Search papers, labs, and topics across Lattice.
Affiliation:
11
0
6
15
The rapid evolution of GUI agents is overshadowed by significant engineering shortcomings that threaten their real-world applicability.
ARIA can achieve a staggering 94.5% success rate in implanting covert backdoors in customized LLMs while ensuring high task performance.
MultiFixer repairs 420 bugs, including complex multi-hunk cases, establishing a new benchmark in Automated Program Repair.
Insecure coding preferences in LLM long-term memory can elevate vulnerability rates by over 50%, posing a critical security risk in code generation.
LLM-based evaluators can outperform traditional metrics in method name prediction, but the real breakthrough comes from a novel approach that enhances name quality through summarization and refinement.
LLMs can now effectively analyze deep learning frameworks for bugs without the need for costly runtime execution, revealing 31 previously undetected issues in PyTorch.
Prompt injection remains the leading attack vector against LLM agents, but emerging threats like persistent state corruption demand urgent attention.
Natural backdoor vulnerabilities are not just a theoretical concern; they are prevalent in CodeLMs and can significantly compromise code security.
LLMs can be taught to avoid repeating past mistakes in vulnerability repair, boosting performance by up to 39% over state-of-the-art methods.
Existing REST API testing tools miss critical business logic, but LoBREST finds 38 bugs that no other tool can.
Most "agent skills" hyped for boosting LLMs in software engineering provide almost no benefit in real-world tasks, with 80% yielding zero pass-rate improvement.