Search papers, labs, and topics across Lattice.
Affiliation:
9
0
10
20
Many failures in AI agents can only be detected through their reasoning trajectories, revealing a hidden layer of risk in deployment that demands urgent attention.
A staggering 68% of MLLM-generated webpages fail to render correctly across different environments, raising serious concerns about their reliability in real-world applications.
Vulnerabilities spanning multiple functions can be detected more accurately with VulAgentRL, which verifies evidence through a novel Code Property Graph approach.
VisualRepair resolves 196 software issues by intelligently focusing on relevant visual regions, outperforming existing methods and showcasing the power of multimodal understanding in automated repair.
Instruction tuning may enhance LLMs' ability to follow commands, but it compromises their performance in critical code infilling tasks, revealing a hidden cost in AI coding assistants.
Debugging agentic coding systems just got a whole lot easier: TrajAudit pinpoints failures with significantly higher accuracy and lower token cost than existing methods.
LLMs aren't always needed: CelerLog shows you can get SOTA log parsing with a hybrid approach that's up to 18x faster and cuts token costs by 94%.
Automated logging systems may perform unpredictably across programming languages, with framework-anchor matching proving particularly sensitive to language differences.
Log-based anomaly detection models are missing 90% of the picture, but AnomalyGen uses LLMs and static analysis to hallucinate realistic training data and close the gap.