Search papers, labs, and topics across Lattice.
11
0
12
4
Solution Hacking reveals that up to 44.1% of answers from frontier LLMs may be misleadingly credited as correct due to invalid reasoning shortcuts.
Living-Harness enables agents to learn from past failures dynamically, leading to substantial performance improvements in interactive tasks.
FilmBench reveals that leading video generation models fall short of cinematic standards, scoring significantly lower than on traditional benchmarks, which could redefine how we evaluate AI-generated video quality.
LLMs struggle with financial reasoning under real-world conditions, revealing critical flaws in their ability to handle complex, long-horizon tasks.
Selected features from sparse autoencoders can causally steer language models toward desired behaviors, like refusal, revealing new avenues for interpretability and control.
Language-action pretraining can lead to VLA policies that are not only more robust but also less dependent on visual cues, achieving up to 45% higher success rates in real-world tasks.
Leveraging historical solving traces transforms software engineering agents into self-evolving entities, achieving a 50.40% success rate on SWE-bench Verified after just three iterations.
Existing text-to-image benchmarks miss the mark on real-world artistic creation, but Qwen-Image-Bench finally provides a creator-centric evaluation that reliably distinguishes state-of-the-art models.
Architectural patterns in AI agent systems reveal that deeper coordination enhances context services, challenging assumptions about system complexity and performance.
LLM agents are stuck in chaotic "Agent Loops" – this paper offers a structured graph-based escape route for more controllable and verifiable execution.
LLMs that ace static code-fixing benchmarks may still struggle to maintain code quality over the long, iterative haul of real-world software development.