Search papers, labs, and topics across Lattice.
5
0
7
7
Coding agents can now be rigorously evaluated on their ability to navigate ambiguity and build software from incomplete requirements, a critical shift in assessing their real-world applicability.
Current search agents fall short of user expectations, with a new benchmark revealing critical gaps in their performance on everyday tasks.
Code agents struggle with evolving user requirements, revealing a 38-point gap in performance across leading LLMs when faced with iterative feedback.
LongCat-Next shatters the language-centric paradigm by unifying text, vision, and audio into a single autoregressive model with minimal modality-specific design, finally reconciling understanding and generation in discrete vision modeling.
Current memory systems like RAG and long-context LLMs stumble in AMemGym's interactive long-horizon conversations, revealing critical performance gaps in maintaining consistent user state.