Search papers, labs, and topics across Lattice.
7
0
10
7
Coding agents can now be rigorously evaluated on their ability to navigate ambiguity and build software from incomplete requirements, a critical shift in assessing their real-world applicability.
OPERA's intrinsic reward mechanism enables LLMs to achieve new heights in open-ended reasoning, outperforming proprietary models without the pitfalls of biased supervision.
Current search agents fall short of user expectations, with a new benchmark revealing critical gaps in their performance on everyday tasks.
Code agents struggle with evolving user requirements, revealing a 38-point gap in performance across leading LLMs when faced with iterative feedback.
Long-context LLM rankings dramatically reshuffle when evaluated across a range of context lengths and capabilities, proving that a single headline score is misleading.
Interactive world models still have a long way to go: a comprehensive benchmark reveals that even state-of-the-art models struggle to consistently perform well across video quality, interaction adherence, and physics compliance.
LongCat-Next shatters the language-centric paradigm by unifying text, vision, and audio into a single autoregressive model with minimal modality-specific design, finally reconciling understanding and generation in discrete vision modeling.