Search papers, labs, and topics across Lattice.
Institute of Software, Chinese Academy of Sciences
12
0
14
13
Advanced autonomous agents struggle with fundamental document manipulation tasks, revealing critical gaps in their operational capabilities.
Transforming historical sequences into a powerful resource, PraMem significantly improves long-horizon behavior prediction beyond existing methods.
ReasoningLens turns the opaque reasoning of large models into clear, actionable insights, enabling researchers to pinpoint errors and optimize performance like never before.
Current AI agents falter in autonomous development, revealing critical gaps in robustness and alignment as they struggle against human-engineered solutions.
ADR transforms the landscape of code task generation, enabling LLMs to tackle genuinely novel and challenging coding problems that enhance their performance.
Forget scraping – this work shows you can generate high-quality, executable terminal environments from scratch to train language agents that outperform models trained on scraped data.
MLLMs can't grasp metaphors in videos, revealing a surprising gap in their high-order cognitive abilities compared to humans.
LLMs trained with ScaleBox, a new high-fidelity code verification system, substantially outperform those trained with heuristic matching, suggesting current RLHF methods are bottlenecked by verification quality.
Multilingual RAG systems are systematically suppressing "answer-critical" documents in non-English languages, crippling their ability to leverage global knowledge.
Forget text-dominance: Today's Omni-modal LLMs surprisingly favor visual inputs, creating new challenges for cross-modal reasoning.
LLMs exhibit a "Utopian bias" when simulating human behavior, converging towards an unrealistic "positive average person" and failing to capture individual differences and long-tail behaviors.
LLMs trained with reinforcement learning from verifiable rewards (RLVR) become overconfident in incorrect answers, but a simple fix—decoupling reasoning and calibration objectives—can restore proper calibration without sacrificing accuracy.