Search papers, labs, and topics across Lattice.
Nanjing University
10
0
12
3
Game solvers can teach LLMs how to make better decisions in long-horizon tasks by providing actionable turn-level feedback, leading to superior performance in complex environments.
Token-level credit allocation using counterfactual replay boosts performance by an average of 4.4 percentage points without the need for an auxiliary scoring model.
Agents trained on static benchmarks falter dramatically in open-world settings, revealing a critical gap in their adaptability to real-world complexities.
MLLMs face severe scalability limitations, with performance dropping by up to 80% on complex visual reasoning tasks, revealing a critical gap in their structural reasoning capabilities.
Multi-agent LLM systems can adapt to new tasks without sacrificing structural integrity, thanks to a novel framework that guarantees role evolution preserves key operational contracts.
LoopLMs don't reliably scale at test time because of an inherent stability vs. effectiveness trade-off, but a new training method can fix that.
Current reward models struggle to distinguish good vs. bad agent behavior in complex tool-using scenarios, especially over long horizons, revealing a critical gap in alignment research.
An 8B parameter model, RideJudge, outperforms 32B baselines in ride-hailing dispute adjudication by aligning visual semantics with evidentiary protocols, achieving 88.41% accuracy.
Current MLLMs struggle with even basic route planning in remote sensing, highlighting a critical gap in their ability to translate perception into action in complex, real-world scenarios.
LLM agents can learn to solve complex, long-horizon tasks much more effectively by using themselves as post-hoc critics to refine Q-values through hindsight reasoning.