Search papers, labs, and topics across Lattice.
Affiliation:
6
0
11
24
This work proposes MATCH, a closed-loop framework for model-aware tool learning with curriculum scheduling and hierarchically gated rewards, and shows consistent improvements across four backbones from two model families.
MERIT-Rank is proposed, a framework that models complementary reasoning trajectories to improve reranking robustness and develops Progressive Rank Policy Optimization (PRPO), a progressive training framework that stabilizes reasoning trajectories while continually improving ranking quality through staged optimization objectives.
Rather than treating agent trajectories as dead post-training demonstrations, researchers can now resurrect thousands of fully executable, verifiable terminal environments directly from tool-execution logs.
LightNav-0 achieves state-of-the-art navigation performance by unifying spatial reasoning and action generation in a single model, eliminating the need for task-specific components.
HBF can significantly boost LLM serving efficiency by enabling more expert replicas and reducing loading times, all while preserving critical execution paths.
MagicSelector achieves unprecedented tool retrieval accuracy by translating vague user instructions into precise subtasks, outperforming state-of-the-art methods.