Search papers, labs, and topics across Lattice.
Affiliation:
3
0
5
Standard single-turn RL produces socially myopic agents, but dynamically scheduling rewards from early relationship-building to mid-dialogue task execution drives a 9.2 percentage point gain in multi-turn goal achievement.
HarnessLens boosts agent performance by up to 13.6% while slashing evaluation costs through smarter, behavior-aware verification.
MiniMax-M2 proves that massive parameter counts don't always translate to better agentic performance; strategic activation of a smaller subset can unlock frontier-level intelligence.