Search papers, labs, and topics across Lattice.
4
0
6
4
Self-improving coding agents can achieve unprecedented efficiency and generalizability by learning from multiple task trajectories simultaneously.
Modality Balance can be harnessed as a powerful form of privileged information, leading to substantial gains in reasoning performance for multimodal models.
Reflecting on failed expert trajectories can boost reasoning performance more than tackling problems directly from scratch.
MADA-RL boosts compact model reasoning accuracy by 2% with 16 times fewer trainable parameters, redefining how critics learn from generators.