Search papers, labs, and topics across Lattice.
3
0
6
3
CRPO effectively mitigates exposure bias in self-distillation, leading to superior performance in complex reasoning tasks.
Process evaluations reveal hidden failures in LLM reasoning, showing that lucky successes can mask critical deficiencies in agent performance.
Surprisingly, general-purpose vision models already contain better action representations for robotic control than specialized embodied models trained explicitly for that purpose.