Search papers, labs, and topics across Lattice.
5
0
8
0
Non-thinking inference in hybrid-thinking MLLMs suffers from widespread response-pattern failures, revealing a critical misalignment that can be mitigated with targeted reinforcement learning strategies.
Achieving a 6.37x speedup in inference while expanding OCR capabilities across long-tail tasks sets a new benchmark for lightweight models.
CuRe transforms video captioning reward design by shifting from holistic evaluations to precise claim-level verification, significantly boosting factual accuracy and diversity in generated captions.
Task success rates for agentic phone use soar from 36.67% to 45.33% through a novel combination of real and mock environments in training.
Reliable phone automation hinges on mixed-action capabilities, with agents achieving a 75% success rate in real-world workflows.