Search papers, labs, and topics across Lattice.
Affiliation:
7
0
10
0
A controlled pure-autoregressive testbed is built and track task-specific validation losses during multimodal continual pretraining across text, image, text-to-image (T2I), and image-to-text (I2T) prediction, showing that better reconstruction does not necessarily yield lower task-specific losses or stronger downstream performance, and that image tokenizer choice can affect text modeling under joint optimization.
Top LLMs can achieve medal-equivalent scores on elite science exams, but they falter on visual grounding and long-horizon consistency.
Discarding irrelevant reasoning tokens can make language models up to 3x faster at test time without sacrificing performance.
Current LLMs struggle to effectively manage memory in multimodal, multi-participant settings, revealing critical gaps in their design.
FiberTune boosts VLA policy performance by preserving critical visual structure, resulting in up to 10.7 percentage points higher success rates on complex tasks.
Fine-grained control over reward signals unlocks significant gains in multi-trait essay scoring, outperforming standard policy optimization techniques.
A single normalization step turns Muon into Muon+, delivering consistent perplexity improvements in LLM pre-training.