Search papers, labs, and topics across Lattice.
9
0
11
13
Executable vector graphics enable MLLMs to achieve human-like spatial reasoning through a structured visual workspace.
Code generation models may excel at passing traditional benchmarks but falter dramatically when faced with novel algorithmic challenges, revealing a hidden weakness in their reasoning capabilities.
MOPD achieves superior capability integration in LLMs by distilling knowledge from multiple RL teachers without losing performance, setting a new standard for post-training methods.
Models may score well on benchmarks but often fail to meet strict perceptual requirements, revealing a hidden brittleness in multimodal evaluations.
Achieve detailed tunnel defect inspection without any training by visually recalibrating foundation model proposals to overcome tunnel-specific interference.
RLVR, the dominant training paradigm for audio language models, may be turning them into unfeeling "answering machines" that excel on benchmarks but fail the vibe check.
Forget noisy pseudo-labels: SpatialEvo unlocks self-supervised 3D spatial reasoning by generating perfectly accurate training data directly from scene geometry.
VLMs can achieve superior visual reasoning by dynamically decomposing queries, extracting premise-conditioned visual latents, and reasoning through grounded rationales, outperforming even multimodal CoT methods.
Even reward models that get the right answer can be dangerously wrong in their reasoning, leading to worse RLHF outcomes, but R-Align fixes this by explicitly aligning rationales with gold standard judgments.