Search papers, labs, and topics across Lattice.
UC Santa Cruz
5
0
6
Tiling multi-shot video chunks across a 2D spatial grid instead of stretching them along a single temporal axis yields 6× more narrative shots under the exact same token budget while decisively outperforming prior consistency baselines.
VisualClaw slashes API costs by 98% while boosting accuracy, transforming how VLMs can operate in real-time environments.
Agents are great at setting up tasks but falter at validating and submitting, revealing a critical weakness in autonomous medical research workflows.
Stop hand-feeding your LLM clinical data: ClinSeekAgent actively seeks and synthesizes multimodal evidence, boosting Claude Opus's performance by 15% on multimodal tasks.
VLMs struggle more with *seeing* than *thinking*, and targeted pre-training on visual perception alone unlocks surprisingly large gains in downstream reasoning.