Search papers, labs, and topics across Lattice.
9
0
12
5
FACET achieves unprecedented task synthesis quality by preserving source intent and ensuring executable state consistency, leading to more reliable terminal agents.
Video-DeepResearch shatters the imitation-learning ceiling, achieving a 64% accuracy in complex video-based question answering, far surpassing existing models.
Executable Blender code transforms text-to-video generation, enabling unprecedented control over scene dynamics and visual fidelity.
Filtering out noise from reasoning traces can drastically improve hallucination detection in large reasoning models, leading to more reliable AI outputs.
LLM agents struggle to juggle multiple tasks when tool use involves realistic delays, revealing critical weaknesses in temporal reasoning and coordination.
Text-based prototypes in vision-language models are fundamentally misaligned with visual data for out-of-distribution detection, but this can be overcome with a novel online pseudo-supervised approach.
Unlock long-context reasoning in LLMs by turning agent trajectories into gold-standard QA pairs, outperforming models 8x larger on challenging reasoning tasks.
Even the best LLMs struggle to effectively discover, refine, and reuse skills over a lifetime of experience, suggesting current benchmarks significantly overestimate real-world agentic capabilities.
Current LLM efficiency metrics fail to capture the true cost of tool use, as measured by wall-clock latency, but a new hardware-aware metric closes the gap.