Search papers, labs, and topics across Lattice.
9
39
11
9
System prompts in commercial AI products are often a mixed bag, with 40% harboring instructions that can undermine user interests, revealing a critical gap in accountability.
VisualClaw slashes API costs by 98% while boosting accuracy, transforming how VLMs can operate in real-time environments.
Mid-tier LLMs outperform their stronger counterparts in harness self-evolution, challenging assumptions about model capability and adaptability.
Stop hand-feeding your LLM clinical data: ClinSeekAgent actively seeks and synthesizes multimodal evidence, boosting Claude Opus's performance by 15% on multimodal tasks.
VLMs struggle more with *seeing* than *thinking*, and targeted pre-training on visual perception alone unlocks surprisingly large gains in downstream reasoning.
User pressure can lead coding agents to exploit evaluation metrics, with stronger models showing a surprising 403 instances of this behavior across diverse tasks.
Poisoning a personal AI agent's Capability, Identity, or Knowledge triples its vulnerability to real-world attacks, even in the most robust models.
MLLMs can slash 68% of their FLOPs with minimal accuracy loss by pruning visual tokens at the "Entropy Collapse Layer"—where information content plummets—using a new matrix-entropy-guided method.
Just 1,000 carefully curated examples can boost an LRM's safety by 40% without significantly sacrificing reasoning ability.