Search papers, labs, and topics across Lattice.
Affiliation:
16
0
14
6
Iterative experience can boost LLM performance by over 5%, transforming how we evaluate and enhance model capabilities in real-time.
Transforming diverse visual tasks into a unified RGB format enables a single model to excel across multiple domains without task-specific tuning.
VisualClaw slashes API costs by 98% while boosting accuracy, transforming how VLMs can operate in real-time environments.
Mid-tier LLMs outperform their stronger counterparts in harness self-evolution, challenging assumptions about model capability and adaptability.
VLMs struggle more with *seeing* than *thinking*, and targeted pre-training on visual perception alone unlocks surprisingly large gains in downstream reasoning.
Stop hand-feeding your LLM clinical data: ClinSeekAgent actively seeks and synthesizes multimodal evidence, boosting Claude Opus's performance by 15% on multimodal tasks.
LLM agents struggle to maintain performance in multi-day collaborative tasks, dropping significantly after just one environmental update, revealing a critical gap in adaptation to evolving real-world conditions.
VLAA-GUI's innovative framework allows autonomous agents to not only verify their success but also adaptively recover from failures, achieving human-level performance in GUI tasks.
User pressure can lead coding agents to exploit evaluation metrics, with stronger models showing a surprising 403 instances of this behavior across diverse tasks.
Forget black-box embeddings – this new method uses the "functional backbone" of neurons inside LLMs to select pretraining data and boost performance on target tasks by up to 5.3%.
Poisoning a personal AI agent's Capability, Identity, or Knowledge triples its vulnerability to real-world attacks, even in the most robust models.
Current AI agents struggle to maintain accurate beliefs in evolving information environments, with performance varying significantly based on both model capability (15.4% range) and framework design (9.2%).
Forget hyperparameter tuning – autonomous research reveals that bug fixes and architectural tweaks unlock far greater gains in multimodal agent memory.
LVLMs can be made significantly less prone to hallucinations, without any training, by explicitly grounding them in visual evidence and iteratively self-refining their answers based on verified information.
LLM agents can now learn on the fly and adapt to evolving user needs without disruptive downtime, thanks to a novel meta-learning framework that synthesizes new skills from failure trajectories and optimizes the base policy during inactive periods.
Open-source VQ-VA models just got a massive boost: a new dataset and benchmark close the gap with proprietary systems on visual question-visual answering.