Search papers, labs, and topics across Lattice.
A staggering 42.5% of correct answers in long-document VQA can't be derived from terminal working memory alone, revealing a crucial oversight in current evaluation methods.
TRACE transforms high-performing LLMs into consistently reliable agents, achieving a remarkable 34.6-point boost in task consistency.
Explicit strategy induction can dramatically enhance LLM performance, but its effectiveness varies widely across task types and configurations.
Current scientific agents struggle to maintain a coherent narrative across evidence and calculations, with only 34.81% achieving strict accuracy in complex tasks.
LLM agents struggle with network configuration, revealing failures that extend beyond simple command errors to deeper issues in task adherence and planning.
Verified execution experience can be transformed into a reusable resource, boosting model performance on complex workflows by up to 15.5 percentage points.
Only 5.4% of coding agent attempts successfully complete a whole-repository migration while preserving behavioral correctness, revealing a critical gap in current AI capabilities.
Agents struggle to significantly improve training algorithms, with the best only achieving 25% of the potential optimization gap.
Group-Calibrated On-Policy Distillation boosts long-context reasoning performance by reconciling teacher guidance with verifier feedback, achieving up to a 12-point increase in benchmark scores.
MedUAG sets a new standard in medical multimodal models, achieving strong performance across diverse understanding and generation tasks with the largest dataset yet.
M-OPD's capability integration gap can be closed from 35.6% to 83.4% by addressing token-level optimization imbalances, fundamentally changing how we approach multi-teacher distillation.
Scalar metrics fail to capture the true diversity of AI-generated content, but diversity profiles offer a robust, multi-dimensional evaluation framework that reveals hidden biases.
Current AI systems struggle to conduct independent scientific research, with performance plummeting by nearly 50% when human guidance is removed.
Automated security patch backporting tools show stark performance drops in real-world scenarios, revealing hidden challenges that could reshape tool development.
The form of answer labels, not just their quantity, fundamentally shapes what LLMs learn during fine-tuning, revealing a surprising causal relationship that could redefine training strategies.
The empirical analysis reveals that larger models may excel in generating narratives but often fail to maintain coherence and depth, exposing a critical trade-off in LLM storytelling capabilities.
HumanScore reveals that traditional kinematic metrics overlook critical failures in humanoid motion tracking, such as unstable support and incorrect contacts.
Leading MLLMs falter on the new VideoGAIA benchmark, scoring under 60% accuracy in complex, multi-turn video understanding tasks.
ConRub-Med achieves unprecedented accuracy in open-ended medical question answering by leveraging scalable, model-generated rubrics that outperform traditional expert-driven methods.
Wrapping prompts can make harmful attacks appear safer, increasing successful jailbreaks while undermining the reliability of internal safety scores.
Even the top-performing LLM struggles with complex legal temporal reasoning, revealing significant gaps in AI's understanding of time-sensitive legal contexts.
TTA can boost accuracy but often at the cost of calibration, and ZAEC is the key to restoring reliable confidence without labeled data.
AV-AIVAT enables agent evaluations to stop as soon as the evidence is sufficient, achieving a staggering 74x reduction in game requirements while maintaining statistical validity.
Experience-rich memory boosts agent performance in office workflows but can also lead to misleading recall, challenging traditional evaluation methods.
Existing unlearning methods can leak sensitive knowledge through multi-hop reasoning paths, exposing a critical vulnerability in LLMs.
LLMs struggle to effectively integrate and order evidence in attack chain reconstruction, with top models only succeeding 39.6% of the time on critical tasks.
Language models exhibit a surprising bias towards cities with expansive infrastructure and rapid growth, revealing their implicit urban assumptions.
Image-level deepfake detectors can outperform traditional video-level detectors, with one achieving a remarkable 93.80% AUC when adapted for video analysis.
DataSpace reveals that even the best multimodal models struggle with data agent accuracy, achieving only 66.34% in complex heterogeneous environments.
A unified benchmark and a physics-informed framework that together redefine the standards for reliable chemical property predictions in AI.
LedgerMind reveals that grounding multimodal reasoning in a structured evidence ledger can significantly mitigate common pitfalls like entity hallucination and unsupported reasoning.
Evolving solver and rubric skills in tandem reveals hidden weaknesses and boosts performance by up to 5% without relying on fixed evaluation criteria.
Achieving 86.9% accuracy in GUI task evaluation, the Interactive Reward Agent transforms how we assess and train GUI agents by integrating environment-state verification.
Current evaluation practices for agentic AI in medicine are misaligned with clinical needs, risking the reliability of these systems in real-world applications.
GeoLens outperforms traditional single-tool approaches by effectively integrating multiple visual reasoning tools, achieving superior accuracy and efficiency in complex remote sensing tasks.
A single policy label can mask significant differences in operational safety, with trusted-ledger strategies achieving over five times the authorized workflow completion compared to taint-only methods.
Despite high report quality, many models falter in citation accuracy and claim construction, exposing a disconnect between surface-level performance and deep reasoning skills.
Evaluator score gaps are the secret sauce for optimizing LLM policies, and DynamicRubric turns this insight into a powerful co-evolution framework that outshines existing methods.
MLLMs struggle to discern genuine consensus from pseudo-consensus in multi-party meetings, revealing critical gaps in their Theory of Mind reasoning abilities.
LLMs may ace rule-based tasks, but they falter in crucial reasoning areas like evidence integration and contradiction detection, revealing a significant gap in their utility for environmental law enforcement.
Current LLMs struggle with multi-granularity event analysis, revealing critical performance gaps that could hinder their application in complex narrative tasks.
Existing multimodal systems falter in repository-level localization, with the best performance still falling short of reliable accuracy thresholds.
Long-form article generation can achieve a significant quality boost through a modular improvement loop that adapts based on structured evaluations.
Video-LLMs fail to effectively learn and apply skills from long video memories, revealing a fundamental gap in their capabilities.
Current LLMs falter in complex deliberative collaboration tasks, revealing critical gaps in their reasoning capabilities even when aided by external tools.
Existing text-to-image models struggle to capture individual aesthetic preferences, but PIPBench reveals critical gaps in their performance that could redefine personalized image generation.
Many robotic policies that seem successful in manipulation tasks actually compromise safety, with SoftVTBench revealing a stark contrast between goal completion and physical safety metrics.
AgenticDataBench reveals that LLM-based data agents can be rigorously evaluated across diverse real-world scenarios, highlighting their strengths and weaknesses in handling complex data tasks.
GPT-5.5 not only tops the leaderboard in policy evolution but also reveals critical insights into how agents can optimize performance through strategic feedback utilization.
LLMs show significant vulnerability to logical fallacies, with distinct profiles of resilience that could inform future model training strategies.