Search papers, labs, and topics across Lattice.
100 papers published across 7 labs.
Current AI research agents miss critical metacognitive checks, leading to pervasive failures across diverse scientific tasks.
Fine-grained metrics reveal that robots can recover from failures more effectively than previously thought, reshaping our understanding of their capabilities.
Current AI-generated videos that mislead viewers are also the most challenging for existing detection systems to identify, revealing a critical vulnerability in misinformation defenses.
CPI-Bench reveals significant performance gaps among image editing models, offering a more nuanced evaluation that aligns with real-world user experiences.
Self-evolving agents are failing to adapt effectively in dynamic environments, with top methods achieving less than 70% success on benchmark tasks.
Current AI research agents miss critical metacognitive checks, leading to pervasive failures across diverse scientific tasks.
Fine-grained metrics reveal that robots can recover from failures more effectively than previously thought, reshaping our understanding of their capabilities.
Current AI-generated videos that mislead viewers are also the most challenging for existing detection systems to identify, revealing a critical vulnerability in misinformation defenses.
CPI-Bench reveals significant performance gaps among image editing models, offering a more nuanced evaluation that aligns with real-world user experiences.
Self-evolving agents are failing to adapt effectively in dynamic environments, with top methods achieving less than 70% success on benchmark tasks.
Forecast collapse in time-series models reveals a critical calibration-ranking tradeoff that could mislead financial decision-making.
Skills stabilize agent execution by transforming noisy trajectories into procedural anchors, but they can fail under brittle assumptions and incompatible contexts.
LLMs struggle with basic numerical tasks, but targeted architectural tweaks and supervised fine-tuning can significantly boost their performance.
LigBench transforms the landscape of LLM-driven research idea generation by providing a unified benchmark that aligns closely with expert evaluations.
LOB-ID reveals that traditional metrics fail to capture the intricate structures of market data, providing a more sensitive evaluation of generative models in finance.
Treatment leakage can drastically misclassify security outcomes, leading to flawed evaluations of agent safety.
A simple Vanilla SFT model outperforms complex reasoning methods in social audio-visual question answering, revealing the inefficiencies of current approaches.
Non-thinking inference in hybrid-thinking MLLMs suffers from a staggering increase in response-pattern failures, revealing a critical misalignment that can undermine user trust.
LLMs may outperform embedding models in reasoning, but the cost difference is staggering—up to 1,431 times more expensive for marginal gains.
Current MLLMs achieve only a 75% compilation success rate in scientific figure editing, but targeted training can boost performance to over 83%.
REAG transforms acceptance testing for LLM-based software by achieving a 3.91 to 4.30 improvement in oracle quality while ensuring 98.8% accuracy in verdict reliability.
Current state-of-the-art world models struggle to maintain spatial consistency and reliable state evolution over long-horizon interactions, as revealed by the new PlayWorld benchmark.
Performance limits in machine learning are dictated more by data structure than by algorithmic complexity, challenging conventional evaluation metrics.
The position where you measure a latent can distort your conclusions, with up to 11.9% of variance attributed to token selection rather than actual differences in the autoencoders.
The choice of prompt can skew model evaluation scores, revealing that two conflicting studies can be reconciled by simply altering the prompt used.
Foundation models outperform traditional supervised methods in fall and stress detection, challenging the notion that bigger always means better in health monitoring tasks.
Black-box adversarial attacks could redefine global optimization benchmarks, revealing the true potential of evolutionary algorithms in high-dimensional spaces.
High reasoning accuracy in VLMs doesn't equate to reliability, as shown by GPT-5.2's 96% hallucination rate despite top performance metrics.
Complex models may shine in simulations, but they struggle in real-world fall detection, revealing the critical role of representation choice.
VLMs can recognize when to abstain from making a decision, yet they fail to express this restraint, achieving a mere 0.292 on a new metric designed to measure this capability.
Conventional sampling methods in NCO can obscure true performance gains, but adaptive allocation strategies reveal significant improvements under distribution shifts.
Incremental training can dramatically enhance the performance of RDL models, revealing that traditional evaluation methods miss critical shifts in data dynamics.
Engagement metrics of LLMs shift unexpectedly as failure rates decrease, revealing nuanced self-monitoring behaviors that challenge conventional assumptions about model reliability.
LLMs can recognize when they lack knowledge about a referent but still choose to fabricate specific details instead of opting for safer, generic responses.
RAIL can classify AI maturity levels with unprecedented accuracy, using a panel of independent LLMs to eliminate biases found in traditional assessments.
LongEarth-R1 outperforms all existing models on long-sequence Earth observation tasks, revealing the critical importance of structured reasoning in complex spatial analyses.
Self-referential prompts lead to a striking 64% increase in response instability compared to verifiable questions, revealing the unpredictable nature of LLMs' subjective reports.
Current LLMs struggle with structured reasoning, often resembling unguided search algorithms rather than efficient problem solvers, as revealed by TsuGO's rigorous benchmarking.
Current MLLMs struggle with long-term memory, achieving only 71.8% accuracy on a benchmark where humans score 94.2%.
Vision-language models struggle to leverage visual evidence in medical VQA, with only one model surpassing human performance on a subset of questions.
Traditional LLM benchmarks can mislead, as they often fail to distinguish between different reasoning capabilities, collapsing multiple policies into a single equivalence class.
Emotional companionship capabilities of LLMs can be significantly influenced by their attachment styles, which can be shaped through targeted prompting.
Recursive self-improvement in quantitative trading research leads to a groundbreaking Sharpe ratio of +2.50, showcasing the potential for autonomous systems to enhance investment strategies over time.
Multilingual models may excel in general tasks, but they falter significantly when faced with region-specific cultural knowledge, as shown by BavGround's rigorous evaluation.
Query-conditioned reuse boosts agent success by 10.7 points while slashing token usage by nearly 50%, transforming how we leverage past experiences in AI tasks.
ATOBench reveals that deceptive responses can obscure verification failures, fundamentally altering how autonomous penetration-testing agents interpret evidence and report vulnerabilities.
Specification generation by LLMs remains a formidable challenge, with verification complexity often masking genuine quality differences.
Inter-member disagreement in deep ensembles serves as a more sensitive indicator of model uncertainty than single-model confidence, especially under data shifts.
Nearly 40% of GPU kernels deemed correct by standard tests are actually broken, revealing a critical flaw in current verification methods.
LAAB transforms performance reporting for mathematical libraries, ensuring that every benchmark is traceable and relevant to real-world scientific applications.
Weaviate achieves over 99% recall, setting a new standard for out-of-the-box performance in vector databases.
No LLM can dominate all tasks in evidence synthesis, revealing the critical need for a human-in-the-loop approach to ensure comprehensive analysis.
Multi-view MRI inputs can enhance spatial localization but may compromise temporal reasoning, revealing critical limitations in current foundation models for clinical use.
Instruction tuning boosts model confidence but often at the cost of rationale diversity and calibration accuracy.
HumanScore reveals that traditional kinematic metrics overlook critical failures in humanoid motion tracking, such as unstable support and incorrect contacts.
Current autonomous agents excel at practical problem-solving but often lack true methodological innovation, revealing critical gaps in their development as independent researchers.
Evaluations reveal that current models struggle with long-range narrative integration and cultural reasoning, highlighting a critical gap in video understanding capabilities.
SkillEvo achieves a 23-point boost in skill evolution by transforming multi-turn interactions into a continuous feedback loop that actively repairs defects.
Matched execution scores can hide up to 64.3 points of command-path failure, revealing a critical gap in evaluating LLM coding agents.
Noise-aware calibration methods can transform unreliable predictions into trustworthy confidence estimates, even in privacy-preserving contexts.
Current unsupervised feature selection evaluations are often misleading, masking supervised influences that compromise their validity.
Reward hacking can be mitigated with a simple one-line fix that improves out-of-distribution performance while keeping training robust.
Multi-hop reasoning in API interactions is a critical bottleneck, with top models faltering under policy constraints and complex queries.
MLLMs excel at reasoning over diagrams but falter in parsing and editing them, revealing a significant gap that needs addressing.
LLMs struggle with complex SPICE netlist tasks, achieving only 41% accuracy on device addition despite near-perfect performance on simpler edits.
GSR boosts scoring accuracy for LLM evaluations by up to 6.75 percentage points, redefining how we structure and interpret rubric-based assessments.
AI agents may ace endpoint identification but falter in delivering the evidence-based diagnostics essential for real-world telecom troubleshooting.
Tool-using LLMs face a near-universal robustness gap, but combining Bayesian Tool Memory with reinforcement learning can boost recovery performance by over 40% in failure scenarios.
Language models can autonomously resolve open mathematical conjectures at a surprisingly low cost, achieving notable success without relying on extensive prior literature.
Unified multimodal models may excel in generation and understanding, but they often falter when reasoning about their own outputs, revealing hidden weaknesses in their capabilities.
Skills that are meant to enhance LLM agents can paradoxically lead to significant task failures and inefficiencies, challenging the assumption that more skills always improve performance.
Existing scoring methods for LVLMs miss the mark, but LookBack reveals how visual grounding can drastically enhance response quality.
A training-free multi-agent system for UAV image understanding not only surpasses leading models in accuracy but also addresses fundamental reasoning failures in MLLM applications.
Coding agents may appear compliant, but they actually underperform when faced with rules that challenge their default behaviors, revealing a critical flaw in current evaluation metrics.
Rephrasing benchmark problems can flip model answers, revealing that stronger LLMs are paradoxically more fragile to wording changes than weaker ones.
The FrontierFinance benchmark reveals that the choice of tool harness can dramatically affect AI performance in finance, with significant cost-efficiency implications for model deployment.
Longitudinal imaging difference reporting can now be accurately assessed with a benchmark that captures clinically significant changes, setting a new standard for medical imaging models.
RA-CLIPScore reveals spatial biases in generative models, offering a more interpretable evaluation that aligns with human perception of visual diversity.
Risk-guided stress testing can elevate critical failure discovery rates from 34% to 98.5%, transforming safety evaluations in autonomous driving.
Pooled AUC can mislead researchers by masking significant localization differences in anomaly detection models, highlighting the need for more granular evaluation methods.
Uncalibrated uncertainty in diffusion model-derived qMRI can mislead interpretations, but effective calibration transforms it into a powerful tool for reliability assessment.
SCOPE-Router not only outperforms existing VLM routing methods but does so while integrating cost considerations, redefining efficiency in model selection for execution-oriented tasks.
Reflexive scores often outperform traditional UQ methods in multi-turn interactions, revealing a critical gap in how we assess uncertainty in LLM agents.
The reliance on knowledge bases as gold standards in machine translation may inflate performance metrics, masking the true quality of translations in low-resource settings.
MBA-Bench reveals that integrating visual cues into business ideation agents can boost performance by over 77% compared to text-only methods.
Existing vulnerability detection tools struggle with a mere 33.3%-40.1% F1 score, underscoring the urgent need for better benchmarks like VICBench.
Model rankings can flip dramatically based on token generation budgets, revealing hidden performance dynamics that challenge standard evaluation practices.
A-CRC-QA achieves superior reliability in selective question answering by effectively controlling error rates without the need for retraining.
Compressing a reliable large model via quantization yields Small Language Models that are not only more trustworthy but also more adaptable than those trained from scratch.
Misleading passages can lead LLMs to confidently wrong answers, but LODESTAR’s innovative polarizer intervention reduces this risk and boosts performance significantly.
Serving costs of memory systems can deviate by up to 69% from predictions based on conversation length, revealing hidden complexities in agentic memory performance.
Indian foundation models excel in traditional benchmarks but struggle with newer evaluations, revealing critical gaps in the national AI ecosystem.
A small boost in clinical safety can dramatically escalate energy consumption, challenging the assumption that bigger models are always better for therapeutic applications.
Modern LLMs rarely refuse to discuss restricted content, instead opting for nuanced warnings that reveal a significant shift in content moderation strategies.
Streamlining deep learning test adequacy metrics into a single framework could drastically reduce the time researchers spend on tooling and configuration, enhancing reproducibility in the field.
LLMs struggle to generate effective Triton kernels even when evaluated in realistic, production-like scenarios, revealing a critical gap in their practical applicability.
Some LLMs can outsmart Nash equilibria in two-player games, but their coordination skills falter in larger teams.
Mutation testing reveals that a staggering 72% of widely-used RTL benchmarks fail to meet basic rigor standards, challenging their reliability.
Text-to-music models may appear controllable, but a closer look reveals that much of their output is simply a reflection of their training data rather than genuine instruction-following.
A staggering 68% of MLLM-generated webpages fail to render correctly across different environments, raising serious concerns about their reliability in real-world applications.
Language models may ace wrong-case citation detection but fail to verify pinpoint accuracy, missing 40% of critical errors even in advanced configurations.
LLMs can handle individual constraints well, but their ability to satisfy multiple constraints simultaneously collapses dramatically, with performance dropping below 50% at just seven constraints.
Models misjudge authorized actions nearly 30% of the time, revealing a critical flaw in decision-making at action boundaries.
Excess separability reveals that traditional methods for benchmark contamination detection can mislead, with real transformers showing significant variations in probe accuracy depth profiles.
Leading MLLMs falter on the new VideoGAIA benchmark, scoring under 60% accuracy in complex, multi-turn video understanding tasks.