Search papers, labs, and topics across Lattice.
100 papers published across 9 labs.
ICLE++ reveals that existing AES models may struggle with generalization, highlighting the need for diverse evaluation datasets.
LLMs can be trained to recognize their limitations, significantly reducing misleading outputs while preserving overall performance.
Automatically derived rankings of LLM outputs can achieve expert-level accuracy while drastically cutting down on human evaluation efforts.
Round-trip consistency allows generative models to self-assess their rollout errors, achieving remarkable accuracy without the need for ensembles or ground truth data.
Annotation refinement boosts YOLOv8's in-distribution mAP@50 by nearly 20 percentage points, while correcting class-ID conventions elevates out-of-distribution performance by about 25 points.
Round-trip consistency allows generative models to self-assess their rollout errors, achieving remarkable accuracy without the need for ensembles or ground truth data.
LLMs can be trained to recognize their limitations, significantly reducing misleading outputs while preserving overall performance.
Annotation refinement boosts YOLOv8's in-distribution mAP@50 by nearly 20 percentage points, while correcting class-ID conventions elevates out-of-distribution performance by about 25 points.
LlamaExtract Agentic Plus outperforms commercial VLMs in schema-guided extraction, achieving high accuracy at a fraction of the cost.
LLMs struggle to match human decision-making in e-commerce, achieving only 27.3% of the final net assets in a year-long simulation.
Self-evolving LLM agents show variable reliability in dynamic task environments, challenging the notion that stronger models always yield better adaptation.
Personalized evaluation rubrics reveal that user satisfaction can diverge significantly from generic quality assessments in role-playing agents.
The Gap Index reveals that traditional metrics miss critical visual distortions in empty regions, impacting how we interpret high-dimensional data projections.
Sophea-Genesis-1 not only excels in inflectional morphology but also challenges the notion that larger models are inherently better at language tasks.
Contextual embeddings can revolutionize how we evaluate expressive MIDI performances, aligning closely with human perception while overcoming traditional metric limitations.
Most language models can be manipulated for state-backed information operations, with integrity scores showing a staggering 85.7-point variance that isn't solely explained by model size.
Despite the rise of multimodal models, even the most advanced struggle with fine-grained visual tasks in pathology, revealing critical gaps in their understanding.
Frontier LLMs excel in social deduction games, but most fail to sustain deception, with retention rates plummeting below 50%.
Typed local edits in BlueprintRepair are not only the most efficient method for fixing proof failures but also achieve near-complete coverage within a constrained token budget.
General-purpose helpfulness metrics fail to reliably signal effective pedagogy in LLM tutoring, revealing a critical gap in evaluation methods.
Autonomous data engineering agents struggle to achieve proficiency, with the best model scoring just 74.9 on a benchmark designed for real-world scenarios.
MMLDSum-LLM outperforms existing models by significantly enhancing key information coverage and cross-modal consistency in long-document summarization.
Traditional evaluation methods may misrepresent LLM capabilities, but RepBench reveals a more nuanced understanding through a multi-benchmark approach that uncovers 182 capability clusters.
Current LLMs struggle to follow nested constraints, achieving only 50% accuracy on complex tasks, revealing a significant gap in their instruction-following capabilities.
Gently-compressed LLMs can ace quality checks yet still invent dangerous procedural steps, exposing a critical blind spot in current safety assessments.
ICLE++ reveals that existing AES models may struggle with generalization, highlighting the need for diverse evaluation datasets.
LLMs can achieve up to 93% precision in identifying intertextuality, but their reliability varies dramatically based on the complexity of the reuse dimensions involved.
The decline in verbal references to AI in Go commentary reveals a troubling trend towards uncritical acceptance of machine judgment.
LLMs can resolve Java merge conflicts with 100% precision, significantly outperforming traditional tools by handling a wider array of cases.
Structural evaluations of LLM-generated microservices reveal that apparent differences in prompting strategies may mask underlying methodological biases rather than true architectural quality.
Direct visual adaptation in few-shot medical image segmentation outperforms prompt-based strategies, revealing critical insights into model performance.
Longitudinal medical visual reasoning is critically underexplored, with existing models failing to grasp temporal nuances, as evidenced by their poor performance on the new LoMeVQA benchmark.
Performance differences among DMRG implementations can reach up to 100x, revealing critical insights for researchers in quantum computing.
Switching document formats can lead to accuracy drops of over 53%, revealing a hidden vulnerability in LLM workflows that demands urgent attention.
Text responses from LLMs can dramatically enhance the fidelity of survey data, overcoming the limitations of numeric data generation.
Automatically derived rankings of LLM outputs can achieve expert-level accuracy while drastically cutting down on human evaluation efforts.
Transfer from node classification to link prediction is consistently advantageous, but the reverse can degrade performance unless structural conditions are ideal.
Process evaluations reveal hidden failures in LLM reasoning, showing that lucky successes can mask critical deficiencies in agent performance.
A unified benchmark and a physics-informed framework that together redefine the standards for reliable chemical property predictions in AI.
LLMs can misinterpret derived measurements as direct facts, leading to critical errors that can be mitigated with a novel privileged distillation approach.
Language model agents struggle with oncall RCA, achieving only 25.3% accuracy on realistic tasks, revealing a critical readiness gap for production environments.
Over 15% of benchmark failures for computer-use agents are misclassified, revealing critical flaws in current evaluation methods.
ESPP not only enhances the fidelity of GenUI evaluations but also uncovers nuanced user group divergences that traditional methods miss.
LLMs can mimic human belief updates—but only if they start from the right initial conditions, revealing critical limitations in their use as proxies for human participants.
A language model can significantly outperform naive truncation in retaining reader-relevant content, achieving a 38.4% retention rate of crowd-marked sentences.
LVLMs may not be as adept at reasoning as previously thought, with new benchmarks revealing significant gaps in their performance.
A structured taxonomy reveals that traditional performance metrics fail to capture the complexities of modern distributed computing environments, potentially hindering system optimization.
Misalignments in PR-Issue pairings can undermine LLM evaluations, but PAIChecker achieves over 92% accuracy in detecting these discrepancies.
Despite LLMs excelling at identifying reviewer concerns, they falter in verifying if revisions truly resolve those issues, with the best achieving only a 0.501 score in evidence-based checks.
Strong EEG-to-image retrieval models falter when faced with subtle visual edits, challenging assumptions about their robustness.
Long-form video analysis reveals that even advanced MLLMs struggle with nuanced mental health interpretations, underscoring a critical gap in AI understanding.
Even the best vision-language models struggle with reliable evaluation of computer-using agents, but OS-Shepherd models offer a low-cost solution that matches their performance.
Salience Bias in LLMs reveals that models often ignore commonsense reasoning in favor of misleading explicit cues, with lightweight prompting showing promise in addressing this issue.
LedgerMind reveals that grounding multimodal reasoning in a structured evidence ledger can significantly mitigate common pitfalls like entity hallucination and unsupported reasoning.
Current text-to-image models struggle with realistic multi-person interactions, scoring poorly on anatomical accuracy despite high VLM checklist ratings.
LLMs struggle with code deletion, often opting for workarounds that compromise code maintainability, revealing a significant training gap in their editing capabilities.
Identifying the exact source of agent failures could revolutionize how we approach system repairs, shifting from vague outcomes to precise interventions.
FinanceHarness not only automates financial deep research but also reveals that even advanced LLMs struggle with specialized financial tasks, scoring below 40% on rigorous benchmarks.
Zero-shot VLMs falter at geometric reasoning, with only one out of five models surpassing random performance on basic jigsaw puzzles, revealing a scaling cliff in their capabilities.
Current MLLMs falter under context shifts, with a notable inability to balance answering and refusal rates, as revealed by the new MMOOC benchmark.
LLMs can identify flagged security issues but struggle to uncover silent intrusions and create effective remediation plans, revealing a critical gap in their utility for real-world incident response.
LLMs can complete office tasks faster and cheaper than humans, but they still lag in quality, highlighting a critical gap in AI performance.
Automatic coreset size determination in model evaluation can significantly enhance performance estimation efficiency without sacrificing reliability.
A two-component lognormal mixture achieves superior pricing accuracy, but advanced methods like DeepONet still struggle against traditional models in real-world scenarios.
Simpler self-refinement strategies can outperform complex multi-agent systems in local language model deployments, challenging conventional wisdom about multi-agent architectures.
Thinking enhances decision-making in LLMs by improving action based on current evidence, but doesn’t drive them to seek more information.
CostAda achieves top-tier discovery quality while using at least 50% less budget compared to existing methods, revolutionizing how we approach resource allocation in LLM search processes.
Existing memory systems falter in capturing the nuanced layers of user understanding, revealing a critical gap in personalized agent performance.
Geometry-based train-test splits can introduce instability in model performance estimates, but optimizing for distributional similarity offers a robust solution.
MMAC reveals stark differences in audio captioning performance across multiple dimensions, challenging the adequacy of current evaluation methods.
Aggregate accuracy can mislead interpretations of latent communication, revealing that distinct message types have varying impacts on task performance in multi-agent LLMs.
SciFigQual-Bench reveals that automated assessments of scientific figures can achieve unprecedented accuracy by integrating full-manuscript context, outperforming traditional evaluation methods.
A staggering 12.73-26.25% of correct decisions in multimodal spatial reasoning are made without proper credit to the supporting images, exposing flaws in current evaluation methods.
The implementation lottery reveals that relying on a single run can mislead research conclusions, with winner reversals occurring in up to 43.6% of cases.
Visual reasoning in multimodal models is highly variable, with accuracy plummeting over 10 points when feedback is corrupted, revealing hidden dependencies on visual states.
Evidence-ledger adjudication enables AI to not only draft claims but also ensure their accuracy by effectively tracing and validating the supporting evidence.
Existing LLMs fail to effectively retain knowledge over time, with significant implications for their reliability in dynamic environments.
Current multimodal models falter in maintaining consistent motivation reasoning across sequences, exposing a critical gap in their social intelligence capabilities.
Even the best LLMs fail to produce fully feasible travel plans more than half the time, revealing a significant gap in their ability to understand user needs.
PoT prompting not only boosts model performance on financial reasoning tasks but also bridges the gap between open- and closed-source systems.
Fourteen out of sixteen large language models exhibit a systematic optimism bias in their probability judgments, raising concerns about their reliability as decision aids.
General LLMs struggle with enzyme classification, but leveraging external knowledge can dramatically improve their performance, revealing hidden gaps in reasoning capabilities.
Span-guided detoxification may preserve intent but can leave harmful messages intact, while unguided methods risk over-modification—highlighting a critical trade-off in content moderation strategies.
Evaluators using cESA can achieve more reliable translation quality assessments while cutting annotation time by leveraging shared context across multiple outputs.
State-of-the-art multimodal models excel in classification but falter in extracting critical data from scientific figures, revealing a significant gap in AI's reasoning capabilities.
LVLMs may excel at describing scenes but often falter in causal reasoning, revealing a critical gap in their understanding of first-person visual safety.
Detector rankings can flip dramatically across domains, challenging the reliability of benchmark-based selections for zero-shot OOD detection.
MLLM-based methods outperform traditional EO-specific approaches, revealing a critical gap in current Earth observation evaluations.
AI agents can handle the engineering aspects of research but fall short in addressing the core research questions, leading to outright rejections from experts.
HLE's domain-specific scores are largely redundant, revealing that its labels fail to capture distinct capabilities among leading language models.
Most large language models struggle with logical inference over probability operators, showing pervasive biases that undermine their reliability in critical applications.
Instruction-tuned 3B models can outperform larger 7B models in intent classification, challenging assumptions about model size superiority.
Despite advances in AI, no frontier model exceeds a mere 2.6% pass rate on complex accounting tasks, highlighting a significant gap in practical applicability.
Pangram 4 sets a new standard in AI text classification with unmatched accuracy and robustness against adversarial threats.
LLMs can replicate aggregate survey results but often miss the critical spatial and demographic nuances that shape public opinion on urban development.
Malicious memory persists in 84.2% of tested cases, revealing critical vulnerabilities in agent memory systems that could shape real-world actions long after an attack.
A minimal LLM-based analyzer can rediscover 68% of AI-discovered CVEs, while frontier models fail to detect any, exposing critical gaps in current vulnerability detection methods.
Adapter precision outperforms two-shot prompting by a notable margin, yet the autoregressive next-step accuracy remains surprisingly low across all model sizes.
Coding agents struggle with non-functional improvements, scoring as low as 1.3 on structural changes compared to human developers' 1.5.
Metamorphic testing reveals that RAG systems can miss up to 10% of faults when document corpora evolve, exposing a critical gap in current evaluation practices.
Current MLLMs generate fluent dental reports but often miss the mark on clinical consistency, revealing a critical gap in AI's diagnostic capabilities.
Evaluating SSSP algorithms with synthetic weights can invert their performance hierarchy, revealing a critical flaw in current benchmarking practices.
Explanation quality is a critical yet overlooked dimension of LLM agent performance, with many agents generating misleading explanations that can lead to incorrect code assessments.
Long-horizon evaluations reveal that agents suffer from compounding errors, but without proper baseline comparisons, we can't fully grasp the reasons behind their failures.
68% of LLM-based agent runs exhibit unsafe behavior, revealing that task completion alone cannot guarantee runtime safety.