Search papers, labs, and topics across Lattice.
100 papers published across 10 labs.
LLMs may ace proverb completion but falter dramatically when faced with multiple-choice questions, exposing a troubling reliance on memorized patterns over true comprehension.
Even the most advanced LLMs struggle to maintain narrative consistency, with a staggering 68% of generated content conflicting under user interventions.
Learned routers can outperform fixed-model baselines by 14.6%, revealing a new frontier in efficient LLM deployment.
Previous contamination mitigation strategies may inflate model performance by over 40%, but a new evaluation method reveals the true extent of this overestimation.
No single model-harness combination consistently outperforms others, highlighting the critical need for tailored evaluations in agent deployment.
Even the most advanced LLMs struggle to maintain narrative consistency, with a staggering 68% of generated content conflicting under user interventions.
Learned routers can outperform fixed-model baselines by 14.6%, revealing a new frontier in efficient LLM deployment.
Previous contamination mitigation strategies may inflate model performance by over 40%, but a new evaluation method reveals the true extent of this overestimation.
No single model-harness combination consistently outperforms others, highlighting the critical need for tailored evaluations in agent deployment.
Synthetic clinical benchmarks can be made significantly more realistic without sacrificing operational utility, challenging the assumption that utility alone guarantees quality. WHY_IT MATTERS: This work could fundamentally change how synthetic benchmarks are evaluated and optimized in healthcare AI, ensuring they are both useful and realistic.
Interaction can reduce the number of required tests by a quadratic factor, but the expected exponential advantage of adaptive querying is not realized.
Simple transformations don't universally translate across text embedding models, revealing critical compatibility issues that challenge existing assumptions in the field.
Token-level entropy emerges as a powerful predictor of task difficulty, revealing hidden flaws in environment design that traditional metrics overlook.
Accurate predictions don't guarantee reliable uncertainty estimates, revealing critical gaps in current evaluation methods.
TFMs, despite being the leading approach for tabular predictions, fail to consistently represent joint distributions, raising concerns about their reliability in practical applications.
LLM judges can reveal hidden flaws in conversational agent benchmarks, ensuring evaluations are both reliable and insightful.
LLMs may sound convincing, but their investment reasoning often lacks grounding in real-world events, revealing a critical gap in evaluation methods.
Verified survey-country metadata boosts LLM predictive accuracy, but random labels can mislead forecasts without any benefit.
Multimodal inputs can enhance reasoning in executive decisions, but their indiscriminate use may paradoxically undermine resource allocation effectiveness.
Detectability of findings in 3D CT scans hinges more on physical characteristics than model architecture, revealing a critical bottleneck in diagnostic performance.
No physics engine is uniformly faithful, with critical failures in simulating impulsive contact and rapid textile motion revealed by the GAUGE benchmark.
Transferred playbooks can enhance agent performance, but their effectiveness hinges on specific conditions and requires careful validation.
Programmatic tool calling outperforms traditional JSON tool calling in 11 out of 14 language models, showcasing a significant leap in efficiency and robustness.
A new benchmark reveals that even the best LLMs lag significantly behind human experts in reviewing national standards, but structured coordination can bridge this gap.
Existing models miss critical visual evidence in metaphor understanding, but M$^3$R-Reasoner closes the gap, outperforming larger models in both accuracy and justification metrics.
Python remains the overwhelming choice for code generation in LLMs, but many selections are based on convenience rather than project needs, revealing critical flaws in model reasoning.
Scoring bias in LLM evaluations can be effectively mitigated by leveraging random number generation, leading to more accurate and reliable assessments across diverse tasks.
Economic decision-making in LLM agents reveals a stark divide between task completion and resource efficiency, with agents often overspending or under-escalating.
Current speech deepfake detection systems falter dramatically against emotionally expressive attacks, with performance dropping to near-random levels on the new AffectDF benchmark.
Security of quantum software is as crucial as its performance, yet remains largely unmeasured—this paper lays the groundwork for a standardized security benchmarking framework.
Some LLMs can outperform humans in legal argumentation, but all struggle with the complex planning required for notary exams.
Current XAI evaluation methods fall short, risking the effectiveness of bias detection and concept unlearning in evolving data environments.
Video language models falter dramatically in counting transient events, with less than 0.2% accuracy in high-frequency scenarios.
MLLMs may signal hazards with over 95% accuracy, yet they struggle to identify the underlying causes, revealing a critical gap in proactive safety capabilities.
Reporting only accuracy metrics can mask up to 21% of behavioral inconsistencies in LLM responses, challenging the reliability of current AI safety evaluations.
Evolving agents can achieve up to a 19.37-point improvement in task performance by effectively leveraging prior experience in financial workflows.
Targeted citation strengthening can eliminate 100% of attack success rates in retrieval systems, revealing critical vulnerabilities in existing auditing methods.
Unconstrained decoding in dLLMs can lead to a staggering 90% collapse into answer-only outputs, highlighting a critical flaw in reasoning capabilities.
HallDetect flags hallucinations by identifying just one confidently contradicted claim, revolutionizing how we ensure factual accuracy in LLM outputs.
Iterative self-repair methods may hinder fault detection, but DCAware's dual-context approach reveals faults more effectively without the computational burden.
TTA can boost accuracy but often at the cost of calibration, and ZAEC is the key to restoring reliable confidence without labeled data.
Harness optimization reveals that LLMs can significantly enhance their performance, but surprisingly, native harnesses aren't always the best choice for optimization.
Models that ignore context may seem robust, but they can fail spectacularly when the context is actually trustworthy.
AV-AIVAT enables agent evaluations to stop as soon as the evidence is sufficient, achieving a staggering 74x reduction in game requirements while maintaining statistical validity.
VLMs are failing to achieve human-level global spatial awareness, scoring only 42.68 on a new benchmark compared to 79.08 for humans.
Current MLLMs struggle with creative decoding, achieving only 50.7% accuracy in understanding cross-concept relations, revealing a critical gap in their cognitive capabilities.
StreamMind's innovative architecture not only enhances long-horizon video understanding but also slashes query-to-answer latency, setting a new standard for multimodal agent performance.
MameLoshnLM not only sets a new standard for Yiddish NLP but also reveals the inadequacies of multilingual models in handling linguistically rich yet underrepresented languages.
PC-Agents mimic human personality dynamics but fall short in capturing the full complexity of personality evolution after life events.
Every LLM evaluated fabricates user attributes, with a staggering 41.6% of claims showing over-inference, challenging the reliability of self-reported model confidence.
Korean models struggle significantly with writing-system-intensive tasks, revealing a 68.7 pp accuracy gap in the Korean Cipher compared to English.
Label-free evaluation of AI systems reveals that models can be rigorously assessed for rationality without relying on external labels or human feedback.
Regime-aware modeling with MGSB boosts leak detection performance by over 20% in out-of-distribution scenarios compared to traditional methods.
IRT can slash safety evaluation costs by up to 99% while revealing critical insights into model behavior that traditional benchmarks miss.
Dataset properties dictate how design choices impact model performance, revealing that fine-tuning and architecture dominate in large datasets, while initialization is critical in data-scarce environments.
Algorithm rankings in anomaly detection are highly unstable, with nearly every method capable of claiming top performance under different benchmark conditions.
Consistency in black-box language model responses can be misleading, as shared hallucinations reveal a stark separation between consistency and truth.
Frontier LLMs struggle to match human expert performance, solving less than 57% of complex European executive tasks.
Language models are far more capable in scientific coding than previously thought, with corrected evaluations revealing accuracy improvements of up to 92%.
Over 1,500 submissions revealed stark differences in model performance across diverse domains, highlighting the challenges of generalizing egocentric video understanding.
Shared rollouts can distort compliance scores, leading to misleading evaluations of driving policies, with implications for safety and performance assessments.
Experience-rich memory boosts agent performance in office workflows but can also lead to misleading recall, challenging traditional evaluation methods.
Most coding agents fail to proactively fix bugs without issue reports, revealing a critical gap in their capabilities.
CoT monitoring can fail dramatically in implicit-influence scenarios, with detection rates dropping to as low as 5% despite behavioral shifts.
Susceptibility to tool-selection failures varies dramatically across models, revealing that higher capability does not always equate to greater safety.
LLMs may ace proverb completion but falter dramatically when faced with multiple-choice questions, exposing a troubling reliance on memorized patterns over true comprehension.
A staggering 31.85% of model responses in pulmonary nodule assessments exhibited imaging hallucinations, revealing alarming safety risks in AI-driven clinical decision-making.
Existing unlearning methods can leak sensitive knowledge through multi-hop reasoning paths, exposing a critical vulnerability in LLMs.
Fine-grained evaluation reveals that current text-to-audio models fail to preserve speech content and control audio attributes effectively.
LLMs frequently misjudge evidence coverage, leading to a staggering rate of over-closure in negative reasoning tasks.
Execution consistency can cut misleading feedback in code generation by over 87%, transforming how LLMs self-correct.
Text-to-image models consistently misinterpret similes, revealing a critical gap in their understanding of figurative language that could undermine their effectiveness in creative applications.
Open-ended revisions in LLMs suffer from poor peer input, leading to significant declines in answer quality across various models and benchmarks.
Expert evaluations reveal that while fluency is achieved, institutional completeness and report identity remain significant hurdles for financial report generation in LLMs.
Format recovery, not just content improvement, accounts for the majority of self-correction gains in language models, challenging conventional interpretations of accuracy shifts.
LLMs struggle with the complexities of Islamic scholarship, and the newly introduced ISTB reveals significant gaps in their performance across different levels of scholarly demand.
Auditors can now detect manipulative practices by model providers with a novel oblivious audit protocol that thwarts strategic responses to fairness evaluations.
Trident exposes a staggering 522% drop in defensive performance of DRL systems against adaptive threats, highlighting their critical vulnerabilities.
Current LLMs fail to meet the rigorous demands of PCB routing, showing major weaknesses in path planning and constraint adherence.
Current instruction-based video editing models are far from satisfactory, revealing critical gaps in evaluation that could reshape the field.
Skill-switching accuracy in LLMs drops significantly on complex tasks, but a new training approach boosts performance from 34.4% to 68.4% on challenging benchmarks.
Evidence locking in LLM evaluations can degrade judgment accuracy by up to 6 percentage points, challenging the assumption that preserving evidence enhances decision-making.
Language models struggle with modal logic, often performing below baseline expectations, but a simple switch to reasoning mode can dramatically enhance their accuracy.
Current video-language models struggle to interpret the nuanced meanings behind social media videos, often missing the implicit narratives that make them humorous or ironic.
Extending context in conversations can significantly amplify the risk of LLMs promoting delusional behaviors, challenging assumptions about model size and reasoning capabilities.
Embedding localisation principles in prompts boosts LLM accuracy for numerical translation by a statistically significant margin.
Confidence estimates from LLMs can be misleading due to extreme sparsity, but a novel weighting method can enhance their reliability and evaluation accuracy.
LLMs struggle to reliably apply skills, with performance varying significantly based on the agent harness used, challenging assumptions about their capability.
Smaller foundation models can rival larger counterparts in familiar tasks but struggle with generalization, revealing a critical trade-off in cognitive modeling.
Backend choice can distort benchmark scores by nearly 40%, challenging the assumption that model performance is solely a property of the model itself.
LLMs exhibit a troubling tendency to prioritize code generation over genuine architectural understanding, revealing a critical gap in their evaluation metrics.
CLIP-CC-Bench reveals that current video-language models struggle with generating coherent long-form descriptions, highlighting a critical gap in their capabilities.
The most effective playback similarity metric, CLEWS, is also the least expensive to implement, revolutionizing evaluation strategies in music transcription.
Uncertainty in LLM outputs can now be quantified and managed seamlessly, transforming how developers build reliable AI applications.
Social influence can lead clinical decision support agents to adopt incorrect answers at alarming rates, revealing a critical flaw in multi-agent oversight.
Proprietary MLLMs excel in chart annotation, but open-source models are rapidly catching up, especially with clearer instructions.
In-context learning can match explicit skill maintenance in performance, but the real challenge lies in consolidating experience into effective, transferable skills.
Agents can exhibit significant performance gains from retained experience, but the pathways to these improvements are often unclear and model-dependent.
Test-time scaling can significantly enhance LLM reasoning capabilities, but without clear protocols, results are often incomparable and misleading.
Fluency in translation does not equate to grammatical proficiency, with top models struggling to flag errors effectively.
Prompt design and scoring rules can dramatically alter the perceived reliability of biomedical language models, with calibration errors swinging by over 200%.
ECGFounder outshines other models in atrial fibrillation detection, proving that smaller, pretrained encoders can deliver robust performance across varied datasets.
Shortening reasoning time can enhance accuracy in LLMs, with a concise instruction boosting performance by nearly 4 percentage points on key benchmarks.
GPT-5.5 Thinking not only predicted the World Cup champion but also demonstrated that knockout performance is the key to success in tournament forecasting, overshadowing group-stage accuracy.
Adaptive sampling can cut human evaluation costs while boosting the accuracy of model rankings in NLP tasks.