Search papers, labs, and topics across Lattice.
Evaluation methodology for AI systems, benchmark design, capability measurement, and safety evaluations.
#5 of 24
5
Synthetic clinical benchmarks can be made significantly more realistic without sacrificing operational utility, challenging the assumption that utility alone guarantees quality. WHY_IT MATTERS: This work could fundamentally change how synthetic benchmarks are evaluated and optimized in healthcare AI, ensuring they are both useful and realistic.
Interaction can reduce the number of required tests by a quadratic factor, but the expected exponential advantage of adaptive querying is not realized.
Simple transformations don't universally translate across text embedding models, revealing critical compatibility issues that challenge existing assumptions in the field.
Token-level entropy emerges as a powerful predictor of task difficulty, revealing hidden flaws in environment design that traditional metrics overlook.
Accurate predictions don't guarantee reliable uncertainty estimates, revealing critical gaps in current evaluation methods.
TFMs, despite being the leading approach for tabular predictions, fail to consistently represent joint distributions, raising concerns about their reliability in practical applications.
LLM judges can reveal hidden flaws in conversational agent benchmarks, ensuring evaluations are both reliable and insightful.
LLMs may sound convincing, but their investment reasoning often lacks grounding in real-world events, revealing a critical gap in evaluation methods.
Verified survey-country metadata boosts LLM predictive accuracy, but random labels can mislead forecasts without any benefit.
Multimodal inputs can enhance reasoning in executive decisions, but their indiscriminate use may paradoxically undermine resource allocation effectiveness.
Detectability of findings in 3D CT scans hinges more on physical characteristics than model architecture, revealing a critical bottleneck in diagnostic performance.
No physics engine is uniformly faithful, with critical failures in simulating impulsive contact and rapid textile motion revealed by the GAUGE benchmark.
Transferred playbooks can enhance agent performance, but their effectiveness hinges on specific conditions and requires careful validation.
Programmatic tool calling outperforms traditional JSON tool calling in 11 out of 14 language models, showcasing a significant leap in efficiency and robustness.
A new benchmark reveals that even the best LLMs lag significantly behind human experts in reviewing national standards, but structured coordination can bridge this gap.
Existing models miss critical visual evidence in metaphor understanding, but M$^3$R-Reasoner closes the gap, outperforming larger models in both accuracy and justification metrics.
Python remains the overwhelming choice for code generation in LLMs, but many selections are based on convenience rather than project needs, revealing critical flaws in model reasoning.
Scoring bias in LLM evaluations can be effectively mitigated by leveraging random number generation, leading to more accurate and reliable assessments across diverse tasks.
Economic decision-making in LLM agents reveals a stark divide between task completion and resource efficiency, with agents often overspending or under-escalating.
Current speech deepfake detection systems falter dramatically against emotionally expressive attacks, with performance dropping to near-random levels on the new AffectDF benchmark.
Security of quantum software is as crucial as its performance, yet remains largely unmeasured—this paper lays the groundwork for a standardized security benchmarking framework.
MameLoshnLM not only excels in performance but also reveals the shortcomings of existing multilingual models in capturing the nuances of low-resource languages like Yiddish.
VLMs can achieve high local spatial accuracy but fail to develop a coherent global spatial understanding, highlighting a critical gap in current AI capabilities.
Some LLMs can outperform humans in legal argumentation, but all struggle with the complex planning required for notary exams.
Current XAI evaluation methods fall short, risking the effectiveness of bias detection and concept unlearning in evolving data environments.
Video language models falter dramatically in counting transient events, with less than 0.2% accuracy in high-frequency scenarios.
MLLMs may signal hazards with over 95% accuracy, yet they struggle to identify the underlying causes, revealing a critical gap in proactive safety capabilities.
Reporting only accuracy metrics can mask up to 21% of behavioral inconsistencies in LLM responses, challenging the reliability of current AI safety evaluations.
Evolving agents can achieve up to a 19.37-point improvement in task performance by effectively leveraging prior experience in financial workflows.
Targeted citation strengthening can eliminate 100% of attack success rates in retrieval systems, revealing critical vulnerabilities in existing auditing methods.
Unconstrained decoding in dLLMs can lead to a staggering 90% collapse into answer-only outputs, highlighting a critical flaw in reasoning capabilities.
HallDetect flags hallucinations by identifying just one confidently contradicted claim, revolutionizing how we ensure factual accuracy in LLM outputs.
Iterative self-repair methods may hinder fault detection, but DCAware's dual-context approach reveals faults more effectively without the computational burden.
TTA can boost accuracy but often at the cost of calibration, and ZAEC is the key to restoring reliable confidence without labeled data.
StreamMind achieves superior performance in streaming video understanding by effectively managing the trade-off between real-time interaction and long-term memory retention.
Harness optimization emerges as a critical, yet underexplored capability of LLMs, revealing that the choice of coding harness can significantly impact performance outcomes.
Models that ignore context may seem robust, but they can fail spectacularly when the context is actually trustworthy.
AV-AIVAT enables agent evaluations to stop as soon as the evidence is sufficient, achieving a staggering 74x reduction in game requirements while maintaining statistical validity.
Every LLM evaluated fabricates user attributes, with a staggering 41.6% of claims showing over-inference, challenging the reliability of self-reported model confidence.
Korean models struggle significantly with writing-system-intensive tasks, revealing a 68.7 pp accuracy gap in the Korean Cipher compared to English.
Label-free evaluation of AI systems reveals that models can be rigorously assessed for rationality without relying on external labels or human feedback.
Regime-aware modeling with MGSB boosts leak detection performance by over 20% in out-of-distribution scenarios compared to traditional methods.
IRT can slash safety evaluation costs by up to 99% while revealing critical insights into model behavior that traditional benchmarks miss.
Dataset properties dictate how design choices impact model performance, revealing that fine-tuning and architecture dominate in large datasets, while initialization is critical in data-scarce environments.
Algorithm rankings in anomaly detection are highly unstable, with nearly every method capable of claiming top performance under different benchmark conditions.
Consistency in black-box language model responses can be misleading, as shared hallucinations reveal a stark separation between consistency and truth.
Frontier LLMs struggle to match human expert performance, solving less than 57% of complex European executive tasks.
Language models are far more capable in scientific coding than previously thought, with corrected evaluations revealing accuracy improvements of up to 92%.
Over 1,500 submissions revealed stark differences in model performance across diverse domains, highlighting the challenges of generalizing egocentric video understanding.
Shared rollouts can distort compliance scores, leading to misleading evaluations of driving policies, with implications for safety and performance assessments.