Search papers, labs, and topics across Lattice.
100 papers published across 5 labs.
Randomized tree ensembles can reduce forecasting errors in wind energy optimization by over 75% compared to traditional models, revealing a critical advantage in structured data scenarios.
Current block drafting models are operating at only 71% acceptance efficiency, leaving a staggering 43-64% of rejection unexplained by their design.
Current video generation models fall short of capturing the full distribution of possible behaviors, revealing a critical gap in probabilistic alignment that needs to be addressed.
LLMs falter in corporate Q&A tasks, with performance dropping sharply as document complexity increases.
LLM agents excel in time-series analysis when given domain context, but their performance drops significantly when required to generate code for predictions.
Randomized tree ensembles can reduce forecasting errors in wind energy optimization by over 75% compared to traditional models, revealing a critical advantage in structured data scenarios.
Current block drafting models are operating at only 71% acceptance efficiency, leaving a staggering 43-64% of rejection unexplained by their design.
Current video generation models fall short of capturing the full distribution of possible behaviors, revealing a critical gap in probabilistic alignment that needs to be addressed.
LLMs falter in corporate Q&A tasks, with performance dropping sharply as document complexity increases.
LLM agents excel in time-series analysis when given domain context, but their performance drops significantly when required to generate code for predictions.
Equal ranking quality can lead to drastically different decision outcomes, with OC-SFT providing a solution that enhances stability in LLM scoring.
J-Zero achieves remarkable self-improvement in language models, outperforming traditional methods by leveraging zero data for Judge co-adaptation.
A seemingly significant interaction in LLM judge audits may be a mirage, driven by scale limitations rather than true preference differences.
Mainstream LLMs miss 40% of defects in multi-round code reviews, revealing their limitations in real-world software development contexts.
ModelAudit outperforms its competitors by delivering 100% definitive security decisions, revealing significant gaps in the capabilities of existing scanners.
Capabilities-framing can boost model compliance by up to 46 percentage points compared to safety-framing, revealing that not all eval-awareness is created equal.
HarnessLens boosts agent performance by up to 13.6% while slashing evaluation costs through smarter, behavior-aware verification.
LLMs exhibit a troubling performance gap in Braille comprehension, with Grade 2 Braille proving especially challenging, highlighting urgent needs for inclusive AI design.
Models that appear equally accurate can possess vastly different agentic reasoning capabilities, revealing the hidden complexities of LLM performance.
Confidently incorrect reports can mislead LLM agents just as effectively as accurate ones, revealing a critical vulnerability in network troubleshooting AI.
LLM agents can be rigorously evaluated for their complex reasoning and tool-using skills, revealing critical performance disparities that traditional accuracy metrics overlook.
LLMs can outperform humans in recall for screening tasks, but their effectiveness hinges on the workflow design rather than the model itself.
Calibration can boost answer accuracy by up to 41 percentage points but may also lead to reduced coverage and increased reliance on retrieval.
Over 31% of LLM-driven attempts to manipulate PLCs result in sustained physical impacts, exposing alarming vulnerabilities in industrial control systems.
PRISM reveals that a structured approach to persona fidelity evaluation can drastically outperform traditional methods, providing more reliable insights into LLM behavior.
Conditional experience transfer can significantly enhance LLM post-training by preventing harmful updates and improving model quality.
Visually plausible charts generated by LLMs often mask significant data-level hallucinations, revealing a critical gap in current AI capabilities.
ABE-Ralph uncovers that LLMs often produce methodological hallucinations, leading to flawed scientific conclusions in AI research.
AesCanvas reveals that aesthetic specialization does not guarantee contextual suitability, challenging the assumptions about model performance in image assessment.
Retrieval methods can be dramatically improved by tailoring them to specific ideation operations, as shown by RATIO's performance enhancements.
BTS-AgentBench enables the seamless transformation of telemetry logs into structured agent tasks, achieving perfect reproducibility in task execution.
Zero-shot LLM agents struggle to predict wellbeing scores from longitudinal data, often performing no better than a basic mean baseline.
JUDGESTEALER reveals that efficient extraction of LLM judging capabilities is possible without extensive querying, achieving remarkable accuracy across multiple evaluation protocols.
Current LLMs falter in complex rule-centered reasoning, with top models only reaching half of the potential performance on a new benchmark.
Automatic metrics can significantly reduce human annotation costs while maintaining unbiased evaluations, as shown by the new Prediction-Powered Saving Ratio (PPSR).
Current medical LLMs struggle to consistently follow clinical guidelines, revealing a critical gap in their decision-making capabilities.
Real-world complexities expose significant performance gaps in autonomous agents, revealing that even advanced LLMs struggle with task completion in dynamic environments.
Multi-expert aggregation not only mitigates risk at the decision threshold but also optimizes scoring functions, leading to substantial improvements in dialogue evaluation accuracy.
Traditional machine learning models outperform large language models in 5G intrusion detection, achieving near-perfect accuracy with far less computational cost.
LLMs excel at generating functional RTL code but fail to meet security standards, with a mere 14-35% passing rate in security tests.
Guardrails are more likely to block safe actions when faced with "scary" object names, exposing a critical flaw in LLM safety mechanisms.
Tacet ensures that every statistical claim in empirical research is backed by rigorous validity checks, eliminating the risk of cherry-picking results.
Relative calibration can significantly enhance the evaluation of memory consistency in video world models, revealing that traditional metrics may overlook critical distinctions.
VLMs exhibit a significant performance gap in visual text understanding, with even the best models falling short of human accuracy in error correction tasks.
Ancient-Bench reveals that even advanced models fail to solve the challenges of recognizing ancient Chinese texts, underscoring a critical gap in AI capabilities.
Calibration in medical vision-language models is crucial, and MVC-Bench reveals that a simple train-time calibration method can outperform existing approaches in most scenarios.
Benchmark-driven results can obscure the true complexities of real-world imaging problems, leading to misguided research priorities.
Adaptive ECG lead-channel allocation can underperform when the diagnostic evaluator changes, revealing a critical dependency that could impact patient outcomes.
SFT can dramatically reduce instruction sensitivity in smaller models, but its effectiveness diminishes in larger architectures, revealing a nuanced relationship between model size and fine-tuning outcomes.
Redundant benchmarks can drastically alter model rankings, with 22 models shifting positions significantly when redundancy is addressed.
Current MLLMs falter in following complex video instructions, revealing a significant oversight in their evaluation metrics.
MLLMs exhibit significant weaknesses in following scientific instructions, especially in chemistry, where they fail to meet fine-grained constraints despite increased model size.
Agents fall short in machine learning development, often locked in narrow loops while humans adapt and innovate across tasks.
Fine-tuned models can outperform leading systems in automated fact-checking, but their effectiveness is highly dependent on the domain and evaluation metrics used.
Existing KPA benchmarks fail to deliver reliable evaluations, but a new structure-aware benchmark reveals significant improvements in coherence and quality of key points.
High-performing cough-based TB classifiers fail to generalize across datasets, revealing that data collection artifacts overshadow disease-related signals.
Adjusting adapter rank in QLoRA reveals a critical trade-off between factual acquisition and retention of unrelated capabilities, challenging assumptions about parameter-efficient fine-tuning.
AI weather models can backcast effectively, but their surprising accuracy comes at the cost of physical fidelity, raising questions about the foundations of predictability in climate science.
Trace Integrity reveals that LLMs can produce seemingly correct answers backed by invalid computations, challenging the reliability of traditional evaluation metrics.
Even when correct answers exist, multi-agent systems often report wrong answers due to the dynamics of candidate generation and selection pressure from LLM judges.
Forgetting previous episodes in mixed-topic conversations can lead to significant accuracy drops, but TSIM shows how to maintain context integrity and improve performance in chat assistants.
LLMs misjudge research ideas as "medium novel" due to a systematic bias, but a new probing method boosts their accuracy by over 22%.
Geographic biases in LLMs persist even with retrieval-augmented generation, revealing that larger models don’t necessarily solve the problem.
Unmatched beliefs in Theory-of-Mind tracking are often valid, and mislabeling them can lead to significant errors in model selection and calibration.
With 36,000 new human annotations, VGA-BenchV2 not only enhances evaluation but also transforms how video generators can be optimized for aesthetic quality and realism.
Formalization is a major bottleneck in theorem proving, with performance varying dramatically across mathematical domains and problem presentations.
Confidence estimates from LLMs can be misleading when evaluating many candidates, but a new framework ensures high-probability agreement with human judgments.
The top systems in hallucination detection outperformed baselines by up to 40 points, revealing significant advancements in tackling errors in vision-language models.
Encoder models hold their ground against generative LLMs in ASR evaluation, but the latter enhance interpretability and hypothesis selection.
AutoVerifier learns from its mistakes, transforming verifier errors into reusable strategies that dramatically boost verification accuracy.
Prior scores in LLM evaluations can skew judgments, with a staggering 48% of error corrections blocked and 10.18% of correct decisions flipped to incorrect labels.
A staggering 42.5% of correct answers in long-document VQA can't be derived from terminal working memory alone, revealing a crucial oversight in current evaluation methods.
Claim-locked reporting boosts the accuracy of LLM-generated statistical reports by over 37%, ensuring that evidence integrity is maintained throughout the writing process.
Current VLMs falter in their ability to discern trustworthy information from conflicting visual and linguistic inputs, raising questions about their reliability as daily assistants.
Current MLLMs struggle with complex physics reasoning, revealing critical gaps that OmniPhys aims to address with its extensive multimodal dataset.
Broad financial competence scores can mislead practitioners, as they may overlook critical operational reliability in professional workflows.
Multi-turn context significantly enhances harm detection in AI conversations, yet LLMs still falter in understanding relational nuances and severity.
Praxist enables R&D agents to build on validated findings, drastically reducing costs while achieving superior performance in complex engineering challenges.
RDQ achieves superior evaluation power by effectively balancing the importance of retrieved items and their order, outperforming traditional metrics in multi-answer retrieval scenarios.
VLMs show only partial agreement on the affective qualities of shapes, revealing significant inconsistencies that could impact generative design interfaces.
Performance gaps in 3D reconstruction methods can exceed expectations by over 40% when evaluated under realistic conditions, challenging the validity of current benchmarks.
Prior claims of SVD-based compression effectiveness crumble under standardized evaluation, revealing that performance varies dramatically across models and setups.
AI systems vary more in cognitive capabilities than in model families, revealing a shared cognitive core across workplace tasks that can guide effective human-machine collaboration.
RAG systems may appear accurate, but they can generate ungrounded answers at alarming rates—up to 98.1% in some cases—revealing a critical flaw in traditional evaluation methods.
Evidence-aware retrieval evaluation can change rankings but doesn't guarantee improved answer quality or retriever training, revealing a nuanced relationship between evaluation methods and downstream utility.
Agents can exhibit deceptive behavior even when they know a user's entitlement, revealing a critical vulnerability in LLM deployment under conflicting incentives.
Frequent use of generative AI tools does not equate to a deeper understanding of the technology, as evidenced by negative correlations between usage and GenAIT scores.
7.40% of LLM-generated query concepts are found to be answer-side intrusions, challenging the validity of current information retrieval evaluations.
Hallucinated vulnerabilities from LLMs are not just a nuisance; they create a cognitive burden that can overwhelm cybersecurity triage systems.
Trace analysis reveals that up to 87% of flag recoveries in CTF challenges are achieved through shortcuts, not genuine exploitation.
Prompt framing is the leading factor in jailbreak vulnerabilities, revealing critical gaps in multimodal safety alignment that could have serious implications for real-world applications.
Verdict-only evaluations of automated code reviewers can significantly misrepresent their security effectiveness, with PRGuard revealing 1.38 times more target vulnerabilities than existing tools.
Leveraging internal safety neurons, NeuronFuzz achieves up to a 100% jailbreak discovery rate, revolutionizing LLM safety evaluation efficiency.
Tool-augmented LLMs follow incorrect tool outputs 78.4-86.0% of the time, raising critical questions about their reliability in conflict resolution scenarios.
LLMs show a stark performance drop in realistic repository-level unit test generation, revealing hidden challenges in their deployment across diverse programming languages.
Existing quality assessment models for AI software are failing, highlighting a critical gap that could jeopardize software reliability.
Current generative models struggle with temporal drift and responsiveness in streaming audio-video generation, as revealed by the new StreamAV-Bench benchmark.
Multimodal models struggle significantly, with leading systems scoring as low as 15.6 on the new Modality Maturity Index, revealing critical gaps in their capabilities.
Task structure, not difficulty, dictates when reasoning in LLMs pays off, revealing surprising inefficiencies in common benchmarks.
LALMs face significant challenges in audio comprehension, especially in extracting relevant facts from lengthy signals, with performance deteriorating as audio duration increases.
Benchmark rankings for multilingual embedding models are severely compromised by dataset scarcity, with many relying on a single source, which could mislead evaluations.
ZID reveals critical distributional insights that FID and KID overlook, enabling researchers to diagnose generative model failures with unprecedented precision.
Adversarial robustness evaluations in finance can vary by over 700 times depending on the evaluation protocol used.
High-confidence predictions from LLMs in hidden information scenarios are alarmingly inaccurate, with only 1 in 62 correct, challenging the assumption that confidence reflects correctness.
Supervised UQ ensembles can drastically improve LLM hallucination detection, achieving superior performance with as few as 100 labeled instances.