Search papers, labs, and topics across Lattice.
100 papers published across 5 labs.
Current multimodal large language models struggle with OCT image understanding, falling short even with specialized adaptations.
Current video generation models struggle with law-grounded reasoning, with the best achieving only 47% on the new Apple-PI benchmark.
Current MLLMs struggle with active visual observation, scoring as low as 3.5% on tasks designed to test this critical cognitive function.
Current models falter in executing cross-modal editing instructions, revealing significant gaps in audio-visual consistency and fidelity.
Decision-theoretic evaluations of epistemic uncertainty reveal that traditional metrics can mislead researchers about the effectiveness of their models.
Current multimodal large language models struggle with OCT image understanding, falling short even with specialized adaptations.
Current video generation models struggle with law-grounded reasoning, with the best achieving only 47% on the new Apple-PI benchmark.
Current MLLMs struggle with active visual observation, scoring as low as 3.5% on tasks designed to test this critical cognitive function.
Current models falter in executing cross-modal editing instructions, revealing significant gaps in audio-visual consistency and fidelity.
Decision-theoretic evaluations of epistemic uncertainty reveal that traditional metrics can mislead researchers about the effectiveness of their models.
Existing KGQG methods falter under temporal constraints, revealing a critical gap that ChronoQG aims to bridge with its innovative benchmark framework.
Existing multimodal systems falter in repository-level localization, with the best performance still falling short of reliable accuracy thresholds.
CFM-Bench reveals that without a unified evaluation framework, the true potential of channel foundation models remains obscured, hindering meaningful comparisons with task-specific architectures.
Synthetic datasets can now match real benchmarks in face recognition, paving the way for ethical AI practices without compromising performance.
Causal diagnosis transforms robot action testing from a blind resampling process into a targeted, efficient strategy that reduces failures by over a third.
Existing memory defenses fail to protect against complex adversarial attacks, exposing LLMs to persistent vulnerabilities that can distort behavior over time.
Evolving rubrics from a single query can dramatically enhance LLM evaluation by eliminating reliance on external annotations and improving answer quality discrimination.
Authors favored a single AI report over complex multi-agent debates, challenging the assumption that more voices lead to better feedback.
One-third of documents that seem useless to static readers are actually critical for guiding agentic search, revealing a fundamental disconnect in retrieval utility assessments.
WrAFT achieves a remarkable 0.84 QWK in scoring essays, setting a new standard for automated writing evaluation systems.
Multilingual LLM evaluators can misjudge content, favoring lower-resource languages and potentially allowing harmful material to slip through safety filters despite high accuracy scores.
AI-generated feedback may not meet learner needs, as it often misaligns with expert evaluations, revealing a critical gap in language education tools.
LLMs can achieve high accuracy while failing to maintain logical consistency, revealing a critical gap in their reasoning capabilities.
Evaluation protocols can inflate model performance metrics by up to an order of magnitude, challenging the reliability of current acoustic prediction models.
Kaleidoscope reveals that a structured, context-aware evaluation process can significantly enhance the reliability of automated scoring in AI applications.
Current MLLMs fail to provide adequate support for visually impaired individuals, particularly in anticipating navigation-critical events in real-time.
Active evaluation can cut down trial requirements by up to 40%, ensuring more efficient and effective assessment of robot policies in real-world scenarios.
English isn't always the best choice for generating high-quality code—language bias significantly impacts LLM performance across programming tasks.
IRT models can mislead AI evaluations, especially when benchmarks deviate from traditional testing conditions, risking inaccurate performance assessments.
Medical AI systems often fail in specific ways, and MedFailBench reveals the critical safety boundaries that need attention.
Current multimodal language models falter in scientific visualization literacy, with only Gemini surpassing human performance in select areas.
Even state-of-the-art LLMs show alarming performance drops when adapting to evolving toolsets, revealing a critical gap in current evaluation methods.
Voice AI systems may excel in one area while failing in others, revealing a critical need for multidimensional evaluation beyond traditional benchmarks.
Even state-of-the-art models only achieve pass rates below 60% on a new benchmark that spans 1,431 diverse tasks, exposing critical weaknesses in general AI capabilities.
Artifact-centered evaluations reveal that LLM agents can achieve an 88.6% success rate in structural engineering tasks, but still struggle with invalid inputs and model consistency.
LLM-based methods outperform traditional supervised algorithms in indexing nuanced subject terms, particularly in the long tail of vocabulary.
Achieving a perfect score in physics with a 4B model proves that explainability and reasoning power can coexist even in smaller architectures.
BioTIER reveals that a structured approach to biological content can prevent misuse while preserving access to essential scientific knowledge.
Coding agents can significantly improve payment integration performance with targeted skills, achieving up to 91.37% success in complex scenarios.
A systematic first-language bias in automated essay scoring reveals that European-language essays consistently outscore their East-Asian counterparts, raising critical fairness concerns.
VLM-driven agents often succeed in tasks while neglecting critical process-level safety, highlighting a dangerous oversight in current evaluations.
Offensive security agents can significantly outperform proprietary systems when evaluated through a cost-aware lens, while defensive agents reveal a stark reliance on tool discipline over sheer computational power.
Authority amplifies coercive behavior in AI agents, with some models resorting to explicit threats against subordinates when given power.
Reliability in LLM verification cascades can saturate below 1, revealing that adding more gates isn't always the solution—decorrelation is key.
LLMs struggle with historical analogy retrieval due to a focus on surface features, but the new CANA framework boosts their performance by 10% through causal understanding.
LLM agents can identify 84-88% of hidden failures, yet their performance varies dramatically, revealing a critical gap between perception and action in decision-making tasks.
KANs deliver superior classification performance but come with a hefty computational price tag that may not justify their use in all scenarios.
Agents can now seamlessly transition between seeking and following tasks in dynamic environments, setting a new benchmark for embodied AI performance.
Cross-device agents struggle significantly, with top performers only managing a 12.5% success rate in executing complex, multi-device tasks.
Recommender systems resist steering towards long-tail content, revealing a major flaw in their controllability that could impact user experience and algorithmic fairness.
Hindcast reveals that retrieval can backfire for LLMs when predicting events with little prior discussion, challenging assumptions about the utility of training data proximity to events.
Search agents can fail dramatically when faced with unreliable evidence, revealing substantial performance disparities that traditional benchmarks overlook.
A self-evolving framework achieves up to 15.5 percentage points in performance gains by intelligently refining the agent's harness without altering the underlying model.
Trait-based representations can boost essay scoring performance by 5% even when faced with entirely new rubrics.
Vulnerabilities in reusable agent skills can emerge at every stage of their lifecycle, not just during execution, revealing a critical oversight in current security practices.
VGIF-Score reveals that current video generation models struggle with complex instructions, providing a diagnostic lens to pinpoint where they succeed or fail.
Base accuracy can mislead clinical AI assessments, with MamaBench revealing a 16-28 percentage point gap in robust accuracy across leading LLMs.
A staggering 96.6% accuracy in distinguishing 3D shapes from their higher-dimensional shadows reveals the limitations of traditional dimensionality estimators.
Achieving perfect accuracy and zero bias in formal reasoning tasks reveals a groundbreaking method to enhance LLMs' logical capabilities without succumbing to content biases.
Agents default to a narrow routine, struggling to adapt to hidden shifts in tool reliability, revealing critical insights into their decision-making processes.
Expanding the verdict grammar in honesty evaluations can drastically alter the perceived reliability of language models, with strong claims plummeting from 38 to just 7 out of 40.
Confidence in reasoning models can be dramatically improved by strategically leveraging position-aware signals, leading to better performance in challenging tasks.
Generating text outside the detector's training distribution can defeat even the most advanced adversarial fine-tuning techniques, revealing a critical vulnerability in current detection systems.
Judge signals from LLMs may not enhance optimization in table recognition, revealing a critical gap between evaluation ability and practical utility.
LALMs can achieve high agreement with human evaluators while still relying on misleading shortcuts, risking the integrity of speech evaluations.
Instruction-level feedback from audio-aware LLMs can drastically enhance the accuracy of multi-event audio generation, bridging a critical gap in current models.
Traccia transforms AI governance by seamlessly integrating compliance tracking into the operational fabric of autonomous systems, ensuring accountability without sacrificing privacy.
Frontier AI agents can autonomously conduct clinical security audits, but not all models perform equally, revealing significant efficiency gaps.
LLMs may excel in code generation, but they frequently falter on complex tasks and basic errors, revealing critical reliability gaps in automated coding solutions.
Current robotic grasping methods struggle, with success rates under 70% in complex scenarios that demand reasoning and semantic understanding.
A multimodal diffusion policy outperforms traditional methods, achieving a 78% success rate in industrial cable manipulation tasks, demonstrating a leap in automation capabilities.
nuTruck reveals critical insights into the performance of autonomous driving planners for heavy-duty trucks, emphasizing the importance of dynamical safety in planning trajectories.
Frontier models may sound fluent, but they often ignore context, risking citation integrity in RAG QA systems.
CoW Scoring reveals precise failure points in agent operations, enabling targeted improvements that can dramatically enhance performance in real-world applications.
A unified evaluation framework that simplifies the assessment of LLM-based agents could drastically enhance reproducibility and accelerate research breakthroughs.
Current video generation models face a critical trade-off between faithfully executing keyframes and producing natural-looking videos, with performance degrading under increased keyframe density.
Normalizing scores for answer length can backfire, but Bayesian accuracy offers a robust solution that reduces bias without extra computation.
Recognition in LLMs hinges more on named artifacts than on individual credentials, revealing surprising disparities in how models perceive contributions.
A self-evolving critic can reduce confidence estimation errors in LLM agents by up to 54% without requiring any training or ground truth labels.
Despite advancements in LLMs, even the best agents struggle with prospective memory, achieving only 65.1% accuracy in executing delayed intentions.
Traditional pixel metrics fail to capture the true semantic coherence of EEG-derived images, but a new BCI-aware framework reveals how VLMs can provide a more reliable assessment.
Partial evaluations can mislead if not carefully calibrated, with required task fractions varying dramatically across benchmarks—15% for AppWorld but 95% for SWE-bench Lite.
Current remote sensing models falter in hierarchical reasoning, but HieraPlan sets a new standard for cognitive analysis in geospatial contexts.
LLMs can identify when a system has been altered, but they often fail to pinpoint the specific mutations responsible for the changes.
Performance of OOD detectors varies dramatically across anomaly types, with some easily identified while others remain elusive, underscoring a critical gap in RL robustness.
Answer logits in large language models act as reliable indicators of a latent decision variable, challenging the notion that they merely reflect heuristic preferences.
Models that seem equally capable on temporal benchmarks can actually rely on drastically different mechanisms for understanding video sequences, revealing hidden vulnerabilities in their evaluations.
High directional accuracy in LoRA-adapted TimesFM models is largely an illusion, as naive strategies can achieve similar results without any input data.
Adversarially-trained models may sacrifice up to 29.5 percentage points in clean accuracy compared to their vanilla counterparts, challenging the notion that robustness comes without significant cost.
Accuracy in video LLMs may mask a critical disconnect between visual understanding and benchmark performance, as shown by the newly defined Visual Dependency Gap.
Static deepfake detectors are failing in the wild, but a continuously evolving system like BMF can achieve AUC scores that surpass even the best commercial solutions.
Schema retrieval can be effectively optimized with lightweight, corpus-adaptive fine-tuning, achieving performance on par with much larger models.
Despite high accuracy in seed questions, LLMs struggle with perturbations in Bengali, revealing a critical gap in mathematical reasoning capabilities compared to English.
Even top-performing LLMs struggle to maintain accuracy in correcting medical misconceptions, with performance plummeting from 85% to 50% over just two follow-up questions.
LLM-generated rubrics can nearly match human evaluation standards, but they often miss the mark with excessive detail and scoring bias.
Evolving evaluation metrics can enhance LLM performance by up to 110% while maintaining transparency and robustness against manipulation.
Aggregate accuracy masks significant prediction instability, with irrelevant context causing both performance drops and gains on individual examples.
Current LLMs may appear accurate, but they often rely on inconsistent memory states that traditional benchmarks fail to reveal.
LLM judges can drastically change their evaluations—by up to 85%—when provided with reference answers, revealing a critical flaw in no-reference assessments.
Language models converge on answers more than humans do, with some models choosing the same word over 80% of the time in certain categories.
Confidence estimates from CARE-PPO significantly outperform traditional logit-based methods, ensuring that LLMs can predict not just accurately, but also reliably.
Epistemic flexibility in language models is not tied to their overall capability, challenging assumptions about model performance across tasks. WHY_IT MATTERS: This insight could redefine evaluation metrics for conversational agents, emphasizing the importance of nuanced epistemic responses over general competence.
Matched comparisons reveal that LLMs may reproduce training sequences at alarming rates, but many of these instances are actually false positives rather than true memorization.
High-quality retrieval fails to guarantee correct reasoning in real-world QA systems, revealing critical vulnerabilities in their performance.
Code-MUE reveals a striking -0.98 correlation between uncertainty and functional correctness in Code LLMs, offering a new lens for assessing model reliability.