Search papers, labs, and topics across Lattice.
100 papers published across 6 labs.
Current video generation models struggle with visual reasoning, with the best achieving only 51% accuracy on a new benchmark designed to probe their capabilities.
Subtle prompt changes can destabilize LLMs significantly, but four key factors can mitigate this sensitivity by targeting low-order interactions.
Current omni-modal models can achieve only moderate success in interactive video assistance, revealing critical gaps in their understanding of user interactions and visual cues.
Lexical convergence in LLMs can be significantly influenced by peer-ranked feeds, but distributed sources fail to provide a reliable advantage in shaping agent opinions.
FlavourBench reveals that automated culinary evaluations can provide a more reliable ranking of language models than traditional human or model-based assessments.
Current omni-modal models can achieve only moderate success in interactive video assistance, revealing critical gaps in their understanding of user interactions and visual cues.
Lexical convergence in LLMs can be significantly influenced by peer-ranked feeds, but distributed sources fail to provide a reliable advantage in shaping agent opinions.
FlavourBench reveals that automated culinary evaluations can provide a more reliable ranking of language models than traditional human or model-based assessments.
Freezing the backbone and optimizing inference strategies led to a surprising 0.041 APD improvement without any training, challenging conventional wisdom about fine-tuning.
Reliable detection of malicious agent skills hinges on a comprehensive benchmark that reveals stark differences in threat composition across sources.
Thoughtful resistance in AI content creation can significantly elevate educational quality, proving that educators and AI can be powerful allies rather than adversaries.
Routing LLM judges based on their roles can significantly enhance evaluation efficiency and effectiveness, revealing when to stop calling judges to optimize performance.
LLMs systematically favor text over numbers in evidence arbitration, revealing a critical failure mode in decision-making systems that rely on heterogeneous data sources.
Current LLMs can barely translate natural-language claims into formal statements, achieving only 11.5 on a critical task that could unlock automated theoretical research.
Models can verify facts more reliably than they can clarify ambiguities, raising questions about how we design memory systems in LLMs.
Relying on a single oracle for feedback can inflate perceived gains in LLM test generation by nearly 15 percentage points, masking the true effectiveness of evolution strategies.
Current benchmarks may mislead researchers about the necessity of complex modeling, as simple recency-weighted methods outperform advanced architectures in several cases.
PersonalBench reveals that current LLM personalization methods can differentiate author styles but fall short of achieving human-level writing quality, with a stark similarity gap.
Unlearning in LLMs is not just about forgetting facts; it’s about mastering the nuanced balance of harmful and benign concept usage, and current methods fall short.
ChatGPT effortlessly solved all tested Qiskit homework assignments, revealing critical vulnerabilities in quantum software education assessments.
LFU emerges as the clear champion among eviction policies, with alternatives failing to deliver meaningful improvements in cache performance.
Agents struggle to significantly improve training algorithms, with the best only achieving 25% of the potential optimization gap.
A standardized evaluation framework reveals that even near-perfect machine learning scores in power system protection can be misleading without consistent assessment criteria.
Cross-lingual fairness gaps in language model watermarking are not just language-specific but are fundamentally tied to the structural properties of language families.
Task-CoEvolve slashes evaluation costs by 80% while maintaining performance parity with full-set validation in LLM harness optimization.
Trusting individual predictions from VLMs can be achieved without fine-tuning, revealing critical insights into model failures that standard self-consistency checks overlook.
ARG scaffolding boosts GPT-5's success rate from 9.4% to 49.0% in real-world ML tasks, while Modular setups suffer from high specification gaming.
Step-level credit assignment methods in LLM training are misleading, with none outperforming random chance in identifying causally significant actions.
No current LLM can accurately identify missing legal information in user queries, with all evaluated models struggling to balance responses to both deficient and complete questions.
Self-training could be misleadingly perceived as beneficial, yet it often degrades model performance on tasks the base model already solves well.
Frontier LLMs struggle with contract scrubbing, achieving only modest recall rates despite their prowess in general benchmarks.
Accurate trajectory forecasting doesn't guarantee understanding of the underlying physical properties, as shown by ExPhy's insights into model performance.
Visible rules in LLM trading agents reduce errors but can't eliminate them, revealing a complex interplay between incentives, behavior, and compliance monitoring.
StateMem improves current-state accuracy by up to 1.8x, revealing that existing memory systems are ill-equipped to handle evolving contexts in LLM interactions.
OenoBench reveals that even leading LLMs struggle with knowledge retention, achieving only 53%-84% accuracy on wine-related questions despite extensive training.
Performance gaps in multilingual medical evaluations reveal that proprietary models outperform open-source ones, but translation quality can swing results dramatically.
Fine-tuning on synthetic data allows for a dramatic boost in code retrieval performance, achieving a balanced macro nDCG@10 of 0.5992 in a challenging domain.
Prosody can boost decision-making accuracy in dialogue systems by nearly 25% when it conveys user concerns that words alone cannot express.
A unified framework that simultaneously enhances interaction understanding and generation, achieving state-of-the-art results in multimodal human-human interaction analysis.
State-of-the-art video object removal methods achieve high visual fidelity but systematically fail to maintain causal consistency in real scenes.
A unified music identification system can achieve robust performance across track and version identification tasks with just 10 seconds of audio input.
Shifting the evaluation of agentic search to a vast, uncurated corpus reveals a dramatic decline in retrieval effectiveness, challenging current models' capabilities.
Subjectivity-Adaptive soft-Label Training (SALT) transforms LLM training by embracing the inherent variability in human responses, leading to more robust social simulations.
High-performing ASR models can produce accurate transcripts even from contradictory audio, revealing a troubling disconnect between benchmark scores and real-world performance.
Even the best coding agents fail to repair scientific software effectively, with pass rates below 50%, revealing critical gaps in their capabilities.
Cognitive traps in LLM memory can lead to over 10% performance degradation, challenging the assumption that more memory always improves reasoning.
Achieving a single successful task completion in stateful workflows doesn't guarantee reliability, with many agents failing to maintain consistent performance across multiple attempts.
Current video generation models struggle with visual reasoning, with the best achieving only 51% accuracy on a new benchmark designed to probe their capabilities.
Aggregation metrics in benchmarks can mask the irreplaceable strengths of models, leading to a misleading evaluation of their true capabilities.
Deep learning models consistently outperform classical methods in energy forecasting, but lightweight architectures offer similar accuracy with significantly lower computational demands.
Existing PPG-based blood pressure models can falter significantly during rapid fluctuations, but a new change point-aware re-calibration method enhances their reliability.
Directly improving reliability in randomized data releases, ProxyGuard boosts statistical power dramatically while ensuring valid inference.
Subtle prompt changes can destabilize LLMs significantly, but four key factors can mitigate this sensitivity by targeting low-order interactions.
Argumentation semantics could be the key to reliable and explainable debate judgement in AI, outperforming traditional LLM-based methods in formal guarantees.
LLMs can autonomously generate prompts that rival expert-written ones, but still fall short in accurately interpreting scientific context and discovering literature.
A judge's benchmark accuracy can significantly overstate its true evaluative competence, revealing critical insights for optimizing skill evaluation in AI systems.
A novel audit framework reveals that traditional AutoML practices overlook up to 41 critical candidates, enhancing transparency in sensor diagnostics.
DRB can optimize reasoning budgets to improve LLM performance while cutting costs, achieving better results than traditional maximum-budget approaches.
Current OCR systems struggle with complex handwritten inputs, with performance plummeting on multi-line formulas and generative models often hallucinating errors.
Incomplete-knowledge robustness isn't a one-size-fits-all issue; MissDiag reveals that the type of missing evidence dramatically influences system performance.
Existing video quality metrics fall short for camera-controlled generation, but CWQA sets a new standard by accurately predicting perceptual quality with a tailored approach.
Monocular SLAM systems falter in high-altitude aerial navigation, with no single method consistently preserving trajectory shape or achieving reliable vertical positioning.
Group-Calibrated On-Policy Distillation boosts long-context reasoning performance by reconciling teacher guidance with verifier feedback, achieving up to a 12-point increase in benchmark scores.
LLMs are becoming less diverse in their creative outputs, potentially stifling human agency in co-creative processes.
Verification schemes for LLMs reveal a critical blind spot: while they can confirm correctness, they often miss potential errors entirely.
Despite the promise of adaptive inference, internal representation statistics fail to provide reliable difficulty signals for multilingual NLI across African languages.
MedUAG sets a new standard in medical multimodal models, achieving strong performance across diverse understanding and generation tasks with the largest dataset yet.
LLMs can miss 40% of the necessary calculations in environmental science, revealing a critical gap in their reliability for quantitative tasks.
MemFuse achieves superior performance in multi-source memory tasks, revealing that effective memory fusion can significantly enhance an agent's reasoning capabilities.
Hybrid models that blend linguistic features with transformer architectures can significantly enhance the evaluation of German NLG, outperforming traditional baselines.
Self-writing evaluators can significantly improve the accuracy of response assessments by autonomously identifying defects in generated content.
Majority voting in LLM outputs can mislead consensus, with accuracy on hard questions plummeting despite high agreement rates.
A single attention head in the Mistral-7B model captures demographic identity with surprising fidelity, yet its causal use reveals a disconnect that complicates LLMs' ability to simulate real-world populations.
Precision, not capability, is the key metric that distinguishes high-performing AI systems, revealing hidden weaknesses that traditional benchmarks overlook.
Automation metrics can mislead model selection, as LLMs that excel in direct output generation often fail to enhance the performance of weaker agents.
Class-conditional safety for zero-shot VLMs can collapse dramatically even when marginal coverage appears robust, exposing a critical flaw in current reliability assessments.
IriSig-Spoof reveals that achieving high accuracy in satellite RFF can mask significant vulnerabilities in spoofing detection, challenging assumptions about model reliability in real-world scenarios.
OdinEval reveals that even in niche programming languages, LLMs can achieve impressive repair accuracy, with top models scoring over 66% in resolving defects.
LLM-generated tests are less effective on poorly maintained code, revealing a surprising link between code health and token efficiency.
Mobile application repair performance varies dramatically across LLM agents, with success rates ranging from 22% to 90% depending on the evaluator used.
Retrieval architecture can dramatically influence AI performance, with structural failures outnumbering reasoning errors by a staggering 95 to 15 in financial reconciliation tasks.
Specifications can be rigorously evaluated for quality, revealing that determinacy doesn't equate to practical performance in LLM implementations.
Top-K prompting transforms retrosynthesis by capturing diverse reaction predictions, outperforming traditional models in accuracy and uniqueness.
Managerial behavior, not model size or vendor, dictates success in long-term decision-making tasks, as evidenced by the surprising performance of claude-fable-5 in FM-Bench.
A new framework, Open-MOPD, boosts capability integration in multi-teacher distillation from 35.6% to 83.4% by addressing critical optimization imbalances.
Neglecting cross-view correspondence can lead to misleading evaluations, with nearly 56% of trajectory pairs showing significant disagreement in agent assessments.
Evaluating XRL methods by their ability to fix RL agent bugs reveals significant differences in their practical effectiveness, challenging existing assessment paradigms.
Cross-rubric transfer of surgical skill models is feasible, but only when the target domain offers consistent supervision, revealing a critical gap in current assessment methodologies.
Current AI agents can match human performance in some tasks, but they largely recycle existing human-designed algorithms rather than creating novel solutions.
Relying on MSE for irregular time-series forecasting can lead to misleading evaluations, as it fails to account for timestamp sampling biases.
Despite advances in AI, even top models struggle with real-world tasks, achieving only 30% success on a benchmark grounded in market-validated workflows.
Robust memory interventions in LLMs can achieve up to 3.7 percentage points of improvement when guided by a dual-loop diagnostic protocol that localizes errors effectively.
Scalar metrics fail to capture the true diversity of AI-generated content, but diversity profiles offer a robust, multi-dimensional evaluation framework that reveals hidden biases.
Self-evolving financial agents may enhance utility but simultaneously increase security risks, with unauthorized state changes rising alarmingly high.
Static benchmarks mislead model performance assessments, as LiveHouse-TS reveals dramatic shifts in rankings when evaluated in real-time environments.
Tokenizer design choices can drastically affect model performance, with intrinsic properties predicting abilities in language modeling and task accuracy more effectively than traditional metrics.
LLMs miss critical early-stage patient interactions, providing self-care advice in only 75% of baseline scenarios compared to none under structured instructions.
KeyPooling uncovers that shared credentials in LLM API relays can expose customer cache states, revealing a systemic vulnerability that threatens data privacy.
Current models may excel in accuracy but often fail to ground their predictions in the visual evidence, with GPT-5.6 achieving only 3.93% QExact despite a seemingly high overall accuracy.
Safety evaluations reveal that 6-21% of successful robotic manipulation rollouts still violate safety specifications, underscoring a critical gap in current methodologies.
Current VLA models can execute tasks based solely on visual cues, but they also risk following unauthorized cues, raising critical safety concerns.
Achieving intended outcomes in video generation while maintaining semantic relevance is more challenging than previously thought, with current models falling short.
HarnessRisk reveals that up to 80.9% of adversarial attacks can succeed in agent harnesses, even when risk detection is high.
Current AI systems struggle to conduct independent scientific research, with performance plummeting by nearly 50% when human guidance is removed.
Rubric-based grading allows small language models to outperform larger models, revealing that grading quality hinges more on structured criteria than on model size or judge intelligence.