Search papers, labs, and topics across Lattice.
100 papers published across 9 labs.
Mixed SFT outperforms next-chunk reasoning RL in reasoning tasks while using over 60 times less compute, reshaping our understanding of effective training strategies.
Discarding irrelevant reasoning tokens can make language models three times faster at test time without sacrificing performance.
Models trained on VBVR-Pro not only excel in native visual reasoning tasks but also reveal critical insights into the effectiveness of different generative modalities.
Free-form language reasoning transforms VLMs into powerful robotic reasoners, significantly boosting performance in complex manipulation tasks.
Imitation learning enables automated theorem provers to solve 46% more problems while drastically reducing proof steps compared to traditional methods.
Discarding irrelevant reasoning tokens can make language models three times faster at test time without sacrificing performance.
Models trained on VBVR-Pro not only excel in native visual reasoning tasks but also reveal critical insights into the effectiveness of different generative modalities.
Free-form language reasoning transforms VLMs into powerful robotic reasoners, significantly boosting performance in complex manipulation tasks.
Imitation learning enables automated theorem provers to solve 46% more problems while drastically reducing proof steps compared to traditional methods.
Operating costs for multi-agent LLM workflows can be slashed without sacrificing performance, thanks to ProgRouter's adaptive routing strategy.
Reflection Steering cuts reasoning token usage by nearly 17% while maintaining accuracy, revolutionizing how LLMs handle reflection during inference.
GRIP achieves superior accuracy-efficiency trade-offs in reasoning tasks by intelligently interpolating parameters from two distinct model types without retraining.
Bridging cognitive islands with a novel framework, PonsRAG boosts long narrative reasoning accuracy by over 11% through innovative evidence integration.
Reusing reasoning from previous queries can boost multi-hop QA accuracy while cutting down on token usage, transforming how we approach graph-based retrieval systems.
Trace Integrity reveals that LLMs can produce seemingly correct answers backed by invalid computations, challenging the reliability of traditional evaluation metrics.
Even when correct answers exist, multi-agent systems often report wrong answers due to the dynamics of candidate generation and selection pressure from LLM judges.
Increasing model size trumps inference compute for grammar-constrained text-to-SQL tasks, with beam search proving superior to sample+vote under matched budgets.
Hallucinated vulnerabilities in AI-driven vulnerability assessments create a cognitive burden that mimics a denial-of-service attack on human triage systems.
ToST empowers students to navigate multiple solution paths simultaneously, transforming the way LLMs can facilitate Socratic teaching.
AdaVDR achieves superior performance in video deep research by dynamically adapting its tool use based on the specific capabilities of the model and the nature of the task.
Formalization is a major bottleneck in theorem proving, with performance varying dramatically across mathematical domains and problem presentations.
DCGC corrects flawed reasoning in LLMs by leveraging imperfect drafts, leading to significant accuracy improvements in complex reasoning tasks.
AutoVerifier learns from its mistakes, transforming verifier errors into reusable strategies that dramatically boost verification accuracy.
ReliableRAG achieves a substantial boost in factual accuracy for multi-hop QA by effectively filtering out deceptive misinformation through a reliability-driven approach.
ClueWeaver transforms how compact language models tackle long narratives, achieving superior evidence retrieval and reasoning transparency.
Current MLLMs struggle with complex physics reasoning, revealing critical gaps that OmniPhys aims to address with its extensive multimodal dataset.
Adaptive triggering can recover lost accuracy in LLM reasoning while cutting down on unnecessary interventions, challenging traditional fixed-interval approaches.
Switching between natural language and structured graphs can boost multi-agent LLM performance by over 12 percentage points while slashing token usage by more than threefold.
Reducing visual token exposure by leveraging selective retrieval can dramatically enhance the performance of 3D medical image question answering systems.
Recursive self-improvement in LLMs can lead to unprecedented performance, with Meta^n surpassing all prior agents on challenging benchmarks by leveraging fixed meta-operations.
A dynamic internal field can govern computation in transformers, but it doesn't enhance cognitive performance—its true value lies in certifiable stability.
Dynamic delegation during reasoning allows LLM agents to outperform static routing strategies, achieving better task success rates with a Bayesian approach.
Conditional memory can either enhance or hinder scientific reasoning, and knowing when to activate it is crucial for optimal performance.
Steering latent dynamics with readout feedback can unlock performance improvements in recurrent models that traditional inference methods fail to achieve.
FedV-KGQA enables multi-hop reasoning across disjoint knowledge graph silos without compromising data sovereignty, achieving near-centralized performance.
Filtering invalid answers with symbolic constraints can boost precision in knowledge graph QA while retaining valuable candidates.
Trial Parallelism accounts for over 65% of reasoning computation in LLMs, and harnessing it can lead to significant speedups in problem-solving.
A staggering 72.9% of medical chain-of-thought rationales fail to influence model answers, raising critical questions about their role in clinical reasoning.
SRD redefines trajectory selection by enabling segment-level intervention, achieving better reasoning efficiency without the need for larger models.
A frozen instruct model can reshape a student's reasoning policy, enabling RL refinement that surpasses traditional methods without costly fine-tuning.
Counterfactual queries can be bounded using a linear programming approach that requires no complete causal graph, revealing insights even with incomplete domain knowledge.
STRIVE transforms longitudinal radiology reporting by integrating structured reasoning and verification, achieving unprecedented accuracy in temporal change detection.
Dynamic adjustment of reasoning paths based on continuous difficulty estimates can save up to 76% in token consumption without losing accuracy in complex reasoning tasks.
DeepRepoQA achieves substantial performance gains in code repository question answering by enabling LLM agents to perform multi-hop reasoning through systematic exploration.
PARTAB achieves significant improvements in table reasoning by intelligently partitioning evidence, outperforming existing methods even in complex scenarios.
SAGE reveals that a multi-agent, evidence-grounded approach can outperform larger models in understanding Chinese ancient documents, challenging the notion that size alone drives performance.
LLMs can dramatically improve taint analysis for Android apps, achieving an F1-score of 0.96 compared to traditional tools' 0.55.
LLMs struggle with raw electromagnetic signal analysis, scoring as low as 21.2% on complex system design tasks, highlighting a critical gap in their reasoning capabilities.
VLMs struggle with multi-agent tactical reasoning, revealing critical weaknesses in spatial understanding and interaction binding that current models fail to address.
The "handoff tax" reveals that switching between models can significantly degrade performance while increasing costs, challenging the assumption that stronger models always yield better outcomes in coding tasks.
AI usage in mathematical research surged from 1.39% to 14.09% of submissions in just five months, with Combinatorics leading the charge.
Newer GPT models show diminishing returns from structured prompting, while Qwen models thrive on Few-Shot techniques—highlighting a critical shift in how LLMs internalize prompting strategies.
Even the best LLMs struggle with Olympiad-level physics, achieving only 33.7% accuracy on a new benchmark that challenges their reasoning capabilities.
GaussVLA achieves a 19.7% improvement in spatial manipulation success rates while being more parameter-efficient than previous models.
Transforming long-form audio meeting comprehension, the GRGA model leverages graph-based planning to overcome acoustic loss and memory challenges, achieving superior QA performance.
LLMs can significantly boost forecasting accuracy, but their effectiveness is hampered by measurement limitations and sensitivity to input changes.
Nonsensical inputs can infiltrate LLM-generated knowledge graphs, but DARKSIDE offers a novel method to track and mitigate these risks effectively.
DRAgent achieves superior localization accuracy by transforming the RES task into a discriminative reasoning challenge, significantly reducing bias and errors in object segmentation.
OaK transforms LLM agents by grounding their decision-making processes in dynamically constructed ontologies, leading to enhanced reliability in multi-step reasoning.
Agents using ParallelWorld can efficiently evaluate multiple future trajectories, leading to superior decision-making in complex environments.
LLMs can now dynamically re-attend to procedural rules, boosting their reasoning accuracy by up to 19 points in complex scenarios.
Explicit strategy induction can dramatically enhance LLM performance, but its effectiveness varies widely across task types and configurations.
Thinking types in LRMs directly correlate with correctness, revealing that cognitive profiling can enhance reasoning performance.
Humor comprehension in AI just got a boost, with CaRGo-T improving performance by up to 20% through innovative graph-based reasoning.
Hierarchical supervision can boost uncertainty estimation in LLMs, enhancing decision-making in biodiversity monitoring by improving prediction accuracy significantly.
SRPO enables LLMs to self-reflect and transform sparse feedback into dense learning signals, achieving state-of-the-art performance with drastically reduced training costs.
Penalizing shifts in safety representations during reasoning fine-tuning can restore LLM safety without sacrificing performance, revealing a critical interplay between reasoning and safety in model training.
SAT solver metrics fail to predict human difficulty in Nonograms, revealing a disconnect between algorithmic and human solving strategies.
Mixed SFT outperforms next-chunk reasoning RL in reasoning tasks while using over 60 times less compute, reshaping our understanding of effective training strategies.
AI's role in mathematics could evolve from mere assistance to a transformative partnership, redefining how we engage with mathematical concepts.
Routing before reasoning can boost function-calling success rates in language models by over 12%, transforming how we approach tool interaction.
A hybrid LLM architecture achieved 99.2% accuracy in O-RADS classification, outperforming traditional methods and highlighting the potential for reliable clinical automation.
Qwen3-32B may be the most reliable at verdicts, but GPT-5 outshines it in mimicking human reasoning paths during scientific fact-checking.
Pruning 64.58% of reasoning tokens not only streamlines MLLM performance but also revitalizes the model's reliance on visual evidence, enhancing task accuracy.
A learnable moderator can transform multi-agent debates into more efficient reasoning processes, outperforming traditional methods by reducing redundancy and enhancing evidence aggregation.
Missing verbal evidence in VLM outputs can lead to significant reasoning errors, but SAVER recovers accuracy by enforcing explicit articulation of visual changes.
Students using Hazel Prover showed significant improvement in inductive reasoning, but initial excessive support hindered their ability to transfer skills to traditional assessments.
DIAG reshapes practice distribution to maximize informative supervision, leading to significantly improved reasoning performance in LLMs.
Fine-grained preference optimization at critical decision points can dramatically enhance the reliability of SQL query generation.
DG-Mem achieves substantial gains in reasoning tasks by leveraging a novel memory architecture that adapts dynamically without altering model parameters.
Spatial reasoning in VLMs operates on coarse object localization rather than precise boundaries, revealing a surprising disconnect between knowing where objects are and how they relate.
IntentQA reveals that understanding intent in videos requires more than just visual recognition, highlighting the critical role of cognitive context in achieving robust performance.
Fine-tuning compact vision-language models on WADE boosts detection performance but still leaves over 75% of floating waste instances undetected.
Post-training with LoRA can boost accuracy in some models while hindering others, revealing the nuanced interplay between architecture and adaptation in audio-dependent tasks.
VLMs blend real visual reasoning with a heavy reliance on language shortcuts, raising questions about their true understanding of visual relations.
Cross-language feature sharing in multilingual LLMs is model-specific and does not always translate to functional equivalence, challenging assumptions about multilingual reasoning.
FormuEvo can accelerate solver performance by up to 5.5× by intelligently evolving MIP formulations through LLM-guided optimization.
MLLMs can achieve precise interior design reasoning without fine-tuning, drastically reducing hallucinations and spatial errors.
Improving descriptive reasoning trace quality can actually hinder recommendation effectiveness, challenging assumptions about the benefits of interpretability in AI systems.
ReAct-SQL matches the accuracy of complex text-to-SQL systems while being up to 8 times faster and simpler.
Training LLMs with expert ICU reasoning leads to significant improvements in clinical decision-making across diverse medical tasks.
CFM achieves unprecedented stability in interdisciplinary climate reasoning, outperforming traditional models in handling complex trade-offs and uncertainties.
Fully-non-leaking wait-free implementations are possible for some concurrent objects, but not all—revealing critical limitations in information security for concurrent systems.
Multi-agent LLMs can drastically improve reasoning accuracy by redesigning agent interactions based on quantitative ethnographic insights.
A novel Tree-of-Thought framework achieves a 62% success rate in repairing smart contracts, significantly outpacing traditional linear methods.
Go programs can now be verified directly with formal methods, enhancing educational tools and developer productivity without altering the original code structure.
QAH enables 4-bit LLMs to outperform their bfloat16 counterparts while reducing memory usage by four times and achieving peak performance seven times faster than traditional methods.
Achieving 94.1% accuracy in spatial relation verification, this modular agent outperforms traditional models by a staggering 42.5 percentage points, highlighting the power of structured reasoning in medical imaging.
The exact form of discarded information in concentration inequalities reveals hidden structures in martingale trajectories that can optimize probabilistic testing strategies.
Adaptive reasoning in language models can reduce token usage by 41% while maintaining high accuracy, reshaping efficiency in AI computations.
Extracting hidden reasoning traces from black-box models is not only feasible but poses a substantial security risk, with EchoCoT achieving over 66% accuracy in retrieval.
BGCMs resolve the ambiguity of interventions in cyclic causal systems, allowing for precise causal reasoning where traditional models fail.
Integrating LLMs into autonomous driving systems can enhance decision-making without sacrificing control or safety, even in unpredictable environments.
LLMs systematically favor text over numbers in evidence arbitration, revealing a critical failure mode in decision-making systems that rely on heterogeneous data sources.
Current LLMs can barely translate natural-language claims into formal statements, achieving only 11.5 on a critical task that could unlock automated theoretical research.