Search papers, labs, and topics across Lattice.
100 papers published across 5 labs.
The integration of LLMs, knowledge bases, and reasoning capabilities could redefine how AI agents learn and operate in dynamic physical environments.
The exact form of discarded information in concentration inequalities reveals hidden structures in martingale trajectories that can optimize probabilistic testing strategies.
Adaptive reasoning in language models can reduce token usage by 41% while maintaining high accuracy, reshaping efficiency in AI computations.
Extracting hidden reasoning traces from black-box models is not only feasible but poses a substantial security risk, with EchoCoT achieving over 66% accuracy in retrieval.
BGCMs resolve the ambiguity of interventions in cyclic causal systems, allowing for precise causal reasoning where traditional models fail.
The exact form of discarded information in concentration inequalities reveals hidden structures in martingale trajectories that can optimize probabilistic testing strategies.
Adaptive reasoning in language models can reduce token usage by 41% while maintaining high accuracy, reshaping efficiency in AI computations.
Extracting hidden reasoning traces from black-box models is not only feasible but poses a substantial security risk, with EchoCoT achieving over 66% accuracy in retrieval.
BGCMs resolve the ambiguity of interventions in cyclic causal systems, allowing for precise causal reasoning where traditional models fail.
Integrating LLMs into autonomous driving systems can enhance decision-making without sacrificing control or safety, even in unpredictable environments.
LLMs systematically favor text over numbers in evidence arbitration, revealing a critical failure mode in decision-making systems that rely on heterogeneous data sources.
Current LLMs can barely translate natural-language claims into formal statements, achieving only 11.5 on a critical task that could unlock automated theoretical research.
G-MARK reveals that grounding multi-agent reasoning in provenance-aware knowledge graphs can drastically enhance occlusion reasoning and decision-making in cooperative driving.
Iterative perception can boost document VQA accuracy by over 25%, proving that smarter evidence acquisition trumps mere model size.
Trusting individual predictions from VLMs can be achieved without fine-tuning, revealing critical insights into model failures that standard self-consistency checks overlook.
Optimizing latent visual representations can boost multimodal reasoning performance by over 9% on complex tasks.
Uncertainty in LLM reasoning is largely a sampling artifact, allowing for more efficient analyses that cut costs without sacrificing accuracy.
SABET-QA outperforms existing methods by effectively refining reasoning states across multiple hops, making it a game-changer for complex temporal queries.
Software 3.0 could redefine the entire software engineering landscape by merging reasoning and context into a unified architecture.
The integration of LLMs, knowledge bases, and reasoning capabilities could redefine how AI agents learn and operate in dynamic physical environments.
Memory Correlation Bias can mislead multi-agent systems into false majorities, but CAMA effectively counters this by recovering independent evidence from correlated memories.
Cognitive traps in LLM memory can lead to over 10% performance degradation, challenging the assumption that more memory always improves reasoning.
A novel deductive verification framework for weighted programming reveals how to express and automate verification conditions for complex quantitative models.
LLMs are miscalibrated in their reasoning, failing to distinguish between scenarios where valid solutions exist and where they do not, which could undermine their effectiveness in real-world applications.
VAKE reveals that over 80% of the knowledge activated through explicit priming is crucial for answering questions, showcasing a new pathway to enhance LLM factual accuracy.
Shared reasoning structures can significantly boost performance in continual reinforcement learning, with a novel replay mechanism achieving parity with multitask training.
Pair-Aware Discriminative Reasoning in UMER reveals critical distinctions between semantically similar candidates, elevating retrieval accuracy in multimodal tasks.
Cost-bounded self-verification allows LLMs to self-correct efficiently, cutting down on response generations while maintaining accuracy.
ORBITER significantly boosts decision-making reliability in last-mile delivery, outperforming existing models by up to 9.2% through enhanced reasoning about spatiotemporal cues.
Simplifying OWL class expressions can enhance reasoning efficiency by reducing complexity without sacrificing formal semantics.
DentAgent outperforms senior specialists by 17.3 percentage points in multi-label diagnosis, revolutionizing multimodal dental reasoning with traceable evidence integration.
PWAL outperforms traditional methods by boosting accuracy by up to 30.86 percentage points while providing a transparent trace of logical decision-making in enthymeme completion.
DRB can optimize reasoning budgets to improve LLM performance while cutting costs, achieving better results than traditional maximum-budget approaches.
AFANet achieves high accuracy in agent failure attribution with a fraction of the computational resources required by traditional LLM-based methods.
A contract-aware proof-repair tool can restore verification integrity in complex theorem proving, but challenges in operational end-to-end verification remain.
A nuanced correctness boundary in Romanov's Triplet Logic reveals that a satisfying set does not guarantee a non-empty intersection, challenging existing assumptions in Boolean satisfiability.
StateTrace reveals that enhancing memory structures in VLMs can drastically improve reasoning about invisible objects, achieving a remarkable 24.6-point performance increase on hidden-state tasks.
Current vision-language models can identify context in driving videos but fail to accurately assign fault in accidents, revealing a critical gap in autonomous driving AI capabilities.
Teacher rewards can mislead learning, but R2-OPD filters out conflicting feedback to boost reasoning performance in language models.
Group-Calibrated On-Policy Distillation boosts long-context reasoning performance by reconciling teacher guidance with verifier feedback, achieving up to a 12-point increase in benchmark scores.
Verification schemes for LLMs reveal a critical blind spot: while they can confirm correctness, they often miss potential errors entirely.
A single lightweight model can now achieve superior recommendation performance by leveraging a structured memory of reasoning, eliminating the need for costly repeated computations.
LLMs can miss 40% of the necessary calculations in environmental science, revealing a critical gap in their reliability for quantitative tasks.
Leveraging SMT conflict counts, SMTrap achieves unprecedented denial-of-service effects against large reasoning models without the need for model feedback or GPU resources.
A neuro-symbolic approach reveals how LLMs can effectively generate implicit premises, transforming the landscape of argument reconstruction.
Majority voting in LLM outputs can mislead consensus, with accuracy on hard questions plummeting despite high agreement rates.
Optimizing the condition number in quantum algorithms could enable a dramatic reduction in circuit complexity, making quantum solutions for Boolean systems more practical than ever.
Grounded reasoning in educational VQA can boost accuracy by over 2.5% when leveraging structured pedagogical cues.
Bridging the intent gap in e-commerce, TTP boosts order volume by 0.46% through reasoning-driven personalized retrieval.
Retrieval architecture can dramatically influence AI performance, with structural failures outnumbering reasoning errors by a staggering 95 to 15 in financial reconciliation tasks.
GUPO reveals that accounting for gradient uncertainty can dramatically improve policy optimization in post-training LLMs, leading to more effective reasoning capabilities.
Understanding the dynamics of knowledge transfer reveals that a tailored curriculum can dramatically enhance performance across reasoning tasks in large language models.
Decodable information in LLMs doesn't guarantee actionable outputs, revealing a critical gap in how these models handle geometric constraints.
Baobab reveals that a mixture indexed by query justifications can outperform independent perception in reasoning tasks, achieving Bayes-optimal performance where others fail.
J64 reveals hidden reasoning states that can significantly boost model accuracy and decision-making, while R64 provides a lightweight, effective proxy for deployment.
SIRLM significantly boosts the reasoning fidelity of LLMs in Knowledge Graph tasks, outperforming existing methods by tightly coupling structural knowledge with parametric learning.
Achieving over three times the accuracy of an untrained model, this research uncovers the transformative potential of domain-specific training in signal mathematical reasoning.
Recirculation enables foundation models to achieve a 23% reduction in perplexity and a 21% increase in accuracy without any added generation latency.
OGA reveals that the G\"odel sentence can be embraced as a consistent fiction, fundamentally altering our understanding of provability and incompleteness in arithmetic systems.
REChart slashes reasoning token usage by 79% while achieving state-of-the-art chart-editing performance, tackling the "overthinking" problem in large reasoning models.
Current models may excel in accuracy but often fail to ground their predictions in the visual evidence, with GPT-5.6 achieving only 3.93% QExact despite a seemingly high overall accuracy.
Candidate conditioning can boost accuracy by nearly 30% when multiple correct answers are present, but surprisingly, it can hinder performance when all candidates are incorrect. WHY_IT MATTERS: This insight challenges existing assumptions about candidate conditioning, potentially reshaping strategies for efficient test-time reasoning in AI systems.
LLM-derived preference judgments reveal significant inconsistencies, undermining the reliability of using a single utility function for decision-making.
Smaller models can outperform larger counterparts in linguistic reasoning tasks, challenging assumptions about model size and capability.
Deep LLMs compress vast semantic distances into navigable pathways, revealing a surprising topological phase transition that enhances reasoning capabilities.
MLLMs struggle with reasoning on complex documents, with even the best models showing significant performance gaps on BEAR-Bench.
CoAL-RAG achieves a remarkable 42.5% boost in BLEU scores for legal question retrieval, redefining efficiency and interpretability in complex legal consultations.
Conditional branching in navigation tasks reveals hidden failures in agent decision-making that standard metrics overlook.
Settling operators can significantly boost accuracy on unseen tasks, transforming additional iterations into a performance advantage rather than a risk of degradation.
Cooperative multi-agent training can unlock unsupervised reasoning capabilities in RL, yielding performance gains that rival supervised methods without the need for costly annotations.
Traditional accuracy metrics can obscure critical improvements in reasoning fluency, as shown by fine-tuned models that reason correctly in low-resource languages despite initial benchmark null results.
Iterative experience can boost LLM performance by over 5%, transforming how we evaluate and enhance model capabilities in real-time.
Uncertainty in multimodal models isn't just a number; it can fundamentally alter decision-making quality under complex evidence conditions.
Temporal confidence allows for adaptive computation in parallel reasoning, leading to a 32% reduction in latency without sacrificing accuracy.
VLMs exhibit a surprising reversal in modality reliance, favoring degraded visuals over text in chart-related tasks, challenging assumptions about their fixed behavior.
LLMs can generate structurally sound legal analyses but often fall short in substantive reasoning, raising questions about their reliability in legal contexts.
HarnessEval-W transforms world model evaluation from mere scoring to a transparent reasoning process that mirrors human judgment.
LLMs consistently underperform in shared-budget reasoning tasks, revealing a critical gap between their single-task capabilities and multi-task resource allocation.
MIRROR exposes the critical flaws in radiology AI reporting by ensuring that generated text can only reflect what the model has actually classified, not what it might imply.
PCs can perform exact probabilistic inference in polynomial time, addressing the NP-hard challenges that have long plagued traditional probabilistic models.
PertMind reveals that leveraging cellular perturbation data can significantly enhance LLMs' biological reasoning capabilities without extensive task-specific retraining.
GRIP cuts query-latent mutual information by 30x while slashing hallucinations by 73%, revolutionizing how we leverage retrieved evidence in reasoning tasks.
Fuzzy semantics for LTLf can dramatically enhance predictive performance and scalability, outperforming traditional methods without relying on automata.
Automating the analysis selection in Robustness Validation can cut design time and enhance thoroughness, transforming how automotive components are validated.
STAIR achieves a remarkable 16.57% improvement in F1 scores by separating semantic interpretation from temporal inference, making temporal reasoning both interpretable and verifiable.
Models trained with ACA-RL not only outperform on missing-premise tasks but also redefine how we evaluate reasoning under uncertainty in NLP.
Latent-OPD reveals that distilling latent representations at trajectory endpoints can dramatically enhance video reasoning efficiency in LMMs.
Spec-Driven Test Generation boosts bug detection rates by nearly 10% by making LLMs reason about code contracts before generating tests.
TDD-Agent transforms how LLMs generate code by making tests integral to the development process, leading to higher correctness and more effective tests.
RGA not only defines its own truth predicate but also proves that internal truth and provability are equivalent, challenging foundational limitations in classical arithmetic.
TRACE can verify previously unverifiable circuits, pushing the boundaries of formal hardware verification for complex arithmetic operations.
By shifting from problem selection to research direction, this framework uncovers a wealth of mathematical conjectures that would otherwise remain hidden.
RUPA reveals that accurately modeling relational dependencies can dramatically enhance uncertainty quantification for LLM agents, leading to earlier failure detection and improved task execution.
Looping mechanisms in language models can dramatically improve the accuracy and efficiency of multi-step API interactions, reshaping agentic tool use.
OPD transfers reasoning skills rather than answers, revealing a complex interplay between teacher-student origins that can either enhance or hinder model capabilities.
Selective visual-text compression in SEER boosts extraction precision while slashing token usage, outperforming leading models in long-context reasoning tasks.
Selective navigation in hybrid knowledge graphs allows agents to outperform full-context models, achieving better results with less information.
Recovering 87% of accuracy lost to KV eviction could redefine efficiency in reasoning tasks for large language models.
Language models struggle to consistently encode the current year, with associative and declarative representations diverging in their update responses.
Multi-3DLLM not only excels at multi-object reasoning but also boosts performance on single-object classification tasks, revealing a surprising synergy in geometric understanding.
Internalized Visual Thinking enables models to reason about future video frames directly, slashing inference time while boosting accuracy.
Achieving similar performance to larger models with significantly less data and faster inference speeds could redefine efficiency benchmarks in foundation models.
Short-context models can achieve superior reasoning performance by leveraging long-context teacher models through innovative token alignment and training strategies.
MathForm-8B not only outperforms specialized autoformalizers but also sets a new standard for verified mathematical formalization with an impressive 88.06% pass rate on syntax checks.
MARC's multi-agent orchestration allows for precise clinical AI reasoning while eliminating the need for manual prompt engineering.