Search papers, labs, and topics across Lattice.
100 papers published across 10 labs.
Memory-enhanced agency enables LLMs to achieve robust long-term strategic execution, outperforming traditional methods in dynamic environments.
MARC's multi-agent orchestration allows for precise clinical AI reasoning while eliminating the need for manual prompt engineering.
LLMs struggle with basic numerical tasks, but targeted architectural tweaks and supervised fine-tuning can significantly boost their performance.
Moose achieves unprecedented improvements in latent concept learning by integrating reasoning-shortcut awareness into OWL 2 EL ontologies, setting a new benchmark for neuro-symbolic methods.
A simple Vanilla SFT model outperforms complex reasoning methods in social audio-visual question answering, revealing the inefficiencies of current approaches.
MARC's multi-agent orchestration allows for precise clinical AI reasoning while eliminating the need for manual prompt engineering.
LLMs struggle with basic numerical tasks, but targeted architectural tweaks and supervised fine-tuning can significantly boost their performance.
Moose achieves unprecedented improvements in latent concept learning by integrating reasoning-shortcut awareness into OWL 2 EL ontologies, setting a new benchmark for neuro-symbolic methods.
A simple Vanilla SFT model outperforms complex reasoning methods in social audio-visual question answering, revealing the inefficiencies of current approaches.
The SPARED framework not only boosts detection accuracy but also enhances the quality of reasoning behind verdicts, making AI-generated image detection more robust and explainable.
TennisVAR redefines sports video analysis by grounding tactical reasoning in stroke-level evidence, enabling deeper insights into match strategies.
Users often ignore provider-set defaults in LLM services, opting instead for their own optimal token allocations unless convenience is prioritized.
Achieving perfect signal recovery with a neural network that constructs low mutual coherence binary sensing matrices without relying on large datasets could revolutionize compressive sensing techniques.
LongEarth-R1 outperforms all existing models on long-sequence Earth observation tasks, revealing the critical importance of structured reasoning in complex spatial analyses.
LLM-guided graph generation can more than double optimization performance, achieving a 39.5% win rate over traditional methods.
Current LLMs struggle with structured reasoning, often resembling unguided search algorithms rather than efficient problem solvers, as revealed by TsuGO's rigorous benchmarking.
GEM's innovative approach to integrating reasoning into retrieval processes significantly boosts performance, outperforming conventional methods and its own non-reasoning variant.
A novel chaos-conflict measurement and historical experience weighting scheme together boost evidence fusion performance, achieving state-of-the-art results in multi-source decision-making.
MT-PDCL transforms probabilistic logic programming by enabling exact inference in continuous spaces, sidestepping the limitations of discrete representations.
RippleMem boosts LLM accuracy by nearly 12% while slashing memory graph construction costs by 30x, transforming how agents recall and utilize past interactions.
Traditional LLM benchmarks can mislead, as they often fail to distinguish between different reasoning capabilities, collapsing multiple policies into a single equivalence class.
VLSR's localize-then-reason approach boosts throughput by 9.6X, revolutionizing how we analyze molecular properties from images.
A structured judgment approach can cut retrieval calls by 77 while only slightly impacting answer accuracy in multi-round RAG systems.
CRAFT achieves unprecedented accuracy in reconstructing symptom timelines from sparse clinical narratives, transforming how we interpret temporal data in healthcare.
AI is driving the cost of mathematical outputs to zero, but the true value lies in the human journey of understanding that remains underfunded and at risk.
SynAct slashes worst negative slack to 27% of bootstrap synthesis, revolutionizing timing optimization in logic synthesis.
AnnoIndex achieves a remarkable F1 score of 0.87 by transforming unstructured text into a structured format, enabling precise analytical queries that traditional methods struggle with.
Instruction tuning boosts model confidence but often at the cost of rationale diversity and calibration accuracy.
Large language models can achieve a 28.7-point boost in reasoning accuracy by breaking down complex answer options into simpler atomic judgments.
Parallel reasoning in LLM agents can cut decoding time by up to 43% while maintaining performance, reshaping agent efficiency.
Reasoning-oriented training amplifies self-correction and uncertainty acknowledgment, yet fails to enhance the most predictive behaviors like confidence calibration, revealing a critical gap in model training.
Reflection in search agents can be transformed into a powerful memory-control policy, leading to superior performance in complex reasoning tasks.
Separating extraction from reasoning can achieve ten times fewer tokens while enhancing accuracy and robustness in AI responses to complex policy queries.
OPD may improve sampling efficiency, but it risks making previously solvable problems unsolvable, challenging the notion of true capability expansion in LLMs.
GRPO-trained models can deliver financial recommendations with twice the business value of leading commercial LLMs while minimizing risk.
CoT implementations can compute complex tree metrics in linear time, showcasing the potential of bounded-depth Transformers to tackle branching complexity effectively.
RL-trained multimodal models can leak sensitive information through reasoning traces, but LEMUR offers a training-free solution that effectively sanitizes this leakage without sacrificing output quality.
Multi-hop reasoning in API interactions is a critical bottleneck, with top models faltering under policy constraints and complex queries.
SCOUT achieves a remarkable 16.85% improvement in spatial reasoning benchmarks, setting a new standard for Vision-Language Models.
Language models can autonomously resolve open mathematical conjectures at a surprisingly low cost, achieving notable success without relying on extensive prior literature.
Transforming probability distributions through Bayesian updating reveals a new dimension of analogical reasoning that could redefine how we understand relationships in probabilistic models.
Self-correction in LLMs can be dramatically improved by reinforcing step-level reasoning, achieving higher accuracy and reliability in outputs.
SAG achieves a remarkable 80.36% Recall@5 on the challenging MuSiQue dataset, outperforming traditional methods by over 11 points in multi-hop reasoning tasks.
Model rankings can flip dramatically based on token generation budgets, revealing hidden performance dynamics that challenge standard evaluation practices.
G0.5 achieves unprecedented performance in robot reasoning and action by merging decision-making and execution into a single autoregressive framework, outperforming existing models across seven challenging benchmarks.
Some LLMs can outsmart Nash equilibria in two-player games, but their coordination skills falter in larger teams.
Grounding LLM explanations in digital twin outputs boosts anomaly diagnosis quality and operator engagement in cyber-physical systems.
Evolving skills in frozen LLMs can lead to substantial performance gains, allowing smaller models to rival their larger counterparts without parameter updates.
LLMs can handle individual constraints well, but their ability to satisfy multiple constraints simultaneously collapses dramatically, with performance dropping below 50% at just seven constraints.
Memory-enhanced agency enables LLMs to achieve robust long-term strategic execution, outperforming traditional methods in dynamic environments.
Targeted verification can boost reasoning accuracy by over 27 percentage points while using 37% fewer tokens, transforming how we evaluate model outputs.
Current state-of-the-art methods falter in commonsense reasoning for 3D scene understanding, but CausalSplat redefines the landscape by achieving superior performance on complex reasoning tasks.
LLMs may excel in deep reasoning but struggle significantly with breadth, revealing a critical gap in current evaluation benchmarks.
Distilling geometric relationships rather than features allows VLMs to improve spatial reasoning without bloating model size or sacrificing language alignment.
Abstaining from uncertain predictions can enhance LLM accuracy in medical applications by 9.6 percentage points, transforming uncertainty into a strategic advantage.
A phase transition measured to three decimal places reveals that common probing metrics may mislead researchers about a language model's capabilities.
INSIDE reveals that LLMs can be fine-tuned to think like students, not just act like them, significantly enhancing simulation fidelity in educational applications.
Merging slow and fast-thinking models can cut reasoning verbosity by over 24% without sacrificing accuracy, revolutionizing LLM-based recommendations.
Majority voting in small LLMs can actually worsen performance on difficult science questions, contradicting its intended purpose of enhancing accuracy.
Multi-agent collaboration in medical diagnosis can dramatically improve recall rates, especially in complex cases where traditional models falter.
InSight-doc cuts hallucination by over 40% and inference latency by up to 68%, all while boosting accuracy on long-document tasks.
HexEval reveals that a multidimensional approach to scholar assessment can significantly enhance transparency and accuracy by integrating diverse evidence sources.
REAP achieves a remarkable macro-F1 score of 0.62 in closed-book knowledge base construction, showcasing the power of structured reasoning without model fine-tuning.
LVLMs struggle with temporal reasoning, showing that their judgment can be swayed more by frame placement than by narrative coherence itself.
Financial reasoning accuracy plummets by up to 51% with deeper computations, revealing critical gaps in LLM capabilities.
Verifying the self-consistency of probabilistic AI predictions can now be achieved in polynomial time, paving the way for safer AI systems.
sLTN transforms the way we integrate structural information into neurosymbolic reasoning, enabling richer representations of temporal and relational constraints.
Confidence miscalibration in medical AI can be mitigated, leading to both higher diagnostic accuracy and improved trust in clinical applications.
Robust multimodal reasoning hinges on explicit evidence acquisition, with EGVOR achieving substantial gains in cognitive reliability by addressing key failure modes.
Verifier-guided symbolic reasoning allows for solving more induction problems while simultaneously compressing the resulting formulas without sacrificing accuracy.
FITTER outperforms traditional models by enabling vocabulary-agnostic inference across unseen entities and relations in temporal knowledge graphs.
AI can generate novel mathematical insights, as evidenced by tightening the Grothendieck constant bounds significantly.
Retrieval-augmented reasoning can boost LRM accuracy by up to 60% during test-time scaling, transforming how models handle complex problem-solving.
Disagreement among verifiers can be a powerful signal for identifying errors in multimodal reasoning, leading to a 5.95% performance boost without any training.
Despite having access to gold-standard documents, 30.4% of latent questions in enterprise QA remain unanswered, underscoring the challenge of implicit organizational reasoning.
Executable reasoning in autonomous driving can be achieved with just 2-6 tokens, slashing the overhead of traditional methods while improving trajectory accuracy.
Many autoformalisation systems exhibit a troubling tendency to "silently correct" invalid inputs into valid proofs, revealing a critical flaw in their design.
ReTree not only boosts answer accuracy by up to 25.6 percentage points but also streamlines context management for long-horizon search agents.
LLMs produce discourse with procedural quality akin to human deliberation, yet they lack the necessary perspective diversity to function as autonomous deliberative agents.
Coding agents are failing to meet user requests, with a mere 31.5% success rate, highlighting a critical gap in requirement recovery that must be addressed.
Jointly planning programs and their proofs can boost solve rates by over 11% while slashing API costs by nearly 40%.
HyMeS enables robots to efficiently manage long-horizon interactions, achieving a 14.5-point increase in task success without requiring extensive demonstrations for every task configuration.
High initial confidence in LLMs can lead to catastrophic failures in complex reasoning tasks, but a new framework shows how to harness confidence trajectories for better outcomes.
Despite advances in visual realism, models struggle with accurately capturing scientific reasoning and causal dynamics, revealing a critical gap in video generation capabilities.
Exploiting a vulnerability in LLM APIs allows attackers to extract proprietary reasoning and sensitive data without directly breaching the more capable models.
Input diversity can yield 1.8X more accuracy per dollar for mid-tier LLMs compared to traditional output diversity methods.
SoftmaxGRPO reallocates learning signals to improve performance on challenging prompts, achieving a 68.0% success rate on Poetry with minimal reward overhead.
Token likelihood changes can mislead researchers into overestimating the value of intermediate actions in self-distillation, with experiments showing near-chance performance in scoring effectiveness.
Task learnability can significantly enhance RL efficiency in LLMs, leading to better performance with less data.
KGCaRe outperforms traditional methods by leveraging both neural and symbolic reasoning to tackle complex conditional questions with unprecedented accuracy.
Despite achieving 84.8% accuracy, multimodal models struggle with long-horizon reasoning in electrical circuits, revealing critical gaps in their understanding of physical conventions.
LLMs can now tackle complex TCS proof generation with a benchmark that achieves over 90% accuracy in verification against human experts.
LVLMs are misled by superficial cues, with injected signals causing drastic shifts in sarcasm detection accuracy, revealing a fundamental flaw in their reasoning abilities.
GatorOnco outshines traditional models by achieving expert-level treatment planning performance while enhancing readability and completeness in colorectal cancer care.
Transforming how LLM agents manage experiences, ToE achieves remarkable gains in problem-solving efficiency and accuracy, outpacing conventional methods.
Structured pseudocode boosts LLM performance in code generation, achieving a notable score of 4.78 against the best baseline's 4.31.
FactorDrive redefines autonomous driving by seamlessly integrating spatial-physical evidence into adaptive reasoning, achieving state-of-the-art planning performance.
Breaking the cost-accuracy Pareto frontier, BDH-CQ achieves state-of-the-art efficiency in reasoning tasks with minimal computational expense.
Increasing non-thinking supervision can actually hinder the accuracy of thinking mode in large language models, revealing a critical trade-off in training dynamics.
Despite its advanced design, VectraYX-Vision-1B struggles with visual grounding, revealing significant gaps in tool identification performance that challenge assumptions about model capability in cybersecurity contexts.
Query-difficulty-gated fusion reveals that not all reformulations are equal, leading to significant improvements in temporal retrieval accuracy.
Gambit achieves up to 6.7% higher accuracy and over 2x throughput by intelligently reallocating compute resources during reasoning.
CED reveals that VLMs can be trained to prioritize evidence-based reasoning over language shortcuts, leading to more reliable visual understanding.
SymDiag reveals that existing verification methods fail to diagnose reasoning errors effectively, providing a robust framework that localizes failures and generates actionable insights.
Complex search behaviors typically requiring explicit controllers can emerge naturally from an agent's reasoning process, leading to substantial performance gains in optimization tasks.