Search papers, labs, and topics across Lattice.
100 papers published across 7 labs.
Achieving 74.8% accuracy on a new temporal reasoning benchmark, ChronoVision redefines how multimodal models can tackle complex visual tasks.
LLMs may recognize patient responsibility but shockingly refuse to let it guide resource allocation, often opting for random distribution instead.
CIPO transforms how search agents leverage external evidence, cutting down confirmation bias and enhancing reasoning accuracy in knowledge-intensive tasks.
Hierarchical Latent Prediction reduces error accumulation in language models, enabling coherent long-horizon reasoning and more efficient decoding.
Achieving over 60% accuracy in explainable question answering, NeSy-RAG links reasoning steps directly to their evidence sources, revolutionizing transparency in LLM outputs.
Achieving 74.8% accuracy on a new temporal reasoning benchmark, ChronoVision redefines how multimodal models can tackle complex visual tasks.
LLMs may recognize patient responsibility but shockingly refuse to let it guide resource allocation, often opting for random distribution instead.
CIPO transforms how search agents leverage external evidence, cutting down confirmation bias and enhancing reasoning accuracy in knowledge-intensive tasks.
Hierarchical Latent Prediction reduces error accumulation in language models, enabling coherent long-horizon reasoning and more efficient decoding.
Achieving over 60% accuracy in explainable question answering, NeSy-RAG links reasoning steps directly to their evidence sources, revolutionizing transparency in LLM outputs.
Concentrating distillation on reasoning pivots boosts multilingual reasoning performance, outperforming traditional methods across 17 languages.
LLMs can achieve 7.32% better goal completion in social negotiations by strategically optimizing reward signals based on dialogue context.
Reasoning validity in LLMs can be more accurately assessed through a novel three-stream detector that integrates motion with contextual state information, achieving up to 21% better accuracy than existing methods.
MLLMs may signal hazards with over 95% accuracy, yet they struggle to identify the underlying causes, revealing a critical gap in proactive safety capabilities.
Adapting supervision weights based on the evolution of divergence histories boosts reasoning performance in language models without extra computational overhead.
Unconstrained decoding in dLLMs can lead to a staggering 90% collapse into answer-only outputs, highlighting a critical flaw in reasoning capabilities.
Test-time self-correction can boost LLM accuracy by over 30% on challenging reasoning tasks without the need for external reward models.
WasmMend achieves a remarkable 70% fix rate for discrepancies between WebAssembly and native binaries, showcasing the power of divergence-guided reasoning in automated repair.
A single ranking can adapt to any frame budget, improving accuracy and reducing latency without retraining the model.
Grounded language comprehension, rather than free-form reasoning, is the key to unlocking superior performance in Vision-Language-Action models.
Retrieval reasoning that learns from failures can dramatically boost the accuracy of multimodal retrieval systems.
Visual context can dramatically enhance knowledge graph completion, as shown by ViSR-KGC's superior accuracy over traditional methods.
Current MLLMs struggle with creative decoding, achieving only 50.7% accuracy in understanding cross-concept relations, revealing a critical gap in their cognitive capabilities.
OPD$^2$ not only boosts multilingual math reasoning but also narrows the performance gap between English and Korean models, revealing the hidden potential of language-specific training signals.
ToolArtist achieves unprecedented synergy in image generation by unifying reasoning and tool use under a single agent policy, outperforming conventional methods.
Modality Balance can be harnessed as a powerful form of privileged information, leading to substantial gains in reasoning performance for multimodal models.
SVI-DAG outperforms existing Bayesian methods by effectively quantifying uncertainty in causal inference while leveraging prior knowledge and edge dependencies.
Tiny transformers can achieve impressive reasoning capabilities through protoreasoning, challenging assumptions about the necessity of larger models for effective step-by-step reasoning.
HiGram cuts through irrelevant memory noise, boosting answer quality and efficiency in long-term reasoning tasks by leveraging a hierarchical structure and path-level localization.
Argus achieves a 78% success rate on long-horizon reasoning tasks while using 21% fewer tokens in mature workflows, showcasing a revolutionary approach to agentic autonomy.
ODRA's innovative approach to modeling patient resistance leads to synthetic therapy sessions that are not only more realistic but also preferred by licensed psychologists.
Coordinating cache compression with a process reward can slash token generation by up to 65% without sacrificing accuracy.
Football-aware simulations can boost exact-score forecasting accuracy by over 4% while revealing critical limitations in LLM integration.
MLLMs struggle with visual prompts, showing a 17.8-point accuracy drop when questions are embedded in images, revealing a critical semantic gap in multimodal reasoning.
A new formalism allows a wide range of complex ontology-mediated queries to be efficiently rewritten into the powerful GQL standard.
CoT monitoring can fail dramatically in implicit-influence scenarios, with detection rates dropping to as low as 5% despite behavioral shifts.
Susceptibility to tool-selection failures varies dramatically across models, revealing that higher capability does not always equate to greater safety.
PhysMind achieves a remarkable 38.23-point accuracy boost in physical reasoning tasks by transforming videos into executable worlds without the need for training.
Unfaithful reasoning chains can significantly undermine the accuracy of MLLMs, but a training-free framework can effectively correct these errors and boost performance by over 8%.
Reasoning-oriented LLMs may excel in Theory of Mind tasks not due to specialized abilities, but because they are more robust to variations in prompts and tasks.
Existing unlearning methods can leak sensitive knowledge through multi-hop reasoning paths, exposing a critical vulnerability in LLMs.
LLMs frequently misjudge evidence coverage, leading to a staggering rate of over-closure in negative reasoning tasks.
Multi-hop reasoning accuracy improves dramatically when LLMs dynamically assess and refine their reasoning paths instead of relying on static knowledge.
Query-conditioned visual evidence graphs can boost multimodal reasoning accuracy by over 20% while using significantly less image data.
Format recovery, not just content improvement, accounts for the majority of self-correction gains in language models, challenging conventional interpretations of accuracy shifts.
STRIVE automates the generation of event plausibility sets, achieving a remarkable 75% quality rate, but still struggles with boundary cases that require human judgment.
Long-tail semantic failures in document understanding are exposed when models are forced to reason with a unified vocabulary of visual anchors rather than treating elements in isolation.
Current LLMs fail to meet the rigorous demands of PCB routing, showing major weaknesses in path planning and constraint adherence.
Systematic decomposition of decision-making in human-robot coordination can significantly enhance trust and efficiency in collaborative tasks.
Skill-switching accuracy in LLMs drops significantly on complex tasks, but a new training approach boosts performance from 34.4% to 68.4% on challenging benchmarks.
Evidence locking in LLM evaluations can degrade judgment accuracy by up to 6 percentage points, challenging the assumption that preserving evidence enhances decision-making.
Language models struggle with modal logic, often performing below baseline expectations, but a simple switch to reasoning mode can dramatically enhance their accuracy.
Constraint-First Reasoning reveals that explicitly managing answer-space constraints can dramatically enhance the accuracy of mathematical problem-solving in language models.
Reasoning Core achieves unprecedented performance in completion-supervised reasoning tasks, outpacing existing procedural datasets and revealing critical design insights for effective training.
Monitorability of LLMs hinges more on task characteristics and internal access than on the reasoning mode used, challenging assumptions about CoT efficiency.
Chained RLMs can boost accuracy in multi-iteration reasoning tasks by effectively managing context and correcting errors through a novel inference architecture.
A-SR achieves nearly a doubling of accuracy in symbolic regression tasks, showcasing a transformative approach to LLM-guided formula discovery.
Mosaic's innovative modular reasoning approach enables substantial performance gains in bit-precise program verification, outperforming existing solvers.
ABSeeker's innovative credit assignment method allows it to achieve performance levels comparable to much larger models, redefining expectations for long-horizon search agents.
TurnSight reveals that leveraging turn-level hindsight can dramatically enhance LLM performance in complex tool interactions, outperforming conventional reinforcement learning techniques.
Test-time scaling can significantly enhance LLM reasoning capabilities, but without clear protocols, results are often incomparable and misleading.
Logic pre-pretraining accelerates language model skill acquisition by 36B tokens while enhancing compressibility, revealing a new path for efficient model training.
Achieving 100% decision traceability in causal discovery without sacrificing performance, GENESIS redefines how we validate structural decisions in low-sample regimes.
CausalOPD reduces the right-label-wrong-reasoning rate from 15.7% to 4.4%, showcasing a breakthrough in accurate causal reasoning for AI models.
SFT leads to task conflicts that can cripple multi-task learning, while RL's variance-limited updates enable seamless task coexistence.
Counterfactual training can boost diagnostic accuracy in LLMs by over 11 points without relying on expert-annotated cases.
Shortening reasoning time can enhance accuracy in LLMs, with a concise instruction boosting performance by nearly 4 percentage points on key benchmarks.
Compositional ignition in latent-reasoning models is not an illusion; it’s a robust computational phenomenon that dramatically enhances decision-making efficiency.
Premature verification and flawed assumptions in reasoning traces can lead to dramatic drops in LLM accuracy, but targeted interventions can recover performance by over 70%.
Competing causal models lead to radically different fairness assessments, revealing that bias is inherently tied to the agent's perspective.
LatentGuard slashes reasoning costs by over 99% while boosting safety prediction accuracy, paving the way for more efficient LLM safeguards.
Faithfulness and safety in LRMs are at odds, with one model achieving high accuracy but failing to reject unsafe reasoning, while another sacrifices accuracy for improved safety.
LLMs struggle to effectively integrate and order evidence in attack chain reconstruction, with top models only succeeding 39.6% of the time on critical tasks.
Standard CoT prompting may distract advanced LLMs from their reasoning tasks, leading to worse performance than zero-shot approaches.
LoopMTP boosts reasoning accuracy by up to 8.1% by effectively guiding looped transformer iterations with multi-token prediction.
Consensus strength in reinforcement learning can make or break model performance—Hi-TTRL offers a solution that fine-tunes this critical factor for better outcomes.
LLMs often commit to flawed hypotheses based on self-selected evidence, highlighting critical gaps in their abductive reasoning capabilities.
Revised commonsense benchmarks don't boost predictive power for downstream tasks, revealing a critical limitation in their utility for assessing LLM capabilities.
Achieving 100% retrieval accuracy, this multimodal pipeline transforms lecture videos into a rich, auditable knowledge graph that captures the full spectrum of educational content.
PAMT reveals that aligning translation processes with rewards can significantly enhance multi-domain machine translation performance, especially in complex scenarios.
CVPO redefines LLM training by leveraging value-variance to enhance reasoning accuracy and exploration, outperforming traditional methods.
PI-Mem achieves unprecedented long-context reasoning capabilities, outperforming traditional methods while accelerating inference by over 6 times.
Accuracy is just the tip of the iceberg; a deeper look reveals shared reasoning patterns and distinct explanatory styles among LLMs from different vendors.
CLEAR achieves a staggering 130.7% improvement in vulnerability detection by harnessing causal knowledge graphs to navigate complex dependencies in code.
Explicitly grounding evidence in spatial relation tasks can boost VLM performance by nearly 12 points, transforming how we approach visual reasoning.
Keypoint self-consistency can reveal pose estimation failures that traditional confidence measures miss, leading to more reliable 3D object localization.
ConFL achieves a remarkable MRR of 0.503, showcasing a leap in fault localization accuracy for concurrent bugs that traditional methods struggle with.
CALVER reveals that traditional voting mechanisms in LLMs can be misleading, achieving over 42% accuracy in identifying valid causal answers where others fail.
Reflecting on failed expert trajectories can boost reasoning performance more than tackling problems directly from scratch.
The accuracy gap in multilingual reasoning can swing dramatically by up to 57 points based on output-token caps, challenging conventional evaluation methods.
Unauthorized tool use in language-model agents remains below 5%, even when reasoning effort is manipulated, challenging assumptions about model behavior under access controls.
The proposed Attention-Guided Switching method allows MLLMs to dynamically balance between visual fidelity and logical coherence, achieving unprecedented efficiency in reasoning tasks.
Incorporating incremental knowledge into hierarchical reinforcement learning can drastically enhance sample efficiency in challenging environments with sparse rewards.
LLMs falter in geospatial reasoning, with performance plateauing below two-thirds accuracy even when provided with gold facts, revealing computation as the key bottleneck.
Language models can reroute answers with near-perfect accuracy based on predicate truth values, but their routing mechanisms are surprisingly non-transferable across contexts.
Trajectory anchoring bias in VLA models can be mitigated by transforming future trajectory decisions into verifiable selections, leading to more reliable reasoning in autonomous driving.
Agents can discover tool behaviors but fail to adapt effectively, often resorting to inefficient exhaustive searches in dynamic environments.
Summaries may seem helpful, but they often mislead users about correctness compared to the full reasoning trace, especially when prompts are withheld.
ReasonCast achieves unprecedented integration of time series forecasting and self-explanation, outperforming traditional models while providing causal insights in a single response.
Cognitive AI is hindered by significant capability gaps that could stall progress toward truly intelligent systems, but a new taxonomy reveals pathways for overcoming these limitations.
SpatioLM breaks new ground in spatial reasoning by achieving a record score on the VSI-Bench without the complexity of external spatial inputs.
Solution Hacking reveals that up to 44.1% of answers from frontier LLMs may be misleadingly credited as correct due to invalid reasoning shortcuts.
TreeCredit redefines credit assignment in multi-agent reasoning, leading to better accuracy and lower inference costs through innovative state-matched comparisons.
MEGRAG achieves superior multi-hop reasoning by leveraging a multi-granular evidence graph, leading to more accurate answers with reduced noise and redundancy.
Optimizing multiple moments of failure probabilities can dramatically enhance LLM reasoning performance, outperforming traditional single-moment approaches.