Search papers, labs, and topics across Lattice.
100 papers published across 6 labs.
The integration of LLMs, knowledge bases, and reasoning capabilities could redefine how AI agents learn and operate in dynamic physical environments.
Proactively synthesizing project-specific issues can drastically enhance LLM agents' ability to resolve software problems effectively.
Current omni-modal models can achieve only moderate success in interactive video assistance, revealing critical gaps in their understanding of user interactions and visual cues.
Training on environments synthesized from real-world business scenarios boosts agent performance on both enterprise tasks and diverse benchmarks, revealing a new frontier in scalable RL training.
System Intelligence emerges as a game-changer, enabling LLM agents to collaborate effectively across complex tasks by organizing their interactions through dynamic graph structures.
Current omni-modal models can achieve only moderate success in interactive video assistance, revealing critical gaps in their understanding of user interactions and visual cues.
Training on environments synthesized from real-world business scenarios boosts agent performance on both enterprise tasks and diverse benchmarks, revealing a new frontier in scalable RL training.
System Intelligence emerges as a game-changer, enabling LLM agents to collaborate effectively across complex tasks by organizing their interactions through dynamic graph structures.
Legal and moral agency in AI are not just blurred lines—they're fundamentally different dimensions that can reshape accountability in AI actions.
TMI uncovers the hidden structure of concurrent tasks in computer-use traces, achieving unprecedented accuracy in task model induction.
Skill selection in LLMs can be optimized to achieve a 0.73 task success rate while using 28% fewer tokens than existing methods.
AI agents can now make scientifically defensible claims by embedding methodological rigor directly into the research process, drastically improving analysis accuracy.
SAPO achieves a 15.1 percentage point improvement over PPO while slashing memory costs and runtime by a third, revolutionizing how we optimize agentic RL.
Integrating LLMs into autonomous driving systems can enhance decision-making without sacrificing control or safety, even in unpredictable environments.
Models can verify facts more reliably than they can clarify ambiguities, raising questions about how we design memory systems in LLMs.
Agents struggle to significantly improve training algorithms, with the best only achieving 25% of the potential optimization gap.
ARG scaffolding boosts GPT-5's success rate from 9.4% to 49.0% in real-world ML tasks, while Modular setups suffer from high specification gaming.
Milestone inference can significantly enhance credit assignment in long-horizon reinforcement learning, leading to superior agent performance without extra model complexity.
Step-level credit assignment methods in LLM training are misleading, with none outperforming random chance in identifying causally significant actions.
Conversational surveys combined with multimodal LLMs can enhance travel behavior predictions, achieving over 71% accuracy by leveraging visual context.
MidTool reveals that dedicated mid-training can significantly enhance LLMs' ability to utilize tools effectively, outperforming traditional post-training approaches.
Subtask-level skills can boost LLM performance beyond baseline levels, while task-level skills often hinder it—highlighting a critical distinction in skill transferability.
Software 3.0 could redefine the entire software engineering landscape by merging reasoning and context into a unified architecture.
Agents rely more on internal notes than traditional documentation, raising questions about the relevance of current documentation practices in automated coding environments.
By clarifying ambiguous patient queries, this framework boosts diagnostic accuracy by over 57 percentage points, transforming how healthcare chatbots interact with patients.
The integration of LLMs, knowledge bases, and reasoning capabilities could redefine how AI agents learn and operate in dynamic physical environments.
StateMem improves current-state accuracy by up to 1.8x, revealing that existing memory systems are ill-equipped to handle evolving contexts in LLM interactions.
SciDSK transforms how AI agents interact with scientific datasets, enabling more effective discovery and interpretation through a structured, reusable skill representation.
LLMs only prove their worth in task scheduling when faced with unpredictable surges of safety-critical demands, revealing their limitations in stable environments.
Shifting the evaluation of agentic search to a vast, uncurated corpus reveals a dramatic decline in retrieval effectiveness, challenging current models' capabilities.
ReCache achieves a staggering 92.43% reduction in KV-tensor memory while maintaining nearly identical performance in tool-augmented language models.
Static environments can be transformed on-the-fly to better suit agent learning, resulting in up to a 9.0-point performance boost with fewer execution steps.
PolicyGuide boosts compliance rates in LLM agents by transforming policy checks into a proactive, workflow-guided system that adapts to user interactions.
Achieving a single successful task completion in stateful workflows doesn't guarantee reliability, with many agents failing to maintain consistent performance across multiple attempts.
RISE not only boosts planning efficiency but also redefines how imagination budgets can be adaptively allocated in complex environments.
Usability testing reveals that DesCartes Builder empowers domain experts to create real-time digital twins with unprecedented ease and reliability.
Agents can now evolve their behavior without altering their core model, achieving over 10% performance gains while retaining learned skills.
DART-SD reveals that leveraging diamond-topology awareness can drastically improve policy diversity and performance in multi-turn tool-calling agents.
Achieving a 95.3% functional accuracy and a 93.0% hallucination-free rate, this multi-agent platform redefines the standards for conversational business intelligence.
Eureka's innovative orchestration of Meta-Agents not only completes complex scientific tasks flawlessly but also uncovers new insights in quantum theory and the Riemann Hypothesis.
ORBITER significantly boosts decision-making reliability in last-mile delivery, outperforming existing models by up to 9.2% through enhanced reasoning about spatiotemporal cues.
RTPO eliminates critical instability in multi-turn RL training, achieving over 21% performance improvement compared to traditional methods.
AFANet achieves high accuracy in agent failure attribution with a fraction of the computational resources required by traditional LLM-based methods.
A single agentic framework can seamlessly tackle multiple intelligent document processing tasks, outperforming specialized models in the process.
By decoupling detection from explanation, PATE-Forensics achieves unprecedented accuracy in deepfake forensics, setting a new benchmark for explainability in AI.
The rise of agentic systems in computational chemistry suggests a looming paradigm shift towards fully autonomous AI scientists, challenging the relevance of human expertise.
LLM agents are stuck in local adjustment loops, unable to adapt their training strategies despite having the resources to do so.
Medical QA systems can achieve unprecedented accuracy by leveraging a multi-agent framework that integrates adaptive memory and structured reasoning.
MemFuse achieves superior performance in multi-source memory tasks, revealing that effective memory fusion can significantly enhance an agent's reasoning capabilities.
AI can enrich art interpretation by enabling diverse narratives, challenging the notion that generative models lead to uniformity in understanding.
Automation metrics can mislead model selection, as LLMs that excel in direct output generation often fail to enhance the performance of weaker agents.
Over half of the bounty listings on RentAHuman impose high proof burdens, often requiring sensitive personal information or physical actions from workers.
Swapping to CTIFoundry allows smaller models to outperform flagship models, achieving higher accuracy with fewer tool calls in cyber threat intelligence investigations.
PILOT's proactive framework boosts recommendation system performance by over 40% in search efficiency and achieves up to 1.60% gains in core metrics, all without human oversight.
Runtime for complex project-scheduling simulations can be slashed from over 1,200 seconds to under 200 seconds using agentic AI optimizations, saving researchers significant computational resources.
Tool failures can be effectively managed with Outcome Monitors, boosting task completion rates by over 150% in critical scenarios.
Automating Assurance Case generation could drastically reduce the resource burden on SMEs striving for compliance with the EU Cyber Resilience Act.
FACET achieves unprecedented task synthesis quality by preserving source intent and ensuring executable state consistency, leading to more reliable terminal agents.
SkillGate achieves a 30% boost in trial success for long-horizon agents by fundamentally rethinking how skill selection is rewarded during execution.
Managerial behavior, not model size or vendor, dictates success in long-term decision-making tasks, as evidenced by the surprising performance of claude-fable-5 in FM-Bench.
Proactively synthesizing project-specific issues can drastically enhance LLM agents' ability to resolve software problems effectively.
Current AI agents can match human performance in some tasks, but they largely recycle existing human-designed algorithms rather than creating novel solutions.
Tying every view of a digital artifact to its version can boost knowledge-work agent performance by over 12 points on critical tasks.
EvoTS-Agent not only outperforms existing models in financial change-point detection but does so with a flawless execution success rate across multiple datasets.
Despite advances in AI, even top models struggle with real-world tasks, achieving only 30% success on a benchmark grounded in market-validated workflows.
Evolution strategies can optimize large language models for long-horizon tasks with minimal GPU resources, outperforming traditional reinforcement learning approaches.
Real-time monitoring and steering of long-running data analyses can now be achieved through a unified storyline interface, transforming how analysts interact with autonomous workflows.
RGE reveals that long-horizon agents can drift significantly from their intended tasks, even while appearing compliant at each step, highlighting the need for deeper oversight mechanisms.
Robust memory interventions in LLMs can achieve up to 3.7 percentage points of improvement when guided by a dual-loop diagnostic protocol that localizes errors effectively.
Self-evolving financial agents may enhance utility but simultaneously increase security risks, with unauthorized state changes rising alarmingly high.
Next-turn user reactions can boost multi-turn agent performance by over 10 percentage points, revealing the critical role of local feedback in reinforcement learning.
Harness provisioning can be optimized to improve LLM agent accuracy by up to 10% while using 48% fewer tokens in liquid cooling tasks.
Wuying-Browser-Agent achieves a groundbreaking 80.6% on WebVoyager, revealing that real-world browser agents can excel beyond short, clean demonstrations.
Human expertise remains essential in the agentic coding process, even as productivity in simulation library development accelerates significantly.
Achieving over 99% precision in structural component detection from framing-plan PDFs could revolutionize the drafting process in structural engineering.
The shift from recoverable failures to irreversible losses in AI agent security on blockchains could redefine our understanding of attack vectors in Web3 environments.
Harnessed agentic RL can boost agent performance significantly, as shown by a 14.6-point improvement in coding tasks with minimal training data.
HarnessRisk reveals that up to 80.9% of adversarial attacks can succeed in agent harnesses, even when risk detection is high.
Bridging semantic logic with geometric constraints, aDSL enables agents to create complex 3D structures more reliably than traditional LLM approaches.
Skill Optimizers trained through execution feedback can outperform traditional models by over 9 points, revealing a critical gap in agent learning methodologies.
Task-conditioned authority selection reduces excess-authority errors in tool-using agents from 4.56% to 0.79%, showcasing a powerful new layer of control.
PACE achieves a flawless safety record in DeFi transactions, eliminating unsafe executions while leveraging LLMs for complex financial operations.
TRUSS reveals that incorporating execution evidence can dramatically improve both the safety and effectiveness of automated agent skills, achieving unprecedented vulnerability detection rates.
DAS achieves a remarkable average score of 4.34 in academic survey automation, surpassing its closest competitor by a significant margin.
An LLM coding agent achieved a 100% success rate in robot manipulation with 46% fewer steps than traditional methods, revolutionizing how we approach task learning without human demonstrations.
GAPL achieves a remarkable reduction in collision rates and displacement errors, showcasing the potential of LLMs in trajectory planning for autonomous driving.
Shifts in observation and action spaces can drastically alter task success rates by over 30%, revealing critical vulnerabilities in current AI agents' performance on UI tasks.
Remediation order in multi-gate AI systems can fundamentally alter decision outcomes, revealing critical vulnerabilities in evidence trustworthiness.
Task difficulty in software issue resolution can be predicted with remarkable accuracy, revealing critical structural features that influence agent performance.
LEGO-RL boosts coding agent performance by up to 8.4% while ensuring robust training signals and execution reliability.
Palmyra x6 outperforms previous models in enterprise agentic tasks while maintaining a strong safety profile and low bias.
Cold-load latency slashed by up to 90% while achieving a surprising boost in relevance accuracy through innovative token management strategies.
Accident retrieval using LLMs can significantly enhance highway construction safety planning, achieving over 75% accuracy in incident classification.
LLMs consistently underperform in shared-budget reasoning tasks, revealing a critical gap between their single-task capabilities and multi-task resource allocation.
GenRouter cuts execution costs by over 95% and latency by 65%, revolutionizing how we approach agentic image generation workflows.
Coordination among AI agents can be optimized by leveraging shared files, reducing communication overhead by 42% in message-heavy tasks, but simply adding a coordinator offers no advantage. WHY_IT MATTERS: These insights could transform how we design multi-agent systems, emphasizing the importance of task structure and communication strategies in enhancing collaborative efficiency.
BATON transforms long-horizon robot manipulation by making subtask exploration efficient and transition-aware, leading to significant performance gains.
Achieving a 91.7% success rate in obstacle avoidance, Orbit-Planner redefines how satellite agents navigate dynamic environments without relying on fixed maps.
State information in LLM-driven agents can be exploited, turning their task execution capabilities into a potential attack surface.
PDDLCoder achieves nearly 90% applicability in long-horizon planning, setting a new standard for LLM-assisted symbolic planning.
CEFITO's innovative action-conditioned representation space allows for precise inference-time planning by eliminating irrelevant actions, setting a new standard in procedure planning accuracy.
Incremental updates to semantic substrates can be 33.7 times cheaper than full re-computation, challenging the assumption that corpus size dictates maintenance costs.
Collective dynamics of AI agents reveal surprising patterns: while communication boosts accuracy on objective tasks, it can lead to political bias in group opinions.
Public skill version histories can dramatically enhance task-specific skill evolution, leading to significant performance gains in AI agents.
Trust-preserving agentic AI can achieve an impressive 86.9% task completion rate while intervening in nearly all policy violations.