Search papers, labs, and topics across Lattice.
100 papers published across 7 labs.
Training API-calling agents just got easier—synthetic data generation using LLMs eliminates the need for complex environments while boosting performance.
Bridging the NL2Pipeline gap, DataFlow-Harness enables LLMs to create reliable, editable data workflows at a fraction of the cost and time of traditional methods.
DSWorld accelerates RL training by 14x while outperforming top LLMs in predicting data science operation outcomes.
Muon can boost agentic RL performance by up to 88% compared to AdamW, challenging assumptions about optimizer efficacy in sparse-reward settings.
Heterogeneous AI agent networks can evolve to outperform homogeneous counterparts, revealing a new scaling law for collaboration that defies traditional strength metrics.
Training API-calling agents just got easier—synthetic data generation using LLMs eliminates the need for complex environments while boosting performance.
Bridging the NL2Pipeline gap, DataFlow-Harness enables LLMs to create reliable, editable data workflows at a fraction of the cost and time of traditional methods.
DSWorld accelerates RL training by 14x while outperforming top LLMs in predicting data science operation outcomes.
Muon can boost agentic RL performance by up to 88% compared to AdamW, challenging assumptions about optimizer efficacy in sparse-reward settings.
Heterogeneous AI agent networks can evolve to outperform homogeneous counterparts, revealing a new scaling law for collaboration that defies traditional strength metrics.
LQCDMaster automates complex LQCD workflows, achieving expert-level precision while slashing computation time from hours to minutes.
Randomizing initial choices in LLM networks can triple collective payoffs, revealing a critical strategy for enhancing collaborative problem-solving.
One-third of documents that seem useless to static readers are actually critical for guiding agentic search, revealing a fundamental disconnect in retrieval utility assessments.
Evidence-grounded detection in FlowGuard reveals that traditional semantic analysis can miss critical execution-related risks, achieving up to 2.23x faster scanning.
Specialist agents can outperform generalist LLMs by up to 20 percentage points in execution accuracy while slashing costs and errors, making them essential for reliable software development.
A unified communication standard could revolutionize how humans interact with robots, enabling seamless collaboration across diverse interfaces.
SEED transforms past experiences into actionable skills, allowing reinforcement learning policies to evolve and improve in real-time.
SAGA reduces the incidence of empty-result queries by systematically grounding SPARQL generation in schema constraints, outperforming existing methods in accuracy across multiple benchmarks.
AutoSynthesis achieves expert-level meta-analysis with automated precision, making evidence synthesis scalable and accessible.
Many autonomous GUI-agent failures can be repaired when users can see and edit plans in real-time, transforming how we interact with automation systems.
BrainPilot achieves state-of-the-art performance in brain science research automation while maintaining rigorous auditability and cost-effectiveness.
Even state-of-the-art LLMs show alarming performance drops when adapting to evolving toolsets, revealing a critical gap in current evaluation methods.
Even state-of-the-art models only achieve pass rates below 60% on a new benchmark that spans 1,431 diverse tasks, exposing critical weaknesses in general AI capabilities.
Artifact-centered evaluations reveal that LLM agents can achieve an 88.6% success rate in structural engineering tasks, but still struggle with invalid inputs and model consistency.
Large language models can automate complex model transformations in automotive engineering, cutting manual effort significantly while ensuring structural validity.
Untrained structural monitors can reduce security attack detection failures from 11.6% to just 3.5%, paving the way for safer AI agent deployments.
SmartRAG achieves competitive multi-hop reasoning on mobile devices using a lightweight architecture that outperforms models up to 18 times its size.
Current LLMs can tackle basic proofs, but MathCoPilot reveals their limitations when faced with advanced theorems requiring true mathematical understanding.
Coding agents can significantly improve payment integration performance with targeted skills, achieving up to 91.37% success in complex scenarios.
Larger models may be their own worst enemies in survival scenarios, consuming energy faster than they can replenish it, while cooperation can lead to unexpected altruism.
Mobile agents can now navigate complex GUIs with unprecedented efficiency, thanks to a novel data-environment co-scaling framework.
Budget-aware training can boost low-turn performance by over 10% while maintaining scalability across diverse tasks.
RetroAgent revolutionizes retrosynthesis planning by combining LLMs with structured memory, leading to significantly improved decision-making in complex chemical searches.
LongStraw enables RL post-training with over 2 million tokens on a fixed GPU budget, pushing the boundaries of context length in AI applications.
Existing payloads in memory can compromise future agent behavior, revealing a critical vulnerability in memory-based AI systems.
SearchOS turns fragile search progress into a robust, shared state, enabling agents to avoid repetitive failures and significantly improve search efficiency.
Offensive security agents can significantly outperform proprietary systems when evaluated through a cost-aware lens, while defensive agents reveal a stark reliance on tool discipline over sheer computational power.
TRACE transforms how long-horizon agents are trained, leading to a remarkable performance increase on complex tasks without the need for supervised fine-tuning.
A novel structured reinforcement learning approach can enhance persuasive signaling strategies in interactive driving, yielding a 30% boost in cost efficiency for route optimization.
Agents using the Experience Memory Graph can recover from failures in a single execution, eliminating the need for costly trial-and-error loops.
Bot adoption in open-source projects not only enhances collaboration but also reduces conflicts, fundamentally reshaping team dynamics.
Traditional loyalty models miss the mark in an era where AI agents autonomously influence purchasing decisions, but the DVM-HALL model captures the complexities of this new landscape.
HealthClaw boosts answer accuracy for personal health management by over 45% while enhancing privacy protection in AI interactions.
CAVA transforms the chaotic landscape of agentic AI actions into a standardized framework, enabling reliable governance and verification of AI behavior across heterogeneous environments.
MEDA reveals that integrating domain knowledge and mechanistic constraints is crucial for accurately discovering biological models, outperforming traditional numerical fitting methods.
Safety Sentry redefines safety in LLM interactions by enabling context-sensitive decision-making, significantly reducing user interruptions while enhancing safety outcomes.
Cross-device agents struggle significantly, with top performers only managing a 12.5% success rate in executing complex, multi-device tasks.
Real-world deployment of the LEA reveals that while it performs well across courses, its ability to maintain content fidelity diminishes with curriculum distance.
Multimodal agents can achieve superior performance by evolving skills alongside their policies, rather than treating them as static resources or mere rewards.
A self-evolving framework achieves up to 15.5 percentage points in performance gains by intelligently refining the agent's harness without altering the underlying model.
MyAG reveals that a graph-based approach can revolutionize the design of LLM agent systems, enabling unprecedented flexibility and efficiency in execution.
Vulnerabilities in reusable agent skills can emerge at every stage of their lifecycle, not just during execution, revealing a critical oversight in current security practices.
Achieving a 78.3% success rate in real-world mobile manipulation, this framework bridges the reality gap with zero-shot transferability to unseen tasks.
UrbanAgent outperforms traditional methods by leveraging multi-agent reasoning to tackle cross-modal inconsistencies in urban profiling tasks.
Agents default to a narrow routine, struggling to adapt to hidden shifts in tool reliability, revealing critical insights into their decision-making processes.
Optimization gains only compound when regression control is integrated into the agent's learning process, revealing a critical factor for effective continual learning.
AI in HRM is reshaping efficiency goals, but its true potential for strategic talent development remains underexplored.
Despite the promise of agentic coding tools, most GitHub projects see minimal adoption, with only a few exceeding the industry standard for PRs per participant.
A surprising 40% of Skeptics transformed into Cautiously Positive users, while 68% of Champions lost enthusiasm, underscoring the dynamic nature of employee acceptance in AI tool adoption.
AI agents can autonomously verify security software with astonishing efficiency, but their trustworthiness is limited by the strength of their feedback mechanisms.
ProfMalPlus not only achieves a 98.1% F1-score but also uncovers hundreds of previously undetected malicious packages, demonstrating a significant leap in supply-chain security for open-source software.
Git's version control could revolutionize memory management in coding agents, achieving 60x better retrieval performance than traditional methods.
User-level permissions in AI agents are not just a feature; they are essential for mitigating risks like unauthorized transactions and data leaks.
NexForge transforms the landscape of LLM training by synthesizing 43.2K tasks, propelling model performance to unprecedented levels without the need for domain-specific infrastructure.
Language corrections in PhysClaw-0 not only enhance robot autonomy but also boost success rates by over 35% while slashing human oversight time.
Robots can now recover from execution failures with up to 39.2% better success rates by leveraging high-level decision-making in uncertain environments.
Multi-Head Latent Control reduces large model usage by up to 90.7% while enhancing decision-making accuracy in LLM agents.
Adaptive memory management can boost LLM task success by over 15 points while slashing token usage by up to 20%.
CoW Scoring reveals precise failure points in agent operations, enabling targeted improvements that can dramatically enhance performance in real-world applications.
ATLAS transforms LLMs from unreliable code generators to trusted partners in analog design, successfully producing SAR ADCs that meet rigorous simulation standards.
Structured feedback can boost LLM agent success rates by up to 44 percentage points, revealing the critical role of admissible alternatives in the repair process.
Safety-aligned LLMs can override deployment instructions up to 43.4% of the time, raising serious liability concerns for organizations using these tools.
Approval gates in LLM-agent frameworks fail to prevent side effects during pauses, with real-world implications seen in 18% of runs.
BPO achieves up to 6.1% higher success rates in sandbox-native RL tasks while cutting down on the number of required policy updates by 38%.
A unified evaluation framework that simplifies the assessment of LLM-based agents could drastically enhance reproducibility and accelerate research breakthroughs.
Cura 1T outperforms existing models in healthcare tasks by leveraging a unique self-evolution training loop that adapts to specific capabilities without sacrificing overall performance.
Personal singularity could redefine the partnership between humans and AI by enabling users to expand their capabilities in a structured, accountable manner.
Everyday AI interactions may be silently reinforcing negative emotional patterns, but a simple observational practice can reverse this trend and promote healthier neural pathways.
Autonomous AI agents can now seamlessly orchestrate IoT environments, transforming how we manage smart buildings and beyond.
The shift from static software components to adaptive, goal-directed agents demands a new engineering framework to ensure reliability and trust in AI systems.
A self-evolving critic can reduce confidence estimation errors in LLM agents by up to 54% without requiring any training or ground truth labels.
Despite advancements in LLMs, even the best agents struggle with prospective memory, achieving only 65.1% accuracy in executing delayed intentions.
Tool-based reasoning in IQA-T1 leads to significant improvements in image quality assessments, outperforming traditional methods while enhancing interpretability.
Partial evaluations can mislead if not carefully calibrated, with required task fractions varying dramatically across benchmarks—15% for AppWorld but 95% for SWE-bench Lite.
Current remote sensing models falter in hierarchical reasoning, but HieraPlan sets a new standard for cognitive analysis in geospatial contexts.
Structured epistemic memory can boost multi-hop reasoning performance by up to 11 points, revealing that organization trumps sheer model size in tackling complex tasks.
Self-evolving knowledge graphs can dramatically enhance multimodal reasoning by continuously adapting to new information and correcting errors in real-time.
Hy-Embodied-VLM-1.0 outperforms its predecessor by 8.4% while activating only a fraction of the parameters, redefining efficiency in embodied agents.
ReflectVLN's closed-loop mechanism allows navigation agents to dynamically adapt and recover from errors, leading to a notable increase in success rates and efficiency.
The new AI-native insurance framework reveals how to effectively underwrite and price policies for autonomous AI systems, balancing risk and governance.
Task-aware execution can cut costs by 85% while maintaining a 100% success rate in complex workflows.
Mid-training with function-aware fill-in-the-middle boosts coding agent performance while preventing capability erosion in non-agentic tasks.
Induced anger can lock LLMs into poor decision patterns by reducing their sensitivity to penalties, unlike human decision-making.
Memory-augmented speculation boosts LLM prediction accuracy by up to 39% without incurring any additional execution time.
Fin-Analyst's innovative LLM hybrid approach outperformed traditional strategies, achieving a +13.51% return on Tesla while exposing critical flaws in rule-based trading for Bitcoin.
Achieving 100% accuracy in empirical reasoning while maintaining source fidelity, EG-VAR redefines how we can eliminate hallucinations in LLM outputs.
Over 21% of reported x402 settlements are fictitious, challenging the narrative of a thriving AI-driven economy.
SoftBoard revolutionizes MVP development by automating prototype creation and usability evaluation, making UX expertise optional for teams.
Multi-perspective reasoning in CT-Repair leads to a 99-bug improvement over the best individual analysis strategy, showcasing the power of structured evidence in automated program repair.
Research artifacts can now be tracked as interconnected exploration trees, revealing the complexities of autonomous scientific discovery like never before.
KnowAct-GUIClaw achieves a groundbreaking 64.1% success rate in long-horizon task execution, outperforming all existing agent frameworks and closed-source models.
Behavior localization is revolutionized, enabling developers to seamlessly connect high-level modification requests to specific code locations in complex AI harnesses.
PalmClaw achieves a staggering 94.9% reduction in task completion time by transforming how mobile agents interact with device capabilities.
Self-improving agents can evolve autonomously with minimal human input, reshaping our approach to AI adaptability and deployment.
OAT achieves up to 5000 times faster failure attribution for LLM agents without the need for costly step-level supervision.