Search papers, labs, and topics across Lattice.
100 papers published across 9 labs.
VirtualSet not only boosts accuracy in LLM-generated queries but also eliminates the risk of executing hallucinated actions by enforcing type safety before execution.
Safety intentions in AI agents can degrade over time, leading to unsafe actions and operational failures, but a new architectural layer can intercept these issues effectively.
IDEAgent achieves a staggering 3.89x improvement in generating diverse and high-quality research ideas compared to existing methods.
Agents struggle to act effectively in 3D scenes, with none of the eleven evaluated VLMs achieving consistent performance across diverse tasks.
Shifting the focus from visual perception to direct program state access can boost agent success rates while slashing costs by nearly 9x.
IDEAgent achieves a staggering 3.89x improvement in generating diverse and high-quality research ideas compared to existing methods.
Agents struggle to act effectively in 3D scenes, with none of the eleven evaluated VLMs achieving consistent performance across diverse tasks.
Shifting the focus from visual perception to direct program state access can boost agent success rates while slashing costs by nearly 9x.
AREX achieves superior performance in deep research tasks by recursively refining answers through a novel self-improvement mechanism that outpaces traditional search methods.
Real-time health insights from wearables can now be processed locally, ensuring privacy without sacrificing personalization.
HiQC achieves unprecedented performance on long-horizon tasks by combining high-level planning with low-level action chunking, effectively reducing value estimation errors.
Silent degradation of AI agents can be mitigated with a continuous-assurance framework that keeps citizen-created tools operationally ready.
GS-Agent transforms natural language into intricate 4D worlds, showcasing a new era of automated creative content generation that rivals traditional manual methods.
Contrastive explanations reveal why policy-aware agents make specific decisions, enhancing user trust and comprehension in autonomous systems.
Open-weight LLMs can achieve nearly 88% task completion in complex data preparation tasks without ever sending sensitive data to the cloud.
LLMs consistently hallucinate on complex problems, but Euclid-MCP delivers precise, low-latency answers that could redefine logical reasoning in AI applications.
Weak policies can achieve up to 18.6% better performance with PATS, a training framework that dynamically adapts guidance based on evolving policy needs.
Traditional regulatory frameworks are failing to keep pace with the shifting landscape of AI autonomy, necessitating a radical rethink of governance strategies.
Cryptographically verifiable authorization could be the key to ensuring that autonomous AI agents operate securely and in compliance with established policies.
Token-level attribution transforms memory learning, enabling agents to identify and leverage crucial information for better performance in complex tasks.
SafeStep not only predicts potential travel failures for elderly users but also tailors interventions that significantly boost their confidence and safety on journeys.
WML not only identifies failure mechanisms in agent workflows but also optimizes knowledge reuse, achieving state-of-the-art accuracy and efficiency in structured tasks.
MemTools transforms agent memory research by enabling seamless integration and evaluation of diverse memory types across architectures.
Agents forget crucial operational facts unless they are anchored to specific situational cues, revealing a fundamental flaw in current memory architectures.
Causal-AgentIR enables image restoration agents to dynamically evolve their knowledge, leading to more effective handling of diverse degradation scenarios.
OpenForgeRL allows researchers to train AI agents in real-world environments with unprecedented ease and efficiency, revealing that some harnesses are significantly harder to learn than others.
Coding agents can now be evaluated on their ability to navigate fuzzy requirements and interactive workflows, reflecting real-world software development challenges.
Agentic Designer achieves unprecedented structural adherence in interior layouts by employing a multi-agent system that iteratively verifies geometric constraints before placement.
Retaining over 64% of in-context learning gains without further environment interactions could revolutionize how agents learn from their experiences.
Forgetting is not just a flaw in AI agents; it's a systemic issue that can be solved by rethinking context management as a lifecycle process.
Active Inference can be framed as a convex MDP, revealing a surprising connection to performative reinforcement learning with robust policy improvement guarantees.
PRO-LONG boosts LLM performance on long-horizon tasks by 18% while using up to 5.8 times fewer tokens than traditional methods.
Achieving a 73.5% reduction in design rule violations, EvoDRC revolutionizes the automation of DRC closure in advanced-node physical design.
Experience-guided policy adaptation can cut adaptation time by up to 38% in dynamic manufacturing environments.
Eight UX principles for human-AI interactions could redefine how businesses integrate AI agents into their workflows.
A single carrier can dominate freight requests, but revealing daily capacity can significantly reduce market concentration and enhance shipper benefits.
Surface accuracy metrics can mislead researchers about the true reliability of multimodal search systems, with silent failures lurking beneath the surface.
Achieving 86% accuracy in automated symbol generation could revolutionize PCB design by drastically reducing manual errors and time investment.
PRTA outperforms traditional and LLM-based recommendation systems by effectively leveraging multiple models through a central LLM planner, enhancing personalization without the pitfalls of hallucination.
Solar Open 2 outperforms its predecessors and competitors with a groundbreaking 1M-token context window, redefining the capabilities of large language models in agentic tasks.
Strong performance in static evaluations masks a critical flaw: LLMs struggle to adapt to evolving user intent during multi-turn interactions.
Adaptive temporal discounting can dramatically enhance reinforcement learning's efficiency and flexibility, outperforming traditional fixed discount methods.
GenDB achieves significantly better performance than existing query engines by automating code generation tailored to specific workloads, fundamentally changing how we approach query processing.
Vanguard not only blocks unsafe actions before they occur but also boosts benign task completion, achieving a remarkable dual improvement in agent safety.
Over two-thirds of malicious issue requests can exploit vulnerabilities in leading AI coding agents, highlighting a critical gap in current safety measures.
AI agents are vulnerable to indirect prompt injection attacks, and KYA reveals how targeted reconnaissance can significantly enhance pentesting effectiveness.
Autonomous AI agents in offensive security blur the lines of accountability, making it easier for malicious actors to exploit their capabilities without clear attribution of responsibility.
PerfAgent doubles the rate of expert-level code optimizations by leveraging profiler-guided feedback, revealing hidden performance bottlenecks that traditional methods miss.
Clarification need prediction can significantly enhance user experience in conversational search systems, leading to more accurate and relevant interactions.
Human-written policies can boost agent performance significantly, but learning from experience remains a major hurdle for effective text policy generation.
Molt reduces the cognitive load on researchers by offering a clean and compact codebase that maintains high performance, enabling faster iterations in agentic reinforcement learning.
Even the most advanced autonomous agents struggle with fundamental document manipulation tasks, revealing critical vulnerabilities in their operational capabilities.
NOOA revolutionizes agent development by treating agents as first-class Python objects, enabling seamless integration of AI capabilities with standard programming practices.
An investigation agent can misinterpret classifier outputs, leading to more errors despite providing seemingly rational explanations.
LangGraph reveals that the right orchestration framework can significantly enhance the reliability and efficiency of stateful AI systems in complex business workflows.
SCM achieves an impressive 84.87% accuracy on conversational memory tasks, showcasing a novel approach to optimizing agent memory retrieval and synthesis.
PhoenixRepair redefines how software agents explore repair strategies, achieving a 76% resolution rate by leveraging multi-location sampling and iterative refinement.
AgentTrails uncovers hidden dependencies in agent trajectories, enabling unprecedented insights into LLM-powered task execution.
Asynchronous attacks across LLM agents can be effectively linked with a new scoring protocol, revealing a high degree of campaign similarity that traditional methods miss.
LLM agents equipped with unique personas can collaboratively navigate complex travel planning discussions, revealing surprising dynamics in group decision-making.
A client-side keepalive can slash follow-up request costs by up to 12.5x, reshaping the economics of LLM usage in agentic workloads.
By intelligently deciding when to zoom in on audio-video content, OmniReasoner boosts reasoning accuracy while reducing computational costs, transforming how LLMs handle complex multimodal inputs.
Recovery routing can outperform escalation strategies by leveraging execution feedback, achieving a higher solve rate at only 35% of the typical recovery cost.
Automating real-to-sim conversion with vision-language agents could revolutionize how we simulate robotic interactions, making it faster and cheaper than ever before.
DeepDebug achieves a 32% improvement in task recovery accuracy, showcasing a powerful new approach to debugging LLM agent failures.
Twin Agent achieves a superior security-utility balance, allowing LLMs to fend off prompt injection attacks without sacrificing performance.
Real-world deployment of LLM-based agents reveals critical safety and reliability challenges that traditional benchmarks overlook.
Domain-specific fine-tuning boosts RF reasoning in LLMs, especially for smaller models, while semantic retrieval proves superior for context alignment.
All 15 evaluated x402 facilitators have critical security flaws that could lead to severe financial losses for merchants and users alike.
Agents fabricate responses 56.6% of the time when faced with silent failures, revealing a critical blind spot in current auditing practices.
A groundbreaking framework for agentic commerce achieves 14.4x faster event verification while ensuring tamper-evident auditability across diverse domains.
Achieving a 100% reduction in data leakage without disrupting application behavior could redefine security protocols for agentic systems.
Traditional web bot defenses crumble under the pressure of LLM agents and commercial solvers, revealing a critical vulnerability in current security architectures.
Skillware redefines agent skills as independent software artifacts, unlocking new possibilities for their management and evolution in AI systems.
VirtualSet not only boosts accuracy in LLM-generated queries but also eliminates the risk of executing hallucinated actions by enforcing type safety before execution.
ARBITER outperforms traditional autoscalers by reliably selecting the right remediation actions and targets, ensuring SLOs are maintained even in complex failure scenarios.
RIME reveals that existing multimodal LLMs struggle with music post-production, highlighting a critical gap in their capabilities that could redefine how music is produced collaboratively.
Knowledge-centric self-improvement can significantly boost AI performance while slashing maintenance costs and enhancing transferability across tasks.
Achieving real-time, long-horizon interactive world rollouts on a single desktop GPU could revolutionize how we develop and deploy AI-driven simulations and games.
Agentic reasoning tools struggle with complex financial documents, revealing substantial gaps in their capabilities that could impact decision-making in finance.
Models can achieve higher accuracy in nuanced football predictions, even when overall match result accuracy shows only slight improvements over traditional baselines.
Pruning tool outputs directly within the agent leads to a remarkable 39% reduction in token usage without sacrificing performance.
Agents trained with SEE-generated trajectories not only navigate complex multi-step procedures but also achieve unprecedented task success rates in real-world applications.
Performance gaps of up to 86.27% among AI agents in EDA workflows reveal that architecture matters more than just domain-specific skills.
Despite high agreement in predictions, LLMs underperform against betting markets, revealing stark differences in decision-making quality and self-awareness.
LLMs can be transformed from unreliable output generators to trusted decision-makers in smart grids by grounding their results in verified numerical tools.
Smaller LLMs falter in ABM simulations, while larger models reveal new dynamics in agent-based decision-making.
Sparse evidence can lead to more effective misinformation detection, with SIEVE outperforming traditional methods by focusing on critical clues rather than exhaustive analysis.
Achieving 86.7% accuracy on direct commands, AdaHome outperforms LLM-based systems while slashing latency by up to 3x, all within a local deployment framework.
EAR boosts memory retrieval performance by nearly 18% in long-term interactions, setting a new standard for LLM adaptability.
Transforming SHCUA outputs into contract-bound UAV commands could revolutionize real-time UAV control by ensuring safety and responsiveness in dynamic environments.
Current LLMs fail to effectively anticipate user needs, with even the latest models only succeeding in 26.7% of proactive scenarios.
Revealing memory utilization patterns through attention can transform how agents refine their memory, leading to substantial gains in performance and efficiency.
Engineers risk losing cognitive control and compromising test validity by overtrusting AI test agents, which could undermine software quality assurance.
Adaptive handover systems can drastically reduce grasp delays and improve user trust in human-robot interactions, especially with asymmetric tools.
AlayaWorld achieves unprecedented long-horizon video generation with just four sampling steps, revolutionizing how we create interactive virtual environments.
A layered defense can thwart most self-state attacks on AI agents, but a hidden vulnerability persists that challenges current OS security paradigms.
Safety intentions in AI agents can degrade over time, leading to unsafe actions and operational failures, but a new architectural layer can intercept these issues effectively.
SOGR1 achieves state-of-the-art performance in knowledge graph question answering while being significantly more efficient than larger models.
Stopping repairs too soon can lead to a 60% drop in true validity, but VRR-Stop ensures LLM agents know exactly when to commit or repair.
Adversarial instructions can exploit LLM agents in HPC, leading to unauthorized actions even under authenticated user credentials.
Ghost references in LLM-generated code can be eliminated by constraining generation to valid runtime environments, ensuring both grammatical and semantic correctness.
RECEIPT uncovers 24 previously unknown XSS vulnerabilities while ensuring no false positives, setting a new standard for trust in automated vulnerability discovery.
Nearly half of agent-generated pull requests lack adequate test coverage, exposing critical vulnerabilities in autonomous software development.
A robust architecture for agentic commerce that ensures transaction integrity and security, even in the face of state changes and unauthorized access.