Search papers, labs, and topics across Lattice.
100 papers published across 9 labs.
Retrieving skills from a crowded library just got smarter—Capability Pages boost retrieval accuracy by distinguishing between confusable alternatives.
Multi-turn interactions can be effectively optimized in LLMs using tailored RL strategies, overcoming significant challenges in credit assignment and reward design.
A self-developing agent that autonomously improves its own coding capabilities while navigating operational safety challenges has set new performance benchmarks.
Complex search behaviors typically requiring explicit controllers can emerge naturally from an agent's reasoning process, leading to substantial performance gains in optimization tasks.
No single model-harness combination consistently outperforms others, highlighting the critical need for tailored evaluations in agent deployment.
Multi-turn interactions can be effectively optimized in LLMs using tailored RL strategies, overcoming significant challenges in credit assignment and reward design.
A self-developing agent that autonomously improves its own coding capabilities while navigating operational safety challenges has set new performance benchmarks.
Complex search behaviors typically requiring explicit controllers can emerge naturally from an agent's reasoning process, leading to substantial performance gains in optimization tasks.
No single model-harness combination consistently outperforms others, highlighting the critical need for tailored evaluations in agent deployment.
Small LLMs can achieve up to 27.2% accuracy gains by leveraging hierarchical memory from larger teacher agents, reshaping how we think about agent training.
SkillZip achieves a remarkable 3.46x compression ratio while preserving 99.2% of dependencies and 98.7% of verifier reachability, revolutionizing how agent skill libraries can be managed.
Visual tool-use in multimodal LLMs may create an illusion of effectiveness, with many models achieving gains that are not causally justified.
CalibForge reveals that adversarial calibration can dramatically enhance the effectiveness of training data for terminal agents, leading to unprecedented performance improvements on standard benchmarks.
Hardware confinement of cryptographic keys eliminates key exfiltration risks, achieving zero successful attacks in a rigorous evaluation against AI agent vulnerabilities.
Agentic LLM-guided feature selection boosts mortality prediction accuracy in cardiac arrest cases, achieving state-of-the-art results with a fraction of the parameters.
GSE achieves up to 180% improvement in recall for coding agents, revolutionizing how skills are evolved and reused in automated programming tasks.
ECHO achieves a remarkable 94.9% tool-execution pass rate while maintaining patient data privacy, setting a new standard for locally-deployable health assistants.
A hybrid ranker outperforms a knowledge graph in skill retrieval, achieving 73.5% accuracy while the graph fails to extend reach despite its structural advantages.
LLMs may sound convincing, but their investment reasoning often lacks grounding in real-world events, revealing a critical gap in evaluation methods.
CIPO transforms how search agents leverage external evidence, cutting down confirmation bias and enhancing reasoning accuracy in knowledge-intensive tasks.
ChainClaw closes critical gaps in on-chain execution, achieving superior safety and task completion compared to existing frameworks.
Transferred playbooks can enhance agent performance, but their effectiveness hinges on specific conditions and requires careful validation.
Programmatic tool calling outperforms traditional JSON tool calling in 11 out of 14 language models, showcasing a significant leap in efficiency and robustness.
DreamGuard achieves a groundbreaking safety-utility balance by predicting long-term risks, outperforming traditional guardrails that only react to immediate threats.
Economic decision-making in LLM agents reveals a stark divide between task completion and resource efficiency, with agents often overspending or under-escalating.
Gated Hindsight Distillation allows GUI agents to learn from future observations, drastically improving their reasoning capabilities in complex environments.
Bayesian evidence acquisition can dramatically enhance diagnostic accuracy in whole-slide image reasoning by focusing on information gain rather than mere relevance.
Software engineers are increasingly dependent on LLMs, risking overreliance that could undermine traditional practices like peer consultation and documentation.
Recursive belief updates in AgentOPSD reveal pivotal decision points, leading to a 89.1% success rate on complex RL tasks.
Persistent vulnerabilities in AI agents can be effectively managed through a novel framework that links agent behavior to a consistent control posture.
READ outperforms traditional dense retrieval methods by over 40 percentage points in answering complex financial queries, revealing the critical flaws in current top-k approaches.
Misleading historical data can corrupt over 30% of tool-calling decisions, but a new method can restore accuracy by effectively transferring reliable policies from teacher to student models.
iARCS transforms 3D scene generation by ensuring that synthetic environments meet essential functional constraints while maintaining diversity and realism.
Current agentic AI systems fall short, with none demonstrating more than two out of nine critical maturity criteria, exposing a significant gap in their capabilities.
Causal memory can boost execution accuracy in Text-to-SQL tasks, but its effectiveness is context-dependent and not universally superior to other retrieval methods.
Unified Agent outperforms existing multi-agent systems by effectively maintaining a compact state across devices and time, revolutionizing user-agent interactions.
Evolving agents can achieve up to a 19.37-point improvement in task performance by effectively leveraging prior experience in financial workflows.
Achieving the highest fidelity in mobile GUI interaction, AppDeltaWorld redefines how agents predict and interact with app interfaces by focusing on code updates instead of images.
CodeGrep slashes token usage and rounds by over 15% while maintaining high resolve rates, revolutionizing how LLM coding agents handle file retrieval.
Operator-residual feedback slashes the rate of misleading score-only decisions from nearly 40% to under 2%, ensuring that autonomous agents make choices grounded in physical reality.
Routing supervision falters precisely when it's most needed, as weaker agents yield fewer labels, limiting the potential for optimal mode selection.
Autogrammar can automatically learn context-free grammars that boost language model performance, achieving near-perfect precision and tripling execution speed on DSL tasks.
AgentExecutor outperforms existing methods by achieving up to 94% code coverage while slashing execution time by over 80%.
F$^2$Agent achieves over 20% better annualized returns than existing models by dynamically capturing inter-modality dependencies and resisting market noise.
Grounded language comprehension, rather than free-form reasoning, is the key to unlocking superior performance in Vision-Language-Action models.
TrajDebug uncovers the root causes of failures in long-horizon agent trajectories, enabling targeted improvements that could significantly boost agent performance.
Harness optimization reveals that LLMs can significantly enhance their performance, but surprisingly, native harnesses aren't always the best choice for optimization.
Runtime-revealed dependency calls can be strategically managed to boost AI-agent workflow efficiency by up to 10%.
Existing economic simulations are stuck in the past—this blueprint reveals how to build adaptive, self-evolving economic agents that could transform decision-making in AI.
World rehearsal enables LLM agents to internalize environment dynamics, achieving superior performance without costly external interactions.
Agents can now recall user actions with 98.4% accuracy using a deterministic memory compilation method that is 86 times more efficient than traditional LLM summaries.
Sensitive information acquisition by LLM agents is rampant, with most existing privacy measures failing to address this critical vulnerability.
Fine-tuning on planning-aware trajectories can enhance model performance across diverse CLI environments, mitigating the pitfalls of scaffold-specific training.
ASTELD uncovers a critical gap in the autonomous AI landscape: no evaluated systems achieve both local-first deployment and enterprise-grade security.
Training LLM agents without expert supervision can lead to better performance and generalization across diverse environments.
ToolArtist achieves unprecedented synergy in image generation by unifying reasoning and tool use under a single agent policy, outperforming conventional methods.
EvolveNet reveals that decentralized evolution of agent harnesses can lead to substantial performance gains by leveraging localized experience rather than relying on centralized optimization.
Myopic planners can fail spectacularly when faced with the need to acquire capabilities for future experiments, leading to unbounded approximation ratios in goal-directed discovery.
HiGram cuts through irrelevant memory noise, boosting answer quality and efficiency in long-term reasoning tasks by leveraging a hierarchical structure and path-level localization.
OCSD reveals how isolating observation effects can lead to more effective token-level updates in reinforcement learning, outperforming traditional methods.
Argus achieves a 78% success rate on long-horizon reasoning tasks while using 21% fewer tokens in mature workflows, showcasing a revolutionary approach to agentic autonomy.
SuperScout matches the best-performing model's solve rate while slashing costs to one-fifth, revolutionizing how we approach coding agent routing.
SkillSV reveals the intricate value of agent skills, enabling precise optimization and safe pruning of complex skill structures.
Experience-rich memory boosts agent performance in office workflows but can also lead to misleading recall, challenging traditional evaluation methods.
InsightEmb reveals that the geometry of state-insight matching can be effectively transferred across domains, enhancing decision-making in self-improving agents.
EviGraph boosts the reliability of autonomous research agents by ensuring every claim is grounded in a validated evidence chain, leading to a 40% increase in claim support.
EASy achieves superior performance-efficiency trade-offs by intelligently coordinating heterogeneous executors based on their capabilities and costs, reshaping agentic system design.
Susceptibility to tool-selection failures varies dramatically across models, revealing that higher capability does not always equate to greater safety.
Eigenius not only validates scientific conclusions but also uncovers discrepancies in published research, revolutionizing how we ensure data integrity in AI-driven science.
Type-conditioned memory decay can enhance LLM performance by ensuring that only the most relevant and timely information is retrieved, leading to a significant boost in temporal reasoning capabilities.
EmpaAva outperforms traditional chatbots by delivering real-time, emotionally aware interactions through a photorealistic 3D avatar.
LoginTrap exposes a staggering 86% success rate for phishing-style attacks on LLM-based web agents, revealing a gaping hole in authentication security.
Trust in AI agent networks hinges on blockchain, which can redefine how agents interact across diverse platforms and stakeholders.
Trident exposes a staggering 522% drop in defensive performance of DRL systems against adaptive threats, highlighting their critical vulnerabilities.
Reliable automated cooking is now achievable with a framework that transforms user preferences into executable recipes, ensuring transparency and adaptability in real kitchens.
DAC-Pose achieves remarkable fidelity in human generation, maintaining texture and identity consistency even under extreme viewpoint changes.
Mimir achieves an impressive 86.0% success rate on long-horizon tasks, outperforming leading closed-source models by a significant margin.
MemoryCPT achieves a superior cost-performance trade-off for LLM agents by intelligently managing memory without overwhelming downstream models with context.
Retrieving skills from a crowded library just got smarter—Capability Pages boost retrieval accuracy by distinguishing between confusable alternatives.
Long-horizon LLM agents can achieve 96.9% task success by learning to adapt their external execution support through trainable harness policies.
SFC redefines semantic understanding in spoken language tasks, achieving superior accuracy and adaptability in open-domain contexts.
A-SR achieves nearly a doubling of accuracy in symbolic regression tasks, showcasing a transformative approach to LLM-guided formula discovery.
LLMs struggle to reliably apply skills, with performance varying significantly based on the agent harness used, challenging assumptions about their capability.
Task context and requester identity can drastically alter permission grants in mobile GUI agents, with one change reducing approvals from 26 to 0 in a critical task.
A single manipulated search result can dramatically amplify the effectiveness of attacks on LLM-based search agents, revealing critical vulnerabilities in their evidence-gathering processes.
Coding agents are alarmingly susceptible to malicious skill files, with exploitation rates exceeding 95% in some cases.
Agent Plans in open-source repositories reveal critical insights into how AI coding tools can be effectively guided through structured task-oriented artifacts.
Achieving 87.71% negotiation accuracy, this architecture revolutionizes how scientific workloads are managed across diverse computational environments.
Fragmented agentic AI workflows expose significant inefficiencies in conventional server architectures, necessitating a radical rethink of resource allocation strategies.
Memory systems that intelligently filter and prioritize information can drastically improve GUI agent performance, as shown by FocusMem's superior results across multiple benchmarks.
ABSeeker's innovative credit assignment method allows it to achieve performance levels comparable to much larger models, redefining expectations for long-horizon search agents.
PIMiner achieves impressive attack success rates against multiple LLMs with minimal query requirements, revolutionizing prompt injection red teaming.
WorldClaw can generate expansive, editable 3D worlds from text prompts while maintaining both global coherence and intricate local details.
State-Matched Routing and Contextualized Self-Distillation boosts task success rates by over 15% in complex interactive environments by aligning guidance with the agent's actual execution state.
Social influence can lead clinical decision support agents to adopt incorrect answers at alarming rates, revealing a critical flaw in multi-agent oversight.
In-context learning can match explicit skill maintenance in performance, but the real challenge lies in consolidating experience into effective, transferable skills.
TurnSight reveals that leveraging turn-level hindsight can dramatically enhance LLM performance in complex tool interactions, outperforming conventional reinforcement learning techniques.
Agents can exhibit significant performance gains from retained experience, but the pathways to these improvements are often unclear and model-dependent.
Hybrid agents that combine LLM-driven planning with RL optimization achieve superior performance in complex decision-making tasks, outperforming traditional methods.
Foundation model agents can achieve stable cooperation in social dilemmas by inferring behavioral similarities, defying classical game theory's expectations of mutual defection.
A semantic re-keying strategy in speculative decoding can boost accepted draft lengths by up to 29% while achieving 4.4x faster decoding speeds.
CARE-X achieves a remarkable 94.0% accuracy in visual question answering, outperforming existing models by 6 percentage points while also enhancing report quality and spatial localization.
Achieving reliable satellite edge-agent orchestration, SAT-Edge-Agent demonstrates that efficient onboard intelligence can be realized even under stringent operational constraints.
TARL revolutionizes memory management for long-term agents by enabling nuanced updates that significantly enhance state recovery and reduce corruption.
LLM-driven agents can violate verification conditions, but a new canonical wrapper can enforce compliance while preserving behavior.