Search papers, labs, and topics across Lattice.
100 papers published across 10 labs.
Evolving the harness of LLM agents can yield a 17-point performance boost without compromising their ability to generalize across tasks.
LlamaExtract Agentic Plus outperforms commercial VLMs in schema-guided extraction, achieving high accuracy at a fraction of the cost.
LLMs struggle to match human decision-making in e-commerce, achieving only 27.3% of the final net assets in a year-long simulation.
Self-evolving LLM agents show variable reliability in dynamic task environments, challenging the notion that stronger models always yield better adaptation.
Zero-Mem slashes memory operation time costs by over 57% while eliminating unnecessary LLM calls, revolutionizing how LLM agents manage memory.
Evolving the harness of LLM agents can yield a 17-point performance boost without compromising their ability to generalize across tasks.
LlamaExtract Agentic Plus outperforms commercial VLMs in schema-guided extraction, achieving high accuracy at a fraction of the cost.
LLMs struggle to match human decision-making in e-commerce, achieving only 27.3% of the final net assets in a year-long simulation.
Self-evolving LLM agents show variable reliability in dynamic task environments, challenging the notion that stronger models always yield better adaptation.
Zero-Mem slashes memory operation time costs by over 57% while eliminating unnecessary LLM calls, revolutionizing how LLM agents manage memory.
RecHarness boosts recommender system performance by over 2% in real-world applications while minimizing the need for manual engineering interventions.
SemPIC boosts long-context retrieval performance by 13% over traditional KV caching, nearing the efficiency of full recompute without sacrificing model integrity.
Multi-agent collaboration can dramatically enhance open-ended reasoning capabilities, but its effectiveness varies significantly with task type.
MANTA reveals that multi-agent systems can dynamically adapt their communication structures during execution, leading to significant performance gains.
Diminishing returns in inference-time scaling reveal that more computation doesn't always equate to better performance in local computer-use agents.
Tycho's innovative approach to active abstraction allows agents to achieve perfect action efficiency while navigating complex game environments, outperforming human players significantly.
Autonomous data engineering agents struggle to achieve proficiency, with the best model scoring just 74.9 on a benchmark designed for real-world scenarios.
MemTxn not only prevents memory corruption in long-running agents but also restores complete state after faults, achieving a remarkable 24-point improvement over existing systems.
Gently-compressed LLMs can ace quality checks yet still invent dangerous procedural steps, exposing a critical blind spot in current safety assessments.
LLMs can achieve up to 93% precision in identifying intertextuality, but their reliability varies dramatically based on the complexity of the reuse dimensions involved.
VIG-RL achieves a new state-of-the-art in Verified Image Grounding by dynamically integrating visual evidence into text responses, outperforming static retrieval methods.
Transition supervision can dramatically boost LLM agent performance, outperforming standard policy optimization by leveraging environmental feedback.
Switching document formats can lead to accuracy drops of over 53%, revealing a hidden vulnerability in LLM workflows that demands urgent attention.
$\Sigma$-Mem not only tracks agent reliability but also adapts dynamically to feedback, outperforming traditional methods in multi-agent coordination.
Training agents in deep, evolving environments can dramatically enhance their performance, with a 9B model achieving a 30.6% accuracy increase through targeted design.
SpatialCLI enables VLMs to achieve a staggering 84.6% accuracy on spatial reasoning tasks, far surpassing existing models.
Change2Task recovers 29.2% more verified coding tasks than traditional methods, streamlining the training of coding agents.
AAPT boosts GUI agent success rates by 58% in decision-critical moments, eliminating delays without sacrificing accuracy.
Process evaluations reveal hidden failures in LLM reasoning, showing that lucky successes can mask critical deficiencies in agent performance.
GRSD transforms how agents learn from their own experiences, leading to significant improvements in both performance and generalization across tasks.
A safety gate in LLM supervisory control can outperform traditional methods, achieving a 16-fold improvement in disturbance rejection and significantly reducing harmful interventions.
Language model agents struggle with oncall RCA, achieving only 25.3% accuracy on realistic tasks, revealing a critical readiness gap for production environments.
Over 15% of benchmark failures for computer-use agents are misclassified, revealing critical flaws in current evaluation methods.
AgenticASR revolutionizes speech recognition by enabling real-time, intent-preserving transcription that adapts as speech evolves, outperforming traditional methods.
The Locksmith Loop can achieve nearly complete test coverage for legacy code migrations, ensuring that the migrated Java code functions identically to its COBOL predecessor.
Vibe-FDTR achieves near-perfect accuracy in thermal property analysis while slashing computational costs and execution time, revolutionizing how researchers approach FDTR data.
Agentic Metaverse Services could revolutionize how we interact with virtual environments, offering tailored agent capabilities that enhance both personal and business experiences.
MIND cuts memory injection attack success rates by more than half while maintaining performance, redefining the defense landscape for LLM agents.
ChronoMem allows LLM agents to seamlessly roll back their memory states, dramatically enhancing their ability to handle corrections and evolving information.
Substantial IP leakage risks in autonomous multi-agent systems are revealed through a novel framework that extracts harness capabilities dynamically during inference.
HALO retains 100% of relevant components in agentic AI responses, while traditional methods fail completely.
FaithEyes reveals that self-judging mechanisms in VLMs can drastically improve tool use fidelity, leading to more reliable multimodal reasoning.
ARES achieves up to a 27% reduction in optimization costs by intelligently adapting reasoning effort based on progress, outperforming fixed-effort approaches.
Forget static memory—MemHarness reconstructs past experiences to fit the present context, dramatically boosting decision-making performance in LLM agents.
Even the best vision-language models struggle with reliable evaluation of computer-using agents, but OS-Shepherd models offer a low-cost solution that matches their performance.
LedgerMind reveals that grounding multimodal reasoning in a structured evidence ledger can significantly mitigate common pitfalls like entity hallucination and unsupported reasoning.
Qwen-UI-Agent outperforms leading models in mobile and cross-platform tasks, achieving up to 97.5% accuracy on key benchmarks.
Existing models mismanage tool use, but Beacon achieves a balance that enhances performance on complex tasks while preserving accuracy on simpler ones.
Life-science AI agents can now access literature more efficiently, achieving over 16-point improvements in citation accuracy through natural language queries.
Identifying the exact source of agent failures could revolutionize how we approach system repairs, shifting from vague outcomes to precise interventions.
Weak student agents can learn effectively from their failures through targeted skill patches that bridge the gap between their capabilities and those of stronger teacher models.
FinanceHarness not only automates financial deep research but also reveals that even advanced LLMs struggle with specialized financial tasks, scoring below 40% on rigorous benchmarks.
LLMs can identify flagged security issues but struggle to uncover silent intrusions and create effective remediation plans, revealing a critical gap in their utility for real-world incident response.
LLMs can complete office tasks faster and cheaper than humans, but they still lag in quality, highlighting a critical gap in AI performance.
CAM-DF reduces tool acquisition costs by 37% while maintaining task success, challenging the conventional wisdom that more tools always lead to better outcomes.
TSDS cuts edge reasoning compute by up to 73% while ensuring reliable decision-making through smart deferral to cloud models.
Transforming hydrologic modeling, the Mass-Conserving Perceptron achieves state-of-the-art predictive accuracy with fewer parameters than traditional methods.
Simpler self-refinement strategies can outperform complex multi-agent systems in local language model deployments, challenging conventional wisdom about multi-agent architectures.
Rapidly inferring hidden partner capabilities can transform how autonomous agents collaborate in dynamic environments.
Native memory in foundation models can significantly enhance efficiency and performance, as shown by the innovative Metis architecture.
DREvo achieves unprecedented stability and performance in harness self-evolution, outperforming existing methods by over 16% on key benchmarks.
PowerAtlas achieves optimal electricity-computing scheduling, ensuring compliance with grid rules while maximizing efficiency in volatile AI workloads.
Nodes can now autonomously decide how to propagate information, leading to improved adaptability in diverse graph structures.
AgentMap reveals that a unified approach to ontology matching can significantly enhance performance across multiple semantic correspondence tasks.
UrbanDS outperforms traditional data science agents by effectively navigating complex urban datasets through a novel graph-guided multi-agent architecture.
Rethinking AI as a friction agent for reflection could fundamentally change how designers engage with their creative processes.
Achieving up to 6.65× speedup in NPU inference by automating Ascend C operator generation could revolutionize performance optimization in low-corpus environments.
Eco3S can replicate complex economic phenomena while allowing for flexible interventions and iterative refinement, setting a new standard for agent-based modeling in socio-economic research.
AlphaSchema reveals that systematic exploration of trading semantics can significantly enhance alpha mining effectiveness, yielding robust predictive performance across various LLMs.
Fine-tuning a language model for PID tuning can boost first-attempt success rates to an impressive 94%, revolutionizing how chemical processes are optimized.
Evidence-ledger adjudication enables AI to not only draft claims but also ensure their accuracy by effectively tracing and validating the supporting evidence.
Living-Harness enables agents to learn from past failures dynamically, leading to substantial performance improvements in interactive tasks.
EvoPINN autonomously discovers new algorithms for physics-informed neural networks, achieving significant performance improvements while ensuring scientific validity.
PUDA revolutionizes self-driving laboratories by enabling AI agents to autonomously execute experiments with complete data provenance, bypassing the limitations of traditional graphical interfaces.
Even the best LLMs fail to produce fully feasible travel plans more than half the time, revealing a significant gap in their ability to understand user needs.
Jointly optimizing knowledge construction and querying leads to a 6.3-point boost in answer correctness, revolutionizing how agents interact with their knowledge bases.
Coding agents may boost productivity, but they risk diminishing developers' understanding and long-term coding skills.
The rise of LLM agents threatens to erase the clarity of authorship and accountability in collaborative knowledge work, raising urgent questions about intellectual integrity.
Predicting the next token's KV entries can boost long-context LLM throughput by over 2.5 times without sacrificing latency or quality.
SkillRise achieves up to 8.5 percentage points better performance than leading methods by effectively reusing transferable skills across related tasks.
Voice Memory cuts speech recognition error rates by over 10% in noisy environments while ensuring the learned model remains auditable and portable.
AI agents can handle the engineering aspects of research but fall short in addressing the core research questions, leading to outright rejections from experts.
SkillSmith achieves unprecedented performance gains by seamlessly merging textual knowledge with parametric skills, outperforming traditional methods that treat these domains separately.
The shift from viewing psychology as a mere evaluative tool to a proactive design paradigm could redefine the future of human-machine partnerships.
Vulnerabilities spanning multiple functions can be detected more accurately with VulAgentRL, which verifies evidence through a novel Code Property Graph approach.
AgentSnare can absorb nearly 47% of an attacker's tool calls while ensuring that no real targets are compromised, showcasing a new frontier in adaptive cybersecurity defenses.
Binary classifiers misclassify nearly 40% of AI agents as humans, but a simple three-class framework achieves perfect detection across all tested evasion strategies.
Compiled harnesses in SIGIL boost execution of mandated steps to 86%, outperforming prose skills by a staggering 30%.
CodeSpec transforms feature development by ensuring that LLM-based code agents produce reliable and verifiable functional chains, achieving up to 70.7% pass rates on complex benchmarks.
CircuitProver not only automates hardware verification but also distills proof knowledge into reusable libraries, slashing verification time by over 23%.
A memory store design that enables conversational AI agents to recall information across sessions without sacrificing user privacy or context limits, achieving up to 80% retrieval accuracy on knowledge-update questions.
Explanation quality is a critical yet overlooked dimension of LLM agent performance, with many agents generating misleading explanations that can lead to incorrect code assessments.
Long-horizon evaluations reveal that agents suffer from compounding errors, but without proper baseline comparisons, we can't fully grasp the reasons behind their failures.
Trusting AI in cyber defense hinges on a benchmark that captures the real complexity of enterprise environments—OSB fills this critical gap.
FAVA achieves a remarkable 90.5% compliance rate in dynamically authorizing LLM agent actions, showcasing a breakthrough in context-sensitive permission management.
Schema normalization boosts schema-drift success to 91.3%, yet fails to address critical issues like stale evidence and session context errors in retrieval-augmented generation systems.
Executable Blender code transforms text-to-video generation, enabling unprecedented control over scene dynamics and visual fidelity.
Organized memory can halve retrieval costs, but without a strong management agent, LLMs struggle to maintain effective organization and answer quality.
Symptom-guided multimodal inputs can dramatically enhance zero-shot disease classification in veterinary settings, outperforming traditional image-only approaches.
Evaluating tool sets as a whole rather than in isolation allows HYSET to achieve superior retrieval performance, even in unseen domains.
Adaptive Org Routing outperforms fixed collaboration protocols by dynamically selecting the best approach for each task, revealing that organizational design must be revalidated for each model family.
Standardized evaluations reveal that while some AI capabilities are advancing rapidly, others are stagnating, challenging our assumptions about agent performance across domains.
VLMs can identify geometry clipping in games, but they generate significant false positives, highlighting their limitations in standalone bug detection.
Unifying the agent and speculator within a single model boosts next tool-call prediction accuracy by over 17% without sacrificing task performance.
MemLens reveals that a value-aware approach to memory management can drastically enhance the efficiency and personalization of LLM interactions, making memory records more impactful.