Search papers, labs, and topics across Lattice.
100 papers published across 7 labs.
Achieving a staggering 95.5% on the ARC-AGI-3 benchmark, Prime Agent redefines the capabilities of long-horizon coding agents by seamlessly integrating memory and computation.
TAU-Agent leverages a novel retrieval mechanism to enhance traffic anomaly detection, achieving impressive benchmark results that challenge existing paradigms in video analysis.
MAGE transforms how coding agents operate by externalizing engineering intent and establishing governance, paving the way for more reliable software development.
JIT-Agent redefines agent performance by enabling on-the-fly harness evolution, leading to substantial improvements over existing models.
An autonomous AI agent achieved 99.5% of the optimal solution for cell-edge power control while slashing inference costs by 600x, revolutionizing the role of researchers in algorithm design.
JIT-Agent redefines agent performance by enabling on-the-fly harness evolution, leading to substantial improvements over existing models.
An autonomous AI agent achieved 99.5% of the optimal solution for cell-edge power control while slashing inference costs by 600x, revolutionizing the role of researchers in algorithm design.
Imitation learning enables automated theorem provers to solve 46% more problems while drastically reducing proof steps compared to traditional methods.
Agents fall short in machine learning development, often locked in narrow loops while humans adapt and innovate across tasks.
Filtering out fit sequences during fine-tuning can boost post-training performance by up to 17%, reshaping how we approach model training for RL applications.
KOPE's ability to retain and leverage past optimization experiences leads to a 1.54x speedup in kernel optimization tasks compared to existing methods.
Language-model agents in SwarmWorld can self-organize to build resilient technological societies, surpassing isolated search methods in innovation and adaptability.
Operating costs for multi-agent LLM workflows can be slashed without sacrificing performance, thanks to ProgRouter's adaptive routing strategy.
AsymSpec achieves 90% accuracy with 1.7x speedups by leveraging asymmetric context access, redefining efficiency in agentic LLMs.
TAU-Agent leverages a novel retrieval mechanism to enhance traffic anomaly detection, achieving impressive benchmark results that challenge existing paradigms in video analysis.
Existing debugging methods for multi-agent systems fail to reliably reproduce and repair failures, but a new symptom-driven approach shows a dramatic improvement in effectiveness.
LocalLSTC reveals that organizing control information temporally can drastically improve the performance of GUI agents, achieving over 64% success rates where previous models faltered.
A self-evolving RCA harness outperforms traditional methods by reusing general agent capabilities, achieving a remarkable 59.0% accuracy in diagnosis.
Calibrating procedural relations in skill retrieval can boost task performance by over 10 points while streamlining execution efficiency.
AdaVDR achieves superior performance in video deep research by dynamically adapting its tool use based on the specific capabilities of the model and the nature of the task.
Game development offers a powerful alternative to traditional reward signals, enabling RL systems to leverage executable environments for superior feedback and data generation.
Multi-agent collaboration in code generation can boost functional and security correctness by over 19 percentage points, reshaping how we approach secure coding.
VISA's self-evolving framework not only enhances multimodal instruction synthesis but also adapts in real-time to improve training data quality and model performance.
Memory management strategies that adapt to individual users can significantly enhance agent performance, outperforming traditional static approaches.
Switching between natural language and structured graphs can boost multi-agent LLM performance by over 12 percentage points while slashing token usage by more than threefold.
Traditional measures of neighborhood livability fail to account for the real-world challenges faced by residents, revealing hidden burdens for those with limited mobility and caregiving roles.
ReDiR slashes attack success rates to under 8% by embedding trajectory-level safety insights directly into the action generation process of LLM agents.
SkillShield slashes malware-generation severity from 3.37 to 0.58, showcasing a new frontier in prompt-space security for coding agents.
Praxist achieves 60 medals on the MLE-bench with an order of magnitude less spending, revolutionizing how autonomous R&D agents can learn from past experiments.
Teams adopting coding agents may be accelerating development at the cost of doubling cognitive complexity and increasing static-analysis warnings without proper AI configuration.
psRL can boost training throughput by over 5x by effectively sharing redundant prefixes across samples, transforming how we approach agentic AI training efficiency.
Mediation with Metis slashes operation time by over 45%, making software agents more efficient in managing external interactions.
IAPO redefines credit assignment in multi-turn interactions, showing that leveraging influence-dependency graphs can significantly enhance service agent performance.
High-confidence predictions from LLMs in hidden information scenarios are alarmingly inaccurate, with only 1 in 62 correct, challenging the assumption that confidence reflects correctness.
CBPO redefines credit assignment in RLVR, enabling precise decision sensitivity that boosts performance across diverse benchmarks.
Dynamic delegation during reasoning allows LLM agents to outperform static routing strategies, achieving better task success rates with a Bayesian approach.
Standardizing terminal-outcome advantages can significantly boost online learning efficiency in asynchronous reinforcement learning scenarios.
PinSieve filters out over twice the non-actionable content while boosting review productivity and reducing operational costs, all while ensuring human oversight remains intact.
LLM coding agents can now leverage a literate programming environment that enhances context utilization and mimics human IDE capabilities.
Jointly training tool creation and use allows LLMs to achieve unprecedented accuracy on procedural reasoning tasks, outperforming larger models and enhancing smaller ones.
MetaRAG achieves a superior accuracy-efficiency trade-off in agentic RAG by aligning decision-making with the model's internal beliefs, outperforming traditional RL methods.
Evidence Blindness can be mitigated by a novel navigation framework that organizes the corpus, leading to faster and more accurate evidence retrieval.
EviDx reveals that structured scaffolding and dynamic evidence acquisition can significantly enhance the reliability and effectiveness of AI-driven clinical diagnosis.
Simthesizer achieves up to 284.96x faster simulation speeds while maintaining a mere 2.51% throughput error, revolutionizing how we model LLM serving systems.
Continuous skill verification in RL agents leads to a substantial performance boost, outperforming traditional static skill banks.
Transforming static industrial documents into dynamic action-effect relationships leads to significantly improved decision-making in operational processes.
Orchestrator agents in deep research systems are responsible for nearly 85% of citation errors, revealing a critical vulnerability in how these systems generate trustworthy reports.
AgentSpec slashes response times for LLM agents by addressing high rejection rates and optimizing token budgets, outperforming existing methods.
LLMs can dramatically improve taint analysis for Android apps, achieving an F1-score of 0.96 compared to traditional tools' 0.55.
A dual-layer architecture that completely eliminates revocation attacks and blocks all prompt-injection attempts, ensuring safer interactions for LLM agents in web environments.
Directly editing 3D meshes in Blender through a visual-centric, agentic approach could revolutionize how artists interact with complex geometry.
A structured LLM-based multi-agent system achieves 100% success in manufacturing process planning, transforming how design artifacts are interpreted into actionable plans.
SCOUT slashes tool-token consumption by 99%, revolutionizing how LLMs manage context and discover tools in enterprise environments.
The "handoff tax" reveals that switching between models can significantly degrade performance while increasing costs, challenging the assumption that stronger models always yield better outcomes in coding tasks.
Even the strongest Android GUI agents are universally vulnerable to runtime anomalies, revealing critical flaws in their robustness.
Transforming action-constraining states into artifacts can lead to a staggering 100% deactivation of safety blockers in LLM workflows.
Recuris transforms long-horizon task execution by reducing common failures by up to 80% and achieving state-of-the-art success rates across multiple models.
StepGuard slashes attack success rates by over 77% while maintaining nearly all utility, setting a new standard for safety in LLM-based agents.
Evolving agent harnesses can boost performance by up to 35% in enterprise environments, even with fixed model weights.
No trading system architecture is inherently safe; adversarial signals can compromise decisions across all roles, revealing a critical vulnerability in multi-agent setups.
Paritok-4B compresses coding agent context to just 25.7% of its original size while retaining 86.5% of solve quality, outperforming existing models.
OODA-Tool achieves substantial improvements in task success by effectively separating state tracking from action execution, outperforming traditional methods in multi-turn interactions.
MAGE transforms how coding agents operate by externalizing engineering intent and establishing governance, paving the way for more reliable software development.
Strong logical planning in LLM agents doesn't ensure safe execution—PeakBench reveals critical flaws in existing benchmarks that could lead to resource overflows.
Despite diverging philosophies, leading agent harnesses are converging on a shared architecture, highlighting a critical gap in external verifiability for trust in AI systems.
Transforming long-form audio meeting comprehension, the GRGA model leverages graph-based planning to overcome acoustic loss and memory challenges, achieving superior QA performance.
Crase achieves 3× higher recall at a third of the cost compared to existing deep research agents, redefining efficiency in scholarly search.
EviGraph's innovative approach to evidence construction boosts accuracy by over 30% compared to traditional methods, reshaping how agents validate information.
ToolMinimize slashes privacy exposure in LLM tool calls by up to 92% without sacrificing task validity, addressing a major vulnerability in AI systems.
Malicious skills can exploit agent permissions to cause physical harm, but a new authority layer could prevent this without hindering legitimate actions.
A new framework categorizes molecular LLM agents into four levels of autonomy, revealing critical gaps and risks in their deployment for scientific discovery.
MediSkill-Evo not only boosts diagnostic accuracy but also ensures adherence to clinical processes, achieving a remarkable 93.61% recovery rate under patient-behavior pressure.
LLMs can significantly boost forecasting accuracy, but their effectiveness is hampered by measurement limitations and sensitivity to input changes.
VideoRover unifies video reasoning and external knowledge retrieval, achieving competitive performance with fewer resources than larger models.
OaK transforms LLM agents by grounding their decision-making processes in dynamically constructed ontologies, leading to enhanced reliability in multi-step reasoning.
TRACE transforms high-performing LLMs into consistently reliable agents, achieving a remarkable 34.6-point boost in task consistency.
An algorithm that guarantees safe execution edits in agent runtimes, preventing potentially hazardous duplications and omissions in task execution.
Achieving high-quality time series forecasting with just a few examples, MetaCaster revolutionizes the way lightweight forecasters are trained in data-scarce environments.
SkillAlchemy boosts agent skill creation efficiency, achieving nearly 20% higher success rates without relying on human authorship.
Agentic AI could revolutionize software development, but it also introduces severe IT security risks that demand immediate attention.
Over-reliance on AI assistance can significantly undermine long-term skill development, leading to a false sense of competence in problem-solving tasks.
Intermediate task decomposition in LLM-agent systems can improve accuracy but fails to consistently outperform a single agent in VAT determination tasks.
LLM agents struggle with network configuration, revealing failures that extend beyond simple command errors to deeper issues in task adherence and planning.
Routing before reasoning can boost function-calling success rates in language models by over 12%, transforming how we approach tool interaction.
Pruning 64.58% of reasoning tokens not only streamlines MLLM performance but also revitalizes the model's reliance on visual evidence, enhancing task accuracy.
AI can now autonomously detect agency and reconstruct policies from mere observations, paving the way for more sophisticated cooperative behaviors.
AgentFlow eliminates data compromise in LLM agents, achieving a remarkable reduction from 33% to 0% in confirmed unsafe actions while boosting utility.
Automating the entire medical imaging model-development pipeline can yield competitive baselines with minimal engineering effort, challenging traditional iterative approaches.
DG-Mem achieves substantial gains in reasoning tasks by leveraging a novel memory architecture that adapts dynamically without altering model parameters.
Reranking tools based on risk exposure can drastically improve safety in LLM agent interactions without sacrificing utility.
ARGUS reliably identifies root causes in Kubernetes incidents but struggles to gain trust for its prescriptive recommendations, revealing a critical gap in automated incident response systems.
AutoSaddler achieves up to 10% performance improvement in LLM agents by automatically optimizing harnesses based on execution failure signals.
Verified execution experience can be transformed into a reusable resource, boosting model performance on complex workflows by up to 15.5 percentage points.
Current LLMs falter in mobile environments, with performance plummeting under real-world constraints, revealing critical gaps in their planning capabilities.
Achieving a staggering 95.5% on the ARC-AGI-3 benchmark, Prime Agent redefines the capabilities of long-horizon coding agents by seamlessly integrating memory and computation.
Achieving top-tier performance in complex professional tasks with a model significantly smaller than its competitors reveals a new frontier in agentic intelligence.
A decentralized bidding system for LLM agents not only enhances efficiency but also reduces manipulation risks, outperforming traditional orchestration methods.
Valid mandate signatures can't protect against manipulated transaction contexts, exposing critical vulnerabilities in LLM-driven payment systems.
TrustShift attacks can manipulate agent trust dynamics, achieving a staggering 69.5% success rate before being mitigated by a novel defense framework.
Achieving an 86.17% success rate in reproduction test generation, DPIAgent reveals that structured task separation can dramatically enhance performance in automated software engineering.
Pufibara outperforms existing agent harnesses by achieving higher task success rates and significantly lower resource consumption in physical system modeling.
Failed tool calls can increase the likelihood of repeating errors by over 800%, revealing a critical flaw in how language models process failure information.
LLMs can now be held accountable in real-time, with a verification system that translates natural language policies into executable obligations, drastically reducing the risk of critical errors.
"Prompt-form collapse" can drastically reduce success rates, but TOWN-VLA's controlled intervention boosts performance by ensuring only meaningful prompts are executed.
In a world of rapidly expanding agent capabilities, a new retrieval-based approach keeps accuracy high and costs low, outperforming traditional methods by a significant margin.