Search papers, labs, and topics across Lattice.
LLM-based autonomous agents, tool-augmented language models, function calling, and agentic workflows.
#7 of 24
3
JIT-Agent redefines agent performance by enabling on-the-fly harness evolution, leading to substantial improvements over existing models.
An autonomous AI agent achieved 99.5% of the optimal solution for cell-edge power control while slashing inference costs by 600x, revolutionizing the role of researchers in algorithm design.
Imitation learning enables automated theorem provers to solve 46% more problems while drastically reducing proof steps compared to traditional methods.
Agents fall short in machine learning development, often locked in narrow loops while humans adapt and innovate across tasks.
Filtering out fit sequences during fine-tuning can boost post-training performance by up to 17%, reshaping how we approach model training for RL applications.
KOPE's ability to retain and leverage past optimization experiences leads to a 1.54x speedup in kernel optimization tasks compared to existing methods.
Language-model agents in SwarmWorld can self-organize to build resilient technological societies, surpassing isolated search methods in innovation and adaptability.
Operating costs for multi-agent LLM workflows can be slashed without sacrificing performance, thanks to ProgRouter's adaptive routing strategy.
AsymSpec achieves 90% accuracy with 1.7x speedups by leveraging asymmetric context access, redefining efficiency in agentic LLMs.
TAU-Agent leverages a novel retrieval mechanism to enhance traffic anomaly detection, achieving impressive benchmark results that challenge existing paradigms in video analysis.
Existing debugging methods for multi-agent systems fail to reliably reproduce and repair failures, but a new symptom-driven approach shows a dramatic improvement in effectiveness.
LocalLSTC reveals that organizing control information temporally can drastically improve the performance of GUI agents, achieving over 64% success rates where previous models faltered.
A self-evolving RCA harness outperforms traditional methods by reusing general agent capabilities, achieving a remarkable 59.0% accuracy in diagnosis.
Calibrating procedural relations in skill retrieval can boost task performance by over 10 points while streamlining execution efficiency.
AdaVDR achieves superior performance in video deep research by dynamically adapting its tool use based on the specific capabilities of the model and the nature of the task.
Game development offers a powerful alternative to traditional reward signals, enabling RL systems to leverage executable environments for superior feedback and data generation.
Multi-agent collaboration in code generation can boost functional and security correctness by over 19 percentage points, reshaping how we approach secure coding.
VISA's self-evolving framework not only enhances multimodal instruction synthesis but also adapts in real-time to improve training data quality and model performance.
Memory management strategies that adapt to individual users can significantly enhance agent performance, outperforming traditional static approaches.
Switching between natural language and structured graphs can boost multi-agent LLM performance by over 12 percentage points while slashing token usage by more than threefold.
Traditional measures of neighborhood livability fail to account for the real-world challenges faced by residents, revealing hidden burdens for those with limited mobility and caregiving roles.
ReDiR slashes attack success rates to under 8% by embedding trajectory-level safety insights directly into the action generation process of LLM agents.
SkillShield slashes malware-generation severity from 3.37 to 0.58, showcasing a new frontier in prompt-space security for coding agents.
Praxist achieves 60 medals on the MLE-bench with an order of magnitude less spending, revolutionizing how autonomous R&D agents can learn from past experiments.
Teams adopting coding agents may be accelerating development at the cost of doubling cognitive complexity and increasing static-analysis warnings without proper AI configuration.
psRL can boost training throughput by over 5x by effectively sharing redundant prefixes across samples, transforming how we approach agentic AI training efficiency.
Mediation with Metis slashes operation time by over 45%, making software agents more efficient in managing external interactions.
IAPO redefines credit assignment in multi-turn interactions, showing that leveraging influence-dependency graphs can significantly enhance service agent performance.
High-confidence predictions from LLMs in hidden information scenarios are alarmingly inaccurate, with only 1 in 62 correct, challenging the assumption that confidence reflects correctness.
CBPO redefines credit assignment in RLVR, enabling precise decision sensitivity that boosts performance across diverse benchmarks.
Dynamic delegation during reasoning allows LLM agents to outperform static routing strategies, achieving better task success rates with a Bayesian approach.
Standardizing terminal-outcome advantages can significantly boost online learning efficiency in asynchronous reinforcement learning scenarios.
PinSieve filters out over twice the non-actionable content while boosting review productivity and reducing operational costs, all while ensuring human oversight remains intact.
LLM coding agents can now leverage a literate programming environment that enhances context utilization and mimics human IDE capabilities.
Jointly training tool creation and use allows LLMs to achieve unprecedented accuracy on procedural reasoning tasks, outperforming larger models and enhancing smaller ones.
MetaRAG achieves a superior accuracy-efficiency trade-off in agentic RAG by aligning decision-making with the model's internal beliefs, outperforming traditional RL methods.
Evidence Blindness can be mitigated by a novel navigation framework that organizes the corpus, leading to faster and more accurate evidence retrieval.
EviDx reveals that structured scaffolding and dynamic evidence acquisition can significantly enhance the reliability and effectiveness of AI-driven clinical diagnosis.
Simthesizer achieves up to 284.96x faster simulation speeds while maintaining a mere 2.51% throughput error, revolutionizing how we model LLM serving systems.
Continuous skill verification in RL agents leads to a substantial performance boost, outperforming traditional static skill banks.
Transforming static industrial documents into dynamic action-effect relationships leads to significantly improved decision-making in operational processes.
Orchestrator agents in deep research systems are responsible for nearly 85% of citation errors, revealing a critical vulnerability in how these systems generate trustworthy reports.
AgentSpec slashes response times for LLM agents by addressing high rejection rates and optimizing token budgets, outperforming existing methods.
LLMs can dramatically improve taint analysis for Android apps, achieving an F1-score of 0.96 compared to traditional tools' 0.55.
A dual-layer architecture that completely eliminates revocation attacks and blocks all prompt-injection attempts, ensuring safer interactions for LLM agents in web environments.
Directly editing 3D meshes in Blender through a visual-centric, agentic approach could revolutionize how artists interact with complex geometry.
A structured LLM-based multi-agent system achieves 100% success in manufacturing process planning, transforming how design artifacts are interpreted into actionable plans.
SCOUT slashes tool-token consumption by 99%, revolutionizing how LLMs manage context and discover tools in enterprise environments.
The "handoff tax" reveals that switching between models can significantly degrade performance while increasing costs, challenging the assumption that stronger models always yield better outcomes in coding tasks.
Even the strongest Android GUI agents are universally vulnerable to runtime anomalies, revealing critical flaws in their robustness.