Search papers, labs, and topics across Lattice.
100 papers published across 3 labs.
Memory-enhanced agency enables LLMs to achieve robust long-term strategic execution, outperforming traditional methods in dynamic environments.
Proactive computing could redefine user interaction by enabling systems to anticipate needs rather than merely responding to them.
Current AI research agents miss critical metacognitive checks, leading to pervasive failures across diverse scientific tasks.
Agentic transactions could redefine how autonomous systems manage reliability and consistency in complex, multi-step workflows.
Personalization in AI co-scientists could be the key to unlocking novel research insights that generic systems overlook.
Current AI research agents miss critical metacognitive checks, leading to pervasive failures across diverse scientific tasks.
Agentic transactions could redefine how autonomous systems manage reliability and consistency in complex, multi-step workflows.
Personalization in AI co-scientists could be the key to unlocking novel research insights that generic systems overlook.
Skills stabilize agent execution by transforming noisy trajectories into procedural anchors, but they can fail under brittle assumptions and incompatible contexts.
SSPO enables deep search agents to learn more effectively by transforming teacher-student disagreements into actionable insights, leading to superior performance with fewer resources.
MARC's multi-agent orchestration allows for precise clinical AI reasoning while eliminating the need for manual prompt engineering.
Efficient belief synchronization in 6G networks can be achieved without requiring homogeneous AI models, preserving privacy and reducing costs.
Treatment leakage can drastically misclassify security outcomes, leading to flawed evaluations of agent safety.
FUSE achieves superior functional grounding performance while cutting computation costs, revolutionizing how agents interact with their environments.
Governing workflows with verifiable provenance can prevent unsupported claims and ensure accountability in agentic decision-making.
ERSkill boosts LLM performance by over 31% through dynamic, skill-guided memory retrieval that evolves with experience.
Segment-level consolidation in LycheeMemory V2 slashes construction costs by up to 86% while maintaining high accuracy in long-term memory tasks.
AI agents can only correctly synthesize implementations and proofs for less than two-thirds of tested multi-module software repositories, revealing significant gaps in current capabilities.
Reducing candidate repair options by half while maintaining performance challenges the notion that complex models always yield better results in agent harness repair.
VALG autonomously navigates the complexities of ML theory research, producing viable theorem candidates while preserving mathematical integrity across proof attempts.
StateBridge reveals that training-free hidden-state alignment can significantly enhance communication efficiency in LLM multi-agent systems, outperforming traditional methods.
CrEST redefines credit assignment in RL by shifting the teacher's role from directing updates to modulating their magnitude, leading to substantial performance improvements in multi-turn agent training.
SkillShapley reveals that not all steps in agent skills are created equal, enabling precise identification of high-impact actions that can enhance performance.
Human intervention in decision-making can enhance multi-agent systems' adaptability, with BoardroomAI achieving 62.11% selective repair of decisions while preserving context.
Legacy bioinformatics code can be transformed into efficient Rust implementations, slashing size by 80x and build time by 10x while boosting performance over threefold.
RippleMem boosts LLM accuracy by nearly 12% while slashing memory graph construction costs by 30x, transforming how agents recall and utilize past interactions.
Recursive self-improvement in quantitative trading research leads to a groundbreaking Sharpe ratio of +2.50, showcasing the potential for autonomous systems to enhance investment strategies over time.
Query-conditioned reuse boosts agent success by 10.7 points while slashing token usage by nearly 50%, transforming how we leverage past experiences in AI tasks.
No existing protocol has successfully unified persistent identity, capability-aware discovery, trust negotiation, and accountability for secure agent interoperability—until now.
ATOBench reveals that deceptive responses can obscure verification failures, fundamentally altering how autonomous penetration-testing agents interpret evidence and report vulnerabilities.
State-corruption attacks can be drastically reduced from 84.7% to just 2.3% with PIPES, while still preserving agent performance.
SynAct slashes worst negative slack to 27% of bootstrap synthesis, revolutionizing timing optimization in logic synthesis.
Task progress in vision-language-action models can be read directly from their internal representations, even before task-specific training, revealing a surprising depth of interpretability.
Intern-S2-Preview-397B not only excels in multimodal scientific reasoning but also enhances biological instruction performance without altering its foundational architecture.
Parallel reasoning in LLM agents can cut decoding time by up to 43% while maintaining performance, reshaping agent efficiency.
Current autonomous agents excel at practical problem-solving but often lack true methodological innovation, revealing critical gaps in their development as independent researchers.
SMA achieves superior spatial reasoning in frozen VLMs by converting verified experiences into transferable lessons, outperforming traditional methods without the need for parameter updates.
AutoDesign outperforms existing design systems by aligning with human design principles and achieving superior poster generation quality through recursive self-improvement.
Matched execution scores can hide up to 64.3 points of command-path failure, revealing a critical gap in evaluating LLM coding agents.
Reflection in search agents can be transformed into a powerful memory-control policy, leading to superior performance in complex reasoning tasks.
Memory systems can now be quantitatively assessed through a utility-capacity frontier, revealing how to optimize agent memory for better performance.
Multi-hop reasoning in API interactions is a critical bottleneck, with top models faltering under policy constraints and complex queries.
An innovative agentic workflow can fully modernize a 48-year-old quantum chemistry codebase without introducing any errors, achieving perfect validation across extensive tests.
Larger, R&D-focused firms are not just adopting ChatGPT Enterprise faster; they’re also leveraging it more intensely across diverse job functions, particularly among early-career employees.
GUIDE slashes document processing time from days to under two hours while maintaining a remarkable 96% success rate in generating deployment-ready artifacts.
AI agents may ace endpoint identification but falter in delivering the evidence-based diagnostics essential for real-world telecom troubleshooting.
Tool-using LLMs face a near-universal robustness gap, but combining Bayesian Tool Memory with reinforcement learning can boost recovery performance by over 40% in failure scenarios.
Skills that are meant to enhance LLM agents can paradoxically lead to significant task failures and inefficiencies, challenging the assumption that more skills always improve performance.
Supervised fine-tuning outperforms complex reinforcement learning techniques in ensuring multilingual API reliability, challenging the notion that more sophisticated methods are always necessary.
Reflexive scores often outperform traditional UQ methods in multi-turn interactions, revealing a critical gap in how we assess uncertainty in LLM agents.
Transforming agent failures into actionable recovery strategies, DARC enhances performance without bloating context, proving that less can be more in self-correction.
MBA-Bench reveals that integrating visual cues into business ideation agents can boost performance by over 77% compared to text-only methods.
ToolHazard reveals that injection timing and placement are critical factors in exploiting vulnerabilities of LLM-based agents, leading to new insights in adversarial robustness.
Serving costs of memory systems can deviate by up to 69% from predictions based on conversation length, revealing hidden complexities in agentic memory performance.
Task completion in LLMs can be achieved at a staggering 92% increase in execution time due to subtle manipulation of skill selection and instruction processes.
Explicit role coordination among LLM-based tools can significantly enhance software quality, but it requires careful human oversight to manage deviations.
Despite the promise of multi-agent systems in software engineering, key frameworks still lack advanced features and show no significant performance difference in summarization tasks.
Videos generated through a novel agentic optimization framework are preferred by users over traditional methods, achieving a striking 69% win rate in preference studies.
Clinician-interactive AI can boost diagnostic accuracy and consistency in thyroid ultrasound reporting while reducing time spent on segmentation and reporting tasks.
Behavioral diversity boosts multiagent performance, yielding up to 48% better outcomes than traditional optimization methods.
Descriptive metadata can boost job execution success rates in multi-cluster AI workflows from 48% to 87%, revolutionizing operational efficiency.
MindMemOS enables AI agents to autonomously refine their memories and skills, achieving state-of-the-art accuracy in dynamic environments.
Grounding LLM explanations in digital twin outputs boosts anomaly diagnosis quality and operator engagement in cyber-physical systems.
By integrating structured website exploration with task-trajectory synthesis, SynWeaver enables web agents to achieve unprecedented levels of generalization across diverse websites.
Proactive computing could redefine user interaction by enabling systems to anticipate needs rather than merely responding to them.
EgoCITE achieves a remarkable 36x reduction in cost while boosting answer accuracy by over 14% in egocentric memory tasks.
Memory-enhanced agency enables LLMs to achieve robust long-term strategic execution, outperforming traditional methods in dynamic environments.
Models misjudge authorized actions nearly 30% of the time, revealing a critical flaw in decision-making at action boundaries.
Keeping GPU-computed decisions on-device can boost performance by up to 2.39x, revolutionizing LLM-agent control efficiency.
Leading MLLMs falter on the new VideoGAIA benchmark, scoring under 60% accuracy in complex, multi-turn video understanding tasks.
AVA-Encoder achieves a 73.1% relative improvement in video representation learning, enabling agents to produce cinematic-grade videos with far fewer resources.
LLM-based agents can negotiate contracts effectively in low uncertainty, but their performance collapses under high uncertainty, revealing critical flaws in their cooperative capabilities.
SkillZip achieves unprecedented skill compression efficiency by transforming repetitive actions into reusable structures, eliminating the need for costly evaluations.
Identifying key traits in human-AI collaboration could redefine how we understand and teach the skills necessary for successful teamwork with AI agents.
Models can now autonomously adapt to new GUI interfaces post-deployment, achieving a 7.4% accuracy boost without human intervention.
A centralized MCP gateway can transform fragmented enterprise authentication into a cohesive and secure identity management system.
A staggering 66% of vulnerabilities in agentic LLMs stem from perception-layer issues, while action-layer risks remain alarmingly underexplored.
Common governance semantics can unify diverse agentic systems, ensuring reproducibility and interoperability across frameworks.
MISA-T boosts rollout throughput by over 53% while preserving workload integrity, revolutionizing how RL pipelines manage heterogeneous demands.
TideRL boosts RL training goodput by up to 5.6 times, transforming how we approach efficiency in multi-turn agentic workloads.
Tool architecture can enhance coding agent performance, with structured interfaces yielding up to 4.7 times more consistency and 41.6% fewer steps in task execution.
Reverse engineering binary software remains largely unsolved, with leading AI models only achieving a mere 31.5% success rate on complex tasks.
The Verifiability Gap reveals that bigger AI models in FinTech may compromise auditability, challenging assumptions about their superiority in governance.
Programmatic skill learning can slash agent costs while enhancing performance, with SpeedRunner leading the charge in cost-efficient adaptation.
A simple screenshot-based judge often outperforms complex LLM evaluation methods, challenging assumptions about the necessity of intricate judging pipelines.
Runtime contracts for AI safety could fundamentally change how we ensure the reliability of autonomous agents in real-world applications.
InSight-doc cuts hallucination by over 40% and inference latency by up to 68%, all while boosting accuracy on long-document tasks.
Long-horizon software development can thrive without relying on persistent agents, as demonstrated by Genesis's ability to evolve complex systems through finite-lived contributors.
SHAPER enables embodied agents to self-evolve their skills and context without retraining, unlocking new possibilities for adaptation in fixed-interface scenarios.
MobileMem shifts the paradigm from static information retrieval to dynamic experiential learning, enabling AI agents to evolve alongside their users.
SKILLER achieves up to 20.4 percentage points improvement in skill generation for small language models, making high-quality task execution accessible without the prohibitive costs of closed-source solutions.
The Signal Rail transforms how users perceive conversational agents by enabling intuitive, visual communication of their internal states through motion rather than text.
Automating data selection with DataMaster not only reduces manual effort but also enhances performance across diverse applications, challenging traditional heuristic methods.
Context interference can significantly degrade the performance of search agents, but a novel context refiner shows how to enhance their reliability and efficiency dramatically.
Action policies in multilingual tool-using agents are more consistent than previously thought, but only for models above a certain size threshold.
LLMs struggle with personal information retrieval, with the best model only achieving 57.3% accuracy on a new benchmark designed to evaluate mobile assistant capabilities.
Despite advances in AI, top-performing agents struggle with real-world data science, achieving only 56.70% success in automating complex workflows.
Current LLM agents struggle to keep pace with the complexities of real-life assistance, scoring low on a benchmark designed to test their proactive and persistent capabilities.
Catastrophic remembering leads to an explosive growth of agentic prompts, but simple prompt comments can reverse this trend and enhance performance dramatically.
Real-time hallucination detection in LLMs is revolutionized by a lightweight adapter that transforms uncertainty into actionable feedback, preventing undesired actions before they occur.
Even the best voice agents struggle with real-world conversational tasks, scoring below 50% in effective assistance across diverse scenarios.
Autoresearch agents can waste compute resolving the same issues repeatedly, but targeted interventions can dramatically enhance their efficiency and performance.
Validation rewards increased by 76% as SINKFLEX-RL tackles the memory limitations of long-horizon reinforcement learning tasks.
AI can generate novel mathematical insights, as evidenced by tightening the Grothendieck constant bounds significantly.
Over 3.7 million agent skills on GitHub reveal how developers are innovating in the absence of formal registries or type checks.