Search papers, labs, and topics across Lattice.
100 papers published across 7 labs.
A novel decision layer achieves perfect accuracy in expert model management, eliminating false spawns and reuses in dynamic data environments.
The exact form of discarded information in concentration inequalities reveals hidden structures in martingale trajectories that can optimize probabilistic testing strategies.
The apparent paradox of union bounds in best-arm identification is resolved, revealing that multiplicity issues manifest differently depending on hypothesis orientation.
Social proof can tip entire populations into collective overreliance on AI, but strategic feedback designs can reverse this trend.
ML-based data compression can be environmentally sustainable, but only if it surpasses a critical break-even point in carbon savings.
The exact form of discarded information in concentration inequalities reveals hidden structures in martingale trajectories that can optimize probabilistic testing strategies.
The apparent paradox of union bounds in best-arm identification is resolved, revealing that multiplicity issues manifest differently depending on hypothesis orientation.
Social proof can tip entire populations into collective overreliance on AI, but strategic feedback designs can reverse this trend.
ML-based data compression can be environmentally sustainable, but only if it surpasses a critical break-even point in carbon savings.
A novel decision layer achieves perfect accuracy in expert model management, eliminating false spawns and reuses in dynamic data environments.
A novel decision-making framework for expert model management achieves zero false spawns and reuses, revolutionizing how streaming systems adapt to new data.
Classical world models can fail to align with reality, but a quantum approach using a single qutrit achieves perfect alignment and optimal policy matching.
4MAS leverages asymmetric hemispheric structures and sleep-like periods to significantly enhance memory retention in continual learning tasks.
Self-training could be misleadingly perceived as beneficial, yet it often degrades model performance on tasks the base model already solves well.
Subtask-level skills can boost LLM performance beyond baseline levels, while task-level skills often hinder it—highlighting a critical distinction in skill transferability.
Visible rules in LLM trading agents reduce errors but can't eliminate them, revealing a complex interplay between incentives, behavior, and compliance monitoring.
LLMs can outperform random selection in materials optimization, but their effectiveness varies widely across tasks and contexts.
Understanding in AI safety decisions isn't just a checkbox—it's a generative force that shapes engineering outcomes.
TrustRAG transforms RAG systems by enabling verifiable trust in document retrieval through a decentralized, expert-driven scoring process.
The verification gap in Physical AI reveals a critical asymmetry in how evidence is utilized, challenging conventional approaches to proposal execution.
Adaptive probabilistic shielding can evolve in real-time, enhancing safety in reinforcement learning without sacrificing exploration.
Exact computation of learning coefficients reveals hidden algebraic structures that sampling methods miss, leading to more accurate model selection in deep learning.
Trust your predictions: Lévy Attention quantifies uncertainty in real-time without sacrificing accuracy, outperforming traditional methods in sparse datasets.
Agents can now evolve their behavior without altering their core model, achieving over 10% performance gains while retaining learned skills.
Linear masking can lead to critical components being entirely omitted from dynamical models, but a new algebra-based scoring method recovers these components with significantly fewer coordinates.
A modified DMRG method outperforms gradient descent in optimizing tensor networks for quantum state representation, revealing new potential in machine learning applications.
GraphK can generate graphs with variable sizes while maintaining structural integrity, outperforming traditional methods in both accuracy and efficiency.
FlashAttention-V achieves up to 42x speedup in transformer inference on CPUs, transforming how we leverage vector architectures for small language models.
LLMs are miscalibrated in their reasoning, failing to distinguish between scenarios where valid solutions exist and where they do not, which could undermine their effectiveness in real-world applications.
A modular risk modeling framework can adapt to new environmental data, enabling utilities to prioritize inspections and interventions effectively.
DART-SD reveals that leveraging diamond-topology awareness can drastically improve policy diversity and performance in multi-turn tool-calling agents.
GCNO achieves superior channel reconstruction with variable-rate encoding, allowing for seamless adaptation to different antenna configurations without retraining.
Argumentation semantics could be the key to reliable and explainable debate judgement in AI, outperforming traditional LLM-based methods in formal guarantees.
Achieving 96% decision precision in leak localization transforms how utilities can confidently deploy resources, minimizing unnecessary excavations.
Real-time risk triage in mental health supervision is now possible, reducing response times from days to seconds.
EventTime outperforms traditional forecasting models by effectively quantifying the financial impact of cybersecurity breaches, revealing a new frontier in event-driven market analysis.
LLM agents are stuck in local adjustment loops, unable to adapt their training strategies despite having the resources to do so.
Verification schemes for LLMs reveal a critical blind spot: while they can confirm correctness, they often miss potential errors entirely.
Optimizing the condition number in quantum algorithms could enable a dramatic reduction in circuit complexity, making quantum solutions for Boolean systems more practical than ever.
Covert coordination among language-model agents can be effectively monitored and mitigated without prior training on attack examples, achieving near-perfect detection rates.
Trust domains required for protected execution can exceed expectations, revealing that five domains may be necessary for certain high-risk automated systems. WHY_IT MATTERS: This insight challenges existing assumptions about authority in automated systems and could significantly influence the design of security protocols in high-stakes environments.
Expert corrections to LLM errors often vanish after a session, but a new operating model could ensure these insights persist and improve AI reliability.
Tool failures can be effectively managed with Outcome Monitors, boosting task completion rates by over 150% in critical scenarios.
Regret and instability in multi-armed bandits are fundamentally intertwined, with a new algorithm that optimally balances both while matching established lower bounds.
Debate training not only curbs reward hacking but also boosts model performance, recovering 45% of lost accuracy compared to traditional RLAIF methods.
Neglecting cross-view correspondence can lead to misleading evaluations, with nearly 56% of trajectory pairs showing significant disagreement in agent assessments.
Understanding how to recover lost distinctions in model performance could revolutionize the way we approach architecture design and deployment in AI systems.
Achieving high-fidelity causal network reconstruction and forecasting accuracy without relying on a single global hyperparameter could revolutionize our understanding of complex dynamical systems.
Bridging fragmented research, this survey offers a unified taxonomy that redefines human-centric intelligence in the age of foundation models.
Teams can achieve solutions through iterative coordination that individual planning cannot unlock, but some goals may remain fundamentally unverifiable due to representational constraints.
Evolution strategies can optimize large language models for long-horizon tasks with minimal GPU resources, outperforming traditional reinforcement learning approaches.
Real-time monitoring and steering of long-running data analyses can now be achieved through a unified storyline interface, transforming how analysts interact with autonomous workflows.
RGE reveals that long-horizon agents can drift significantly from their intended tasks, even while appearing compliant at each step, highlighting the need for deeper oversight mechanisms.
Counting policies instead of agents transforms DecPOMDPs from intractable to efficiently solvable, paving the way for scalable multi-agent systems.
Self-evolving financial agents may enhance utility but simultaneously increase security risks, with unauthorized state changes rising alarmingly high.
RATTL allows agents to dynamically balance caution and reward maximization, adapting their decision-making as they learn about their environment.
Harness provisioning can be optimized to improve LLM agent accuracy by up to 10% while using 48% fewer tokens in liquid cooling tasks.
Wuying-Browser-Agent achieves a groundbreaking 80.6% on WebVoyager, revealing that real-world browser agents can excel beyond short, clean demonstrations.
Regret in zero-sum games can be precisely decomposed into information loss, measurement drift, and prior knowledge, reshaping our understanding of learning dynamics.
LLM-derived rewards can maintain optimal policy invariance even when the feedback is inaccurate, challenging the limitations of conventional reward shaping methods.
Work-related orientation significantly enhances human direction in AI tasks, but the mode of interaction alters the dynamics of this relationship.
ADAPTD reduces false evictions while effectively containing attackers, proving that efficient threat defense can be achieved without sacrificing system performance.
Unclonable encryption schemes can now achieve the gold standard of indistinguishability through a novel simultaneous Goldreich-Levin reduction.
Current AI systems struggle to conduct independent scientific research, with performance plummeting by nearly 50% when human guidance is removed.
Task order can dramatically skew the performance of self-improving agents, revealing hidden dependencies that could undermine their reliability in real-world applications.
A novel two-threshold framework allows LLMs to judge outputs with formal control over reliability, achieving higher coverage without compromising error rates.
Generative AI is reshaping workplace interactions by creating a veil of opacity around human effort, complicating trust and collaboration.
Attackers can bypass defenses by cleverly fragmenting harmful tasks, rendering current stateful defenses ineffective in real-world scenarios.
Replacing causal mechanisms with constants in structural models reveals a surprising equivalence to graph surgery, clarifying how interventions affect dependencies.
Fair-ordering protocols can ensure that conflicting requests are invalidated, even when the policy state is hidden and unrecoverable.
Remediation order in multi-gate AI systems can fundamentally alter decision outcomes, revealing critical vulnerabilities in evidence trustworthiness.
Autonomous scientific agents can now be audited more effectively by linking claims directly to their evidence and verification, transforming how we ensure scientific integrity.
Agents can significantly reduce communication needs by leveraging memory more effectively, revealing a critical balance between remembering and signaling.
Sharpening the affinity matrix in t-SNE can significantly enhance the preservation of nearest neighbors, while smoothing broadens local neighborhood fidelity—outperforming traditional multiscale methods.
Potential theory could unlock new levels of sample efficiency in reinforcement learning algorithms.
Reallocating optimization effort based on reward saturation can boost performance by up to 9.2% in complex reasoning tasks.
Achieving a new upper bound of $\omega < 2.371177 could redefine our understanding of matrix multiplication efficiency.
Power systems could become the gold standard for testing graph machine learning, yet face a reproducibility crisis due to scarce benchmarks and datasets.
Achieving improved regret bounds in parallel GP bandit optimization without the need for an initial uncertainty sampling phase could redefine efficiency in practical applications.
Lifelong learning just got a major upgrade: SoftModel's dynamic topology adapts in real-time, breaking free from the constraints of fixed neural architectures.
Randomly subsampling edges can yield correlation clustering approximations that rival those of complete graphs, challenging existing lower bounds for general graphs.
Enforcing latent independence in representation learning can compromise the alignment with semantically meaningful covariates, revealing a critical trade-off in model design.
Identical delay summaries can yield vastly different regret outcomes, highlighting the crucial impact of timing in bandit optimization.
Oversmoothing in graph neural networks can be effectively tackled by a new index-theoretic criterion that reveals the true discriminative power of sheaf configurations.
Competing dealers can increase market instability by over 3 times, revealing the critical need for a pre-deployment stability margin in machine learning models for trading.
Collective dynamics of AI agents reveal surprising patterns: while communication boosts accuracy on objective tasks, it can lead to political bias in group opinions.
Reducing clinical label disagreement by 36% reveals how LLMs can transform rare-disease modeling by embedding expert knowledge directly into representation learning.
A prompt's value can exponentially reduce the cost of generating complex artifacts, revealing a new dimension of efficiency in LLM interactions.
Uncertainty in AI decision-making can now be visualized and translated into actionable oversight responses, making it explicit rather than implicit.
Trust-preserving agentic AI can achieve an impressive 86.9% task completion rate while intervening in nearly all policy violations.
Gamifying debate practice with AI opponents not only enhances engagement but also provides structured, multimodal feedback that traditional text-based systems lack.
HyperSkill's hypergraph memory structure enables LLM agents to leverage relational knowledge, resulting in up to 11.51% performance gains over traditional memory systems.
Subtle prompt cues can exert powerful control over AI models, revealing a vulnerability that challenges our understanding of AI behavior.
Expert-guided policy revisions in language models can drastically improve diagnostic accuracy, achieving up to a 32.7 percentage point increase in Recall@1 for rare diseases.
Black-box RL can boost agent performance by nearly 15 points on complex tasks, revealing a new frontier for scalable optimization.
Timing the entry of preference dimensions can lead to substantial performance gains in multi-preference alignment for LLMs.
Mint-Agent models not only outperform existing benchmarks in financial reasoning but also integrate long-horizon execution with auditable evidence trails, setting a new standard for financial intelligence.
A feedback loop of user evaluation and AI action can transform how personal AIs adapt and improve their observational capabilities.
Users can identify AI responses as "Claudish" even without knowing the underlying model, revealing a deeper layer of AI character recognition that transcends mere identification.
Achieving over 27,000 requests per second, this system redefines the durability-latency trade-off in AI audit records, but raises questions about security and compliance.
Covert channels in LLM traffic can leak private information without any malicious intent, revealing a surprising vulnerability in seemingly benign interactions.
Skill, not lineage, is the key determinant of success in trusted-monitor ensembles, with implications for how we build and evaluate these systems.
As AI automates coding, the real challenge shifts to ensuring that human specifications are accurate and verifiable, revealing a critical paradox in software development.
Agile teams can now integrate EU AI Act compliance into their workflows without sacrificing agility, thanks to a practical, evaluated guideline.
Uncovering a coherent long-tail structure in urban navigation reveals critical safety scenarios that traditional data approaches often miss.