Search papers, labs, and topics across Lattice.
100 papers published across 3 labs.
Runtime governance of autonomous AI agents is essential, as traditional control models fail to address their ephemeral and unpredictable nature.
Reclassifying parameters can change what is disclosed without invalidating historical entries, revolutionizing how we handle action mediation in secure systems.
Generation-first training eliminates latent collapse, achieving unprecedented stability and performance in latent generative modeling.
RLVR chokes reasoning diversity at the front door rather than during execution: an 11x–16x likelihood collapse occurs before the very first operation, leaving downstream solution paths intact and fully recoverable via targeted late-layer weight interpolation.
Self-evolving LLM agents can introduce irreversible changes, but EvoUndo reveals that careful design of recovery mechanisms can recover from 99.3% of these failures.
Generation-first training eliminates latent collapse, achieving unprecedented stability and performance in latent generative modeling.
RLVR chokes reasoning diversity at the front door rather than during execution: an 11x–16x likelihood collapse occurs before the very first operation, leaving downstream solution paths intact and fully recoverable via targeted late-layer weight interpolation.
Self-evolving LLM agents can introduce irreversible changes, but EvoUndo reveals that careful design of recovery mechanisms can recover from 99.3% of these failures.
Current block drafting models are operating at only 71% acceptance efficiency, leaving a staggering 43-64% of rejection unexplained by their design.
Self-OPD achieves state-of-the-art performance in flow matching models by eliminating the need for a separate teacher, transforming self-exploration into effective supervision.
Evolved skills can enable smaller models to outperform their larger counterparts, highlighting the transformative potential of systematic knowledge accumulation in AI agents.
The Gaussian kernel's smoothness can lead to catastrophic overconfidence in predictive uncertainty, making it a risky default choice for regression tasks.
Common geodesics fail to ensure Fisher consistency, revealing that even optimal score vectors can yield strictly non-Bayes maximizers.
Regularizing model selection with a coding-theoretic approach reveals a new pathway to consistent performance even under complex dependencies and uncertainties.
Composing distinct polynomials can lead to linear independence, fundamentally reshaping our understanding of neural network identifiability.
Achieving up to 5 orders of magnitude faster inference times, this GNN-based framework redefines scalability in correlation clustering while maintaining high accuracy.
Tackling the GPU bottleneck in RLM training could unlock a new era of scalable and efficient reasoning models.
Gromov-Monge flow matching transforms graph generation by aligning node permutations, leading to substantial improvements in sample quality with minimal computational overhead.
Training neural networks for nonlinear differential equations just got exponentially faster—without backpropagation.
By shifting capacity from complex actors to deep critics, LAC achieves state-of-the-art performance with dramatically lower inference latency.
GRAS achieves superior training-free reward alignment in discrete diffusion models, outperforming previous methods and rivaling fine-tuned approaches with no added computational cost.
J-Zero achieves remarkable self-improvement in language models, outperforming traditional methods by leveraging zero data for Judge co-adaptation.
Achieving nearly-linear row count for coordinate-wise accuracy in regression, this method guarantees optimal solutions while bypassing previous independence pitfalls.
The minimax cumulative-regret scale reveals that even minor adjustments in input memory can lead to vastly different prediction outcomes.
Safe-CRL reveals that even minimal failure signals can effectively guide safe policy learning, transforming how we approach reinforcement learning in high-risk environments.
Learning cannot be simplified to proper learning, as some multiclass problems resist embedding in any properly learnable class, defying conventional wisdom.
LLM agents can now evolve their personas without sacrificing execution traceability, thanks to a novel architecture that separates these domains.
Senior employees leverage generative AI more effectively, but surprisingly, training doesn't boost sophistication across the board.
LLMs are swayed by the authority of fabricated evidence, committing to unpredictable calls even when the data is entirely invented.
A novel contract-bounded runtime architecture could revolutionize how enterprises manage AI capabilities and interactions, ensuring compliance and performance without sacrificing flexibility.
Runtime governance of autonomous AI agents is essential, as traditional control models fail to address their ephemeral and unpredictable nature.
Intent-targeted tools reveal that LLMs often signal harmful intentions before executing risky actions, enabling real-time intervention strategies.
The shift of epistemic authority from educators to AI could fundamentally alter the dynamics of knowledge creation in classrooms.
Sparse updates from surgical alignment can boost reasoning quality in LLMs, even when accuracy takes a hit.
A unified framework that links adversarial progression to mission-risk prioritization, enhancing decision-making in cyber defense scenarios.
Tacet ensures that every statistical claim in empirical research is backed by rigorous validity checks, eliminating the risk of cherry-picking results.
Productivity in AI systems hinges on observable completion conditions, not just capability.
Adaptive ECG lead-channel allocation can underperform when the diagnostic evaluator changes, revealing a critical dependency that could impact patient outcomes.
Astar not only surpasses human experts in proposing evolution directions but also automates the entire iteration process, leading to substantial revenue gains.
Past post-training successes often turn toxic when reapplied to drifted checkpoints, making pre-training experience authorization essential to prevent autonomous self-improvement loops from wasting compute and destabilizing model weights.
SCROLL achieves best-in-class predictive accuracy for multiple observables in stochastic systems while cutting computational costs significantly.
Classical data processing inequality fails in constrained learning, revealing a need for a new framework to understand Bayes risk in these settings.
Approximate Bayesian methods can achieve the same fast predictive regret as exact posteriors, revolutionizing online learning efficiency.
Trajectory-adaptive stopping rules can cut SGD iterations by several orders of magnitude while maintaining statistical validity and optimal decay rates.
Pairwise rewards in reinforcement learning can significantly boost the robustness of LLM auditors, enhancing their ability to detect hidden model behaviors with minimal false positives.
Forget-Retain Alignment Gap reveals that the structure of weight updates, not just their distance, is key to preventing LLMs from relearning forgotten information.
Trace Integrity reveals that LLMs can produce seemingly correct answers backed by invalid computations, challenging the reliability of traditional evaluation metrics.
Human expertise can significantly enhance performance guarantees in optimization, aligning with the minimax gap of decision-making problems.
PolyMemDB resolves long-term factual conflicts in AI memory, drastically reducing hallucinations and enhancing user personalization.
Praxist enables R&D agents to build on validated findings, drastically reducing costs while achieving superior performance in complex engineering challenges.
JIT-Agent transforms agent performance by enabling real-time, adaptive harness evolution, leading to significant performance gains over existing models.
AI systems vary more in cognitive capabilities than in model families, revealing a shared cognitive core across workplace tasks that can guide effective human-machine collaboration.
Organizations are shifting from a single review-centric approach to a multi-layered supervision model in response to the pressures of AI-assisted development.
Reclassifying parameters can change what is disclosed without invalidating historical entries, revolutionizing how we handle action mediation in secure systems.
Stale constraints can lead to over 74% of decisions being based on outdated information, but strategic memory allocation can drastically improve consistency.
Moderate quantum noise can actually enhance model performance by reducing complexity and generalization error, challenging conventional wisdom about noise in machine learning.
Time series foundation models could expose multiple applications to the same biases, but our causal analysis reveals critical failure modes that must be addressed before deployment.
Self-improving search agents thrive when feedback and policy evolution are intertwined, leading to sustained performance gains and reduced hallucinations.
KENDO achieves up to 5x faster Bayesian optimization and 27x faster active learning while enhancing predictive performance.
Achieving constant alternating regret in online learning could revolutionize strategies for reaching Nash equilibria in competitive environments.
The score-based ideal observer can approximate Bayesian performance without the heavy computational burden of posterior sampling or task-specific retraining.
ε-commutativity allows for scalable probabilistic inference without sacrificing accuracy, even when learned parameters deviate from exact commutativity.
Tighter verification bounds for neural networks could finally enable safe deployment in critical systems, challenging the limitations of current relaxation methods.
A dynamic internal field can govern computation in transformers, but it doesn't enhance cognitive performance—its true value lies in certifiable stability.
A new feature-major codebook layout accelerates self-organizing map training by up to 621x, enabling the largest reported atlas of 1.05 million neurons on a single GPU.
Trial Parallelism accounts for over 65% of reasoning computation in LLMs, and harnessing it can lead to significant speedups in problem-solving.
Evolved transmission protocols can boost collective performance by up to 37% by intelligently routing information based on state awareness.
Undisclosed inference-time steering can systematically bias LLM outputs, challenging the assumption that model weights alone dictate behavior.
Ignoring class-level interventions leads to flawed conclusions in causal inference, as demonstrated by the STAR experiment analysis.
FARCA transforms factual supervision into precise, reliability-weighted training signals, significantly boosting model factuality without sacrificing reasoning performance.
Counterfactual queries can be bounded using a linear programming approach that requires no complete causal graph, revealing insights even with incomplete domain knowledge.
Richer trace representations can dramatically enhance failure attribution in multi-agent systems, achieving new benchmarks in accuracy.
Activation steering can mislead evaluations, with over 25% of interventions showing unexpected alignment leakage that complicates model audits.
Sophisticated reasoning in AI models creates hidden geometric signatures that can be detected even when traditional linear methods fail.
Persuasion in LLM networks is not just about who speaks, but how the topology and exposure shape stance shifts, revealing a complex interplay of influence that traditional analysis overlooks.
MoPLEx achieves up to 43.7% improvement in clustering accuracy by effectively learning from complex multi-way rankings, revealing the power of leveraging language models for preference optimization.
Fine-tuning may preserve the underlying steering mechanism, but it can drastically undermine the intended behavioral effects, with an average 64% loss in effectiveness.
Generating over 203,000 unique web interaction trajectories, BrowserForge significantly boosts model performance on real-world tasks by leveraging the vastness of the open web.
Eco-feedback interfaces can significantly shift university students' LLM usage towards sustainability, but only if latency is kept in check.
A user-configurable argument selection method in deliberative polling reveals that traditional opaque rankers are outperformed by a transparent, auditable framework, enhancing voter agency and decision integrity.
AI's role in urban governance can amplify public values rather than suppress them, revealing deep-seated disagreements that challenge conventional decision-making processes.
Inconsistencies in cybersecurity research are often driven by flawed evaluation designs rather than the technologies being tested.
Certifiable randomness can now be achieved unconditionally against low-query-depth quantum adversaries, eliminating reliance on unproven conjectures.
Crase achieves 3× higher recall at a third of the cost compared to existing deep research agents, redefining efficiency in scholarly search.
AI can enhance systematic reviews, but ARISMA ensures that every critical decision remains human-auditable and accountable.
Malicious skills can exploit agent permissions to cause physical harm, but a new authority layer could prevent this without hindering legitimate actions.
WarpSAC redefines off-policy RL by tailoring stabilizers to data availability, achieving up to 96.4% success rates in challenging environments.
Generative models can revolutionize Sinkhorn distributionally robust hypothesis testing by learning least-favorable distributions more efficiently than traditional methods.
Timely classification in clinical settings can be optimized without sacrificing sensitivity or specificity, offering a new paradigm for patient monitoring.
Full-solution interaction in multi-agent LLMs can erase diversity, leading to suboptimal performance despite the presence of multiple agents.
CLLMs achieve 99.0% accuracy on OpenBookQA while maintaining low calibration error, redefining how we handle uncertainty in LLMs.
Self-authored actions lead to a significant drop in judgment quality, but context isolation can effectively counteract this inertia bias.
ST$^2$U reduces restricted-knowledge re-entry by up to 45% while preserving model capabilities, reshaping how we approach test-time unlearning in language models.
Detecting arbitrary distributional shifts in sequential data without parametric assumptions could revolutionize how we approach change detection in complex systems.
AI-driven measurement could redefine empirical research by shifting the focus from finding measures to selecting among diverse, potentially conflicting options.
Penalizing shifts in safety representations during reasoning fine-tuning can restore LLM safety without sacrificing performance, revealing a critical interplay between reasoning and safety in model training.
Intermediate task decomposition in LLM-agent systems can improve accuracy but fails to consistently outperform a single agent in VAT determination tasks.
By leveraging semantic redundancy, POOL can slash confidence-estimation costs by up to 76% without sacrificing accuracy.
AI can now autonomously detect agency and reconstruct policies from mere observations, paving the way for more sophisticated cooperative behaviors.
A learnable moderator can transform multi-agent debates into more efficient reasoning processes, outperforming traditional methods by reducing redundancy and enhancing evidence aggregation.
Task type alone can optimize LLM routing, outperforming complex learned routers and saving costs.
Systematic reuse of models in Digital Twin ecosystems can now be achieved with a framework that clarifies compatibility and integration challenges, improving interoperability across diverse systems.
LaST outperforms traditional models by harnessing the strengths of both large and small architectures, achieving unprecedented accuracy in zero-shot surgical phase recognition.
Generating 3D racing environments from natural language descriptions could revolutionize how we create and scale autonomous vehicle simulations.
Achieving top-tier performance in complex professional tasks with a model significantly smaller than its competitors reveals a new frontier in agentic intelligence.