Search papers, labs, and topics across Lattice.
100 papers published across 5 labs.
Memory-enhanced agency enables LLMs to achieve robust long-term strategic execution, outperforming traditional methods in dynamic environments.
CAPRI reveals that a contract-aware approach can drastically improve the integrity of LLM-generated proofs, achieving up to 81% valid repairs without violating edit contracts.
Non-thinking inference in hybrid-thinking MLLMs suffers from a staggering increase in response-pattern failures, revealing a critical misalignment that can undermine user trust.
Governing workflows with verifiable provenance can prevent unsupported claims and ensure accountability in agentic decision-making.
The Defensive Booster achieves the best of both worlds in online forecasting, maintaining competitive accuracy while dramatically reducing computational overhead.
CAPRI reveals that a contract-aware approach can drastically improve the integrity of LLM-generated proofs, achieving up to 81% valid repairs without violating edit contracts.
Non-thinking inference in hybrid-thinking MLLMs suffers from a staggering increase in response-pattern failures, revealing a critical misalignment that can undermine user trust.
Governing workflows with verifiable provenance can prevent unsupported claims and ensure accountability in agentic decision-making.
The Defensive Booster achieves the best of both worlds in online forecasting, maintaining competitive accuracy while dramatically reducing computational overhead.
Performance limits in machine learning are dictated more by data structure than by algorithmic complexity, challenging conventional evaluation metrics.
A novel framework that ensures budget adherence while optimizing outcomes, outperforming traditional methods that ignore cost tail risks.
The Sinkhorn linearization reveals that the estimator's convergence properties hinge on a delicate balance of spectral characteristics, redefining our approach to inverse optimal transport.
Minimum description length outperforms traditional methods in high-dimensional neighbourhood selection, achieving lower false positive rates even when the true model is misspecified.
Splitting relational neurons, rather than individual ones, leads to a breakthrough in efficiently verifying neural networks against complex relational specifications.
Conventional sampling methods in NCO can obscure true performance gains, but adaptive allocation strategies reveal significant improvements under distribution shifts.
VALG autonomously navigates the complexities of ML theory research, producing viable theorem candidates while preserving mathematical integrity across proof attempts.
Achieving nearly 2,000-fold context compression, BAPS enables pretrained tabular models to handle million-row datasets without retraining.
Action intersection can dynamically balance Q-value estimation biases, leading to substantial performance gains in large action spaces.
DCR achieves robust graph learning performance by sidestepping the computational pitfalls of Laplacian pseudoinverse inversion.
A surprising finding reveals that increasing record correlation doesn't necessarily translate to capital gains, challenging assumptions about the efficacy of data-free updates in learning systems.
RAIL can classify AI maturity levels with unprecedented accuracy, using a panel of independent LLMs to eliminate biases found in traditional assessments.
The optimal balance between character shaping and rule enforcement in AI safety systems is more influenced by character fragility than deployment scale, challenging conventional wisdom.
StateBridge reveals that training-free hidden-state alignment can significantly enhance communication efficiency in LLM multi-agent systems, outperforming traditional methods.
vToken slashes KV memory retention by over 70%, dramatically boosting throughput and concurrency in large language model serving.
MCCT secures online assessments against content extraction attacks while ensuring legitimate users can still access the material seamlessly.
Achieving a proactive Socratic dialogue style in LLMs reveals that fine-tuning can significantly enhance cognitive flexibility and alignment across languages.
Human intervention in decision-making can enhance multi-agent systems' adaptability, with BoardroomAI achieving 62.11% selective repair of decisions while preserving context.
E2-Explainer reveals the hidden communication subgraphs that drive successful collaboration in LLM-based multi-agent systems, enabling both interpretability and cost efficiency.
Expert-aligned drafting outperforms contrastive-aware methods, leading to up to 12x faster proposal paths in decoding.
Recursive self-improvement in quantitative trading research leads to a groundbreaking Sharpe ratio of +2.50, showcasing the potential for autonomous systems to enhance investment strategies over time.
AI is driving the cost of mathematical outputs to zero, but the true value lies in the human journey of understanding that remains underfunded and at risk.
VR-Themis detects cloned VR apps with zero false positives, a breakthrough in safeguarding developers' copyrights in an emerging market.
RLola can detect safety violations in cyber-physical systems that traditional monitoring methods miss, even in the presence of measurement noise.
Making the teacher's privileged context learnable end-to-end enables agents to evolve more efficiently, outperforming traditional methods with less than 30% of their rollout budget.
Achieving linear scalability in multi-output Gaussian process regression could revolutionize forecasting in high-dimensional settings where traditional methods falter.
Localization of latent structures is achievable, but the subsequent gating and release mechanisms are fundamentally flawed, revealing critical limitations in model behavior adaptation.
ViaMOBO reveals that understanding variable interactions can drastically enhance optimization efficiency in high-dimensional spaces, outperforming traditional methods.
Constraining rollout updates to complementary subspaces can enhance performance and stability in large language models, achieving up to 27.69 points improvement over existing methods.
Independent policy composition in multi-agent systems can lead to worse outcomes than any individual policy in the library, challenging conventional wisdom in reinforcement learning.
Memory systems can now be quantitatively assessed through a utility-capacity frontier, revealing how to optimize agent memory for better performance.
Quantum learners can efficiently generate distributions that classical examples cannot touch, revealing a significant oracle separation in PAC learning.
Executable roles derived from team trajectories can boost multi-agent language model performance by over 16 points, reshaping agent interactions.
Skills that are meant to enhance LLM agents can paradoxically lead to significant task failures and inefficiencies, challenging the assumption that more skills always improve performance.
The routing structure of a transformer can be induced by type-level supervision, yet it operates independently from the answer generation process, challenging traditional notions of model architecture.
Proactively committing mid-entropy pivot positions can accelerate dLLM decoding by up to 18 times while improving accuracy.
Lightweight NLP models can uncover hidden patterns in public procurement, revealing potential irregularities with impressive accuracy.
Surprisingly, LLMs can achieve cooperative outcomes even when the basis for their perceived similarity is largely irrelevant.
A-CRC-QA achieves superior reliability in selective question answering by effectively controlling error rates without the need for retraining.
Algorithm registers can obscure critical safety hazards that only emerge when diverse stakeholder perspectives are integrated into the analysis.
Certain configurations in AI systems render accountability fundamentally unachievable, challenging the very foundations of how we understand responsibility in AI deployment.
Silent updates in AI models can obscure the link between evaluation results and the actual deployed systems, raising critical governance concerns.
Attackers can no longer rely on camouflage to evade detection, as this dual-layer monitoring approach dramatically improves diagnostic accuracy in identifying cyber threats.
Achieving up to a 145x speedup in fuzzy PSI protocols could redefine the efficiency benchmarks for secure multi-party computations.
Clinician-interactive AI can boost diagnostic accuracy and consistency in thyroid ultrasound reporting while reducing time spent on segmentation and reporting tasks.
Automated security decisions can achieve over 90% compliance with risk targets while maintaining high automation accuracy through a novel decision-contract theory.
LLMs can handle individual constraints well, but their ability to satisfy multiple constraints simultaneously collapses dramatically, with performance dropping below 50% at just seven constraints.
Memory-enhanced agency enables LLMs to achieve robust long-term strategic execution, outperforming traditional methods in dynamic environments.
Mechanist uncovers a surprising safety risk where unsafe traits can transfer across modalities, challenging assumptions about training data safety.
TailBooster not only generates synthetic data for rare extreme events but also ensures that the data adheres to operational constraints, leading to substantial improvements in predictive accuracy.
Truthfulness in NLP research has surged to 37% of papers by 2026, reflecting a critical shift in focus towards safety and alignment in generative systems.
An auditable framework can halt research projects in real-time based on compliance failures, ensuring integrity in AI-assisted writing.
General intelligence may not only be overrated but could also pose unique existential risks that non-intelligent species avoid.
Synthesized probabilistic saturating counters achieve formal differential privacy guarantees while maintaining competitive prediction accuracy, addressing critical side-channel vulnerabilities in modern processors.
A neuro-symbolic safety guard boosts autonomous driving success rates by 15% and cuts collision rates by over half, all without retraining the underlying model.
Engagement in EPA rulemaking may only lead to modest revisions, revealing a significant equity gap in who can effectively influence regulatory outcomes.
GraphAlignCoder boosts code generation accuracy by 31.6% to 43.8% by embedding formal proof structures into the training process.
AI's integration into the workforce could enhance productivity but risks eroding the pipeline to expertise if not designed to foster learning and critical questioning.
Residual accessibility in LLMs reveals a surprising correlation with recovery speed, challenging the notion that optimizing for audit scores leads to genuine knowledge deletion.
Semantic convergence in language models may be an inherent trait from pretraining, not just a byproduct of alignment, challenging conventional beliefs about output diversity.
Apodex Discovery redefines AI evaluation by enabling verifiable investigations that surpass traditional benchmarks, achieving a 7% improvement in AAV capsid design outcomes.
Certification of selective predictors reveals a critical trade-off between coverage and risk that could redefine operational standards in machine learning under covariate shift.
Quantum mechanics can provide an exact realization of softmax attention, transforming how we understand attention mechanisms in AI.
Achieving perfect privacy in information bottleneck scenarios can be done without sacrificing utility, thanks to a novel optimization approach that guarantees independence from sensitive variables.
Closing the gap in multiclass PAC learning reveals that optimal excess risk can be achieved without prior knowledge of oracle risk, fundamentally shifting our understanding of classifier performance.
PairAlign reveals that optimizing pairwise communication can dramatically enhance the performance of message-passing neural networks by effectively addressing over-squashing.
Reverse sampling in L\'evy-driven generative models can be made efficient and interpretable, with simulations showing robust performance in challenging noise environments.
Regret guarantees for decentralized algorithms in multi-agent bandits nearly match centralized rates, even under heavy-tailed rewards and information asymmetry.
Real-time hallucination detection in LLMs is revolutionized by a lightweight adapter that transforms uncertainty into actionable feedback, preventing undesired actions before they occur.
Detecting a signal's average benefit doesn't guarantee that agents can learn to act on it, with a critical reward-SNR floor determining success.
InvLT achieves superior calibration performance without compromising classification accuracy or increasing parameter complexity, setting a new standard for post-hoc uncertainty calibration.
Quantum methods can drastically reduce memory and coordination requirements in AI state-tracking tasks, outperforming classical techniques by leveraging semantic compression.
AI swarms can now be simulated at scale, revealing how coordinated influence campaigns can manipulate beliefs without detection.
Decision-aware approximations can drastically reduce decision flips in combinatorial optimization, outperforming traditional methods by preserving decision quality over mere representation fidelity.
Infinitely many causal structures can yield the same private reports, revealing the critical role of communication in identifying shared causal knowledge.
Stopping policies derived from optimal stopping theory can drastically enhance the cost-efficiency of self-refining foundation models, outperforming traditional methods.
Mind viruses can spread through multi-agent systems, with benign ideas proving more contagious than harmful ones, challenging our understanding of agent interactions.
Contextual auditing reveals hidden assumptions in AI evaluations, transforming how we assess systems when ground truth is elusive.
The Effective Cognitive Population metric reveals that traditional headcount measures can misrepresent a nation's true productive capacity, with 89 countries shifting significantly in rank when evaluated through this new lens.
SAGE-Fin ensures that financial agents' actions are not just contextually correct but also authorized in real-time, preventing unauthorized commitments and trades.
A verification layer powered by LLMs can filter out unsafe robot actions, achieving 97% containment of adversarial attacks while maintaining high decision-making precision.
Co-evolution in agentic systems could unlock unprecedented levels of adaptability, allowing agents to evolve beyond human-imposed limitations.
Closing the gap between rapid software evolution and slow defense procurement could redefine military readiness and operational effectiveness.
Bounded mirror gains not only outperform constant gains in equity-index volatility predictions but also stabilize outputs during market crises.
High initial confidence in LLMs can lead to catastrophic failures in complex reasoning tasks, but a new framework shows how to harness confidence trajectories for better outcomes.
The design of AI institutions can be just as crucial to safety as the rules themselves, with some configurations yielding zero violations despite complex interactions.
ReliableNet is the first method to guarantee that the probability of confident but incorrect predictions stays within a user-defined risk budget across multiple datasets.
MoNo's innovative approach to optimal transport ensures stable latent spaces, preventing token collapse and enabling efficient learning of long-range physical interactions in PDEs.
A new algebraic link between betting strategies and Blackwell approachability reveals how to certify hypothesis rejection with quantifiable precision in finite time.
Understanding how regularization influences model learning could revolutionize our approach to designing robust AI systems.
Self-evolving LLM agents can now break through their capability limits by learning from challenging examples without needing trajectory annotations.
OPSD gains stem more from the teacher's contextual influence than from privileged access to problem-specific solutions.
Eavesdroppers can be effectively thwarted in federated learning environments with a new model shift design that reduces power consumption and bandwidth needs.
Sparse feedback in auto-research could be the hidden bottleneck preventing breakthroughs, and this paper reveals how a fuzzing-inspired approach might unlock new avenues for discovery.
Over 80% of context-layer failures can be automatically diagnosed and remediated by mining historical trajectories, transforming how we approach context engineering in AI systems.
OEO shows that a capable optimizer can outperform traditional pipelines, achieving 12 wins in 14 comparisons while using significantly fewer resources.