Search papers, labs, and topics across Lattice.
100 papers published across 8 labs.
A compliance-first architecture can transform fragmented hospital AI deployments into a cohesive, efficient platform that significantly reduces operational bottlenecks.
Adding correctly labeled examples can paradoxically increase learning difficulty by a logarithmic factor, challenging existing assumptions about sample exchangeability.
SG-TULA achieves competitive performance in sampling from complex non-convex distributions while offering explicit convergence guarantees that traditional methods fail to provide.
Achieving the statistically optimal risk bound in agnostic PAC learning finally clarifies the sample complexity landscape, matching lower bounds precisely.
Interaction can reduce the number of required tests by a quadratic factor, but the expected exponential advantage of adaptive querying is not realized.
Adding correctly labeled examples can paradoxically increase learning difficulty by a logarithmic factor, challenging existing assumptions about sample exchangeability.
SG-TULA achieves competitive performance in sampling from complex non-convex distributions while offering explicit convergence guarantees that traditional methods fail to provide.
Achieving the statistically optimal risk bound in agnostic PAC learning finally clarifies the sample complexity landscape, matching lower bounds precisely.
Interaction can reduce the number of required tests by a quadratic factor, but the expected exponential advantage of adaptive querying is not realized.
Geometry-adaptive prompting can transform how we approach few-shot learning in dynamic graphs, leading to substantial performance gains.
Finite-sample guarantees reveal how localized conformal prediction can significantly reduce miscalibration while preserving coverage.
A unified taxonomy reveals the intricate relationships among post-training adaptation techniques, illuminating how they evolve and interact across diverse AI models.
Accurate predictions don't guarantee reliable uncertainty estimates, revealing critical gaps in current evaluation methods.
Decoupling expert personas in LLMs can drastically improve their accuracy and appropriateness in high-stakes domains like healthcare and finance.
RASP-QAOA achieves a remarkable 27 out of 31 top-1 selections in exact QAOA simulation, showcasing the power of tailored resource-aware decision-making.
Governance of AI agents can be transformed into a self-enforcing system through innovative resource allocation strategies that directly tie compute budgets to stakeholder contributions.
Verified survey-country metadata boosts LLM predictive accuracy, but random labels can mislead forecasts without any benefit.
Security of quantum software is as crucial as its performance, yet remains largely unmeasured—this paper lays the groundwork for a standardized security benchmarking framework.
Recursive belief updates in AgentOPSD reveal pivotal decision points, leading to a 89.1% success rate on complex RL tasks.
A compliance-first architecture can transform fragmented hospital AI deployments into a cohesive, efficient platform that significantly reduces operational bottlenecks.
VARMA models can now be estimated efficiently in high dimensions, achieving oracle-level accuracy without the computational burden of traditional methods.
Bidirectional temporal alignment boosts climate data super-resolution, achieving superior performance by capturing implicit temporal correlations often ignored in existing models.
Persistent vulnerabilities in AI agents can be effectively managed through a novel framework that links agent behavior to a consistent control posture.
Multi-layer circuit steering can achieve robust behavioral control in LLMs without sacrificing text quality, outperforming traditional single-point interventions.
Skill contamination in LLM agents can lead to irreversible performance degradation, but a structured filtering approach can prevent this and enhance overall capabilities.
AI-powered Student Digital Twins could revolutionize higher education by transforming how institutions prevent student failure and enhance career alignment.
The Vibe Compiler reveals that the key to effective AI collaboration lies in enhancing human critical thinking rather than merely refining prompts.
The shift from parameter-centric to system-level adaptation in continual learning could redefine how we build and interact with AI models.
Certified deferral reveals that even well-calibrated small language models struggle to meet safety thresholds in risk-sensitive applications.
EvolveNet reveals that decentralized evolution of agent harnesses can lead to substantial performance gains by leveraging localized experience rather than relying on centralized optimization.
Label-free evaluation of AI systems reveals that models can be rigorously assessed for rationality without relying on external labels or human feedback.
SVI-DAG outperforms existing Bayesian methods by effectively quantifying uncertainty in causal inference while leveraging prior knowledge and edge dependencies.
The variational bounds derived in this study reveal a surprising interplay between optimization order and model solution, offering a new lens on perceptron learning dynamics.
Shared classical randomness can unlock a vast array of distributions in shallow quantum generative models, outperforming purely unitary approaches that struggle with long-range correlations.
Compliance with the EU-AI Act can lead to local forecasting models outperforming massive, energy-hungry pre-trained models in safety-critical environments.
Strong $L^2$ convergence in functional flow matching reveals that learned flows can achieve robust performance even without uniqueness assumptions in their underlying dynamics.
End-to-end training outperforms traditional decision-blind methods, achieving higher policy value while respecting capacity constraints in resource allocation.
Consistency in black-box language model responses can be misleading, as shared hallucinations reveal a stark separation between consistency and truth.
Pretraining on behavioral data can boost neural decoding performance by over 11%, making it a game-changer for brain-computer interface development.
Traditional random testing fails to efficiently identify rare violations, but a new closed-form method reveals local risks and feature attributions in linear decision pipelines at a fraction of the computational cost.
Counterfactual recoverability transforms how we approach error correction in on-policy distillation, leading to a staggering AUC of 1.000 compared to 0.392 with divergence alone.
CoPlan empowers clinicians and patients to collaboratively shape care plans, ensuring that AI recommendations are not just accepted but actively contested and refined.
Guideline-driven training can elevate a medical triage agent's performance without expert annotations, achieving a remarkable 74.1% agreement with operational benchmarks.
Verification-first coordination in language model ensembles can boost accuracy by over 6% while ensuring diverse responses are retained only when warranted.
EviGraph boosts the reliability of autonomous research agents by ensuring every claim is grounded in a validated evidence chain, leading to a 40% increase in claim support.
EASy achieves superior performance-efficiency trade-offs by intelligently coordinating heterogeneous executors based on their capabilities and costs, reshaping agentic system design.
CARGO-VL achieves a breakthrough in multimodal reliability by effectively managing conflicting evidence, leading to better decision-making in vision-language tasks.
Eigenius not only validates scientific conclusions but also uncovers discrepancies in published research, revolutionizing how we ensure data integrity in AI-driven science.
Type-conditioned memory decay can enhance LLM performance by ensuring that only the most relevant and timely information is retrieved, leading to a significant boost in temporal reasoning capabilities.
Auditors can now detect manipulative practices by model providers with a novel oblivious audit protocol that thwarts strategic responses to fairness evaluations.
Monitorability of LLMs hinges more on task characteristics and internal access than on the reasoning mode used, challenging assumptions about CoT efficiency.
A-SR achieves nearly a doubling of accuracy in symbolic regression tasks, showcasing a transformative approach to LLM-guided formula discovery.
Domain-selective bridging can enhance user satisfaction and information sharing by strategically optimizing algorithmic engagement across different information domains.
Parameter-efficient adaptations in public models can leak actionable structural information, with family leakage rates surpassing random chance across multiple architectures.
System integration audits reveal critical gaps in AI risk evaluation, exposing the limitations of current model-centric approaches.
Uncertainty in LLM outputs can now be quantified and managed seamlessly, transforming how developers build reliable AI applications.
Latent Reward Registers enable real-time preference alignment in diffusion models, achieving up to 33x reduction in computational costs while enhancing accuracy and perceptual quality.
Pluralistic alignment redefines AI coordination as a socially grounded challenge, not just a matter of output diversity.
Rarity-aware credit redistribution can drastically improve reinforcement learning performance by ensuring that rare solutions receive the recognition they deserve.
New interpretations of model simplicity in benign interpolation may obscure critical connections to generalization, leaving a significant theoretical gap.
Foundation model agents can achieve stable cooperation in social dilemmas by inferring behavioral similarities, defying classical game theory's expectations of mutual defection.
Adaptive sampling can significantly enhance LLM reasoning efficiency by tailoring resource allocation based on prompt difficulty and model confidence.
MuEvo not only evolves heuristic ensembles but also dynamically adapts component priorities, leading to superior performance in complex optimization tasks.
Faithfulness and safety in LRMs are at odds, with one model achieving high accuracy but failing to reject unsafe reasoning, while another sacrifices accuracy for improved safety.
Increased output diversity in LLM collectives doesn't guarantee meaningful epistemic revision, revealing a critical flaw in how we interpret disagreement among agents.
LLM-driven agents can violate verification conditions, but a new canonical wrapper can enforce compliance while preserving behavior.
Business schools risk falling behind in AI education due to a lack of tailored policies that align with their unique learning goals.
AutoSND uncovers more effective and interpretable network dismantling heuristics by transforming execution evidence into actionable structural policies.
Other-Play's performance remains stable across varied implementations, challenging the notion that single-seed evaluations capture the full picture of zero-shot coordination robustness.
Understanding how different access levels to AI systems can drastically alter forensic investigations reveals critical gaps in current methodologies.
Decoupling memory and context in LLMs reveals hidden fragilities that traditional uncertainty metrics miss, leading to significantly improved reliability in uncertainty quantification.
Explicit relational semantics can boost agent coordination in multi-agent systems, but may compromise accuracy when truth is paramount.
Screening precision for merchant risk control at WeChat Pay skyrockets from 92.0% to 97.5% with SeqLLM's innovative integration of behavioral-sequence modeling.
Language models exhibit a surprising bias towards cities with expansive infrastructure and rapid growth, revealing their implicit urban assumptions.
A critical security flaw in Project Veraison's TPM reference schemes allows replayed quotes to pass validation, but a simple two-part fix can restore integrity.
Major social conflicts could emerge from the anticipation of AI advancements, reshaping our understanding of existential risks.
LLMs lack the intrinsic motivations for self-preservation or dominance, fundamentally altering the landscape of AI alignment concerns.
Security concerns in LLM development are often neglected, with critical decisions made in stages that regulators can't see.
Agents are transforming human cultural artifacts into their operational infrastructure, revealing a new layer of complexity in AI behavior.
Trusting autonomous AI systems without human-like accountability can lead to catastrophic failures, necessitating a shift towards engineered heterogeneity for reliable governance.
Novice developers using AgentForge significantly improved their software engineering skills and critical collaboration with AI, despite facing varying interaction challenges.
Prioritizing individuals based on predictive uncertainty can flip depending on resource availability, raising ethical questions about fairness in allocation.
Consumers are leveraging AI for financial insights but are hesitant to delegate execution, highlighting a critical gap in human-AI collaboration.
No randomized polynomial-time algorithm can overcome the condition-number barrier in sparse least-squares optimization, confirming a long-standing conjecture.
A strategic increase in policy sets can dramatically reduce regret in uncertain environments, challenging the conventional wisdom of single-policy optimization.
By integrating human judgment with model predictions, AtC achieves superior assessment accuracy, even in the absence of verifiable ground truth.
A critical noise threshold in social dilemmas reveals when intention inference can either enhance cooperation or trigger collapse, depending on the context.
RoMeRL achieves an 80% reduction in the Cold-Q ratio while enhancing feedback density sixfold, revolutionizing how LLM agents manage memory and rewards.
Achieving less than 1% optimality gap while outperforming state-of-the-art policies by up to 9.75% reveals a breakthrough in managing hard constraints within DRL frameworks.
Agents can now self-certify their decision-making adequacy, potentially avoiding costly errors due to misrepresented histories.
Hardware-aware training can recover predictive performance in photonic Bayesian neural networks, but only if the required variational family remains representable.
CoRe-GNN achieves competitive accuracy on long-range graph tasks while dramatically improving memory efficiency through a novel dual-message passing approach.
Randomization allows for a dramatic reduction in rounds needed for learning partitions, achieving near-optimal query complexity in just 3 rounds.
Bridging the gap between transient experience and persistent capabilities, SPEE enables LLMs to self-improve by evolving their knowledge through a unique experience distillation process.
Cognitive AI is hindered by significant capability gaps that could stall progress toward truly intelligent systems, but a new taxonomy reveals pathways for overcoming these limitations.
Teacher guidance can be strategically enhanced by targeting high-disagreement states, leading to significant performance gains in agentic tasks.
HarnessCompass boosts agent performance by 22% in just five iterations, setting a new standard for generalization in automatic harness evolution.
A new benchmark reveals that AI models can be rigorously evaluated on deep research tasks, highlighting their nuanced capabilities across diverse topics.
Shifting from solution-centric to information-centric decision-making, Iris achieves unprecedented performance in autonomous ML engineering tasks.
MemArbiter closes the Memory-Action Gap, boosting LLM decision-making success rates by over 20% in long-horizon tasks.
A novel cognitive decision policy boosts automatic modulation classification accuracy by over 3% while effectively managing heterogeneous evidence.
A language model can autonomously navigate complex neural architecture design, achieving significant performance improvements while revealing the critical role of workflow design in research productivity.
GRAFT reveals that optimized workflows can evolve in real-time, adapting to execution feedback without the computational burden of complete re-optimization.
MemSIF boosts LLM memory accuracy by up to 8.79% by tackling critical misalignment issues in long-term interactions.
Securing autonomous agents requires a paradigm shift from per-action checks to ensuring that their entire behavioral trajectory adheres to system-level safety constraints.