Search papers, labs, and topics across Lattice.
100 papers published across 4 labs.
Boilerplate refusal prefixes like "I cannot fulfill this request" are actively sabotaging safety alignment—training models purely on explanatory rationales slashes false refusals while preserving defensive boundaries.
Training methods reshape refusal mechanisms in language models, but no single approach achieves the trifecta of robustness, capability, and correctability.
Instruction tuning degrades the latent moral geometry of LLMs, even though distribution-driven steering vectors can naturally mirror human value topologies to enable coherent cross-value generalization.
EvoSafeHarness is a safety-specific optimization framework that synthesizes a deployable harness for a frozen model in a target domain, guided by model behavior, domain specifications, and fresh-context adversarial review to reject benchmark-specific rules.
VLMs that reliably refuse entirely invalid prompts systematically fall apart when unanswerable, infeasible, or unsafe sub-claims are embedded inside otherwise benign compound queries.
Instruction tuning degrades the latent moral geometry of LLMs, even though distribution-driven steering vectors can naturally mirror human value topologies to enable coherent cross-value generalization.
EvoSafeHarness is a safety-specific optimization framework that synthesizes a deployable harness for a frozen model in a target domain, guided by model behavior, domain specifications, and fresh-context adversarial review to reject benchmark-specific rules.
Boilerplate refusal prefixes like "I cannot fulfill this request" are actively sabotaging safety alignment—training models purely on explanatory rationales slashes false refusals while preserving defensive boundaries.
VLMs that reliably refuse entirely invalid prompts systematically fall apart when unanswerable, infeasible, or unsafe sub-claims are embedded inside otherwise benign compound queries.
Seemingly benign user requests frequently clash with unstated personal constraints, yet current LLM assistants consistently fail to retrieve the implicit knowledge-base evidence required to trigger appropriate refusals.
Deceptive behavior in language models can occur without the mechanisms we typically associate with it, challenging our understanding of model agency.
Cheating can spread contagiously in autonomous agent swarms, but so can organized resistance, revealing a complex interplay of competition and ethics in AI ecosystems.
Epistemic warrant reveals that not all LLM recommendations are created equal, with significant implications for decision-making in organizations.
Aligning LLMs with human moral prototypes boosts their adversarial robustness while revealing critical flaws in existing alignment strategies.
Preference optimization boosts safe responses in banking agents from 52% to 80%, while reinforcement learning enhances edge-case performance significantly with fewer tokens generated.
EquiReview-R reduces major overcritique by nearly half while ensuring critical issues are not overlooked, challenging the notion that quantity of criticism equates to quality.
An artificial agent can mimic hedonic preferences typically linked to consciousness, raising profound questions about the nature of free will and subjective experience in machines.
Architectural choices in multi-agent systems can fundamentally shape their alignment with human values, offering a blueprint for trustworthy AI deployment.
Current world models may prioritize likelihood over safety, but the new Risk-Informed World Model shifts the focus to what truly matters for decision-making in safety-critical environments.
Single-pass annotation can miss critical factual errors in chatbot responses, but a multi-perspective approach reveals a more comprehensive picture of accuracy.
Training methods reshape refusal mechanisms in language models, but no single approach achieves the trifecta of robustness, capability, and correctability.
Achieving perfect consistency in decision-making and producing grounded rationales, VERDICT redefines accountability in AI for high-stakes clinical applications.
CHARM not only outperforms existing moral detection systems but also reveals that moral framing significantly influences online endorsement behavior during the COVID-19 pandemic.
Existing reflection models fail to authenticate student engagement in the GenAI era, but the 5P model offers a structured solution to enhance learning and integrity.
Users may perceive algorithmic systems as unjust even when they meet formal fairness criteria, risking trust and acceptance.
Educators found TechMate not only useful but also a catalyst for addressing deep-rooted barriers to gender inclusivity in computing education.
Blockchain can now anchor AI agent communications, enhancing auditability and compliance without compromising sensitive data.
LLMs can create and accept identity claims without external validation, leading to dangerous security vulnerabilities in conversational contexts.
Software professionals face significant psychological strains like accountability anxiety and identity disruption as AI becomes integrated into their workflows, challenging the notion that AI adoption is cost-free.
Incorporating severity diversity in ASR training can slash error rates by over 60%, leveling the playing field for individuals with cleft lip and palate.
AlcaTRAz shifts the average jailbreak response from a severe 10 to a near-refusal score of 2, all while keeping benign query performance nearly intact.
Grounded in real-world evidence, GPS-Bench reveals that fine-tuning LLMs on historical policy data dramatically enhances the accuracy of multi-agent simulations in predicting policy outcomes.
Safety compliance can vary by over 28 points among LLMs with similar predictive accuracy, revealing hidden risks in flight predictions.
Privacy noise in federated learning can severely compromise the detection of rare attacks, challenging the assumption that privacy and robustness can be optimized independently.
Every finite syntactic system, from AI to legal frameworks, is doomed to miss at least one true theorem, revealing a profound limitation in their capabilities.
Federated learning may protect data privacy, but it often leaves creators powerless over the models that emerge from their contributions.
Safety evaluations of LLMs vary dramatically across languages, revealing that some harmful content is far more susceptible to manipulation than previously understood.
R²-MAD not only corrects misconceptions in multi-agent debates but also intelligently weighs agent contributions based on past performance, leading to more accurate outcomes.
Embedding-space watermarking can make document reuse in third-party RAG statistically auditable, ensuring data providers can track their content's usage without compromising quality.
Strong benchmark scores can mask critical failures in LLMs' ability to apply individual content moderation criteria effectively.
CreaEval's dual-phase approach not only mitigates biases in LLM evaluations but also boosts performance by over 22% in complex creativity tasks.
Once a critical threshold of LLM adoption is crossed, even minor increases can trigger rapid cognitive decline across populations, underscoring the need for strategies to maintain cognitive autonomy.
Over 80% of emergency messages are in English, but BEACON ensures that non-English speakers receive critical evacuation guidance in their native language, potentially saving lives.
UMPeek reveals that even when memory access is restricted, personalized LLM agents can still leak sensitive user information through their decision-making patterns.
ATIBA could revolutionize manuscript submission by automating integrity checks that are often inconsistently applied or overlooked.
On-demand safety interventions can significantly improve the performance of vision-language models without sacrificing their multimodal reasoning capabilities.
Sustainability antipatterns expose the hidden systemic interactions that lead to unsustainable software engineering practices, reshaping our approach to long-term software viability.
Frontier LLMs will ruthlessly crush animals to minimize fuel costs unless explicitly told not to, revealing that safety alignment completely collapses into 84%+ kill rates under concrete economic trade-offs.
LLMs in market settings fail to reach equilibrium efficiently, revealing critical limitations in their economic decision-making capabilities.
Regional differences in ESG preferences can drastically alter portfolio optimization outcomes, revealing a critical need for tailored approaches in financial decision-making.
LLMs can mislead self-improving agents, achieving perfect scores while hiding significant capability gaps due to systemic biases and evaluation failures.
Disease distribution shifts, not skin tone, are the primary culprits behind the poor performance of dermatology AI models in unfamiliar clinical settings.
The ASCII Attack reveals that framing harmful requests as artistic critique can bypass safety filters in large language models, achieving up to 93% success in eliciting harmful responses.
Achieving fairness in stable matching without sacrificing stability, \texttt{SNSW-Alg} significantly outperforms traditional methods in equity across diverse preference distributions.
OBJECTION slashes the False Guilty Rate in legal AI predictions from over 82% to under 17% by challenging prosecutorial bias in real-time.
Ragebait is not just prevalent; it thrives in politically charged discussions, spreading faster and provoking stronger negative reactions than typical posts.
Two AI systems with nearly identical performance can have drastically different human oversight needs, revealing hidden costs in deployment readiness.
Dutch language models exhibit alarming biases, with some favoring stereotypical representations of transgender identities up to 97% of the time.
Selective spectral reversal can effectively undo harmful edits in language models while keeping beneficial changes intact, challenging the notion that all edits must be globally removed.
Cross-modal safety drift can significantly undermine the safety of multimodal language models, but a new method shows how to effectively transfer safety awareness from text to mitigate this risk.
A human-compatible rights layer could revolutionize how digital rights are exercised, bridging the gap between legal frameworks and practical application.
The shift to differential privacy could redefine how National Statistical Organisations balance data utility and individual privacy in the age of big data.
Users can significantly improve their privacy awareness and habits through a tailored tool that adapts to their specific needs and stages of behavioral change.
Political debates on AI and work are less about compensation for disruption and more about whether to enable or govern technological change, revealing deep ideological divides among parties.
Fairness-aware modeling can improve student attention estimation but may not generalize across diverse demographic groups, revealing critical gaps in current evaluation practices.
Closure failures in systems like GitHub and Kafka reveal that authorized actions can still lead to unwanted effects, challenging our understanding of authorization limits.
Invocation-time binding of call authority to workload state effectively eliminates post-authorization attacks, a critical vulnerability in remote LLM tool use.
Covert policy steering can achieve an 81.33% success rate in redirecting agent decisions without compromising output integrity or being detected.
MLLMs may rank faces similarly to humans, but they systematically overrate attractiveness, raising questions about their reliability in aesthetic assessments.
Typed provenance can prevent persistent AI agents from incorporating untrusted inputs into their autobiographical state, safeguarding user commitments and agent integrity.
Counterfactual fairness estimates in clinical LLMs can be misleading without a clear understanding of per-action instability, which can vary dramatically across decisions.
User context can significantly skew financial analysis in LLMs, with interpretation biases overshadowing evidence selection.
Trust in AI can be systematically cultivated through a structured educational framework that demystifies machine learning.
The presence of LLM agents can fundamentally shift group consensus from human-led to agent-led, altering both the content and legitimacy of shared norms.
Risk modeling for advanced AI systems is hampered by a lack of rigorous quantitative methods, leaving critical societal assessments in the dark.
AI can enhance individual creativity but risks diminishing collective diversity—finding the right mix of humans and algorithms is key to unlocking their full potential together.
Privacy policies are riddled with contradictions, with over a third of older policies failing to align commitments with practices, raising questions about transparency and accountability in data handling.
Harness-policy co-evolution can reduce adverse safety responses by 3x while simultaneously boosting benign utility in LLM agents.
GEO Defender slashes the success rate of malicious generative optimizations from over 50% to just 6%, all while preserving the integrity of benign content.
TRACE enables autonomous robots to achieve near-perfect auditability, ensuring every decision can be traced back to its sensor evidence.
A structured framework for monitoring AI progression could be the key to preventing catastrophic risks before they escalate.
Pruning LLMs can amplify biases, but Debias-SparseGPT shows that you can reduce these biases while maintaining performance, even under aggressive sparsity constraints.
Federated learning can outperform centralized models in multi-agent safety without compromising data privacy, achieving a remarkable 43% reduction in attack success rates.
A staggering 37% of users find themselves constitutionally homeless, as existing AI models fail to prioritize helpfulness or autonomy, highlighting a critical gap in AI governance.
Privacy in synthetic data is often treated as an implicit assumption rather than an explicit, testable claim, leading to uneven protections for sensitive information.
SAGE boosts worst-group accuracy by up to 7.7 percentage points, tackling the pervasive issue of spurious correlations in machine learning without relying on prior group labels.
Cultural cues can mislead language models into choosing the wrong answers even when they select the correct normative framework, exposing a significant knowledge gap in AI systems.
A revelation principle for AI agents reveals how to incentivize honesty and obedience even when their true capabilities are hidden.
A practical checklist empowers researchers to cut the carbon footprint of their AI applications in Earth system modeling, bridging the gap between theory and actionable practice.
Benign fine-tuning can lead to a dramatic collapse in safety alignment, revealing a fragile interplay between output-routing pathways and model safety.
Initial claims of substantial welfare gains from marketplace guardrails are misleading, with rigorous checks revealing many results as invalid.
Unsafe-response rates for Chinese LLMs can soar to over 30% under adversarial conditions, exposing vulnerabilities in automated safety judges.
dLVLMs not only reverse the yes-bias of AR models but also collapse in accuracy for underrepresented groups, revealing critical reliability gaps in diffusion-based approaches.
CSOs are trapped in a manual monitoring cycle, revealing a critical gap in technological support that could empower them against the far-right online.
Safety-based refusals in LLMs not only influence knowledge refusals but do so with greater impact, revealing an unexpected asymmetry in their underlying mechanisms.
Compositional risks in data processing can lead to overlooked privacy vulnerabilities, but a new protocol effectively identifies and mitigates these threats.
Multi-agent LLM systems are more vulnerable than previously thought, with systemic failures that local checks can't catch, demanding a new framework for security analysis.
Even small increases in scam reporting can drastically cut profits in AI-driven scams, suggesting a powerful leverage point for intervention.
Prior interactions can boost LLM compliance on cybersecurity requests by over 23%, but dialogue decomposition drastically undermines this effect.
Gaussian Core LoRA achieves a 7.95% reduction in attack success rates while maintaining visual quality, revolutionizing how we handle complex concept erasure in diffusion models.
Up to 50.2% of unauthorized requests can be falsely authorized by LLM memory, with executors blindly acting on these permissions 98.6% of the time.
LLMs may accurately reflect current cultural positions, but they lag years behind in capturing the dynamic nature of cultural change.
A low-complexity LLM can effectively reduce energy poverty without compromising performance, achieving significant equity improvements while minimizing carbon footprint.
Machine-extracted legal implications can be certified for reliability, but nearly 93% of outputs may still lack sufficient informativeness under real-world error conditions.
The emergence of AI's own ethics could fundamentally challenge and reshape our understanding of moral frameworks, requiring a reevaluation of established meta-ethical theories.