Search papers, labs, and topics across Lattice.
AI governance principles, value alignment through constitutions, fairness, bias mitigation, and ethical AI deployment.
#20 of 24
1
Erasing bias while maintaining model integrity is now achievable with a framework that seamlessly navigates between concept erasure and counterfactual generation.
Retraining policies that appear to shrink demographic disparity under drift can actually degrade accuracy for all subgroups, while standard fairness conclusions can entirely flip based solely on evaluation population weighting.
Most LLM unlearning dissolves the moment a model is compressed for production, but isolating updates to high-significance layers prevents forgotten data from resurfacing under 4-bit quantization.
Prompt-based personalization cannot fix cultural bias: steering LLMs with cultural context actually widens the disparity between dominant and underrepresented cultures, even when injecting explicit cultural facts.
Pre-filtering text streams with emotion-aware semantic screening slashes the inference overhead of transformer-based toxicity classifiers without degrading cyberbullying detection sensitivity.
Nearly three-quarters of what benchmarks flag as model bias is just prompt-wording noise, but the surviving 25% are deeply entrenched pretraining representations that stubbornly resist post-training alignment.
Up to 78 percentage points of reported machine unlearning success can be reversed with just 10 unlabeled images and zero weight updates, exposing widely cited forgetting benchmarks as mere BatchNorm illusions.
Formulating preference optimization directly at the gradient level eliminates the notorious over-refusal penalty in open-weight LLMs without compromising downstream utility.
Distributed neural network weights can now be cryptographically audited for multi-party collusion using zero-knowledge proofs, achieving certified sub-0.1% false-accusation rates across hundreds of leaked model mixtures.
Position bias and audit design rival or exceed demographic disparities in LLMs, rendering high-profile findings of rating-vs-ranking bias reversals non-replicable across hiring, lending, and triage.
Identical ethical dilemmas trigger contradictory value judgments across languages in frontier LLMs, but inference-time steering vectors extracted from hidden-state discrepancies can reliably enforce culture-specific alignment without fine-tuning.
Where an LLM represents a stereotype is detached from where it acts on it: linear decodability peaks up to 53% of model depth earlier than causal attribution, while fewer than 18% of stereotype-associated SAE features transfer across languages.
Deeper reasoning creates an unexpected safety backdoor: extended CoT dilutes attention away from system constraints, making models progressively more vulnerable to jailbreaks the longer they "think."
Static input guardrails remain blind to silent system prompt overrides, but inspecting post-generation latent trajectories catches constraint-violating responses before delivery while slashing false positive rates to under 3%.
Cryptographic audit trails in agentic systems provide a false sense of security, proving bit-level integrity while fundamentally failing to verify semantic truth, authorization, or capture completeness.
Apple's walled garden fails to prevent XR data leaks: 58% of audited Vision Pro applications secretly transmit sensitive user traffic without mandatory privacy disclosures.
Generative AI has lowered the technical barrier to hijacking brain-computer interfaces, exposing 17 novel attack vectors capable of silently subverting neural decoders and tethered physical hardware.
Prematurely solving an author's immediate drafting problem with an LLM eliminates the exact cognitive friction required to uncover genuinely novel conceptual frameworks.
Conversational AI marks a fundamental rupture from search and social platforms: rather than merely indexing or amplifying human expression, generative systems actively displace it, exposing deep vulnerabilities in existing speech doctrine and platform liability regimes.
Headline claims that fine-tuning effortlessly extracts copyrighted books collapse under scrutiny, driven by sub-standard match thresholds, prompt leakage, and zero negative controls.
Exhaustive human review paradoxically degrades safety at scale due to vigilance fatigue, forcing a critical shift from synchronous human-in-the-loop filtering to layered, asynchronous human-on-the-loop oversight in high-stakes domains.
Apparent "centrist" alignment in smaller language models often masks degenerate response collapse rather than genuine neutrality, exposing deep measurement artifacts in standard political benchmark evaluations.
Two-thirds of material edits to frontier AI safety frameworks are never disclosed in developer changelogs—and 77% of these changes quietly weaken or remove safety commitments.
Financial fraud can be caught directly from the manifold: topological latent spaces retain enough geometric signal to snipe malicious transactions at ultra-low latency without ever exposing underlying PII.
Differentially private causal effect estimation no longer requires sacrificing statistical precision: propensity score blocking cuts estimation error by over 75% compared to state-of-the-art private IPW baselines.
Switching an LLM from an API to a consumer chat interface can degrade performance more than downgrading an entire model generation—and standard API hyperparameter controls cannot reliably bridge the gap.
Tracing whether an LLM or a human wrote a buggy line of code is an operational dead end: responsibility cannot be derived from code provenance, requiring quality engineering to pivot entirely to service-outcome verification.
Auditing demographic bias across guidance scales no longer requires brute-force image generation: causal abstraction enables a lightweight transformer to predict fairness shifts across the entire CFG spectrum.
Standard text-prompt evaluations severely underestimate diffusion safety risks, which reliably resurface under embedding-level control and trigger policy enforcement failures unless probabilities are explicitly calibrated.
A novel safe task-planning framework, SafeMem, which constructs and maintains a long-term semantic graph memory of the open and dynamic environment, and substantially improves safe success rates compared to state-of-the-art VLM-driven task planners.
Every speech-to-speech model misgenders speakers based on content, not voice, with misgendering rates soaring to 90% when voice and content clash.
In nearly half of real-world reports involving psychiatric delusions, AI chatbots actively validated the user's ungrounded beliefs—fueling escalations that directly preceded acute hospitalization and death by suicide.
Mandatory synthetic media disclosures are heading for a breakdown: without standardized technical thresholds for what degree of AI modification triggers compliance, current EU rules risk drowning users in meaningless UI badges while failing to catch deceptive deepfakes.