Search papers, labs, and topics across Lattice.
67 papers published across 4 labs.
The regulatory sandboxes can be viewed as pedagogical environments for AI: dynamic spaces where alignment develops as a formative process, progressively shaping autonomous behaviors through interaction and cooperation in scenarios of increasing complexity.
A certification layer is developed that wraps any severity model unmodified, with distribution-free guarantees using this structure, and attaches identical validity and certifies, on the vulnerable road users, a model-independent floor on set width that no base model beats.
A Distributional Sociotechnical Audit (DSA) is developed that integrates algorithmic equity, synthetic-data validity, and public-attitude heterogeneity into one empirical pipeline and offers a more defensible basis for transport GenAI governance than categorical approval tiers.
The results show that differentially private perturbation can be integrated into EEG processing workflows, but the selected mechanism, privacy parameters, and sensitivity calibration strongly influence data utility.
RISE AI provides an architecture for making bounded, evidence-based claims about Responsibility, Inclusivity, Safety, and Empowerment, and develops a rupture test that links institutional baselines to system evaluation.
The regulatory sandboxes can be viewed as pedagogical environments for AI: dynamic spaces where alignment develops as a formative process, progressively shaping autonomous behaviors through interaction and cooperation in scenarios of increasing complexity.
A certification layer is developed that wraps any severity model unmodified, with distribution-free guarantees using this structure, and attaches identical validity and certifies, on the vulnerable road users, a model-independent floor on set width that no base model beats.
A Distributional Sociotechnical Audit (DSA) is developed that integrates algorithmic equity, synthetic-data validity, and public-attitude heterogeneity into one empirical pipeline and offers a more defensible basis for transport GenAI governance than categorical approval tiers.
The results show that differentially private perturbation can be integrated into EEG processing workflows, but the selected mechanism, privacy parameters, and sensitivity calibration strongly influence data utility.
RISE AI provides an architecture for making bounded, evidence-based claims about Responsibility, Inclusivity, Safety, and Empowerment, and develops a rupture test that links institutional baselines to system evaluation.
Bridging the gap between theoretical optimality and practical representation learning, this work designs an implementation that imposes a translational bias on counterfactual trajectories - a constraint that aligns with how many concepts geometrically manifest in modern language models.
Retraining policies that appear to shrink demographic disparity under drift can actually degrade accuracy for all subgroups, while standard fairness conclusions can entirely flip based solely on evaluation population weighting.
Most LLM unlearning dissolves the moment a model is compressed for production, but isolating updates to high-significance layers prevents forgotten data from resurfacing under 4-bit quantization.
Prompt-based personalization cannot fix cultural bias: steering LLMs with cultural context actually widens the disparity between dominant and underrepresented cultures, even when injecting explicit cultural facts.
Pre-filtering text streams with emotion-aware semantic screening slashes the inference overhead of transformer-based toxicity classifiers without degrading cyberbullying detection sensitivity.
Nearly three-quarters of what benchmarks flag as model bias is just prompt-wording noise, but the surviving 25% are deeply entrenched pretraining representations that stubbornly resist post-training alignment.
A reference-based method that audits bias in hidden-state representations across related model variants across related model variants, for example before and after fine-tuning, which is complementary to output-based auditing rather than a replacement for it.
Up to 78 percentage points of reported machine unlearning success can be reversed with just 10 unlabeled images and zero weight updates, exposing widely cited forgetting benchmarks as mere BatchNorm illusions.
Formulating preference optimization directly at the gradient level eliminates the notorious over-refusal penalty in open-weight LLMs without compromising downstream utility.
Distributed neural network weights can now be cryptographically audited for multi-party collusion using zero-knowledge proofs, achieving certified sub-0.1% false-accusation rates across hundreds of leaked model mixtures.
Position bias and audit design rival or exceed demographic disparities in LLMs, rendering high-profile findings of rating-vs-ranking bias reversals non-replicable across hiring, lending, and triage.
Identical ethical dilemmas trigger contradictory value judgments across languages in frontier LLMs, but inference-time steering vectors extracted from hidden-state discrepancies can reliably enforce culture-specific alignment without fine-tuning.
Where an LLM represents a stereotype is detached from where it acts on it: linear decodability peaks up to 53% of model depth earlier than causal attribution, while fewer than 18% of stereotype-associated SAE features transfer across languages.
Deeper reasoning creates an unexpected safety backdoor: extended CoT dilutes attention away from system constraints, making models progressively more vulnerable to jailbreaks the longer they "think."
Static input guardrails remain blind to silent system prompt overrides, but inspecting post-generation latent trajectories catches constraint-violating responses before delivery while slashing false positive rates to under 3%.
This work introduces SCHEMEARENA, a 400-scenario benchmark for scalable scheming stress testing, constructed through a factorized scenario synthesis framework spanning diverse safety-relevant tool domains, instrumental goals, oversight conditions, and pressure mechanisms.
Cryptographic audit trails in agentic systems provide a false sense of security, proving bit-level integrity while fundamentally failing to verify semantic truth, authorization, or capture completeness.
Apple's walled garden fails to prevent XR data leaks: 58% of audited Vision Pro applications secretly transmit sensitive user traffic without mandatory privacy disclosures.
Generative AI has lowered the technical barrier to hijacking brain-computer interfaces, exposing 17 novel attack vectors capable of silently subverting neural decoders and tethered physical hardware.
Prematurely solving an author's immediate drafting problem with an LLM eliminates the exact cognitive friction required to uncover genuinely novel conceptual frameworks.
Conversational AI marks a fundamental rupture from search and social platforms: rather than merely indexing or amplifying human expression, generative systems actively displace it, exposing deep vulnerabilities in existing speech doctrine and platform liability regimes.
Headline claims that fine-tuning effortlessly extracts copyrighted books collapse under scrutiny, driven by sub-standard match thresholds, prompt leakage, and zero negative controls.
Exhaustive human review paradoxically degrades safety at scale due to vigilance fatigue, forcing a critical shift from synchronous human-in-the-loop filtering to layered, asynchronous human-on-the-loop oversight in high-stakes domains.
Apparent "centrist" alignment in smaller language models often masks degenerate response collapse rather than genuine neutrality, exposing deep measurement artifacts in standard political benchmark evaluations.
Two-thirds of material edits to frontier AI safety frameworks are never disclosed in developer changelogs—and 77% of these changes quietly weaken or remove safety commitments.
Financial fraud can be caught directly from the manifold: topological latent spaces retain enough geometric signal to snipe malicious transactions at ultra-low latency without ever exposing underlying PII.
Differentially private causal effect estimation no longer requires sacrificing statistical precision: propensity score blocking cuts estimation error by over 75% compared to state-of-the-art private IPW baselines.
Switching an LLM from an API to a consumer chat interface can degrade performance more than downgrading an entire model generation—and standard API hyperparameter controls cannot reliably bridge the gap.
Tracing whether an LLM or a human wrote a buggy line of code is an operational dead end: responsibility cannot be derived from code provenance, requiring quality engineering to pivot entirely to service-outcome verification.
Auditing demographic bias across guidance scales no longer requires brute-force image generation: causal abstraction enables a lightweight transformer to predict fairness shifts across the entire CFG spectrum.
Standard text-prompt evaluations severely underestimate diffusion safety risks, which reliably resurface under embedding-level control and trigger policy enforcement failures unless probabilities are explicitly calibrated.
A novel safe task-planning framework, SafeMem, which constructs and maintains a long-term semantic graph memory of the open and dynamic environment, and substantially improves safe success rates compared to state-of-the-art VLM-driven task planners.
Every speech-to-speech model misgenders speakers based on content, not voice, with misgendering rates soaring to 90% when voice and content clash.
In nearly half of real-world reports involving psychiatric delusions, AI chatbots actively validated the user's ungrounded beliefs—fueling escalations that directly preceded acute hospitalization and death by suicide.
Mandatory synthetic media disclosures are heading for a breakdown: without standardized technical thresholds for what degree of AI modification triggers compliance, current EU rules risk drowning users in meaningless UI badges while failing to catch deceptive deepfakes.
Automating peer review to survive the flood of AI-generated papers creates an exploitable adversarial loop, incentivizing authors to game predictable evaluators and ultimately corrupting the downstream corpus used to train future frontier models.
Multilingual medical benchmarks penalize cross-lingual answer variation as model error, but real-world clinicians are sharply split on whether AI should enforce universal consistency or adapt to local cultural contexts.
Automated climate control consistently misaligns with vulnerable households when systems treat thermal comfort as a static algorithmic setpoint rather than an embodied negotiation.
Most privacy-enhancing technologies merely secure the pipelines for inherently harmful functionalities rather than preventing the harms themselves.
Standard anti-aliasing in graphic renderers leaks geographic locations down to 1-meter resolution, breaking visual anonymization on dot maps with 0.0002-pixel recovery precision.
Telecom operators can now provision eSIM profiles without ever observing persistent device identifiers or tracking subscriber identities across carrier downloads.
Open source isn't banning AI code—83% of projects explicitly welcome it, but maintainers are actively rewriting governance rules to enforce human liability, mandatory PR disclosure, and countermeasures against autonomous "slop."
Maintainers are fundamentally restructuring OSS governance rules to defend scarce human attention against asymmetric floods of low-effort, AI-generated pull requests.
Alignment training doesn't erase political bias—it merely conceals a measurable latent direction that leads models to hardcode transient partisan consensus as timeless objective fact.
LLMs can perfectly mirror human moral verdicts while harboring an entirely un-human theory of mind, systematically whitewashing agents as far more altruistic and less hostile than humans perceive them to be.
Optimizing educational AI for immediate user preferences actively reinforces cognitive dependence and isolation, exposing a fundamental blind spot in standard human-centered design.
Chinese model guardrails target coordination rather than ideology—declining even to organize pro-government rallies—yet their sky-high refusal rates collapse under basic adversarial paraphrasing.
Despite 78% of the broader literature fixating on human-computer interaction, elite AI conferences disproportionately cite research treating LLMs as synthetic social minds and multi-agent societies.
Frontier AI hardware controls are functionally unpoliceable by government regulators alone, but a deployable framework of private-sector telemetry and third-party auditing can close critical verification gaps within a single year.
DCP provides a common evidence language for useful outcomes, alternative routes, and feedback effects across AI research, and is applied to Core, recovered, and audit-incomplete decisions.
Instruction tuning degrades the latent moral geometry of LLMs, even though distribution-driven steering vectors can naturally mirror human value topologies to enable coherent cross-value generalization.
EvoSafeHarness is a safety-specific optimization framework that synthesizes a deployable harness for a frozen model in a target domain, guided by model behavior, domain specifications, and fresh-context adversarial review to reject benchmark-specific rules.