Search papers, labs, and topics across Lattice.
AI governance principles, value alignment through constitutions, fairness, bias mitigation, and ethical AI deployment.
#10 of 24
2
FairTPT not only improves fairness in vision-language models but also prevents catastrophic forgetting, setting a new standard for test-time adaptation in AI.
FRAME reveals that up to 41% of racial performance differences in medical imaging may stem from sampling variation rather than actual bias, challenging conventional fairness assessments.
Omitted variable bias in deep learning can be effectively controlled, leading to unbiased predictions and improved model performance even in the presence of confounding variables.
LLM-enhanced GNNs may boost performance, but they also expose sensitive information, revealing a troubling vulnerability to privacy attacks.
Individual fairness in hierarchical clustering can be achieved without sacrificing local similarities, revealing a surprising $\Theta(\log n)$ gap between local and global constraints.
Data citation for large language models is not just a verification issue; it’s a multifaceted challenge that could redefine how we acknowledge and trace the origins of AI-generated information.
Diverse refusal prefixes can significantly bolster the stability of refusal mechanisms, making them more resilient against targeted attacks.
Human expertise can significantly enhance performance guarantees in optimization, aligning with the minimax gap of decision-making problems.
Hallucinated vulnerabilities in AI-driven vulnerability assessments create a cognitive burden that mimics a denial-of-service attack on human triage systems.
Geographic biases in LLMs persist even with retrieval-augmented generation, revealing that larger models don’t necessarily solve the problem.
Multi-agent collaboration in code generation can boost functional and security correctness by over 19 percentage points, reshaping how we approach secure coding.
Attack success rates plummet as this self-evolving defense framework learns from each jailbreak attempt, adapting in real-time without any parameter tuning.
Prior scores in LLM evaluations can skew judgments, with a staggering 48% of error corrections blocked and 10.18% of correct decisions flipped to incorrect labels.
Instruction-tuned LLMs can outperform specialized models in hate speech detection, achieving state-of-the-art results across multiple domains and languages.
Adaptive triggering can recover lost accuracy in LLM reasoning while cutting down on unnecessary interventions, challenging traditional fixed-interval approaches.
GGSS reduces demographic bias in generative VLMs while preserving accuracy, outperforming traditional debiasing methods.
Explicitly defining threat actors could transform how we assess risks associated with the release of open-weight AI models.
Non-great-power conflicts could be as critical as great power conflicts in shaping AI risk landscapes, challenging conventional wisdom in the field.
Multi-turn context significantly boosts harm detection in AI conversations, yet LLMs still falter in understanding nuanced relational dynamics.
Frequent use of generative AI tools does not guarantee a solid understanding of their underlying concepts, as shown by the GenAIT test results.
ReDiR slashes attack success rates to under 8% by embedding trajectory-level safety insights directly into the action generation process of LLM agents.
Misplaced buttons in WeChat Mini Programs can lead to involuntary payments and serious security breaches, highlighting a critical gap in mobile app compliance.
GIFT enforces user data isolation in LLM serving with less than 11% throughput overhead, revolutionizing privacy protection in shared infrastructures.
Face-swapping may not be as secure as it seems, with significant identity leakage that can be quantitatively analyzed and predicted.
Reclassifying action parameters can enhance privacy without sacrificing audit integrity, revealing a new dimension in ledger design.
A one-shot intervention can slash demographic disparities in synthetic face generation by up to 98% without retraining models or sacrificing image quality.
Time series foundation models could expose multiple applications to the same biases, but our causal analysis reveals critical failure modes that must be addressed before deployment.
The Multilevel Weighted Round Robin method can ensure fairness in hierarchical resource allocation, but its guarantees are not universal across all preference types.
Neighborhood-based fairness audits can be surprisingly unstable, with small perturbations leading to significant shifts in fairness assessments.
Minority-class recall can be boosted by over 23 percentage points using a novel diffusion approach that mitigates majority-class noise in temporal graphs.
$\texttt{findr}$ achieves the dual goals of interpretability and predictive accuracy in credit risk modeling, revealing when traditional explanations hold and when deeper insights are necessary.
PinSieve filters out over twice the non-actionable content while boosting review productivity and reducing operational costs, all while ensuring human oversight remains intact.
RePolicy achieves superior safety-policy invocation in language model agents, adapting dynamically to changing contexts and unseen trajectories.
Grounding clinical language models in structured physiological knowledge can boost safety scores by over 21 percentage points, surpassing even state-of-the-art models like GPT-4.
High agreement among LLM judges masks a troubling insensitivity to meaningful changes in the constructs they evaluate, with sensitivity scores averaging only 31.9%.
Safety judgments for reasoning traces are far more complex than for final responses, revealing critical gaps in current guardrail models' capabilities.
Label instability in LLM hate speech detection can reach over 31% when comparing Urdu scripts to English translations, revealing a critical gap in safety evaluation for a major global language.
High-quality model outputs can be misleading, with trustworthiness plummeting under degraded conditions, revealing critical flaws in current evaluation methods.
Even top-performing multimodal judges fail to reliably detect errors in specific modalities, revealing hidden blind spots in their evaluations.
Prioritizing questions based on the severity of potential diagnostic errors can drastically reduce high-severity misses in medical dialogue by nearly 30%.
Conventional semantic metrics can mislead safety assessments, revealing a critical gap in operational reliability for language models in air traffic control.
TrustDABench reveals that even the best LLMs struggle with reliability and robustness in structured data analysis, achieving less than 25% accuracy in tracing evidence paths.
Counterfactual explanations can empower individuals to contest algorithmic decisions, but only if they are tailored to effectively reveal underlying errors.
AI-powered gamification can transform cybersecurity education by significantly boosting student engagement and attention to critical topics.
AI psychosis could redefine how we understand the mental health impacts of interacting with chatbots, highlighting a unique feedback loop that amplifies delusional beliefs.
A dual-layer architecture that completely eliminates revocation attacks and blocks all prompt-injection attempts, ensuring safer interactions for LLM agents in web environments.
LLMs prioritize equal resource distribution over patient needs, revealing a troubling bias in ethical decision-making that could impact clinical outcomes.
Students embrace AI more enthusiastically than faculty and staff, but this enthusiasm is tempered by significant concerns over academic integrity.
Scholarly research on generative AI often overlooks the cultural and structural realities of Black communities, reducing their experiences to mere technical variables.
Eco-feedback interfaces can significantly shift university students' LLM usage towards sustainability, but only if latency is kept in check.