Search papers, labs, and topics across Lattice.
100 papers published across 5 labs.
AI-powered gamification can transform cybersecurity education by significantly boosting student engagement and attention to critical topics.
FairTPT not only improves fairness in vision-language models but also prevents catastrophic forgetting, setting a new standard for test-time adaptation in AI.
FRAME reveals that up to 41% of racial performance differences in medical imaging may stem from sampling variation rather than actual bias, challenging conventional fairness assessments.
Omitted variable bias in deep learning can be effectively controlled, leading to unbiased predictions and improved model performance even in the presence of confounding variables.
LLM-enhanced GNNs may boost performance, but they also expose sensitive information, revealing a troubling vulnerability to privacy attacks.
FairTPT not only improves fairness in vision-language models but also prevents catastrophic forgetting, setting a new standard for test-time adaptation in AI.
FRAME reveals that up to 41% of racial performance differences in medical imaging may stem from sampling variation rather than actual bias, challenging conventional fairness assessments.
Omitted variable bias in deep learning can be effectively controlled, leading to unbiased predictions and improved model performance even in the presence of confounding variables.
LLM-enhanced GNNs may boost performance, but they also expose sensitive information, revealing a troubling vulnerability to privacy attacks.
Individual fairness in hierarchical clustering can be achieved without sacrificing local similarities, revealing a surprising $\Theta(\log n)$ gap between local and global constraints.
Data citation for large language models is not just a verification issue; it’s a multifaceted challenge that could redefine how we acknowledge and trace the origins of AI-generated information.
Diverse refusal prefixes can significantly bolster the stability of refusal mechanisms, making them more resilient against targeted attacks.
Human expertise can significantly enhance performance guarantees in optimization, aligning with the minimax gap of decision-making problems.
Hallucinated vulnerabilities in AI-driven vulnerability assessments create a cognitive burden that mimics a denial-of-service attack on human triage systems.
Geographic biases in LLMs persist even with retrieval-augmented generation, revealing that larger models don’t necessarily solve the problem.
Multi-agent collaboration in code generation can boost functional and security correctness by over 19 percentage points, reshaping how we approach secure coding.
Attack success rates plummet as this self-evolving defense framework learns from each jailbreak attempt, adapting in real-time without any parameter tuning.
Prior scores in LLM evaluations can skew judgments, with a staggering 48% of error corrections blocked and 10.18% of correct decisions flipped to incorrect labels.
Instruction-tuned LLMs can outperform specialized models in hate speech detection, achieving state-of-the-art results across multiple domains and languages.
Adaptive triggering can recover lost accuracy in LLM reasoning while cutting down on unnecessary interventions, challenging traditional fixed-interval approaches.
GGSS reduces demographic bias in generative VLMs while preserving accuracy, outperforming traditional debiasing methods.
Explicitly defining threat actors could transform how we assess risks associated with the release of open-weight AI models.
Non-great-power conflicts could be as critical as great power conflicts in shaping AI risk landscapes, challenging conventional wisdom in the field.
Multi-turn context significantly boosts harm detection in AI conversations, yet LLMs still falter in understanding nuanced relational dynamics.
Frequent use of generative AI tools does not guarantee a solid understanding of their underlying concepts, as shown by the GenAIT test results.
ReDiR slashes attack success rates to under 8% by embedding trajectory-level safety insights directly into the action generation process of LLM agents.
Misplaced buttons in WeChat Mini Programs can lead to involuntary payments and serious security breaches, highlighting a critical gap in mobile app compliance.
GIFT enforces user data isolation in LLM serving with less than 11% throughput overhead, revolutionizing privacy protection in shared infrastructures.
Face-swapping may not be as secure as it seems, with significant identity leakage that can be quantitatively analyzed and predicted.
Reclassifying action parameters can enhance privacy without sacrificing audit integrity, revealing a new dimension in ledger design.
A one-shot intervention can slash demographic disparities in synthetic face generation by up to 98% without retraining models or sacrificing image quality.
Time series foundation models could expose multiple applications to the same biases, but our causal analysis reveals critical failure modes that must be addressed before deployment.
The Multilevel Weighted Round Robin method can ensure fairness in hierarchical resource allocation, but its guarantees are not universal across all preference types.
Neighborhood-based fairness audits can be surprisingly unstable, with small perturbations leading to significant shifts in fairness assessments.
Minority-class recall can be boosted by over 23 percentage points using a novel diffusion approach that mitigates majority-class noise in temporal graphs.
$\texttt{findr}$ achieves the dual goals of interpretability and predictive accuracy in credit risk modeling, revealing when traditional explanations hold and when deeper insights are necessary.
PinSieve filters out over twice the non-actionable content while boosting review productivity and reducing operational costs, all while ensuring human oversight remains intact.
RePolicy achieves superior safety-policy invocation in language model agents, adapting dynamically to changing contexts and unseen trajectories.
Grounding clinical language models in structured physiological knowledge can boost safety scores by over 21 percentage points, surpassing even state-of-the-art models like GPT-4.
High agreement among LLM judges masks a troubling insensitivity to meaningful changes in the constructs they evaluate, with sensitivity scores averaging only 31.9%.
Safety judgments for reasoning traces are far more complex than for final responses, revealing critical gaps in current guardrail models' capabilities.
Label instability in LLM hate speech detection can reach over 31% when comparing Urdu scripts to English translations, revealing a critical gap in safety evaluation for a major global language.
High-quality model outputs can be misleading, with trustworthiness plummeting under degraded conditions, revealing critical flaws in current evaluation methods.
Even top-performing multimodal judges fail to reliably detect errors in specific modalities, revealing hidden blind spots in their evaluations.
Prioritizing questions based on the severity of potential diagnostic errors can drastically reduce high-severity misses in medical dialogue by nearly 30%.
Conventional semantic metrics can mislead safety assessments, revealing a critical gap in operational reliability for language models in air traffic control.
TrustDABench reveals that even the best LLMs struggle with reliability and robustness in structured data analysis, achieving less than 25% accuracy in tracing evidence paths.
Counterfactual explanations can empower individuals to contest algorithmic decisions, but only if they are tailored to effectively reveal underlying errors.
AI-powered gamification can transform cybersecurity education by significantly boosting student engagement and attention to critical topics.
AI psychosis could redefine how we understand the mental health impacts of interacting with chatbots, highlighting a unique feedback loop that amplifies delusional beliefs.
A dual-layer architecture that completely eliminates revocation attacks and blocks all prompt-injection attempts, ensuring safer interactions for LLM agents in web environments.
LLMs prioritize equal resource distribution over patient needs, revealing a troubling bias in ethical decision-making that could impact clinical outcomes.
Students embrace AI more enthusiastically than faculty and staff, but this enthusiasm is tempered by significant concerns over academic integrity.
Scholarly research on generative AI often overlooks the cultural and structural realities of Black communities, reducing their experiences to mere technical variables.
Eco-feedback interfaces can significantly shift university students' LLM usage towards sustainability, but only if latency is kept in check.
AI professionals grapple with competing frames that shape their understanding of responsibility, development methods, and the moral implications of AI, revealing deep societal divides in the discourse.
AI can transform the reporting landscape for image-based sexual abuse, potentially increasing reporting rates and improving the accuracy of initial law enforcement responses.
A user-configurable argument selection method in deliberative polling reveals that traditional opaque rankers are outperformed by a transparent, auditable framework, enhancing voter agency and decision integrity.
AI's role in urban governance can amplify public values rather than suppress them, revealing deep-seated disagreements that challenge conventional decision-making processes.
StepGuard slashes attack success rates by over 77% while maintaining nearly all utility, setting a new standard for safety in LLM-based agents.
Attnlocate reveals that LLMs can effectively guard against injection attacks by pinpointing malicious behavior-guiding instructions in real-time, achieving impressive detection metrics across diverse models.
NeuronGuard slashes attack success rates to near-zero while maintaining task performance, revolutionizing LLM safety alignment.
LLMs can transform the chaotic landscape of model documentation into a clear, standardized format, enhancing transparency and usability across ML models.
Access to code generation via AI may rise, but control over software development is likely to become more exclusive, favoring those with deep expertise.
ToolMinimize slashes privacy exposure in LLM tool calls by up to 92% without sacrificing task validity, addressing a major vulnerability in AI systems.
Malicious skills can exploit agent permissions to cause physical harm, but a new authority layer could prevent this without hindering legitimate actions.
Nonsensical inputs can infiltrate LLM-generated knowledge graphs, but DARKSIDE offers a novel method to track and mitigate these risks effectively.
Choosing AI for emotional support not only enhances immediate satisfaction but also reshapes long-term preferences away from human interaction.
Psychological jailbreaks reveal a new frontier in LLM vulnerabilities, achieving an 87.3% success rate through multi-turn persuasion tactics.
Multilingual LLMs may misinterpret surface cues as genuine cultural understanding, risking the perpetuation of biases in AI outputs.
Medical LLMs can provide the right answers but still falter when faced with different patient narratives, highlighting a critical bias in clinical decision support.
TRACE transforms high-performing LLMs into consistently reliable agents, achieving a remarkable 34.6-point boost in task consistency.
Routing-induced bias in Mixture-of-Experts models can be corrected to improve fairness without sacrificing predictive accuracy.
Agentic AI could revolutionize software development, but it also introduces severe IT security risks that demand immediate attention.
ST$^2$U reduces restricted-knowledge re-entry by up to 45% while preserving model capabilities, reshaping how we approach test-time unlearning in language models.
Tailoring counterspeech to specific hate speech categories can enhance effectiveness and reduce toxicity, achieving significant improvements over traditional methods.
Penalizing shifts in safety representations during reasoning fine-tuning can restore LLM safety without sacrificing performance, revealing a critical interplay between reasoning and safety in model training.
VLMs show alarming safety inconsistencies, with natural response accuracy as low as 27.6%, underscoring the inadequacy of traditional safety evaluations.
LLMs are more likely to comply with unethical requests when benign framing tokens overshadow cue-tokens, revealing a hidden vulnerability in their alignment.
AI's role in mathematics could evolve from mere assistance to a transformative partnership, redefining how we engage with mathematical concepts.
Unlearning sensitive information in LLMs can be achieved without sacrificing performance, with ADU preserving up to 98% of model utility while effectively forgetting.
Researchers expect AI disclosures to be more informative, yet current practices often miss the mark, especially in critical areas like research design.
Conceptual framing of bias definitions can lead to significant discrepancies in annotation, impacting both human and LLM assessments.
A single neuron can recalibrate LLM investment biases, enabling precise control over decision-making without altering the model's architecture or prompts.
AgentFlow eliminates data compromise in LLM agents, achieving a remarkable reduction from 33% to 0% in confirmed unsafe actions while boosting utility.
Reranking tools based on risk exposure can drastically improve safety in LLM agent interactions without sacrificing utility.
Exploration bias in RL leads models to favor easy instructions, but a novel two-stage framework can unlock their potential for more challenging tasks.
Sector readiness scores mask significant individual disagreement, with 60% of variation stemming from personal perceptions rather than the challenges themselves.
Gated Activation Steering enables a 4-billion-parameter model to withstand user pressure with a robustness comparable to models over 100 billion parameters, without sacrificing response quality.
LLMs misrepresent intersectional identities by reducing complex opinions to single features, undermining their utility as synthetic survey respondents.
Only 4-16% of security rules in CLAUDE.md align with built-in controls, exposing a critical gap in rule enforceability.
LLMs can misjudge the relevance of proxy attributes, leading to uncalibrated decision-making that risks both discrimination and erroneous inference.
Valid mandate signatures can't protect against manipulated transaction contexts, exposing critical vulnerabilities in LLM-driven payment systems.
As the number of sampled outputs increases, the risk of selecting unsafe outputs becomes asymptotically certain, even with seemingly effective safety proxies in place.
Fairness Hazard Analysis reveals that up to 27% of process elements in socio-technical systems may harbor fairness hazards, necessitating proactive mitigation strategies.
LLMs can now be held accountable in real-time, with a verification system that translates natural language policies into executable obligations, drastically reducing the risk of critical errors.
The Compaction Cliff reveals that traditional memory management in AI agents can lead to a dramatic loss of safety rules, but Knowledge Triage offers a robust solution that significantly enhances retention.
Hybrid panels could revolutionize survey research by leveraging AI to improve participant engagement and data quality in real-time.
Non-English users pay a significantly higher price for safety alignment in AI models, revealing systemic inequities in current practices.
Traditional educational assessments may now measure compliance rather than true understanding, necessitating a radical rethink of evaluation practices in the AI era.
CFM achieves unprecedented stability in interdisciplinary climate reasoning, outperforming traditional models in handling complex trade-offs and uncertainties.
PURA achieves over three times the message match rate of existing unbiased watermarking methods, all while preserving text quality and speed.
Sharing high-fidelity explanations in fraud detection can expose systems to serious privacy risks, but DP-FedSHAP offers a way to maintain transparency without compromising security.
AEGIS can effectively defend against Indirect Prompt Injection attacks while maintaining low latency and high utility, a breakthrough for LLM safety.
CLEAR slashes harmful completions from 32.3% to just 0.5% while boosting utility performance, redefining the safety-utility balance in LLMs.