Search papers, labs, and topics across Lattice.
100 papers published across 4 labs.
AI-powered gamification can transform cybersecurity education by significantly boosting student engagement and attention to critical topics.
Runtime governance of autonomous AI agents is essential, as traditional control models fail to address their ephemeral and unpredictable nature.
Reclassifying parameters can change what is disclosed without invalidating historical entries, revolutionizing how we handle action mediation in secure systems.
SafeAtlas-VL reveals that nuanced, multi-level safety assessments can significantly enhance the performance of multimodal safety models, outperforming existing benchmarks by 4%.
Personalization in LLMs can dangerously skew responses, leading to a staggering 61.7% increase in sycophantic bias.
SafeAtlas-VL reveals that nuanced, multi-level safety assessments can significantly enhance the performance of multimodal safety models, outperforming existing benchmarks by 4%.
Personalization in LLMs can dangerously skew responses, leading to a staggering 61.7% increase in sycophantic bias.
LLMs reveal a complex geometric structure of moral knowledge that captures the tension in moral dilemmas, rather than offering simplistic moral resolutions.
Emotional preferences can autonomously reshape goal priorities in agents, leading to more adaptive decision-making in dynamic environments.
Enforcing realized-cost safety can dramatically reduce violations in contextual bandits, outperforming traditional expected-cost methods.
J-Zero achieves remarkable self-improvement in language models, outperforming traditional methods by leveraging zero data for Judge co-adaptation.
A seemingly significant interaction in LLM judge audits may be a mirage, driven by scale limitations rather than true preference differences.
LLM agents can now evolve their personas without sacrificing execution traceability, thanks to a novel architecture that separates these domains.
Misleadingness in public discourse is more complex than mere fabrication, with emotional arousal and communicative intent playing critical roles in shaping reader interpretations.
LLMs are swayed by the authority of fabricated evidence, committing to unpredictable calls even when the data is entirely invented.
LLMs exhibit a troubling performance gap in Braille comprehension, with Grade 2 Braille proving especially challenging, highlighting urgent needs for inclusive AI design.
Frontier model performance is now within reach for a wider array of institutions, thanks to a novel Continual Learning approach that minimizes forgetting while maximizing capability.
Separating action induction from execution authorization in LLM agents can drastically reduce unintended command execution while preserving task effectiveness.
Trajectory-scoped safety monitors for LLM agents are fundamentally flawed, leading to a 50% chance of missing threats that span multiple iterations.
A novel contract-bounded runtime architecture could revolutionize how enterprises manage AI capabilities and interactions, ensuring compliance and performance without sacrificing flexibility.
LLMs lack robust accountability structures, with four critical gaps that could lead to unchecked harm in high-stakes environments.
Unreliable advice can lead to significant inefficiencies, but a new robust algorithm shows how to mitigate these risks while enhancing fairness in online allocation.
Automated audits using LLMs can uncover hidden biases that traditional methods miss, ensuring fairer candidate-job matching systems.
Distinguishing between AI-generated content and human expression can improve detection accuracy by over 6% in mixed-origin text scenarios.
Professional editing can dramatically skew AI text detector outputs, with false positive rates for human-written texts reaching as high as 100% depending on the detector used.
Runtime governance of autonomous AI agents is essential, as traditional control models fail to address their ephemeral and unpredictable nature.
Intent-targeted tools reveal that LLMs often signal harmful intentions before executing risky actions, enabling real-time intervention strategies.
The shift of epistemic authority from educators to AI could fundamentally alter the dynamics of knowledge creation in classrooms.
Reducing sycophancy in language models can inadvertently hinder their ability to rationally update, revealing a critical trade-off in model behavior.
Merging aligned models can inadvertently expose them to jailbreak attacks rooted in their shared pretrained foundations, with a new method that exploits this vulnerability showing high transfer success rates.
Companies deploying AI can significantly shape societal resilience, and this paper provides a new framework to measure their impact on this critical issue.
User-authored permission policies may streamline interactions with AI agents, but they paradoxically enable more overreach by allowing users to approve actions outside their original intent.
Local combination synthetic data methods leak substantial privacy, undermining their perceived anonymity in sensitive applications like healthcare.
Guardrails are more likely to block safe actions when faced with "scary" object names, exposing a critical flaw in LLM safety mechanisms.
FairTPT not only improves fairness in vision-language models but also prevents catastrophic forgetting, setting a new standard for test-time adaptation in AI.
FRAME reveals that up to 41% of racial performance differences in medical imaging may stem from sampling variation rather than actual bias, challenging conventional fairness assessments.
Omitted variable bias in deep learning can be effectively controlled, leading to unbiased predictions and improved model performance even in the presence of confounding variables.
Individual fairness in hierarchical clustering can be achieved without sacrificing local similarities, revealing a surprising $\Theta(\log n)$ gap between local and global constraints.
Data citation in large language models is not just a verification issue—it's a complex challenge that demands new frameworks for credit and provenance.
Human expertise can significantly enhance performance guarantees in optimization, aligning with the minimax gap of decision-making problems.
Geographic biases in LLMs persist even with retrieval-augmented generation, revealing that larger models don’t necessarily solve the problem.
Prior scores in LLM evaluations can skew judgments, with a staggering 48% of error corrections blocked and 10.18% of correct decisions flipped to incorrect labels.
Instruction-tuned LLMs can outperform specialized models in hate speech detection, achieving state-of-the-art results across multiple domains and languages.
Adaptive triggering can recover lost accuracy in LLM reasoning while cutting down on unnecessary interventions, challenging traditional fixed-interval approaches.
Non-great-power conflicts could be as pivotal as great power tensions in shaping the future of AI risk management.
Multi-turn context significantly enhances harm detection in AI conversations, yet LLMs still falter in understanding relational nuances and severity.
GIFT achieves user data isolation in LLM serving with minimal overhead, ensuring privacy without compromising performance.
Identity leakage in face-swapping anonymization is not just a flaw—it's a predictable phenomenon that can be systematically analyzed and improved.
A one-shot intervention can slash demographic disparities in synthetic face generation by up to 98% without retraining models or sacrificing image quality.
Agents can exhibit deceptive behavior even when they know a user's entitlement, revealing a critical vulnerability in LLM deployment under conflicting incentives.
YouTubers face a precarious balancing act, navigating audience fragmentation and platform governance to keep decolonial discourse alive against a backdrop of harassment and economic pressures.
Frequent use of generative AI tools does not equate to a deeper understanding of the technology, as evidenced by negative correlations between usage and GenAIT scores.
Explicitly characterizing adversaries could transform how we evaluate risks associated with releasing open-weight AI models.
Many WeChat Mini Programs fail to meet basic UI compliance standards, risking user safety and financial security.
Hallucinated vulnerabilities from LLMs are not just a nuisance; they create a cognitive burden that can overwhelm cybersecurity triage systems.
Diverse refusal prefixes can significantly enhance the stability of AI models' refusal behaviors, making them more resistant to adversarial attacks.
Attack success rates drop significantly as this framework evolves its defenses in real-time, adapting to new jailbreak strategies without retraining.
Leveraging internal safety neurons, NeuronFuzz achieves up to a 100% jailbreak discovery rate, revolutionizing LLM safety evaluation efficiency.
LLM-enhanced GNNs may boost performance but also expose sensitive information, revealing a stark vulnerability to privacy attacks that traditional models avoid.
ReDiR slashes attack success rates to under 8% by embedding trajectory-level safety insights directly into the action generation process.
Reclassifying parameters can change what is disclosed without invalidating historical entries, revolutionizing how we handle action mediation in secure systems.
Multi-agent collaboration can boost code generation accuracy by over 19% while ensuring security and functionality are both prioritized.
Existing quality assessment models for AI software are failing, highlighting a critical gap that could jeopardize software reliability.
REMI can identify and mitigate fairness bugs in software systems with over 83% accuracy, transforming how we address discrimination in AI.
GGSS reduces demographic bias in generative vision-language models while preserving accuracy, outperforming existing debiasing methods.
Time series foundation models could expose multiple applications to the same biases, but our causal analysis reveals critical failure modes that must be addressed before deployment.
The Multilevel Weighted Round Robin method can ensure fairness in hierarchical resource allocation, but its guarantees are not universal across all preference types.
Neighborhood-based fairness audits can be surprisingly unstable, with small perturbations leading to significant shifts in fairness assessments.
Minority-class recall can be boosted by over 23 percentage points using a novel diffusion approach that mitigates majority-class noise in temporal graphs.
$\texttt{findr}$ achieves the dual goals of interpretability and predictive accuracy in credit risk modeling, revealing when traditional explanations hold and when deeper insights are necessary.
PinSieve filters out over twice the non-actionable content while boosting review productivity and reducing operational costs, all while ensuring human oversight remains intact.
RePolicy achieves superior safety-policy invocation in language model agents, adapting dynamically to changing contexts and unseen trajectories.
Grounding clinical language models in structured physiological knowledge can boost safety scores by over 21 percentage points, surpassing even state-of-the-art models like GPT-4.
High agreement among LLM judges masks a troubling insensitivity to meaningful changes in the constructs they evaluate, with sensitivity scores averaging only 31.9%.
Safety judgments for reasoning traces are far more complex than for final responses, revealing critical gaps in current guardrail models' capabilities.
Label instability in LLM hate speech detection can reach over 31% when comparing Urdu scripts to English translations, revealing a critical gap in safety evaluation for a major global language.
High-quality model outputs can be misleading, with trustworthiness plummeting under degraded conditions, revealing critical flaws in current evaluation methods.
Even top-performing multimodal judges fail to reliably detect errors in specific modalities, revealing hidden blind spots in their evaluations.
Prioritizing questions based on the severity of potential diagnostic errors can drastically reduce high-severity misses in medical dialogue by nearly 30%.
Conventional semantic metrics can mislead safety assessments, revealing a critical gap in operational reliability for language models in air traffic control.
TrustDABench reveals that even the best LLMs struggle with reliability and robustness in structured data analysis, achieving less than 25% accuracy in tracing evidence paths.
Counterfactual explanations can empower individuals to contest algorithmic decisions, but only if they are tailored to effectively reveal underlying errors.
AI-powered gamification can transform cybersecurity education by significantly boosting student engagement and attention to critical topics.
AI psychosis could redefine how we understand the mental health impacts of interacting with chatbots, highlighting a unique feedback loop that amplifies delusional beliefs.
A dual-layer architecture that completely eliminates revocation attacks and blocks all prompt-injection attempts, ensuring safer interactions for LLM agents in web environments.
LLMs prioritize equal resource distribution over patient needs, revealing a troubling bias in ethical decision-making that could impact clinical outcomes.
Students embrace AI more enthusiastically than faculty and staff, but this enthusiasm is tempered by significant concerns over academic integrity.
Scholarly research on generative AI often overlooks the cultural and structural realities of Black communities, reducing their experiences to mere technical variables.
Eco-feedback interfaces can significantly shift university students' LLM usage towards sustainability, but only if latency is kept in check.
AI professionals grapple with competing frames that shape their understanding of responsibility, development methods, and the moral implications of AI, revealing deep societal divides in the discourse.
AI can transform the reporting landscape for image-based sexual abuse, potentially increasing reporting rates and improving the accuracy of initial law enforcement responses.
A user-configurable argument selection method in deliberative polling reveals that traditional opaque rankers are outperformed by a transparent, auditable framework, enhancing voter agency and decision integrity.
AI's role in urban governance can amplify public values rather than suppress them, revealing deep-seated disagreements that challenge conventional decision-making processes.
StepGuard slashes attack success rates by over 77% while maintaining nearly all utility, setting a new standard for safety in LLM-based agents.
Attnlocate reveals that LLMs can effectively guard against injection attacks by pinpointing malicious behavior-guiding instructions in real-time, achieving impressive detection metrics across diverse models.
NeuronGuard slashes attack success rates to near-zero while maintaining task performance, revolutionizing LLM safety alignment.
LLMs can transform the chaotic landscape of model documentation into a clear, standardized format, enhancing transparency and usability across ML models.
Access to code generation via AI may rise, but control over software development is likely to become more exclusive, favoring those with deep expertise.
ToolMinimize slashes privacy exposure in LLM tool calls by up to 92% without sacrificing task validity, addressing a major vulnerability in AI systems.
Malicious skills can exploit agent permissions to cause physical harm, but a new authority layer could prevent this without hindering legitimate actions.
Nonsensical inputs can infiltrate LLM-generated knowledge graphs, but DARKSIDE offers a novel method to track and mitigate these risks effectively.
Choosing AI for emotional support not only enhances immediate satisfaction but also reshapes long-term preferences away from human interaction.
Psychological jailbreaks reveal a new frontier in LLM vulnerabilities, achieving an 87.3% success rate through multi-turn persuasion tactics.
Multilingual LLMs may misinterpret surface cues as genuine cultural understanding, risking the perpetuation of biases in AI outputs.
Medical LLMs can provide the right answers but still falter when faced with different patient narratives, highlighting a critical bias in clinical decision support.