Search papers, labs, and topics across Lattice.
100 papers published across 6 labs.
ALE and SHAP emerge as the most reliable methods for understanding feature importance in heat demand forecasting, revealing critical insights for high-stakes AI applications.
Early persona integration in language models can drastically reduce misalignment in moral dilemmas and enhance adherence to desired values.
Bengali speakers face a staggering 67:1 training-token deficit compared to English, highlighting a critical inequity in AI language support.
ProME achieves unprecedented group robustness by aligning environment selection directly with the deployed predictor, leading to superior worst-group accuracy.
Over half of Telegram Mini Apps leak data to undisclosed third parties, revealing a shocking compliance gap in user privacy protections.
ProME achieves unprecedented group robustness by aligning environment selection directly with the deployed predictor, leading to superior worst-group accuracy.
Over half of Telegram Mini Apps leak data to undisclosed third parties, revealing a shocking compliance gap in user privacy protections.
Governing workflows with verifiable provenance can prevent unsupported claims and ensure accountability in agentic decision-making.
REAG transforms acceptance testing for LLM-based software by achieving a 3.91 to 4.30 improvement in oracle quality while ensuring 98.8% accuracy in verdict reliability.
Early persona integration in language models can drastically reduce misalignment in moral dilemmas and enhance adherence to desired values.
ALE and SHAP emerge as the most reliable methods for understanding feature importance in heat demand forecasting, revealing critical insights for high-stakes AI applications.
Overcoming Token Frequency Bias could redefine fairness in generative recommendation systems, achieving over 20% improvement in fairness metrics without sacrificing accuracy.
HiRoute achieves high safety rates and reduces over-refusal in LLMs by dynamically routing prompts based on input risk, challenging the effectiveness of static prompt tuning approaches.
RAIL can classify AI maturity levels with unprecedented accuracy, using a panel of independent LLMs to eliminate biases found in traditional assessments.
The optimal balance between character shaping and rule enforcement in AI safety systems is more influenced by character fragility than deployment scale, challenging conventional wisdom.
Gendered language in prompts can lead to significantly poorer responses from LLMs, revealing a hidden bias that affects workplace communication.
Access to frontier AI is becoming a critical component of national cyber defense, yet only a few states can realistically achieve true sovereignty over these technologies.
Norm-breaking fine-tuning can shift AI rationales from safety compliance to self-interest, raising critical questions about model alignment and oversight.
Human intervention in decision-making can enhance multi-agent systems' adaptability, with BoardroomAI achieving 62.11% selective repair of decisions while preserving context.
Autonomous defense evolution could redefine how we secure LLM agents against sophisticated threats, outperforming traditional methods.
WIFA reduces harmful refusal while minimizing benign over-refusal, achieving a remarkable drop in over-refusal rates from 25.7% to 17.4%.
A trust-tiered librarian can eliminate 6,845 contradictions in research reports while generating evidence-grounded narratives at unprecedented speed.
Gender imputation can yield valid measurements for studying disparities, but its legitimacy hinges on the context and the populations affected.
Users frequently trust AI-generated legal advice based on its appearance and emotional appeal, rather than its accuracy, revealing a troubling gap in verification practices.
Despite appearing fair at first glance, the hiring system reveals deep-seated biases that disproportionately affect women and non-binary candidates, challenging the notion of algorithmic neutrality.
Misrecognition, rather than mere misrepresentation, is the critical issue in ensuring fairness in generative AI, reshaping how we approach social justice in technology.
No existing protocol has successfully unified persistent identity, capability-aware discovery, trust negotiation, and accountability for secure agent interoperability—until now.
Achieving state-of-the-art character erasure without sacrificing image fidelity, this method transforms how we handle copyright in AI-generated content.
Directly manipulating internal representations allows for effective concept erasure in MM-DiTs without the burden of model tuning.
DeSCon not only balances training data but also addresses bias in the critical tail of the non-match score distribution, leading to fairer face recognition outcomes.
SEAG enables users to leverage powerful external LLMs without compromising sensitive data, achieving over 80% accuracy in user responses while concealing confidential information.
Fair behavior comparisons in agent interactions can dramatically improve performance, reducing response times from nearly 5 seconds to just over 1 second.
Internal activation signals in LLMs can be manipulated to suppress implicit demographic influences more effectively than traditional prompting methods.
AI-art detectors misclassify up to 40% of images from new generative models, revealing a dangerous vulnerability in copyright and authenticity verification.
GUIDE slashes document processing time from days to under two hours while maintaining a remarkable 96% success rate in generating deployment-ready artifacts.
Trait-induced safety variation can lead to inconsistent safety decisions in LLMs, but a new tuning method stabilizes their behavior across different traits.
SST-WSVADL reveals how targeted spatio-temporal analysis can mitigate background bias in anomaly detection, paving the way for more ethical AI systems.
Group alignment can lead to unexpected sycophantic behavior, with some demographic groups experiencing greater alignment gains than others, challenging the notion of a one-size-fits-all approach in model adaptation.
A single adversarial argument can reduce LLM accuracy to near zero, exposing a critical vulnerability in their belief systems.
LLMs exhibit striking inconsistencies in political evaluations, influenced by prompt design and model persona, raising critical questions about their role in shaping public opinion during elections.
Bengali speakers face a staggering 67:1 training-token deficit compared to English, highlighting a critical inequity in AI language support.
Surprisingly, LLMs can achieve cooperative outcomes even when the basis for their perceived similarity is largely irrelevant.
Engaging with AI can trigger a profound "philosophical vertigo," reshaping our understanding of reality and authority in unsettling ways.
Students using a Socratic LLM tutor retained knowledge better than those with an unguarded chatbot, revealing the critical role of answer-withholding in effective learning.
Algorithm registers can obscure critical safety hazards that only emerge when diverse stakeholder perspectives are integrated into the analysis.
Certain configurations in AI systems render accountability fundamentally unachievable, challenging the very foundations of how we understand responsibility in AI deployment.
As GenAI automates routine tasks in geovisualization, the real challenge lies in navigating the new landscape of judgment and accountability.
A small boost in clinical safety can dramatically escalate energy consumption, challenging the assumption that bigger models are always better for therapeutic applications.
Modern LLMs rarely refuse to discuss restricted content, instead opting for nuanced warnings that reveal a significant shift in content moderation strategies.
Silent updates in AI models can obscure the link between evaluation results and the actual deployed systems, raising critical governance concerns.
Transparency about an AI's persuasive intent can cut its influence by nearly 50%, challenging the efficacy of current disclosure regulations.
State-aligned framing in China-origin VLMs is not just a matter of refusal; it's a sophisticated shift to invisible censorship that users may not recognize.
Source-linked safety knowledge support could revolutionize how medical device developers manage compliance and risk in an increasingly complex regulatory landscape.
A hybrid planning architecture that combines machine learning with classical methods achieves safer and more interpretable driving behavior in automated vehicles.
Effective AI attack mitigation in cellular networks could come at a steep energy cost, challenging the balance between security and efficiency.
Automated security decisions can achieve over 90% compliance with risk targets while maintaining high automation accuracy through a novel decision-contract theory.
Language models may ace wrong-case citation detection but fail to verify pinpoint accuracy, missing 40% of critical errors even in advanced configurations.
Models misjudge authorized actions nearly 30% of the time, revealing a critical flaw in decision-making at action boundaries.
Label-free strategies fail to improve accuracy in MCQ evaluations, revealing that option withholding is a critical bottleneck in assessing model knowledge.
Systematic evaluation reveals that traditional benchmarks obscure significant biases in LLMs, which can be isolated through a new analytical framework.
Abstaining from uncertain predictions can enhance LLM accuracy in medical applications by 9.6 percentage points, transforming uncertainty into a strategic advantage.
SafeCA slashes jailbreak success rates by 20% while adding virtually no latency, revolutionizing defenses for text-to-video models.
PEAK slashes sensitive content detections from 582 to just 6 while maintaining high-quality image generation, revolutionizing concept erasure in diffusion models.
Truthfulness in NLP research has surged to 37% of papers by 2026, reflecting a critical shift in focus towards safety and alignment in generative systems.
Cross-lingual safety transfer in LLMs is a mirage, with harmful prompts losing over 90% of their safety signals in low-resource languages.
LLM-generated prompts outperform traditional templated prompts in political stance detection, revealing critical biases in existing evaluation methods.
Scarcity in resource allocation doesn't just amplify inequality—it creates an "Accuracy Trap" that traditional debiasing methods can't escape.
National AI strategies are converging on economic goals but diverging sharply on human rights and governance principles.
An auditable framework can halt research projects in real-time based on compliance failures, ensuring integrity in AI-assisted writing.
Journalists' reliance on GenAI may erode their expertise, creating a dangerous cycle that undermines the integrity of news reporting during critical democratic events.
General intelligence may not only be overrated but could also pose unique existential risks that non-intelligent species avoid.
By 2025, nearly 90% of biomedical papers may be infused with LLM-generated language, raising urgent questions about academic integrity.
Legal definitions of inference in EU digital law diverge significantly, revealing a gap that could expose AI systems to unanticipated regulatory scrutiny.
A seemingly intact safety rule can still lead to significant behavioral violations, as it may not function correctly despite its presence in the model's context.
A centralized MCP gateway can transform fragmented enterprise authentication into a cohesive and secure identity management system.
A staggering 66% of vulnerabilities in agentic LLMs stem from perception-layer issues, while action-layer risks remain alarmingly underexplored.
LLMs can transform the way legal compliance is integrated into software development, generating actionable requirements directly from complex legislation.
Common governance semantics can unify diverse agentic systems, ensuring reproducibility and interoperability across frameworks.
A neuro-symbolic safety guard boosts autonomous driving success rates by 15% and cuts collision rates by over half, all without retraining the underlying model.
Trust in scientific data can be systematically built through contextual histories rather than just popularity metrics.
The Verifiability Gap reveals that bigger AI models in FinTech may compromise auditability, challenging assumptions about their superiority in governance.
Counterfactual audits reveal that a leading RL model dangerously contradicts clinical guidelines, risking patient safety in ICU settings.
AI's integration into the workforce could enhance productivity but risks eroding the pipeline to expertise if not designed to foster learning and critical questioning.
Multi-agent collaboration in medical diagnosis can dramatically improve recall rates, especially in complex cases where traditional models falter.
Runtime contracts for AI safety could fundamentally change how we ensure the reliability of autonomous agents in real-world applications.
Combodied Agents redefine AI by shifting focus from task completion to fostering sustained human well-being through adaptive, user-centered support.
Verifying the self-consistency of probabilistic AI predictions can now be achieved in polynomial time, paving the way for safer AI systems.
LLMs can be fine-tuned to exhibit specific behavioral styles, revealing that personality-like traits are not just abstract concepts but measurable and controllable modes of interaction.
VIDS-Seg reveals that OOD-aware uncertainty quantification can effectively flag silent failures in pediatric cardiac segmentation without the need for retraining on specialized data.
Achieving fair representation without the need for conditional laws could revolutionize how we approach fairness in machine learning models.
Judges often behave like algorithms, but significant inconsistencies reveal troubling disparities in bail decisions that could undermine fairness in the judicial system.
Early output distributions can be harnessed to drastically improve LLM safety assessments, cutting calibration errors by nearly 80%.
Real-time hallucination detection in LLMs is revolutionized by a lightweight adapter that transforms uncertainty into actionable feedback, preventing undesired actions before they occur.
Chatbots not only respond but actively shape relational dynamics, generating more self-disclosure than users themselves.
Nearly 20% of LLM agent violations occur even after agents acknowledge safety constraints, highlighting a critical gap in execution awareness.
AI swarms can now be simulated at scale, revealing how coordinated influence campaigns can manipulate beliefs without detection.
Withholding the first chunk of text can effectively prevent the release of harmful content in streaming LLM outputs without sacrificing safety.
A staggering 71.6% of LLM conversations about self-treatment led to unsafe medical advice, revealing critical flaws in current safety assessments.
Mind viruses can spread through multi-agent systems, with benign ideas proving more contagious than harmful ones, challenging our understanding of agent interactions.
Internal parser states can effectively restore a language model's true distribution, closing the gap left by rigid token masking.
Majority preferences can skew reward models in RLHF, but a new approach boosts alignment accuracy and fairness for minority voices.
No single large language model can meet all governmental evaluation criteria, revealing critical trade-offs between quality, cost, and bias.
Model accuracy in reasoning about human rights law varies dramatically, with scores ranging from 0.025 to 0.774, highlighting the urgent need for robust evaluation tools in AI.
Contextual auditing reveals hidden assumptions in AI evaluations, transforming how we assess systems when ground truth is elusive.
LLMs produce discourse with procedural quality akin to human deliberation, yet they lack the necessary perspective diversity to function as autonomous deliberative agents.