Search papers, labs, and topics across Lattice.
100 papers published across 5 labs.
Hallucinations in LLMs may blur the line between advanced simulation and true cognitive processes, challenging our understanding of AI consciousness.
CLEAR slashes harmful completions from 32.3% to just 0.5% while boosting utility performance, redefining the safety-utility balance in LLMs.
Legal and moral agency in AI are not just blurred lines—they're fundamentally different dimensions that can reshape accountability in AI actions.
Technical debt in software development is escalating, and without a shift to integrated testing and specification, we risk compounding vulnerabilities in AI systems.
In a post-AGI world, human welfare may hinge entirely on ownership shares in a corporate economy dominated by machine agents, rendering employment policies obsolete.
CLEAR slashes harmful completions from 32.3% to just 0.5% while boosting utility performance, redefining the safety-utility balance in LLMs.
Legal and moral agency in AI are not just blurred lines—they're fundamentally different dimensions that can reshape accountability in AI actions.
Technical debt in software development is escalating, and without a shift to integrated testing and specification, we risk compounding vulnerabilities in AI systems.
In a post-AGI world, human welfare may hinge entirely on ownership shares in a corporate economy dominated by machine agents, rendering employment policies obsolete.
Thoughtful resistance in AI content creation can significantly elevate educational quality, proving that educators and AI can be powerful allies rather than adversaries.
Social proof can tip entire populations into collective overreliance on AI, but strategic feedback designs can reverse this trend.
Unlearning in LLMs is not just about forgetting facts; it’s about mastering the nuanced balance of harmful and benign concept usage, and current methods fall short.
Cross-lingual fairness gaps in language model watermarking are not just language-specific but are fundamentally tied to the structural properties of language families.
Language models can reconstruct sensitive user data from innocuous outputs, with 2-digit secrets retrieved almost perfectly and 4-digit secrets at 82% accuracy, raising serious privacy concerns.
Safety Nets can shrink AI system sizes by three orders of magnitude while guaranteeing 100% correct outputs, making them a game-changer for certifying AI in aviation.
Human-mediated AI guidance can transform how families engage with emergency preparedness, making it more interactive and tailored to children's needs.
Treating specification changes as the primary unit of change can drastically reduce deployment times and defects in data platforms.
SafeBranch enables embodied agents to achieve ten times more safe task completions without sacrificing performance, even in novel environments.
Memory Correlation Bias can mislead multi-agent systems into false majorities, but CAMA effectively counters this by recovering independent evidence from correlated memories.
Compression can lead to a significant loss of critical knowledge while leaving models confidently incorrect, revealing hidden biases that standard metrics fail to detect.
Understanding in AI safety decisions isn't just a checkbox—it's a generative force that shapes engineering outcomes.
Static profiles lead to identity essentialism in LLMs, but a new longitudinal memory framework reveals a path to richer, more diverse social simulations.
The verification gap in Physical AI reveals a critical asymmetry in how evidence is utilized, challenging conventional approaches to proposal execution.
PolicyGuide boosts compliance rates in LLM agents by transforming policy checks into a proactive, workflow-guided system that adapts to user interactions.
Achieving competitive skin disease classification while ensuring fairness across skin tones, MIFR aligns clinical and dermoscopic data in a shared embedding space.
A curriculum that intelligently adapts to diverse user preferences can double population satisfaction while slashing training time.
TestifAI can predict the robustness of deep learning models against complex perturbations with remarkable accuracy while slashing testing costs by up to 80%.
Achieving 96% decision precision in leak localization transforms how utilities can confidently deploy resources, minimizing unnecessary excavations.
Safety alignment in VLMs can suppress grounded visual reasoning, yet visual evidence remains influential during refusal, revealing a complex interplay between safety and perception.
Social.Wiki redefines web ownership by enabling communities to collaboratively create and govern their online spaces without centralized control.
Clients transformed their view of AI from a mere tool for answers to a collaborative partner in problem-solving.
Balancing hate speech detection and user privacy is not just a challenge—it's a necessary trade-off that could redefine online safety standards.
Hallucinations in LLMs may blur the line between advanced simulation and true cognitive processes, challenging our understanding of AI consciousness.
A single attention head in the Mistral-7B model captures demographic identity with surprising fidelity, yet its causal use reveals a disconnect that complicates LLMs' ability to simulate real-world populations.
The review reveals five distinct methods for quantifying the carbon footprint of web advertising, paving the way for more rigorous environmental accountability in digital marketing.
Six incorrect references and three legal misqualifications highlight the critical need for precision in regulatory cross-references within the EU AI Act.
Generative AI doesn't just reflect biases; it embeds a dominant cultural epistemology that marginalizes minority voices at the very foundation of knowledge creation.
A novel framework reveals that psychological safety risks in autonomous vehicles can be systematically assessed and prioritized, bridging the gap between human factors and technology.
Over half of the bounty listings on RentAHuman impose high proof burdens, often requiring sensitive personal information or physical actions from workers.
Trust domains required for protected execution can exceed expectations, revealing that five domains may be necessary for certain high-risk automated systems. WHY_IT MATTERS: This insight challenges existing assumptions about authority in automated systems and could significantly influence the design of security protocols in high-stakes environments.
A striking trade-off emerges: ensuring differential privacy in voting can render computationally simple tasks intractable, especially for STV.
Expert corrections to LLM errors often vanish after a session, but a new operating model could ensure these insights persist and improve AI reliability.
Attention pooling enhances performance but paradoxically amplifies subgroup disparities, challenging assumptions about fairness in model adaptation.
Associative context retrieval can dramatically amplify the effectiveness of white-box attacks on LLMs without compromising their general performance.
RGE reveals that long-horizon agents can drift significantly from their intended tasks, even while appearing compliant at each step, highlighting the need for deeper oversight mechanisms.
Scalar metrics fail to capture the true diversity of AI-generated content, but diversity profiles offer a robust, multi-dimensional evaluation framework that reveals hidden biases.
Evidence representation, not model choice, is the key to improving explanation quality in credit risk decision-making.
Targeted memory manipulation in LLM-agent communities can escalate group polarization, revealing a new vector for social influence attacks.
Work-related orientation significantly enhances human direction in AI tasks, but the mode of interaction alters the dynamics of this relationship.
Complexity in AI and robotics challenges can significantly undermine community readiness, with a one-point increase in complexity correlating to a 0.21-point drop in perceived preparedness.
Achieving the best spurious-mitigation performance in medical imaging, SpurCon reshapes representation geometry to enhance model reliability with minimal retraining.
HarnessRisk reveals that up to 80.9% of adversarial attacks can succeed in agent harnesses, even when risk detection is high.
Current AI systems struggle to conduct independent scientific research, with performance plummeting by nearly 50% when human guidance is removed.
Valid inference from AI-generated data is possible without gold-standard labels, challenging the conventional reliance on costly benchmarks.
A novel two-threshold framework allows LLMs to judge outputs with formal control over reliability, achieving higher coverage without compromising error rates.
No single LLM excels across all governance dimensions, revealing critical trade-offs that public institutions must navigate in model selection.
The phrasing of belief expressions can swing LLM accuracy from a +50% boost to a -14% drop, revealing a critical vulnerability in how models navigate user beliefs and facts.
Keyword searches on TikTok reveal up to 56% harmful content, significantly outpacing passive scrolling results and challenging existing moderation assumptions.
Reflex-Guard filters harmful prompts with 95.9% accuracy in just 37.6 ms, revolutionizing real-time safety for LLMs.
Answer format can dramatically alter gender bias measurements in LLMs, sometimes reversing bias rankings entirely.
Media bias analysis just got a boost—this LLM-based pipeline reveals hidden patterns in how news sources frame events and describe public figures.
Generative AI is reshaping workplace interactions by creating a veil of opacity around human effort, complicating trust and collaboration.
Population-level validation of CGM AI tools hides critical disparities, with T1D patients facing 6 mg/dL higher prediction errors than T2D counterparts.
Trust in AI-driven decisions can be systematically documented, ensuring accountability at the crucial output-to-action boundary in bioscience research.
Cybersecurity education in Australia risks perpetuating inequity by ignoring the unique challenges faced by women and CALD communities, with four key barriers identified that hinder inclusivity.
Task-conditioned authority selection reduces excess-authority errors in tool-using agents from 4.56% to 0.79%, showcasing a powerful new layer of control.
MemCatalyst reveals that targeted data poisoning can drastically boost membership inference accuracy in Vision-Language Models with minimal resource expenditure.
All tested GUI agents are alarmingly vulnerable to environmental injection attacks, with success rates reaching over 66%, revealing a pressing need for improved safety measures.
MLLMs can be tricked into unsafe behavior by seemingly harmless prompts and images, but COMIC effectively mitigates this risk by focusing on the operation-target relationship.
Attack rankings shift dramatically when evaluated under shared target-call budgets, revealing hidden efficiencies in traditional methods.
PACE achieves a flawless safety record in DeFi transactions, eliminating unsafe executions while leveraging LLMs for complex financial operations.
Configurable privacy-preserving MRI processing can balance the trade-off between anatomical detail and privacy, revolutionizing neuroimaging practices.
Fair-ordering protocols can ensure that conflicting requests are invalidated, even when the policy state is hidden and unrecoverable.
Counterfactual recourse in education becomes truly actionable when recommendations are not just model-valid but also semantically feasible and machine-checkable.
Clinical translation standards could revolutionize how we ensure the reliability of machine learning systems.
Redakto achieves PII redaction while preserving the utility of text for LLM applications, addressing critical privacy concerns in the wake of new EU regulations.
Traditional risk assessment fails for AI, but a new capability-based planning framework reveals actionable insights for crisis preparedness.
International crises can synchronize political narratives across media, but national policies still drive significant divergence in coverage.
Accident retrieval using LLMs can significantly enhance highway construction safety planning, achieving over 75% accuracy in incident classification.
LLMs can generate structurally sound legal analyses but often fall short in substantive reasoning, raising questions about their reliability in legal contexts.
Achieving a 15.4% shift in LLM sycophancy control with a method that ensures predictable and gradual adjustments could redefine user interactions with AI.
Reasoning-capable models can significantly reduce the impact of misinformation in RAG systems without the heavy computational costs of isolation.
QVIRL achieves robust apprenticeship learning from raw pixel data while quantifying uncertainty, a breakthrough for safety-critical AI applications.
Human preferences in ethical decision-making for autonomous vehicles reveal a troubling preference for self-sacrifice over minimizing casualties, challenging traditional ethical frameworks.
Quipu achieves zero defects in knowledge representation by enforcing strict governance and trust mechanisms, revolutionizing how agents interact with knowledge graphs.
Collective dynamics of AI agents reveal surprising patterns: while communication boosts accuracy on objective tasks, it can lead to political bias in group opinions.
CUBICS reveals that context-aware performance estimation can significantly improve the reliability of safety-critical ML components by avoiding oversimplified failure models.
Human-LLM collaboration can significantly enhance the accuracy and depth of privacy risk assessments in complex welfare schemes, revealing critical insights that neither could achieve alone.
Mandatory GenAI labeling may be more about regulatory optics than effective governance, risking the very innovation it aims to control.
Uncertainty in AI decision-making can now be visualized and translated into actionable oversight responses, making it explicit rather than implicit.
Trust-preserving agentic AI can achieve an impressive 86.9% task completion rate while intervening in nearly all policy violations.
CAPO enables LLMs to meet stringent operational constraints without sacrificing task performance, achieving feasible prompts in every evaluated domain.
User motivations for keeping GPT-4o reveal that effective model replacement hinges on maintaining established user value, not just technical upgrades.
BabelSteering boosts harmful request refusals across multiple languages by an average of 11 percentage points, all while maintaining task performance.
Analyst bias can be effectively eliminated in qualitative research by using a domain-specialized LLM that structures insights directly from data.
Non-custodial digital assets can achieve regulatory compliance without sacrificing user privacy, fundamentally reshaping the landscape of digital payments.
Users can identify AI responses as "Claudish" even without knowing the underlying model, revealing a deeper layer of AI character recognition that transcends mere identification.
A single query to a large language model could cost the planet $0.4, underscoring the urgent need to consider AI's environmental footprint.
Agentic flooding is already straining government services, with LLMs generating demand surges that could outpace current capacities.
Proprietary LLMs can match human educators in delivering relevant answers from curated lecture videos, potentially transforming how programming education leverages AI.
Identity-conditioned prompts can lead to significant differences in the readability of LLM-generated robot design descriptions, raising critical questions about bias in AI outputs.
A single authorization rule can effectively prevent memory leakage across different audience groups in language agents, ensuring that sensitive information remains compartmentalized.
Stripped of safety, language models can be tricked into generating misleading yet confident responses, with up to 90% of outputs being decoys under attack.
Achieving over 27,000 requests per second, this system redefines the durability-latency trade-off in AI audit records, but raises questions about security and compliance.
GEO-optimized content is more prevalent than expected, with nearly 9% of web pages showing signs of manipulation, raising alarms about the integrity of information in generative search engines.