Search papers, labs, and topics across Lattice.
22 papers from Google Research on Constitutional AI & AI Ethics
Scholarly research on generative AI often overlooks the rich cultural practices and structural conditions affecting Black communities, reducing their experiences to mere data points.
Continuous state optimization can slash movement costs by over 300% while maintaining performance in complex learning systems.
AMS reveals that safety training modifications can significantly alter the activation landscape of language models, impacting their compliance with safety protocols.
Foundation model agents can achieve stable cooperation in social dilemmas by inferring behavioral similarities, defying classical game theory's expectations of mutual defection.
Existing safety evaluations miss critical vulnerabilities in LLMs used in K-12 education, revealing that dynamic interactions pose unique risks that current guardrails fail to mitigate.
Safety fine-tuning in LLMs not only suppresses self-attribution of consciousness but also inadvertently erases culturally significant beliefs about non-human entities and spirituality.
Ignoring the social dynamics that led to past disasters could doom AI systems to repeat history's mistakes.
Jointly verifying user intent and response harm can reduce attack success rates to just 4.1%, setting a new standard for LLM safety.
Ignoring the interplay between engineering, business, and legal perspectives can doom privacy-enhancing technologies to failure in software systems.
A novel auditing framework reveals that synthetic data can leak private information without model access, challenging assumptions about data privacy in generative AI.
Redefining data work through a reparative lens reveals the urgent need to prioritize the voices of those harmed by online systems, challenging existing norms of accountability in AI.
In high-stakes health contexts, stakeholders demand that trustworthiness in AI systems be inspectable, not just asserted, reshaping how we design health information tools.
A strategic messaging shift on Google Search reduced CSAM-related queries by 3.8%, effectively redirecting some users towards therapeutic resources.
Even the best LLM judges miss cultural faux pas that are obvious to locals, achieving only 52% F1 score on a new benchmark.
Multilingual LLMs exhibit a surprising "American bias," even when prompted in other languages, and instruction tuning makes it worse.
GAAP offers a deterministic, trust-minimized approach to AI agent security, safeguarding user data even when models are compromised or prompts are injected.
Ethics interventions in AI development often fail because practitioners don't trust them – here's a breakdown of why, and how to fix it.
Safety fine-tuning might inadvertently be stripping LLMs of their ability to understand non-human minds and entertain spiritual beliefs, even while preserving Theory of Mind.
Despite the effort required, Android developers overwhelmingly support platform-level changes to combat fingerprinting, suggesting a path to enhanced user privacy through collaborative platform-developer initiatives.
LLMs are becoming "epistemic agents" that shape our knowledge environment, so we need a new framework for evaluating and governing them based on trustworthiness, not just performance.
Finally, a framework to quantify AI's cultural intelligence, moving beyond ad-hoc cultural benchmarks to a systematic, extensible, and theoretically grounded approach.
DPO's success isn't just clever engineering—it's deeply rooted in human choice theory, unlocking a surprisingly flexible framework for preference optimization and justifying many DPO extensions.