Search papers, labs, and topics across Lattice.
12
0
11
0
LLMs' personalities are dynamic and layer-dependent, revealing that quantization can significantly disrupt their behavioral consistency.
Identity drift in generative agents reveals that anti-self-deception is a dominant modification behavior, challenging assumptions about agent fidelity under pressure.
Simulating 8.3 billion diverse personas reveals nuanced user interactions that traditional evaluations miss, transforming how we assess AI systems.
TREK transforms the way models tackle challenging prompts by expanding their exploration support, leading to substantial performance gains even in the hardest task scenarios.
Role-typed credit assignment can drastically improve reinforcement learning outcomes by accurately distinguishing between useful exploration and regression in agent actions.
Integrating real images into the GRPO process and using dual-reward guidance allows PortraitGen to achieve unprecedented levels of photorealism while effectively suppressing AI artifacts.
Trajectory mining reveals skill structures but fails to translate these insights into meaningful performance gains for downstream policies.
WeGenBench exposes the hidden deficiencies of text-to-image models, revealing that many leading systems struggle with specific generation tasks despite overall high performance.
LLMs show significant variability in the actionability of their UX critiques, with some models outperforming others across different product categories.
Agent Development Kits vary dramatically in usability, with some enabling agents to outperform general-purpose coding tools at a fraction of the cost.
Explicitly invoking external image tools in vision-language models dramatically reduces jailbreak success rates, even when the tool's output is overridden or unsafe.
MLLMs still struggle to reason about everyday situations when they require identifying and using visual clues, despite excelling at tasks relying on pre-existing knowledge.