Search papers, labs, and topics across Lattice.
This study systematically audits open-weight language models using persona vectors, revealing a comprehensive inventory of 53 traits across four behavioral domains. The findings indicate that while both models exhibit a default helpfulness aligned with expert judgments, steering can significantly enhance traits typically suppressed, such as hyperbole and hallucination. Notably, the research highlights that persona vectors serve as effective probes for understanding behavioral organization rather than merely as controls for model output.
Persona vectors expose the hidden complexities of LLM behavior, revealing that steering can unlock traits like hyperbole and hallucination that models typically suppress.
What a language model will and will not do is largely set during post-training, but which behaviors it expresses, hides, or resists is not revealed by prompting alone. Persona vectors, behavioral directions in activation space, can probe this organization, but prior work covers only a handful of traits. We present the first systematic application of persona vectors at this scale, compiling a 53-trait inventory across four behaviorally distinct domains and labeling every trait in two open-weight models as natural (expressed at baseline), steerable latent but amplifiable, or intractable (resistant to standard extraction). Both models default to helpful, task-oriented behavior: all nine agentic traits are natural, and their default clinician behavior matches a board-certified psychologist's independent desirability judgments on 16 of 17 traits. Steering produces its largest gains on traits these defaults exclude: hyperbole, hallucination, and sycophancy. The same asymmetry holds across all 171 generic-trait pairs: two steerable traits can collapse the composition, but pairs involving a default never do. Where standard extraction fails on a trait like"evil,"a vector transferred from a fine-tuned variant still recovers it, with the residual refusals appearing inside the model's chain-of-thought. Persona vectors are most informative not as a set of controls but as a probe of behavioral organization.