Mar 4, 2026arXiv:2603.03915

Rethinking Role-Playing Evaluation: Anonymous Benchmarking and a Systematic Study of Personality Effects

AI Summary

The paper introduces an anonymous evaluation method for Role-Playing Agents (RPAs) to mitigate the bias introduced by models relying on memorized information associated with famous fictional characters. Experiments show that anonymization significantly degrades role-playing performance, highlighting the impact of name exposure. The authors then demonstrate that incorporating personality information, especially self-generated personalities, can effectively enhance RPA performance in the anonymous setting, achieving results comparable to using human-annotated personalities.

Key Contribution

LLMs role-play worse when they can't rely on character names, but surprisingly, self-generated personalities can restore performance to near human-annotated levels.

Abstract

Large language models (LLMs) have demonstrated significant potential in developing Role-Playing Agents (RPAs). However, current research primarily evaluates RPAs using famous fictional characters, allowing models to rely on memory associated with character names. This dependency creates a bias that limits the generalization of RPAs to unseen personas. To address this issue, we propose an anonymous evaluation method. Experiments across multiple benchmarks reveal that anonymization significantly degrades role-playing performance, confirming that name exposure carries implicit information. Furthermore, we investigate personality augmentation to enhance role fidelity under anonymous setting. We systematically compare the efficacy of personality traits derived from human annotations versus those self-generated by the model. Our results demonstrate that incorporating personality information consistently improves RPA performance. Crucially, self-generated personalities achieve performance comparable to human-annotated ones. This work establishes a fairer evaluation protocol and validates a scalable, personality-enhanced framework for constructing robust RPAs.

Eval Frameworks & Benchmarks Natural Language Processing Tool Use & Agents

Citation Metrics

Citations0

Influential citations0

References0

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

Rethinking Role-Playing Evaluation: Anonymous Benchmarking and a Systematic Study of Personality Effects

Related Papers