Search papers, labs, and topics across Lattice.
This paper introduces a comprehensive framework for assessing the behavioral fidelity of long-horizon human activity simulations, addressing a significant gap in the evaluation of LLM-based human simulators. By analyzing a 43-hour dataset of real office activities, the authors compare various conditioning mechanisms, revealing that while statistical priors yield distributions closest to real behavior, they also lead to fragmented routines and reduced variability. The study emphasizes the need for a multi-faceted evaluation approach that considers different metrics and temporal scales to better capture the complexity of human behavior.
Statistical priors may mimic real-world activity patterns but risk oversimplifying human behavior by fragmenting routines and limiting variability.
As LLM-based human simulators are increasingly used for policy, evaluation, and training, they must faithfully reproduce real behavioral patterns. While prior work has examined behavioral fidelity in survey responses and dialogue, longer-horizon real-world activity remains largely unexplored. We introduce a framework for evaluating behavioral fidelity in long-horizon activity simulations across temporal granularities and levels of analysis. As a case study, we collect a 43-hour multi-camera dataset of in-the-wild office activity and compare trace-derived conditioning mechanisms: persona descriptors, few-shot exemplars, and statistical transition and time-of-day priors. We find that behavioral fidelity is not uniform across metrics: statistical priors bring activity and sequence distributions closest to real behavior, yet over-fragment routines and suppress within-person variability. These findings motivate a more holistic evaluation that spans multiple metrics, temporal granularities, and levels of analysis.