Search papers, labs, and topics across Lattice.
MonitrLLM is an open-source evaluation infrastructure designed to connect full conversation transcripts of large language model (LLM) interactions with user-defined task intents and outcomes. This approach addresses a significant gap in LLM evaluation by treating interaction trajectories and user feedback as primary signals, rather than optional metadata. A pilot study with 26 college students using ChatGPT revealed that while users reported high satisfaction, there was a notable 23.1% failure rate in achieving their goals, highlighting the need for more nuanced evaluation methods.
Despite high user satisfaction with LLM interactions, a staggering 23.1% of goal tasks failed, challenging assumptions about conversational success.
Benchmark suites assess model capability on controlled tasks; large-scale conversation corpora capture naturalistic use without user feedback; and in-interface feedback mechanisms record satisfaction without task purpose. Together, they leave a critical gap in LLM evaluation: no existing infrastructure routinely links interaction trajectories to user-defined outcomes. We introduce MonitrLLM, open-source infrastructure for community-centered LLM evaluations that links full conversation transcripts to user-reported task intent and outcome assessments, treating all three as primary evaluative signals rather than optional metadata. To demonstrate the value of this approach, we conducted a two-week feasibility pilot with 26 college students using ChatGPT, collecting 206 evaluation reports with full conversation transcripts. The findings from our pilot demonstrate the value of connecting conversation trajectories with user-reported outcomes. For instance, despite reporting high average satisfaction (4.19/5) with their LLM interactions, participants also experience a substantial 23.1% failure rate on their goal tasks. We also find that multi-turn conversations are reported as failing at 2.5 times the rate of single-turn exchanges, a pattern that reframes extended interaction as a signal of difficulty rather than engagement. We conclude by discussing the value of incorporating direct user feedback with observational data for robust LLM evaluations, and the possibilities for infrastructure that enables this goal.