Search papers, labs, and topics across Lattice.
The paper introduces user-turn generation as a novel probe to evaluate interaction awareness in LLMs, going beyond the standard assistant-turn paradigm. Experiments across 11 open-weight LLMs and 5 datasets reveal that interaction awareness is decoupled from task accuracy, with even high-performing models showing near-zero follow-up rates under deterministic generation. However, higher temperature sampling and collaboration-oriented post-training can elicit and improve interaction awareness, suggesting it exists latently within the models.
Even the most powerful LLMs often lack genuine interaction awareness, blindly generating responses without anticipating or reacting to subsequent user actions.
Standard LLM benchmarks evaluate the assistant turn: the model generates a response to an input, a verifier scores correctness, and the analysis ends. This paradigm leaves unmeasured whether the LLM encodes any awareness of what follows the assistant response. We propose user-turn generation as a probe of this gap: given a conversation context of user query and assistant response, we let a model generate under the user role. If the model's weights encode interaction awareness, the generated user turn will be a grounded follow-up that reacts to the preceding context. Through experiments across 11 open-weight LLMs (Qwen3.5, gpt-oss, GLM) and 5 datasets (math reasoning, instruction following, conversation), we show that interaction awareness is decoupled from task accuracy. In particular, within the Qwen3.5 family, GSM8K accuracy scales from 41% (0.8B) to 96.8% (397B-A17B), yet genuine follow-up rates under deterministic generation remain near zero. In contrast, higher temperature sampling reveals interaction awareness is latent with follow up rates reaching 22%. Controlled perturbations validate that the proposed probe measures a real property of the model, and collaboration-oriented post-training on Qwen3.5-2B demonstrates an increase in follow-up rates. Our results show that user-turn generation captures a dimension of LLM behavior, interaction awareness, that is unexplored and invisible with current assistant-only benchmarks.