Search papers, labs, and topics across Lattice.
Terminal-Universe resolves the scarcity of executable code sandboxes by inverting static agent trajectory logs to reconstruct functional, interactive terminal workspaces. The framework restores pre-execution file trees via tool-operation replay, infills missing dependencies with an LLM agent, and expands these environments into cross-repository and multi-turn interaction tasks. Applied to public trajectories, the pipeline yielded 37.3k task-sufficient environments whose training data boosted Qwen3.5-27B performance by 11.9 points on Terminal-Bench 2.1 and 13.8 points on EvoCode-Bench v2 MT@4.
Rather than treating agent trajectories as dead post-training demonstrations, researchers can now resurrect thousands of fully executable, verifiable terminal environments directly from tool-execution logs.
As terminal-based code agents become prevalent, agent trajectories have accumulated at scale, while realistic, executable environments remain scarce. However, environments are what agent post-training actually requires: each can be re-queried into many verifiable tasks and provides execution feedback, whereas a trajectory is a single frozen demonstration. Rather than generating environments from scratch, we observe that the tool-execution history in existing trajectories exposes the structure and contents of the environments in which they ran, making it possible to reconstruct those environments from the trajectories themselves. Thus, we introduce Terminal-Universe, a framework which turns each trajectory into a reusable environment and explores it for synthesizing new tasks and continued interactions. Specifically, Terminal-Universe replays the file operations recorded in a trajectory to restore each file before the agent modified it, yielding a partial workspace; a completion agent then supplies the missing files and dependencies. On this recovered workspace, we both reconstruct the original intent task and synthesize entirely new ones. Besides, we also scale the tasks along two complementary axes: breadth and depth. For breadth, we mine directional dependency relations between related environments and synthesize cross-workspace queries spanning multiple codebases, as developers routinely do in real-world development. For depth, we extend the initial single-turn query into a multi-round session that captures iterative user feedback and requirement refinement via a user agent. Applied to public terminal agent trajectories, Terminal-Universe produces 37.3k task-sufficient environments. Supervised fine-tuning of Qwen3.5-27B on this corpus improves single-round performance on Terminal-Bench 2.1 by 11.9 points and multi-round performance on EvoCode-Bench v2 MT@4 by 13.8 points.