Search papers, labs, and topics across Lattice.
This study investigates the linear representations of temporal horizons in the Qwen3-32B language model to manipulate its time-related preferences and recommendations. By employing contrastive linear probes trained on temporal-choice answers, the authors successfully steer the model's preferences in both short-term and long-term directions, demonstrating significant shifts in decision-making thresholds on various tasks. The findings reveal that intertemporal preferences in AI models can be effectively measured and adjusted, which has implications for applications involving delayed rewards and long-term planning safety.
Steering a language model's intertemporal preferences can induce significant shifts in decision-making, impacting how AI systems advise on delayed costs and benefits.
We study linear representations of temporal horizon in the large language model Qwen3-32B and use them to change the model's time-related preferences, recommendations, and capabilities. We train contrastive linear probes on teacher-forced temporal-choice answers to find a short-term versus long-term direction in the model's residual stream, and evaluate contrastive activation-addition steering on a held-out binary temporal-choice task, an out-of-distribution monetary intertemporal-choice task, and a TravelPlanner capability benchmark. The central result is that temporal-horizon directions can be identified with simple contrastive linear probes and then used for steering to induce large, bidirectional preference changes. On an out-of-distribution monetary choice task that varies reward size and delay, steering strongly shifts the model's indifference threshold between smaller-sooner and larger-later rewards in both directions. We further show improvements on a planning-related capability metric under moderate temporal steering. These results suggest that model intertemporal preferences are measurable and steerable, which is relevant for AI systems that give advice involving delayed costs and benefits, and for safety questions about long-horizon planning.