Search papers, labs, and topics across Lattice.
This paper investigates the dynamics of a large language model (LLM) service where providers set token prices and default reasoning-token allocations, while users can either accept, customize, or exit the service. By modeling this interaction as a Stackelberg game, the authors derive the user's optimal customized allocation and characterize the provider's optimal default through a three-regime rule. The findings reveal that defaults only influence user behavior when convenience is valued, indicating that users typically prefer their optimal allocations regardless of the defaults set by providers.
Users often ignore provider-set defaults in LLM services, opting instead for their own optimal token allocations unless convenience is prioritized.
We study a large language model (LLM) service in which a provider chooses a per-token price and a default reasoning-token allocation, while a user may accept the default, customize the allocation, or exit. Larger allocations can improve accuracy but increase token cost and latency. We model this interaction as a Stackelberg game and derive the user's unique optimal customized allocation in closed form. For any price, the acceptable defaults form either an empty set or a compact interval. We characterize the provider's optimal default through a three-regime rule, reduce equilibrium computation to a one-dimensional price optimization, and prove the existence of the equilibrium. We further show that defaults affect the implemented reasoning allocation only when users value the convenience of avoiding customization; otherwise, every service-providing outcome implements the user's optimal customized allocation. Experiments with two compact open-weight reasoning models on five mathematics and science benchmarks support the accuracy-token model and show how model and task characteristics determine equilibrium prices, defaults, and reasoning allocations.