Search papers, labs, and topics across Lattice.
This paper evaluates the coherence of probabilistic forecasts generated by language models using a method based on de Finetti's theorem. By eliciting forecasts from models on stock return events and calculating the maximum Dutch-book profit, the authors quantify the degree of incoherence in these forecasts. The findings reveal significant incoherence, particularly in scenarios with complex logical relationships and irrelevant contextual details, suggesting a need for improved training strategies to enhance probabilistic accuracy.
Language models exhibit substantial incoherence in probabilistic forecasts, with irrelevant details amplifying errors by an order of magnitude.
People increasingly use language models to support life decisions. Many such decisions involve a probabilistic forecast: How likely is a major life event, a natural disaster, or an economic outcome? Users of language models may implicitly trust that these forecasts fall out of a coherent world model. In this paper, we evaluate the coherence of language model probabilistic forecasts through a procedure that builds on a theorem due to de Finetti. We elicit forecasts from language models across events generated from stock returns data. We then use linear programs to compute the largest Dutch-book profit - the profit an arbitrageur could guarantee by betting against model-generated probabilities - which we use as a measure of incoherence. Our procedure does not require outcome labels, so we can evaluate coherence even in settings where outcomes are not observed or have not yet resolved. We find substantial evidence of incoherence in language model forecasts. Such incoherence increases when there are richer logical relationships between events, and irrelevant contextual details can increase incoherence by an order of magnitude. We conclude by discussing how alternative training strategies may improve probabilistic coherence.