Search papers, labs, and topics across Lattice.
This study investigates the impact of survey-country metadata on the performance of large language models (LLMs) in social inference tasks, revealing that while verified country information can significantly enhance predictive accuracy (reducing Brier loss by 0.040), random metadata can mislead forecasts without improving outcomes. A randomized audit across multiple models and countries demonstrated that disclosing the random origin of metadata did not effectively mitigate its influence on model predictions, with both opaque and disclosed-random labels causing similar shifts in country-directed forecasts. The findings highlight the complexities of metadata utilization in LLMs, suggesting that not all cues improve model performance and that random labels can introduce significant noise.
Verified survey-country metadata boosts LLM predictive accuracy, but random labels can mislead forecasts without any benefit.
Survey-country metadata can improve an LLM's forecast of an individual response when informative, yet the same cue may redirect the forecast when assigned at random. A within-record audit tests whether disclosing a random label's uniform, record-independent origin reduces its country-directed uptake, and whether verified survey country lowers held-out Brier loss. Independent population anchors and recorded human answers measure direction and consequence across five fixed API models, six countries, and seven development-selected targets. In the primary post-review 72-record panel, opaque and disclosed-random labels each produced country-direction shifts of 0.214. Paired attenuation was 0.0003 (95% CI [-0.0157, 0.0166]). Verified country reduced Brier loss by 0.040 (95% CI [0.024, 0.056]), while random-label regret included zero. A non-overlapping mixed-coverage consistency panel retained positive disclosed-random movement and verified utility, while attenuation remained uncertain. On the selected targets, verified metadata was useful in both panels, but disclosure did not reliably attenuate random-label uptake. PROV-FORECAST contains 14,400 paired item-level probability distributions from the corrected panel.