Search papers, labs, and topics across Lattice.
This study investigates how large language models (LLMs) resolve conflicts between textual summaries and numerical data when they provide opposing evidence. By creating a controlled synthetic benchmark that manipulates various factors such as modality and temporal recency, the authors reveal that LLMs exhibit systematic arbitration behaviors, favoring text or numbers based on context. The findings indicate that LLMs often resort to heuristic strategies, leading to potential misjudgments in tool-augmented decision-making systems.
LLMs systematically favor text over numbers in evidence arbitration, revealing a critical failure mode in decision-making systems that rely on heterogeneous data sources.
Large language models (LLMs) are increasingly used in settings where textual summaries, numerical observations, and external tool outputs may provide conflicting evidence. We study how LLMs arbitrate between such sources when they support opposing decisions. To do so, we introduce a controlled synthetic benchmark in which latent risk trajectories generate both numerical time series and natural language summaries, allowing us to construct conflicts where exactly one evidence source is aligned with the ground-truth label. This design lets us independently manipulate modality, temporal recency, source reliability, and evidence provenance. Across open-weight instruction-tuned models, we find that arbitration behaviour is systematic rather than random: models exhibit distinct text-versus-number preferences, follow temporal recency more consistently than explicit reliability cues, and can over-rely on external forecasts even when they conflict with direct contextual evidence. These results suggest that current LLMs often rely on heuristic arbitration strategies when integrating heterogeneous evidence, highlighting a failure mode for tool-augmented decision systems.