Search papers, labs, and topics across Lattice.
This study investigates the performance of large language models (LLMs) as economic agents in a double auction market, replacing human participants to assess their effectiveness in resource allocation. The findings reveal that LLMs struggle to achieve market equilibrium, resulting in less efficient outcomes compared to human agents. Additionally, the analysis uncovers significant variability in trading behaviors across different LLMs and market roles, highlighting a shift in decision-making from strategic to urgent considerations during trading.
LLMs in market settings fail to reach equilibrium efficiently, revealing critical limitations in their economic decision-making capabilities.
Large language models (LLMs) are increasingly deployed as economic agents, yet there is little evidence whether LLM agents are suited for participating in market mechanisms designed for humans, and whether these mechanisms deliver desired outcomes when faced with LLM agents. We address this question by replicating seminal economic experiments, replacing human subjects with LLM agents. We place agents in a double auction environment, which is a widely-used market mechanism. We check whether such a market is able to deliver an efficient allocation of resources, thereby testing a novel dimension of alignment of LLM agents -- their compatibility with a fundamental market mechanism. We find that markets populated by LLM agents exhibit slower or no convergence towards market equilibrium, thus providing less efficient allocations than markets populated by humans. We then analyze agents' individual trading decisions and find substantial heterogeneity both across model families and market roles. We also run a lexical analysis of Chain-of-Thought (CoT) traces generated by the agents. We find that the decision to execute a trade rather than continue incrementally adjusting prices is associated with a shift from strategic considerations toward urgency. We publicly release our testing framework, which can be used for future evaluations.