Search papers, labs, and topics across Lattice.
This study evaluates the reasoning capabilities of large language models (LLMs) in understanding and generating Chinese xiehouyu riddles using a novel dataset created by linguists to mitigate data contamination. Through multiple-choice questions and free-form explanations, the research reveals that frontier Chinese models outperform English-centric models in memorization and understanding, with Gemini 3.1 Pro achieving a remarkable 92.6% accuracy on novel xiehouyu. However, the creativity of LLMs in generating xiehouyu remains significantly lower than that of human experts, indicating a gap in their reasoning and creative abilities in this linguistic domain.
Frontier Chinese models excel in understanding novel xiehouyu riddles, but their creative output still lags behind human performance.
In this paper, we push the boundary of LLM reasoning by testing them in a Chinese language game, xiehouyu, with novel xiehouyu created by linguists that had not existed before to avoid data contamination. We use multiple-choice questions (MCQ), free-form explanation generation, and new xiehouyu creation to evaluate LLMs'ability to understand and create xiehouyu. In MCQ, we use the delta of accuracy ($\Delta_{acc}$) between existing but low-frequency xiehouyu and novel ones as an index for memorization. $\Delta_{acc}$ for native speakers is very low, suggesting similar processing mechanisms. However, we found that frontier Chinese models have on average a $\Delta_{acc}$ of 23.6\%, while English-centric models tested have a mean $\Delta_{acc}$ of 5.1\%, suggesting that frontier Chinese models are likely trained with much larger Chinese data, thus memorizing more low-frequency xiehouyu. For novel xiehouyu, Gemini 3.1 Pro demonstrated remarkable ability with acc 92.6, which is 24\% higher than human accuracy. In xiehouyu creation, those created by LLMs receive much worse ratings than those by humans. These results suggest that claims about the reasoning abilities of LLMs may need careful re-examination considering the data contamination issue, and that LLMs'creativity in language-related tasks may still be behind human experts, at least in Chinese xiehouyu.