LLMs Struggle with Novel Chinese Xiehouyu Riddles, Study Finds
A new study from arXiv (2607.23440) tests large language models on Chinese xiehouyu riddles, using novel examples created by linguists to avoid data contamination. The study employs multiple-choice questions, free-form explanation generation, and new riddle creation to evaluate understanding. In multiple-choice tests, the delta accuracy between existing low-frequency xiehouyu and novel ones was measured as an index of memorization. Native speakers showed a very low delta, indicating similar processing. Frontier Chinese models averaged a 23.6% delta, while English-centric models averaged 5.1%, suggesting Chinese models memorize more low-frequency xiehouyu due to larger training data. Gemini 3.1 Pro showed remaining capability on novel xiehouyu.
Key facts
- Study tests LLMs on Chinese xiehouyu riddles
- Novel xiehouyu created by linguists to avoid data contamination
- Evaluation includes multiple-choice, explanation generation, and new riddle creation
- Delta accuracy between existing and novel xiehouyu used as memorization index
- Native speakers have very low delta accuracy
- Frontier Chinese models average 23.6% delta accuracy
- English-centric models average 5.1% delta accuracy
- Gemini 3.1 Pro showed capability on novel xiehouyu
Entities
Institutions
- arXiv