ProverbIT Benchmark Reveals LLM Struggles with Italian Proverbs
A recent research article presents ProverbIT, a set of 100 multiple-choice questions aimed at assessing large language models (LLMs) on Italian proverbs. The investigation analyzes 13 advanced models, which include both large reasoning models (LRMs) and conventional LLMs, through three specific tasks: completing proverbs, answering multiple-choice questions with correct answers, and tackling multiple-choice questions without correct answers. The findings indicate that, although most models excel in completion tasks, their performance significantly declines in multiple-choice scenarios lacking correct answers, even among top reasoning models. This research underscores the limitations of LLMs in grasping culturally rooted linguistic expressions. The paper can be accessed on arXiv under ID 2608.04670.
Key facts
- ProverbIT is a novel Italian benchmark with 100 multiple-choice questions.
- 13 frontier models were assessed, including LRMs and traditional LLMs.
- Three tasks: proverb completion, multiple-choice with correct answers, and multiple-choice without correct answers.
- Most models succeeded at proverb completion.
- Performance dropped dramatically in multiple-choice without correct answers.
- Even state-of-the-art reasoning models struggled in that format.
- The study focuses on culturally embedded linguistic expressions.
- Paper available on arXiv: 2608.04670.
Entities
Institutions
- arXiv