CITBench: New Benchmark for Interactive Tabular Data Processing with LLMs
CITBench has been developed by researchers as a robust benchmark to assess large language models (LLMs) in the realm of interactive tabular data processing tasks. Unlike current benchmarks that emphasize single-turn, explicitly defined instructions for table reasoning, CITBench tackles the intricacies of multi-turn interactions where user needs may change. This benchmark features a classification of four primary categories: table matching, cleaning, augmentation, and transformation, which includes 18 task types and 1,296 examples sourced from various domain datasets. CITBench allows for both offline and online evaluations. The online mode simulates multi-turn interactions with restricted operational procedures and structured task scripts, reflecting essential behavioral traits of LLM-based assistants in practical settings. Further insights into this initiative can be found in a paper on arXiv (arXiv:2608.00018).
Key facts
- CITBench is a new benchmark for evaluating LLMs on interactive tabular data processing.
- It covers four high-level categories: table matching, cleaning, augmentation, and transformation.
- The benchmark includes 18 task types and 1,296 instances.
- Instances are curated from datasets across diverse domains.
- CITBench supports both offline and online evaluation.
- The online setting models multi-turn interactions with evolving user requirements.
- The benchmark aims to address limitations of existing single-turn table reasoning benchmarks.
- The paper is available on arXiv with identifier 2608.00018.
Entities
Institutions
- arXiv