V-FiLLM: New Benchmark for Financial Reasoning in LLMs
A new framework named V-FiLLM has been developed by researchers to create benchmarks for financial reasoning using executable computation trees based on actual tables. This system guarantees the accuracy of answers by evaluating trees symbolically to derive ground truth, subsequently converting them into natural-language questions, thus eliminating any model from the labeling process. Consequently, it enables the generation of items at any scale without incurring annotation costs or adopting a generator's error rate. V-FiLLM features four distinct difficulty axes: computation depth, expression breadth, complexity of financial concepts, and context size. Evaluations of open-source models indicate that accuracy can decrease by as much as 51% with increased reasoning depth and by up to 47 percentage points in adversarial numerical scenarios. This benchmark fills a gap in financial reasoning over structured data, complementing existing STEM-focused benchmarks. The paper can be found on arXiv under the identifier 2608.11047.
Key facts
- V-FiLLM is a framework for generating financial reasoning benchmarks.
- It uses executable computation trees grounded in real tables.
- Answers are correct by construction, with ground truth obtained symbolically.
- No model is involved in labeling, allowing arbitrary scale generation.
- Four difficulty axes: computation depth, expression breadth, financial concept complexity, context size.
- Accuracy drops up to 51% with increased reasoning depth.
- Accuracy drops up to 47 percentage points under adversarial numerical conditions.
- Paper available on arXiv:2608.11047.
Entities
Institutions
- arXiv