ARTFEED — Contemporary Art Intelligence

TKFQA: New Benchmark Tests LLM Factuality and Order Robustness in Heterogeneous Knowledge

ai-technology · 2026-08-11

A new benchmark called TKFQA has been launched to assess the consistency of factuality and the robustness of large language models (LLMs) in multi-hop reasoning across diverse knowledge sources. This benchmark includes 10,130 question-answering (QA) pairs derived from tables, texts, and knowledge graphs (KGs). Each instance is based on a clear counterfactual reasoning chain, allowing for a comprehensive evaluation of answer accuracy, the correctness of reasoning chains, and sensitivity to variations in input order. An in-depth analysis of 14 LLMs, both open and closed-source, demonstrated that leading models struggle with reasoning-chain accuracy and are affected by input order. This research fills a gap in current benchmarks, which inadequately evaluate LLMs' capabilities in multi-hop reasoning across varied knowledge formats. The findings underscore the difficulties in maintaining factuality and consistency when knowledge is presented in different structures. The paper can be found on arXiv with the identifier 2608.07838.

Key facts

  • TKFQA is a new benchmark for factuality consistency and order-robust grounded reasoning in LLMs.
  • The benchmark includes 10,130 QA pairs grounded in tables, texts, and knowledge graphs.
  • Each example is built from an explicit counterfactual reasoning chain.
  • The benchmark evaluates answer correctness, reasoning-chain accuracy, and robustness to input-order variations.
  • 14 open- and closed-source LLMs were evaluated.
  • State-of-the-art models showed limited reasoning-chain accuracy and sensitivity to input order.
  • The paper is available on arXiv with identifier 2608.07838.
  • The benchmark addresses a gap in existing benchmarks for multi-hop reasoning over heterogeneous knowledge.

Entities

Institutions

  • arXiv

Sources