ARTFEED — Contemporary Art Intelligence

KnowHal: New Benchmark for Multimodal Hallucination Evaluation

ai-technology · 2026-08-06

KnowHal has been developed by researchers as a benchmark aimed at thoroughly assessing hallucinations in Multimodal Large Language Models (MLLMs). Unlike current benchmarks that primarily address hallucinations related to entities, attributes, and relationships, KnowHal specifically includes knowledge hallucination, which encompasses four areas: entity, attribute, relation, and knowledge. This benchmark features paired questions—both positive and negative—based on shared images and entities, facilitating controlled evaluations of perceptual mistakes, knowledge errors, and acceptance of false premises. It comprises 1,800 samples spanning 10 domains and 50 categories, created through a semi-automated process that integrates LLM support, CLIP filtering, and human validation. The findings are published in a paper on arXiv (arXiv:2608.03782).

Key facts

  • KnowHal is a new benchmark for evaluating hallucinations in Multimodal Large Language Models (MLLMs).
  • It covers four hallucination dimensions: entity, attribute, relation, and knowledge.
  • The benchmark includes 1,800 samples across 10 domains and 50 categories.
  • It uses paired positive and negative questions over shared images and entities.
  • Construction pipeline combines LLM assistance, CLIP-based filtering, and human verification.
  • The paper is available on arXiv with identifier 2608.03782.

Entities

Sources