ARTFEED — Contemporary Art Intelligence

Framework for Auditing Faithfulness of Multimodal LLMs in Grid Diagnosis

other · 2026-07-29

A novel framework has been introduced for assessing the reliability of multimodal large language models (LLMs) in grid diagnosis. This method evaluates self-reported dependencies, behavioral shifts during controlled modality ablations, and preregistered engineering significance to identify inconsistencies. A mechanism for evidence-gated correction and re-audit reprocesses unsuccessful responses under evidence limitations, confirming enhanced grounding without sacrificing performance. Case studies analyze three LLMs of varying sizes within IEEE 39- and 118-bus scenarios. The goal of the framework is to guarantee the utilization of task-relevant evidence, moving beyond simply measuring answer accuracy.

Key facts

  • Framework audits faithfulness of multimodal LLMs in grid diagnosis.
  • Compares self-reported reliance, behavioral changes, and engineering importance.
  • Uses controlled modality ablations to detect discrepancies.
  • Evidence-gated correction and re-audit mechanism regenerates failed responses.
  • Case studies evaluate three LLMs on IEEE 39- and 118-bus scenarios.
  • Published on arXiv with ID 2607.24539.

Entities

Institutions

  • arXiv

Sources