ARTFEED — Contemporary Art Intelligence

New AI Method Interprets Black-Box LLMs via Energy-Based Models

ai-technology · 2026-08-06

A new method has been proposed by researchers to enhance the interpretability of proprietary Large Language Models (LLMs) accessed solely through closed APIs, tackling the essential issue of responsible AI deployment. This approach, outlined in a paper on arXiv (2608.02879), introduces a model-agnostic, post-hoc attribution interpreter that functions at the sentence level. It employs an Energy-Based Model (EBM) as a surrogate to understand the internal consistency of the LLM between prompts and outputs, utilizing this energy landscape to train a lightweight interpreter network. After training, the interpreter independently quantifies the impact of prompt sentences on a specified target output without additional API calls. By training a local interpreter globally across varied inputs, the framework identifies broader generation trends and reduces instance-specific biases. The paper includes experiments that showcase the EBM's effectiveness in this role, marking a crucial advancement for the AI community by providing insights into the opaque decision-making of black-box models commonly used across various fields.

Key facts

  • The paper is available on arXiv with ID 2608.02879.
  • The method is model-agnostic and post-hoc.
  • It operates at the sentence level.
  • An Energy-Based Model (EBM) is used as a surrogate.
  • The interpreter is lightweight and standalone after training.
  • No further API queries are needed after training.
  • The framework captures broader generation patterns.
  • Experiments demonstrate the EBM's accuracy.

Entities

Institutions

  • arXiv

Sources