ARTFEED — Contemporary Art Intelligence

NxN E-valuation: A New Method for Hypothesis Certification in LLM Systems

ai-technology · 2026-08-10

A new algorithm called NxN E-valuation has been proposed for hypothesis certification, designed to verify hypotheses without constructing case-specific certification procedures, provided a sufficiently large dataset is available. The method is particularly suited to LLM-based exploration systems, where LLMs excel at proposing hypotheses but suffer from hallucination, preventing direct use of their outputs. Existing remedies, such as circular verification and held-out testing, fall short; held-out testing can pass false hypotheses via spurious correlations. NxN E-valuation exploits the naturally existing large training set by letting different samples serve as null hypotheses for one another. The algorithm is based on e-values and conformal prediction, offering a handy and rigorous approach. The paper is available on arXiv under the identifier 2608.06621.

Key facts

  • NxN E-valuation is an e-value-based hypothesis-certification algorithm.
  • It verifies hypotheses without building case-specific certification procedures.
  • Requires a large enough dataset.
  • Suited for LLM-based exploration systems.
  • Addresses hallucination in LLM outputs.
  • Existing remedies like circular verification and held-out testing have limitations.
  • Held-out testing can pass false hypotheses via spurious correlations.
  • The method uses different samples as null hypotheses for one another.
  • Paper available on arXiv: 2608.06621.

Entities

Institutions

  • arXiv

Sources