ARTFEED — Contemporary Art Intelligence

New Method to Reduce Scoring Bias in LLM-as-a-Judge Evaluators

ai-technology · 2026-08-10

A research paper on arXiv (2608.05726) proposes a novel method to mitigate scoring bias in Large Language Models (LLMs) used as evaluators of text quality, a practice known as LLM-as-a-Judge. The method involves instructing the LLM to randomly generate number tokens, then measuring the deviation of the observed distribution from a uniform distribution to identify latent numerical bias. Task-specific bias is measured by adding a definition of the downstream task to the random number generation prompts. During evaluation, token generation probabilities are rectified to account for this bias. Experiments were conducted on four tasks, though details are not provided in the abstract. The paper was announced as a cross-type submission on arXiv.

Key facts

  • Paper arXiv:2608.05726 proposes a method to mitigate scoring bias in LLM-as-a-Judge.
  • LLM-as-a-Judge uses LLMs to evaluate text quality, outperforming conventional metrics.
  • Scoring bias is the tendency of LLM evaluators to generate particular scores regardless of context.
  • The method instructs the LLM to randomly generate number tokens to identify latent numerical bias.
  • Deviation from uniform distribution is used to measure bias.
  • Task-specific bias is measured by adding task definition to random number generation prompts.
  • Token generation probabilities are rectified considering the LLM's latent number bias.
  • Experiments were conducted on four tasks.

Entities

Institutions

  • arXiv

Sources