ARTFEED — Contemporary Art Intelligence

LLM Benchmark Scores Vary Significantly with Inference Backend Choice

ai-technology · 2026-08-06

A new arXiv preprint (2608.04714) reveals that the choice of inference backend—such as HuggingFace, vLLM, or Ollama—can substantially alter large language model (LLM) benchmark scores, even under deterministic greedy decoding. The study, which is a fully-crossed analysis of three instruction-tuned models, five inference frameworks, six benchmarks, and four generation modes, found that backend selection is a non-negligible factor that can significantly change model performance. The effect is structural and strongly model-dependent, with variance decomposition showing that roughly 39% of the variability in scores can be attributed to the backend choice. The authors argue that benchmark scores are often reported as properties of the model alone, while the inference framework used to produce them is considered non-influential and its name and version are rarely disclosed. This oversight could lead to misleading comparisons and reproducibility issues in the AI research community. The paper is available on arXiv under the identifier 2608.04714.

Key facts

  • Study investigates influence of inference backend on LLM benchmark scores.
  • Backends tested include HuggingFace, vLLM, and Ollama.
  • Fully-crossed design: 3 models x 5 frameworks x 6 benchmarks x 4 generation modes.
  • Backend choice can significantly alter performance even under greedy decoding.
  • Effect is structural and strongly model-dependent.
  • Roughly 39% of score variability is attributed to backend choice.
  • Paper available on arXiv with ID 2608.04714.
  • Authors call for disclosure of inference framework in benchmark reporting.

Entities

Institutions

  • HuggingFace
  • vLLM
  • Ollama

Sources