ARTFEED — Contemporary Art Intelligence

Hybrid Evaluation of Four Open-Source Model Routers Across Benchmarks

ai-technology · 2026-08-18

A recent paper on arXiv (2608.14641) introduces a standardized measurement protocol and a hybrid assessment of four open-source model routers: RouterBench, BFCL v4, tau2-bench, and WebArena. This research analyzes 290 fixed tasks against a defined matrix of 2,610 potential outcomes. Among the routers, three maintain consistent tier assignments, while only the vLLM Semantic Router shows significant variation based on prompt content, although it does not achieve the highest success rate across the four benchmarks. Always-Mid aligns perfectly with Aurelio on three benchmarks and is within 0.003 on the fourth. Task-level superiority tests for vLLM reveal no specific advantage over a content-blind allocation, with equivalence confirmed only on WebArena at the five-percentage-point margin. The results underscore the challenges in directly comparing router evaluations and suggest a unified method for future studies.

Key facts

  • arXiv:2608.14641v1
  • Evaluates four open-source routers: RouterBench, BFCL v4, tau2-bench, WebArena
  • 290 frozen tasks
  • 2,610 candidate outcomes
  • Three routers emit constant or near-constant tier assignments
  • vLLM Semantic Router varies with prompt content
  • vLLM has highest success rate on none of the benchmarks
  • Always-Mid matches Aurelio exactly on three benchmarks and within 0.003 on the fourth
  • Equivalence for vLLM established only on WebArena at five-percentage-point margin

Entities

Institutions

  • arXiv
  • RouterBench
  • BFCL v4
  • tau2-bench
  • WebArena
  • vLLM Semantic Router
  • Always-Mid
  • Aurelio

Sources