ARTFEED — Contemporary Art Intelligence

FLEX Benchmark Exposes Routing Saturation in Multimodal Continual Instruction Tuning

ai-technology · 2026-08-04

A recent study published on arXiv (2608.01437) indicates that expert routing in Multimodal Continual Instruction Tuning (MCIT) has reached near saturation on current benchmarks, masking the long-horizon routing challenge. The researchers present FLEX (Fingerprint-reduced Long-horizon Expert eXamination), a benchmark consisting of 34 tasks aimed at minimizing textual fingerprints that reveal task identity. FLEX categorizes tasks with analogous instruction and response formats but varies in visual and knowledge domains, standardizes outer templates, and assesses routing over a significantly extended horizon. This research emphasizes the overlooked task-identification issue that is vital for allowing multimodal large language models to sequentially learn new tasks while preserving previously acquired skills. The results indicate that existing benchmarks may not sufficiently test routing systems, and FLEX seeks to offer a more authentic evaluation framework.

Key facts

  • The paper is available on arXiv with ID 2608.01437.
  • MCIT enables multimodal large language models to acquire new tasks sequentially while retaining prior capabilities.
  • Many recent methods use task-specific LoRA experts and route inputs to one or more experts at inference.
  • The task-identification problem underlying expert routing is under-explored.
  • Routing is nearly saturated on widely used MCIT benchmarks.
  • Textual fingerprints that leak task identity and short 4-10-task sequences obscure the long-horizon routing problem.
  • FLEX is a 34-task long-horizon MCIT benchmark with weakened textual fingerprints.
  • FLEX groups tasks with similar instruction and answer formats but diverse visual and knowledge domains.
  • FLEX normalizes outer templates and evaluates routing over a substantially longer horizon.

Entities

Institutions

  • arXiv

Sources