ARTFEED — Contemporary Art Intelligence

MoCA: A New Approach to Chart-to-Code Generation with Cross-modal Arbitration

ai-technology · 2026-08-18

A new method called MoCA (Mixture of Cross-modal Arbitration) has been proposed for chart-to-code generation, which involves reading fine-grained visual details of a chart and writing executable code that reproduces it. Existing methods either train visual and coding abilities separately or fine-tune on chart-to-code data with the two abilities entangled, neither accounting for the distinct nature of the abilities or interference when optimized together. MoCA separates the two abilities rather than blending them, built on a Cross-modal Arbitration Block (CAB) that maintains a visual branch and a code branch as distinct pathways, with a lightweight arbiter arbitrating their relative contributions at every layer and generated token. Training occurs in two stages: supervised warm-up on self-distilled reasoning trajectories that decompose visual understanding into explicit steps, followed by reinforcement learning. The paper is available on arXiv with identifier 2608.15510.

Key facts

  • MoCA stands for Mixture of Cross-modal Arbitration.
  • It separates visual and coding abilities rather than blending them.
  • The method uses a Cross-modal Arbitration Block (CAB) with distinct visual and code branches.
  • A lightweight arbiter arbitrates contributions at every layer and token.
  • Training involves two stages: supervised warm-up and reinforcement learning.
  • The warm-up uses self-distilled reasoning trajectories to decompose visual understanding.
  • The paper is on arXiv with identifier 2608.15510.
  • The approach addresses interference between visual and coding abilities.

Entities

Institutions

  • arXiv

Sources