ARTFEED — Contemporary Art Intelligence

Rendering Code as Images Reduces Input Tokens by Up to 86.5%

ai-technology · 2026-07-27

A recent study published on arXiv (2607.21672) examines the token counting methods of commercial APIs when handling source code as images compared to raw text for vision-language models. Researchers performed a reproducible measurement case study involving five programming languages, nine different source lengths (ranging from 20 to 2,000 lines), and 15 model variants from Anthropic, OpenAI, and Google Vertex AI, which can be categorized into approximately five unique accounting signatures. Analyzing 675 text/image pairs revealed aggregate image-to-text token ratios of 0.135, 0.194, and 0.242, leading to input-token reductions of 86.5%, 80.6%, and 75.8%. This research addresses a systems question regarding how providers count tokens for code represented as images, which could lower costs for lengthy contexts.

Key facts

  • Study examines input-token accounting for source code as text vs. images
  • Uses five programming languages and nine source lengths from 20 to 2,000 lines
  • Tests 15 model aliases from Anthropic, OpenAI, and Google Vertex AI
  • Aliases collapse to approximately five distinct accounting signatures
  • Aggregate image-to-text ratios: 0.135, 0.194, and 0.242
  • Corresponding token reductions: 86.5%, 80.6%, and 75.8%
  • Based on 675 complete text/image pairs
  • Published on arXiv with ID 2607.21672

Entities

Institutions

  • Anthropic
  • OpenAI
  • Google Vertex AI
  • arXiv

Sources