ARTFEED — Contemporary Art Intelligence

V2N: First Complete Visual Piano Transcription System Achieves State-of-the-Art Results

ai-technology · 2026-08-06

A groundbreaking system called V2N (Video to Notes) has been developed by researchers, marking the first fully realized Visual Piano Transcription (VPT) solution that fills critical voids in current techniques. While traditional audio-based piano transcription excels in detecting onset, pitch, and velocity, the sustain pedal complicates matters by allowing sounds to linger after a key is released, causing audio systems to misinterpret pedal-extended offsets instead of actual key releases. Current VPT systems mainly concentrate on onset detection from brief video segments, leading to less accurate offset predictions and a lack of reported note-level velocity. V2N utilizes a shared temporal backbone with specialized heads for onset, offset, key hold, and velocity, trained with per-frame supervision. Ablation studies reveal that multi-task supervision enhances offset and velocity prediction while boosting onset accuracy, with extended temporal context providing additional benefits. V2N achieves new state-of-the-art performance on the PianoVAM and R3 datasets. Detailed in a paper submitted to arXiv (ID: 2608.03419) in the Computer Science > Sound category, this work signifies a major leap in visual music transcription, promising more precise and thorough analysis of piano performances from video.

Key facts

  • V2N is the first complete Visual Piano Transcription (VPT) system.
  • It predicts onset, offset, key hold, and velocity from video.
  • Uses a shared temporal backbone with task-specific heads.
  • Trained with per-frame supervision.
  • Multi-task supervision improves offset and velocity prediction.
  • Longer temporal context improves performance.
  • Achieves state-of-the-art results on PianoVAM and R3 datasets.
  • Paper available on arXiv with ID 2608.03419.

Entities

Institutions

  • arXiv

Sources