ARTFEED — Contemporary Art Intelligence

UniVVT: Unified End-to-End Framework for Video Virtual Try-On

ai-technology · 2026-08-07

A recent study presents UniVVT, an integrated framework designed for high-fidelity video virtual try-on (VVT). This innovative approach redefines VVT as semantically conditioned video generation, removing the necessity for distinct modules for human parsing, pose estimation, and garment warping during inference. It overcomes the challenges posed by prevalent methods that depend on mask-conditioned video inpainting and multi-stage processes, which can complicate deployment and allow for the propagation of errors in geometric priors into the final videos. UniVVT employs a scene-task perceiver based on a Multimodal Large Language Model to encode the source video, target garment, and task instructions into compact, task-aware latent tokens. A lightweight semantic bridge connects these tokens to the video generation process. The paper is accessible on arXiv under identifier 2608.05745 and is noted as a cross-type abstract. This research holds significance for the digital art and AI technology fields, enhancing the ability to create realistic videos of individuals in various garments, with applications in fashion, e-commerce, and digital content creation.

Key facts

  • UniVVT is a unified end-to-end framework for video virtual try-on.
  • It reframes VVT as semantically conditioned video generation.
  • It eliminates mask, pose, and warping modules at inference.
  • The framework uses a scene-task perceiver built on a Multimodal Large Language Model.
  • It encodes source video, target garment, and task instruction into latent tokens.
  • A lightweight semantic bridge aligns tokens with video generation.
  • The paper is available on arXiv with identifier 2608.05745.
  • The approach aims to reduce error propagation from geometric priors.

Entities

Institutions

  • arXiv

Sources