ARTFEED — Contemporary Art Intelligence

TGFusion: Text-Guided Flow Matching for Image Fusion Under Complex Degradations

ai-technology · 2026-08-04

Researchers have introduced TGFusion, a text-guided latent-space flow matching framework designed to address the challenges of infrared-visible image fusion under realistic degradation scenarios. The framework unifies degradation suppression and cross-modal fusion by encoding task, degradation, and generation cues into structured prompts. Unlike existing methods that inject fixed global text representations, TGFusion adapts textual guidance to spatially varying degradations, local structures, and thermal saliency. The approach leverages text to provide prior information about degradation characteristics, complementing limited evidence from corrupted input images. The paper, available on arXiv (2608.00530), proposes a Prompt-conditioned Multi-stream J (likely a typo for 'Joint') architecture to fully exploit these priors. This work is significant for computer vision applications where image fusion must operate under complex, real-world conditions.

Key facts

  • TGFusion is a text-guided latent-space flow matching framework.
  • It unifies degradation suppression and cross-modal fusion.
  • It encodes task, degradation, and generation cues into structured prompts.
  • It adapts to spatially varying degradations, local structures, and thermal saliency.
  • Existing methods use fixed global text representations.
  • Text provides prior information about degradation characteristics.
  • The paper is on arXiv with ID 2608.00530.
  • The framework uses a Prompt-conditioned Multi-stream J architecture.

Entities

Institutions

  • arXiv

Sources