ARTFEED — Contemporary Art Intelligence

DD-CMD: Dual-Domain Cross-Modal Decoding for Text-Guided Medical Image Segmentation

ai-technology · 2026-08-13

A new framework called Dual-Domain Cross-Modal Decoding (DD-CMD) has been developed by researchers for segmenting pulmonary infections using clinical text guidance. This innovative method employs two types of language assistance during the decoding process: Text-Guided Spatial Cross-Attention (TGSA), which aligns visual tokens of various scales with textual semantics and refines features via gated residual fusion, and Spectral-Text Adaptive Modulation (STAM), which utilizes a 2D DCT to derive learnable band-energy statistics, predicting text-conditioned FiLM parameters to adjust decoder channels for frequency-aware decoding. DD-CMD incorporates TGSA and STAM within a coarse-to-fine decoder (from 7x7 to 56x56) and employs a lightweight two-stage refinement module to restore full-resolution masks. This method overcomes the shortcomings of recent text-guided approaches that focus on spatial alignment but neglect frequency content essential for texture and boundaries. Experiments were performed on the QaTa-CO dataset, although specific findings are not included in the abstract. The paper can be found on arXiv with the identifier 2608.11335.

Key facts

  • DD-CMD integrates spatial and frequency domain language guidance for medical image segmentation.
  • TGSA aligns multi-scale visual tokens with text semantics using gated residual fusion.
  • STAM uses 2D DCT to compute band-energy statistics and predicts FiLM parameters.
  • Decoder operates from 7x7 to 56x56 resolution with a two-stage refinement module.
  • Experiments performed on QaTa-CO dataset.
  • Paper available on arXiv:2608.11335.

Entities

Institutions

  • arXiv

Sources