ARTFEED — Contemporary Art Intelligence

DualDiT: A Diffusion Transformer for Joint OCT Image and Segmentation Mask Generation

ai-technology · 2026-08-03

A new model named DualDiT has been developed by researchers; it is a conditional dual-output Diffusion Transformer aimed at simultaneously generating optical coherence tomography (OCT) B-scans and segmentation masks for upper retinal cell layers in ex vivo mouse retina. This innovation tackles the challenge of limited annotated medical imaging data, especially in OCT of mouse eyes, where manual delineation of retinal layers is both time-consuming and requires specialized skills. While diffusion models have shown potential in medical image creation, most joint image-mask synthesis has relied on U-Net-based denoisers, leaving diffusion transformers underutilized. DualDiT utilizes a pretrained VAE to encode both modalities into a shared latent space, concatenates their representations, and applies conditional diffusion on the combined tensor. The paper can be found on arXiv with the identifier 2607.29337.

Key facts

  • DualDiT is a conditional dual-output Diffusion Transformer for joint OCT image and segmentation mask generation.
  • The model targets OCT B-scans and segmentation masks of upper retinal cell layers in ex vivo mouse retina.
  • It addresses the shortage of annotated data in medical imaging, particularly in OCT of mouse eyes.
  • Manual retinal layer delineation is labor-intensive due to tiny structures and required expertise.
  • Diffusion models perform well in medical image synthesis, but joint image-mask generation has relied mainly on U-Net-based denoisers.
  • DualDiT encodes both modalities into a shared latent space via a pretrained VAE.
  • It concatenates latent representations and performs conditional diffusion over the joint tensor.
  • The paper is available on arXiv with identifier 2607.29337.

Entities

Institutions

  • arXiv

Sources