MIND: Multimodal Intent-Driven Network via Diffusion Transformers
for Medical Image Fusion

Yunzhan Fu1, Xiangyu Shen4, Yifei Sun2,3, Yuhan Chen5, Jian Wu2,†, Hongxia Xu6,†
1Transvascular Implantation Devices Research Institute, Zhejiang University, Hangzhou, China
2Zhejiang University, Hangzhou, China
3Liangzhu Laboratory, Hangzhou, China
4Hangzhou Institute of Technology, Xidian University, Hangzhou, China
5Hangzhou Dianzi University, Hangzhou, China
6Zhejiang Key Laboratory of Medical Imaging Artificial Intelligence, Hangzhou, China
ACM International Conference on Multimedia (ACM MM) 2026
†Corresponding Author

Intent-Driven vs Process-Driven Fusion

🔻 Drag the slider to compare Fusion Result (a) [MIND, intent-driven] and Fusion Result (b) [DiTFuse, process-driven].

MIND is a multimodal intent-driven medical image fusion framework built on Diffusion Transformers (DiTs). It uses BioMedGPT to generate pathology-aware fusion texts that explicitly describe the desired fused image, guiding the generation with clinical diagnostic intent instead of a uniform fusion rule. A Multi-scale Latent Adapter (MLA) preserves 2D spatial continuity that DiTs lose during sequence flattening, and a medical semantic consistency loss with a timestep-truncation mechanism keeps the fused image semantically locked to the fusion texts.

Abstract

Medical image fusion aims to integrate complementary information from diverse imaging modalities to support clinical diagnosis. Existing methods typically apply uniform fusion rules globally, lacking a deep understanding of diagnostic intents and pathological structures. To address these limitations, we propose MIND, a Multimodal Intent-Driven Network via Diffusion Transformers (DiTs) for medical image fusion. Specifically, we utilize BioMedGPT to generate intent-driven fusion texts from source images, guiding the fusion process with pathology-aware diagnostic intents. To combat the loss of 2D spatial continuity caused by 1D sequence flattening in DiTs, we design a Multi-scale Latent Adapter (MLA), which explicitly extracts source image features before serialization and injects them into the network via strict dimensional alignment. To resolve the semantic shift caused by decoupling image outputs from diagnostic intents, we design a medical semantic consistency loss with a timestep-truncation mechanism, ensuring deep semantic locking between fused images and fusion texts while maintaining the stability of the underlying physical manifold reconstruction. Comprehensive experiments on the Harvard, BraTS, and GFP datasets reveal that MIND delivers superior fusion quality, significantly improves downstream brain tumor segmentation accuracy, and enables flexible interactive fusion.

Figures

Results

Quantitative comparison on Harvard

Table 1: Quantitative comparison of MIND with eight state-of-the-art methods on Harvard. Best results are highlighted as first, second and third.

Ablation study

Table 4: Ablation study of our key contributions.

BibTeX

@inproceedings{fu2026mind,
  title={MIND: Multimodal Intent-Driven Network via Diffusion Transformers for Medical Image Fusion},
  author={Fu, Yunzhan and Shen, Xiangyu and Sun, Yifei and Chen, Yuhan and Wu, Jian and Xu, Hongxia},
  booktitle={Proceedings of the ACM International Conference on Multimedia (ACM MM)},
  year={2026}
}