Medical image fusion aims to integrate complementary information from diverse imaging modalities to support clinical diagnosis. Existing methods typically apply uniform fusion rules globally, lacking a deep understanding of diagnostic intents and pathological structures. To address these limitations, we propose MIND, a Multimodal Intent-Driven Network via Diffusion Transformers (DiTs) for medical image fusion. Specifically, we utilize BioMedGPT to generate intent-driven fusion texts from source images, guiding the fusion process with pathology-aware diagnostic intents. To combat the loss of 2D spatial continuity caused by 1D sequence flattening in DiTs, we design a Multi-scale Latent Adapter (MLA), which explicitly extracts source image features before serialization and injects them into the network via strict dimensional alignment. To resolve the semantic shift caused by decoupling image outputs from diagnostic intents, we design a medical semantic consistency loss with a timestep-truncation mechanism, ensuring deep semantic locking between fused images and fusion texts while maintaining the stability of the underlying physical manifold reconstruction. Comprehensive experiments on the Harvard, BraTS, and GFP datasets reveal that MIND delivers superior fusion quality, significantly improves downstream brain tumor segmentation accuracy, and enables flexible interactive fusion.
Table 1: Quantitative comparison of MIND with eight state-of-the-art methods on Harvard. Best results are highlighted as first, second and third.
Table 4: Ablation study of our key contributions.
@inproceedings{fu2026mind,
title={MIND: Multimodal Intent-Driven Network via Diffusion Transformers for Medical Image Fusion},
author={Fu, Yunzhan and Shen, Xiangyu and Sun, Yifei and Chen, Yuhan and Wu, Jian and Xu, Hongxia},
booktitle={Proceedings of the ACM International Conference on Multimedia (ACM MM)},
year={2026}
}