About the journal
Browse
Collections
Multimedia collections
Authors & reviewers
Multimodal triplane diffusion transformer for high fidelity 3D shape generation
Shuang WU , Youtian LIN , Yifei ZENG , Feihu ZHANG , Hao ZHU , Xun CAO , Yao YAO
Front. Comput. Sci. ›› 2026, Vol. 20 ›› Issue (10) : 2010715
Recent advancements in native 3D generation have demonstrated remarkable capabilities in producing high-quality 3D assets from image or text prompts. However, these methods face a critical challenge: insufficient alignment between generated meshes and input conditions. In this paper, we propose Multimodal Triplane Diffusion Transformer to address the issue, featuring two core components: A triplane-based 3D variational autoencoder that compresses point clouds, sampled uniformly from mesh surfaces and concentrated near sharp edges, into a triplane latent space, and a diffusion model trained on this latent space, empowered by multimodal diffusion transformer blocks to establish cross-modality interactions between latent representations and conditions. Extensive experiments demonstrate that our method achieves not only superior generalization capability but also significantly enhanced geometric alignment with input images compared to state-of-the-art approaches in the image-to-3D task.
native 3D generation / diffusion model / multimodal diffusion transformer
| [1] |
|
| [2] |
|
| [3] |
|
| [4] |
|
| [5] |
Nichol A, Jun H, Dhariwal P, Mishkin P, Chen M. Point-E: A system for generating 3D point clouds from complex prompts. 2022, arXiv preprint arXiv: 2212.08751 |
| [6] |
|
| [7] |
|
| [8] |
|
| [9] |
Wang P, Shi Y. ImageDream: Image-prompt multi-view diffusion for 3D generation. 2023, arXiv preprint arXiv: 2312.02201 |
| [10] |
|
| [11] |
|
| [12] |
|
| [13] |
|
| [14] |
|
| [15] |
|
| [16] |
Xu J, Cheng W, Gao Y, Wang X, Gao S, Shan Y. InstantMesh: efficient 3D mesh generation from a single image with sparse-view large reconstruction models. 2024, arXiv preprint arXiv: 2404.07191 |
| [17] |
Shi R, Chen H, Zhang Z, Liu M, Xu C, Wei X, Chen L, Zeng C, Su H. Zero123++: a single image to consistent multi-view diffusion base model. 2023, arXiv preprint arXiv: 2310.15110 |
| [18] |
|
| [19] |
|
| [20] |
|
| [21] |
Hong F, Tang J, Cao Z, Shi M, Wu T, Chen Z, Yang S, Wang T, Pan L, Lin D, Liu Z. 3DTopia: Large text-to-3D generation model with hybrid diffusion priors. 2024, arXiv preprint arXiv: 2403.02234 |
| [22] |
|
| [23] |
|
| [24] |
|
| [25] |
|
| [26] |
|
| [27] |
Jun H, Nichol A. Shap-E: Generating conditional 3D implicit functions. 2023, arXiv preprint arXiv: 2305.02463 |
| [28] |
|
| [29] |
|
| [30] |
|
| [31] |
|
| [32] |
Li Y, Zou Z X, Liu Z, Wang D, Liang Y, Yu Z, Liu X, Guo Y C, Liang D, Ouyang W, Cao Y P. Triposg: High-fidelity 3d shape synthesis using large-scale rectified flow models. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025, doi: 10.1109/TPAMI.2025.3633512 |
| [33] |
|
| [34] |
|
| [35] |
|
| [36] |
|
| [37] |
|
| [38] |
|
| [39] |
|
| [40] |
Wu S, Lin Y, Zhang F, Zeng Y, Yang Y, Bao Y, Qian J, Zhu S, Cao X, Torr P, Yao Y. Direct3D-S2: Gigascale 3D generation made easy with spatial sparse attention. 2025, arXiv preprint arXiv: 2505.17412 |
| [41] |
|
| [42] |
|
| [43] |
|
| [44] |
|
| [45] |
|
| [46] |
|
| [47] |
|
| [48] |
Oquab M, Darcet T, Moutakanni T, Vo H V, Szafraniec M, Khalidov V, Fernandez P, Haziza D, Massa F, El-Nouby A, Assran M, Ballas N, Galuba W, Howes R, Huang P Y, Li S W, Misra I, Rabbat M, Sharma V, Synnaeve G, Xu H, Jégou H, Mairal J, Labatut P, Joulin A, Bojanowski P. DINOv2: Learning robust visual features without supervision. 2023, arXiv preprint arXiv: 2304.07193 |
| [49] |
Ho J, Salimans T. Classifier-free diffusion guidance. 2022, arXiv preprint arXiv: 2207.12598 |
| [50] |
Chang A X, Funkhouser T, Guibas L, Hanrahan P, Huang Q, Li Z, Savarese S, Savva M, Song S, Su H, Xiao J, Yi L, Yu F. ShapeNet: An information-rich 3D model repository. 2015, arXiv preprint arXiv: 1512.03012 |
| [51] |
|
| [52] |
Xue L, Yu N, Zhang S, Panagopoulou A, Li J, Martín-Martín R, Wu J, Xiong C, Xu R, Niebles J C, Savarese S. ULIP-2: towards scalable multimodal pre-training for 3D understanding. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024, 27081−27091 |
| [53] |
|
Higher Education Press
/
| 〈 |
|
〉 |