Multimodal triplane diffusion transformer for high fidelity 3D shape generation

Shuang WU , Youtian LIN , Yifei ZENG , Feihu ZHANG , Hao ZHU , Xun CAO , Yao YAO

Front. Comput. Sci. ›› 2026, Vol. 20 ›› Issue (10) : 2010715

PDF (7564KB)
Front. Comput. Sci. ›› 2026, Vol. 20 ›› Issue (10) :2010715 DOI: 10.1007/s11704-026-50361-3
Image and Graphics
RESEARCH ARTICLE
Multimodal triplane diffusion transformer for high fidelity 3D shape generation
Author information +
History +
PDF (7564KB)

Abstract

Recent advancements in native 3D generation have demonstrated remarkable capabilities in producing high-quality 3D assets from image or text prompts. However, these methods face a critical challenge: insufficient alignment between generated meshes and input conditions. In this paper, we propose Multimodal Triplane Diffusion Transformer to address the issue, featuring two core components: A triplane-based 3D variational autoencoder that compresses point clouds, sampled uniformly from mesh surfaces and concentrated near sharp edges, into a triplane latent space, and a diffusion model trained on this latent space, empowered by multimodal diffusion transformer blocks to establish cross-modality interactions between latent representations and conditions. Extensive experiments demonstrate that our method achieves not only superior generalization capability but also significantly enhanced geometric alignment with input images compared to state-of-the-art approaches in the image-to-3D task.

Graphical abstract

Keywords

native 3D generation / diffusion model / multimodal diffusion transformer

Cite this article

Download citation ▾
Shuang WU, Youtian LIN, Yifei ZENG, Feihu ZHANG, Hao ZHU, Xun CAO, Yao YAO. Multimodal triplane diffusion transformer for high fidelity 3D shape generation. Front. Comput. Sci., 2026, 20 (10) : 2010715 DOI:10.1007/s11704-026-50361-3

登录浏览全文

4963

注册一个新账户 忘记密码

References

[1]

Podell D, English Z, Lacey K, Blattmann A, Dockhorn T, Müller J, Penna J, Rombach R. SDXL: Improving latent diffusion models for high-resolution image synthesis. In: Proceedings of 12th International Conference on Learning Representations. 2024

[2]

Schuhmann C, Beaumont R, Vencu R, Gordon C, Wightman R, Cherti M, Coombes T, Katta A, Mullis C, Wortsman M, Schramowski P, Kundurthy S, Crowson K, Schmidt L, Kaczmarczyk R, Jitsev J. LAION-5B: An open large-scale dataset for training next generation image-text models. In: Proceedings of the 36th International Conference on Neural Information Processing Systems. 2022, 1833

[3]

Deitke M, Schwenk D, Salvador J, Weihs L, Michel O, VanderBilt E, Schmidt L, Ehsani K, Kembhavi A, Farhadi A. Objaverse: A universe of annotated 3D objects. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2023, 13142−13153

[4]

Deitke M, Liu R, Wallingford M, Ngo H, Michel O, Kusupati A, Fan A, Laforte C, Voleti V, Gadre S Y, VanderBilt E, Kembhavi A, Vondrick C, Gkioxari G, Ehsani K, Schmidt L, Farhadi A. Objaverse-XL: A universe of 10M+ 3D objects. In: Proceedings of the 37th International Conference on Neural Information Processing Systems. 2023, 1554

[5]

Nichol A, Jun H, Dhariwal P, Mishkin P, Chen M. Point-E: A system for generating 3D point clouds from complex prompts. 2022, arXiv preprint arXiv: 2212.08751

[6]

Mildenhall B, Srinivasan P P, Tancik M, Barron J T, Ramamoorthi R, Ng R. NeRF: Representing scenes as neural radiance fields for view synthesis. In: Proceedings of the 16th European Conference on Computer Vision. 2020, 405−421

[7]

Poole B, Jain A, Barron J T, Mildenhall B. DreamFusion: Text-to-3D using 2D diffusion. In: Proceedings of the 11th International Conference on Learning Representations. 2023

[8]

Chen Z, Wang F, Wang Y, Liu H. Text-to-3D using Gaussian splatting. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024, 21401−21412

[9]

Wang P, Shi Y. ImageDream: Image-prompt multi-view diffusion for 3D generation. 2023, arXiv preprint arXiv: 2312.02201

[10]

Shi Y, Wang P, Ye J, Long M, Li K, Yang X. MVDream: Multi-view diffusion for 3D generation. In: Proceedings of the 12th International Conference on Learning Representations. 2024

[11]

Qiu L, Chen G, Gu X, Zuo Q, Xu M, Wu Y, Yuan W, Dong Z, Bo L, Han X. RichDreamer: A generalizable normal-depth diffusion model for detail richness in text-to-3D. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024, 9914−9925

[12]

Huang Z, Guo Y C, Wang H, Yi R, Ma L, Cao Y P, Sheng L. Mv-adapter: Multi-view consistent image generation made easy. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. 2025, 16377−16387

[13]

Long X, Guo Y C, Lin C, Liu Y, Dou Z, Liu L, Ma Y, Zhang S H, Habermann M, Theobalt C, Wang W. Wonder3D: Single image to 3D using cross-domain diffusion. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024, 9970−9980

[14]

Li J, Tan H, Zhang K, Xu Z, Luan F, Xu Y, Hong Y, Sunkavalli K, Shakhnarovich G, Bi S. Instant3D: fast text-to-3d with sparse-view generation and large reconstruction model. In: Proceedings of the 12th International Conference on Learning Representations. 2024

[15]

Liu M, Shi R, Chen L, Zhang Z, Xu C, Wei X, Chen H, Zeng C, Gu J, Su H. One-2-3-45++: fast single image to 3D objects with consistent multi-view generation and 3D diffusion. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024, 10072−10083

[16]

Xu J, Cheng W, Gao Y, Wang X, Gao S, Shan Y. InstantMesh: efficient 3D mesh generation from a single image with sparse-view large reconstruction models. 2024, arXiv preprint arXiv: 2404.07191

[17]

Shi R, Chen H, Zhang Z, Liu M, Xu C, Wei X, Chen L, Zeng C, Su H. Zero123++: a single image to consistent multi-view diffusion base model. 2023, arXiv preprint arXiv: 2310.15110

[18]

Lu Y, Zhang J, Li S, Fang T, McKinnon D, Tsin Y, Quarn L, Cao X, Yao Y. Direct2.5: Diverse text-to-3D generation via multi-view 2.5D diffusion. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024, 8744−8753

[19]

Shue J R, Chan E R, Po R, Ankner Z, Wu J, Wetzstein G. 3D neural field generation using triplane diffusion. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2023, 20875−20886

[20]

Zhao Z, Liu W, Chen X, Zeng X, Wang R, Cheng P, Fu B, Chen T, Yu G, Gao S. Michelangelo: Conditional 3D shape generation based on shape-image-text aligned latent representation. In: Proceedings of the 37th International Conference on Neural Information Processing Systems. 2023, 3236

[21]

Hong F, Tang J, Cao Z, Shi M, Wu T, Chen Z, Yang S, Wang T, Pan L, Lin D, Liu Z. 3DTopia: Large text-to-3D generation model with hybrid diffusion priors. 2024, arXiv preprint arXiv: 2403.02234

[22]

Zhang L, Wang Z, Zhang Q, Qiu Q, Pang A, Jiang H, Yang W, Xu L, Yu J . CLAY: a controllable large-scale generative model for creating high-quality 3D assets. ACM Transactions on Graphics, 2024, 43( 4): 120

[23]

Esser P, Kulal S, Blattmann A, Entezari R, Müller J, Saini H, Levi Y, Lorenz D, Sauer A, Boesel F, Podell D, Dockhorn T, English Z, Rombach R. Scaling rectified flow transformers for high-resolution image synthesis. In: Proceedings of the 41st International Conference on Machine Learning. 2024, 12606−12633

[24]

Wu S, Lin Y, Zhang F, Zeng Y, Xu J, Torr P, Cao X, Yao Y. Direct3D: Scalable image-to-3D generation via 3D latent diffusion transformer. In: Proceedings of the 38th International Conference on Neural Information Processing Systems. 2024, 3873

[25]

Peebles W, Xie S. Scalable diffusion models with transformers. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. 2023, 4172−4182

[26]

Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez A N, Kaiser Ł, Polosukhin I. Attention is all you need. In: Proceedings of the 31st International Conference on Neural Information Processing Systems. 2017, 6000−6010

[27]

Jun H, Nichol A. Shap-E: Generating conditional 3D implicit functions. 2023, arXiv preprint arXiv: 2305.02463

[28]

Hong Y, Zhang K, Gu J, Bi S, Zhou Y, Liu D, Liu F, Sunkavalli K, Bui T, Tan H. LRM: large reconstruction model for single image to 3D. In: Proceedings of the 12th International Conference on Learning Representations. 2024

[29]

Lan Y, Hong F, Yang S, Zhou S, Meng X, Dai B, Pan X, Loy C C. LN3DIFF: scalable latent neural fields diffusion for speedy 3D generation. In: Proceedings of the 18th European Conference on Computer Vision. 2025, 112−130

[30]

Xiang X K, Yuan Y J, Hu W B, Liu Y T, Ma Y W, Gao L . PGT-NeuS: Progressive-growing tri-plane representation for neural surface reconstruction. IEEE Transactions on Visualization and Computer Graphics, 2025, 31( 10): 9213–9224

[31]

Zhang B, Tang J, Nießner M, Wonka P . 3DShape2VecSet: a 3D shape representation for neural fields and generative diffusion models. ACM Transactions on Graphics, 2023, 42( 4): 92

[32]

Li Y, Zou Z X, Liu Z, Wang D, Liang Y, Yu Z, Liu X, Guo Y C, Liang D, Ouyang W, Cao Y P. Triposg: High-fidelity 3d shape synthesis using large-scale rectified flow models. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025, doi: 10.1109/TPAMI.2025.3633512

[33]

Chen R, Zhang J, Liang Y, Luo G, Li W, Liu J, Li X, Long X, Feng J, Tan P. Dora: Sampling and benchmarking for 3D shape variational auto-encoders. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2025, 16251−16261

[34]

Gao J, Shen T, Wang Z, Chen W, Yin K, Li D, Litany O, Gojcic Z, Fidler S. GET3D: a generative model of high quality 3D textured shapes learned from images. In: Proceedings of the 36th International Conference on Neural Information Processing Systems. 2022, 2308

[35]

Lin C H, Gao J, Tang L, Takikawa T, Zeng X, Huang X, Kreis K, Fidler S, Liu M Y, Lin T Y. Magic3D: High-resolution text-to-3D content creation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2023, 300−309

[36]

Shen T, Gao J, Yin K, Liu M Y, Fidler S. Deep marching tetrahedra: a hybrid representation for high-resolution 3D shape synthesis. In: Proceedings of the 35th International Conference on Neural Information Processing Systems. 2021, 466

[37]

Yang J, Mo K, Lai Y K, Guibas L J, Gao L . DSG-Net: Learning disentangled structure and geometry for 3D shape generation. ACM Transactions on Graphics, 2023, 42( 1): 1

[38]

Ren X, Huang J, Zeng X, Museth K, Fidler S, Williams F. XCube: Large-scale 3D generative modeling using sparse voxel hierarchies. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024, 4209−4219

[39]

Xiang J, Lv Z, Xu S, Deng Y, Wang R, Zhang B, Chen D, Tong X, Yang J. Structured 3D latents for scalable and versatile 3D generation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2025, 21469−21480

[40]

Wu S, Lin Y, Zhang F, Zeng Y, Yang Y, Bao Y, Qian J, Zhu S, Cao X, Torr P, Yao Y. Direct3D-S2: Gigascale 3D generation made easy with spatial sparse attention. 2025, arXiv preprint arXiv: 2505.17412

[41]

He X, Zou Z X, Chen C H, Guo Y C, Liang D, Yuan C, Ouyang W, Cao Y P, Li Y. SparseFlex: High-resolution and arbitrary-topology 3D shape modeling. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. 2025, 14822−14833

[42]

Liu R, Wu R, Van Hoorick B, Tokmakov P, Zakharov S, Vondrick C. Zero-1-to-3: Zero-shot one image to 3D object. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. 2023, 9264−9275

[43]

Wang P, Liu L, Liu Y, Theobalt C, Komura T, Wang W. NeuS: learning neural implicit surfaces by volume rendering for multi-view reconstruction. In: Proceedings of the 35th International Conference on Neural Information Processing Systems. 2021, 2081

[44]

Wu K, Liu F, Cai Z, Yan R, Wang H, Hu Y, Duan Y, Ma K. Unique3D: High-quality and efficient 3D mesh generation from a single image. In: Proceedings of the 38th International Conference on Neural Information Processing Systems. 2024, 3974

[45]

Li P, Liu Y, Long X, Zhang F, Lin C, Li M, Qi X, Zhang S, Xue W, Luo W, Tan P, Wang W, Liu Q, Guo Y. Era3D: high-resolution multiview diffusion using efficient row-wise attention. In: Proceedings of the 38th International Conference on Neural Information Processing Systems. 2024, 1780

[46]

Li W, Liu J, Yan H, Chen R, Liang Y, Chen X, Tan P, Long X. CraftsMan3D: High-fidelity mesh generation with 3D native diffusion and interactive geometry refiner. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2025, 5307−5317

[47]

He K, Zhang X, Ren S, Sun J. Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2016, 770−778

[48]

Oquab M, Darcet T, Moutakanni T, Vo H V, Szafraniec M, Khalidov V, Fernandez P, Haziza D, Massa F, El-Nouby A, Assran M, Ballas N, Galuba W, Howes R, Huang P Y, Li S W, Misra I, Rabbat M, Sharma V, Synnaeve G, Xu H, Jégou H, Mairal J, Labatut P, Joulin A, Bojanowski P. DINOv2: Learning robust visual features without supervision. 2023, arXiv preprint arXiv: 2304.07193

[49]

Ho J, Salimans T. Classifier-free diffusion guidance. 2022, arXiv preprint arXiv: 2207.12598

[50]

Chang A X, Funkhouser T, Guibas L, Hanrahan P, Huang Q, Li Z, Savarese S, Savva M, Song S, Su H, Xiao J, Yi L, Yu F. ShapeNet: An information-rich 3D model repository. 2015, arXiv preprint arXiv: 1512.03012

[51]

Loshchilov I, Hutter F. Decoupled weight decay regularization. In: Proceedings of the 7th International Conference on Learning Representations. 2019

[52]

Xue L, Yu N, Zhang S, Panagopoulou A, Li J, Martín-Martín R, Wu J, Xiong C, Xu R, Niebles J C, Savarese S. ULIP-2: towards scalable multimodal pre-training for 3D understanding. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2024, 27081−27091

[53]

Zhou J, Wang J, Ma B, Liu Y S, Huang T, Wang X. Uni3D: exploring unified 3D representation at scale. In: Proceedings of the 12th International Conference on Learning Representations. 2024

Rights & permissions

Higher Education Press

PDF (7564KB)

Supplementary files

Highlights

365

Accesses

0

Citation

Detail

Sections
Recommended

/