CDJMP: an adaptive conditional diffusion model for multi-agent joint motion prediction in autonomous driving

Wenwen Zheng , Chao Huang , Hao Zhang , Huilin Yin

Autonomous Intelligent Systems ›› 2026, Vol. 6 ›› Issue (1) : 18

PDF
Autonomous Intelligent Systems ›› 2026, Vol. 6 ›› Issue (1) :18 DOI: 10.1007/s43684-026-00138-z
Original Article
research-article
CDJMP: an adaptive conditional diffusion model for multi-agent joint motion prediction in autonomous driving
Author information +
History +
PDF

Abstract

In autonomous driving trajectory prediction, it is important to generate multi-modal trajectories. As a generative method, diffusion model has been increasingly adopted in the field of trajectory prediction. In this paper, we propose CDJMP (Conditional Diffusion model-based Joint Motion Prediction), an adaptive conditional diffusion framework designed for multi-agent trajectory forecasting in autonomous driving. However, when conventional diffusion models generate multi-modal trajectories, particularly for multi-agent joint prediction, they often fail to balance the diversity of generated trajectories and the accuracy of prediction. At the same time, diffusion models also suffer from time-consuming inference. To address these issues, CDJMP utilizes a two-stage framework to generate diverse and highly accurate trajectories. In the first stage, a probabilistic initializer predicts proposal trajectories together with adaptive denoising steps. In the second stage, a group-aware conditional encoder captures dynamic multi-agent interactions and guides the diffusion process to produce coherent multimodal outcomes. Experiments on the INTERACTION and Argoverse datasets demonstrate that CDJMP achieves state-of-the-art performance, reducing minADE and minFDE by up to 9% and 10%, respectively, on the INTERACTION dataset. Ablation experiments demonstrate that CDJMP can effectively reduce inference time while maintaining prediction accuracy. These results highlight the potential of CDJMP as an efficient and accurate framework for real-time multi-agent trajectory prediction in autonomous driving.

Keywords

Motion prediction / Autonomous driving / Multi-modal trajectory prediction / Interaction modeling / Diffusion model

Cite this article

Download citation ▾
Wenwen Zheng, Chao Huang, Hao Zhang, Huilin Yin. CDJMP: an adaptive conditional diffusion model for multi-agent joint motion prediction in autonomous driving. Autonomous Intelligent Systems, 2026, 6 (1) : 18 DOI:10.1007/s43684-026-00138-z

登录浏览全文

4963

注册一个新账户 忘记密码

References

[1]

Liu J., Mao X., Fang Y., et al.. A survey on deep-learning approaches for vehicle trajectory prediction in autonomous driving. 2021 IEEE International Conference on Robotics and Biomimetics (ROBIO), 2021IEEE, 978-985

[2]

Kamenev A., Wang L., Bohan O.B., et al.. Predictionnet: real-time joint probabilistic traffic prediction for planning, control, and simulation. 2022 International Conference on Robotics and Automation (ICRA), 2022IEEE, 8936-8942

[3]

Phan-Minh T., Grigore E.C., Boulton F.A., et al.. Covernet: multimodal behavior prediction using trajectory sets. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 202014074-14083

[4]

Casas S., Gulino C., Liao R., et al.. Spagnn: spatially-aware graph neural networks for relational behavior forecasting from sensor data. 2020 IEEE International Conference on Robotics and Automation (ICRA), 2020IEEE, 9491-9497

[5]

Gao J., Sun C., Zhao H., et al.. Vectornet: encoding hd maps and agent dynamics from vectorized representation. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 202011525-11533

[6]

A. Cui, S. Casas, K. Wong, et al., GoRela: go relative for viewpoint-invariant motion forecasting. arXiv preprint (2022). arXiv:2211.02545

[7]

Fang L., Jiang Q., Shi J., et al.. Tpnet: trajectory proposal network for motion prediction. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 20206797-6806

[8]

M. Ye, J. Xu, X. Xu, et al., Dcms: motion forecasting with dual consistency and multi-pseudo-target supervision. arXiv preprint (2022). arXiv:2204.05859

[9]

Cui H., Radosavljevic V., Chou F.C., et al.. Multimodal trajectory predictions for autonomous driving using deep convolutional networks. 2019 International Conference on Robotics and Automation (ICRA), 2019IEEE, 2090-2096

[10]

Y. Chai, B. Sapp, M. Bansal, et al., Multipath: multiple probabilistic anchor trajectory hypotheses for behavior prediction. arXiv preprint (2019). arXiv:1910.05449

[11]

Liang M., Yang B., Hu R., et al.. Learning lane graph representations for motion forecasting. European Conference on Computer Vision, 2020, Berlin, Springer, 541-556

[12]

Gu J., Sun C., Zhao H.. Densetnt: end-to-end trajectory prediction from dense goal sets. Proceedings of the IEEE/CVF International Conference on Computer Vision, 202115303-15312

[13]

Varadarajan B., Hefny A., Srivastava A., et al.. Multipath++: efficient information fusion and trajectory aggregation for behavior prediction. 2022 International Conference on Robotics and Automation (ICRA), 2022IEEE, 7814-7821

[14]

Li L.L., Yang B., Liang M., et al.. End-to-end contextual perception and prediction with interaction transformer. 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2020IEEE, 5784-5791

[15]

Liu Y., Zhang J., Fang L., et al.. Multimodal motion prediction with stacked transformers. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 20217577-7586

[16]

Yuan Y., Weng X., Ou Y., et al.. Agentformer: agent-aware transformers for socio-temporal multi-agent forecasting. Proceedings of the IEEE/CVF International Conference on Computer Vision, 20219813-9823

[17]

J. Ngiam, B. Caine, V. Vasudevan, et al., Scene transformer: a unified architecture for predicting multiple agent trajectories. arXiv preprint (2021). arXiv:2106.08417

[18]

Zhou Z., Ye L., Wang J., et al.. Hivt: hierarchical vector transformer for multi-agent motion prediction. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 20228823-8833

[19]

N. Nayakanti, R. Al-Rfou, A. Zhou, et al., Wayformer: motion forecasting via simple & efficient attention networks. arXiv preprint (2022). arXiv:2207.05844

[20]

Shi S., Jiang L., Dai D., et al.. Motion transformer with global intention localization and local movement refinement. Adv. Neural Inf. Process. Syst., 2022, 35: 6531-6543

[21]

Dhariwal P., Nichol A.. Diffusion models beat gans on image synthesis. Adv. Neural Inf. Process. Syst., 2021, 34: 8780-8794

[22]

J. Song, C. Meng, S. Ermon, Denoising diffusion implicit models. arXiv preprint (2020). arXiv:2010.02502

[23]

Ho J., Jain A., Abbeel P.. Denoising diffusion probabilistic models. Adv. Neural Inf. Process. Syst., 2020, 33: 6840-6851

[24]

J. Ho, T. Salimans, Classifier-free diffusion guidance. arXiv preprint (2022). arXiv:2207.12598

[25]

Mao W., Xu C., Zhu Q., et al.. Leapfrog diffusion model for stochastic trajectory prediction. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 20235517-5526

[26]

Hagedorn S., Hallgarten M., Stoll M., et al.. The integration of prediction and planning in deep learning automated driving systems: a review. IEEE Trans. Intell. Veh., 2024, 10: 3626-3643

[27]

Lu Y., Wang W., Hu X., et al.. Vehicle trajectory prediction in connected environments via heterogeneous context-aware graph convolutional networks. IEEE Trans. Intell. Transp. Syst., 2022, 24(88452-8464

[28]

Song H., Luan D., Ding W., et al.. Learning to predict vehicle trajectories with model-based planning. Conference on Robot Learning, 2022PMLR, 1035-1045

[29]

Mo X., Liu H., Huang Z., et al.. Map-adaptive multimodal trajectory prediction via intention-aware unimodal trajectory predictors. IEEE Trans. Intell. Transp. Syst., 2023, 25(6): 5651-5663

[30]

Zhou Z., Wang J., Li Y.H., et al.. Query-centric trajectory prediction. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 202317863-17873

[31]

Jia X., Wu P., Chen L., et al.. Hdgt: heterogeneous driving graph transformer for multi-agent trajectory prediction via scene encoding. IEEE Trans. Pattern Anal. Mach. Intell., 2023, 45(11): 13860-13875

[32]

Wang X., Liu J., Lin H., et al.. A multi-modal spatial–temporal model for accurate motion forecasting with visual fusion. Inf. Fusion, 2024, 102 102046

[33]

Liu J., Lin H., Wang X., et al.. Reliable trajectory prediction in scene fusion based on spatio-temporal structure causal model. Inf. Fusion, 2024, 107 102309

[34]

Lu Y., Wang W., Bai R., et al.. Hyper-relational interaction modeling in multi-modal trajectory prediction for intelligent connected vehicles in smart cites. Inf. Fusion, 2025, 114 102682

[35]

C. Luo, Understanding diffusion models: a unified perspective. arXiv preprint (2022). arXiv:2208.11970

[36]

Gilles T., Sabatini S., Tsishkou D., et al.. Gohome: graph-oriented heatmap output for future motion estimation. 2022 International Conference on Robotics and Automation (ICRA), 2022IEEE, 9107-9114

[37]

Li J., Shen T., Gu Z., et al.. Adm: accelerated diffusion model via estimated priors for robust motion prediction under uncertainties. 2024 IEEE 27th International Conference on Intelligent Transportation Systems (ITSC), 2024IEEE, 2221-2227

[38]

Ronneberger O., Fischer P., Brox T.. U-net: convolutional networks for biomedical image segmentation. International Conference on Medical Image Computing and Computer-Assisted Intervention, 2015, Cham, Springer, 234-241

[39]

W. Zhan, L. Sun, D. Wang, et al., Interaction dataset: an international, adversarial and cooperative motion dataset in interactive driving scenarios with semantic maps. arXiv preprint (2019). arXiv:1910.03088

[40]

Chang M.F., Lambert J., Sangkloy P., et al.. Argoverse: 3d tracking and forecasting with rich maps. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 20198748-8757

[41]

Limeros S.C., Majchrowska S., Johnander J., et al.. Towards explainable motion prediction using heterogeneous graph representations. Transp. Res., Part C, Emerg. Technol., 2023, 157 104405

Funding

National Natural Science Foundation of China(62433014)

Rights & permissions

The Author(s)

PDF

3

Accesses

0

Citation

Detail

Sections
Recommended

/