A generation-based defect detection system for rail transit infrastructure

Xinyu Zheng , Lingfeng Zhang , Yuhao Luo , Tiange Wang

High-speed Railway ›› 2026, Vol. 4 ›› Issue (1) : 1 -9.

PDF (3701KB)
High-speed Railway ›› 2026, Vol. 4 ›› Issue (1) :1 -9. DOI: 10.1016/j.hspr.2025.09.004
Research article
research-article
A generation-based defect detection system for rail transit infrastructure
Author information +
History +
PDF (3701KB)

Abstract

The use of Unmanned Aerial Vehicles (UAVs) for defect detection on railway slopes is becoming increasingly widespread due to their ability to capture high-resolution images over large, inaccessible, and topographically complex areas. However, current UAV-based detection methods face several critical limitations, including constrained deployment frequency, limited availability of annotated defect data, and the lack of mature risk assessment frameworks. To address these challenges, this study introduces a novel approach that integrates diffusion models with Large Language Models (LLMs) to generate high-quality synthetic defect images tailored to railway slope scenarios. Furthermore, an improved transformer-based architecture is proposed, incorporating attention mechanisms and LLM-guided diffusion-generated imagery to enhance defect recognition performance under complex environmental conditions. Experimental evaluations conducted on a dataset of 300 field-collected images from high-risk railway slopes demonstrate that the proposed method significantly outperforms existing baselines in terms of precision, recall, and robustness, indicating strong applicability for real-world railway infrastructure monitoring and disaster prevention.

Keywords

Railway / Large language models / Computer vision / Object detection

Cite this article

Download citation ▾
Xinyu Zheng, Lingfeng Zhang, Yuhao Luo, Tiange Wang. A generation-based defect detection system for rail transit infrastructure. High-speed Railway, 2026, 4 (1) : 1-9 DOI:10.1016/j.hspr.2025.09.004

登录浏览全文

4963

注册一个新账户 忘记密码

References

[1]

J. Liu, Y. Zhang, J. Han, et al., Intelligent hazard-risk prediction model for train control systems, IEEE Trans. Intell. Transp. Syst. 21 (2019) 4693-4704.

[2]

Y. Yürekli, C. Özarpa, İ. Avcı, Real-time railway hazard detection using distributed acoustic sensing and hybrid ensemble learning, Sensors 25 (2025) 3992.

[3]

M. Ezuma, F. Erden, C.K. Anjinappa, et al., Micro-UAV detection and classification from rf fingerprints using machine learning techniques, 2019 IEEE Aerospace Conference, IEEE, 2019, pp. 1-13.

[4]

R. Opromolla, G. Fasano, D. Accardo, A vision-based approach to uav detection and tracking in cooperative applications, Sensors 18 (2018) 3391.

[5]

J. Zhao, J. Zhang, D. Li, et al., Vision-based anti-uav detection and tracking, IEEE Trans. Intell. Transp. Syst. 23 (2022) 25323-25334.

[6]

L.M. Dang, S.I. Hassan, I. Suhyeon, et al., UAV based wilt detection system via convolutional neural networks, Sustain. Comput. Inform. Syst. 28 (2020) 100250.

[7]

Q. Shi, J. Li, Objects detection of uav for anti-uav based on YOLOv4, 2020 IEEE 2nd International Conference on Civil Aviation Safety and Information Technology (ICCASIT), IEEE, 2020, pp. 1048-1052.

[8]

C.F.R. Chen, Q. Fan, R. Panda, Crossvit: Cross-attention multi-scale vision transformer for image classification, Proceedings of the IEEE/CVF international conference on computer vision, IEEE, 2021, pp. 357-366.

[9]

K. Han, Y. Wang, H. Chen, et al., A survey on vision transformer, IEEE Trans. Pattern Anal. Mach. Intell. 45 (2022) 87-110.

[10]

Z. Liu, Y. Lin, Y. Cao, et al., Swin transformer: Hierarchical vision transformer using shifted windows, Proceedings of the IEEE/CVF international conference on computer vision, IEEE, 2021, pp. 10012-10022.

[11]

Z. Xi, W. Chen, X. Guo, et al., The rise and potential of large language model based agents: A survey, Sci. China Inf. Sci. 68 (2025) 121101.

[12]

L. Zhang, X. Hao, Q. Xu, et al., 2025.Mapnav: A novel memory representation via annotated semantic maps for vlm-based vision-and-language navigation. arXiv:2502.13451, 2025.

[13]

L. Zhang, H. Wang, E. Xiao, et al., Multi-floor zero-shot object navigation policy. arXiv:2409.10906, 2024.

[14]

L. Zhang, Q. Zhang, H. Wang, et al., Trihelper: Zero-shot object navigation with dynamic assistance, 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), IEEE, 2024, pp. 10035-10042.

[15]

W.X. Zhao, K. Zhou, J. Li, et al., A survey of large language models. arXiv:2303. 18223, 2023.

[16]

J. Li, A. Hassani, S. Walton, et al., Convmlp: Hierarchical convolutional mlps for vision, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, IEEE, 2023, 6307-6316.

[17]

Ç. Aytekin, Y. Rezaeitabar, S. Dogru, et al., Railway fastener inspection by real- time machine vision, IEEE Trans. Syst. Man Cybern. Syst. 45 (2015) 1101-1107.

[18]

Y. Min, B. Xiao, J. Dang, et al., Real time detection system for rail surface defects based on machine vision, EURASIP J. Image Video Process. (2018) 1-11.

[19]

T. Wang, F. Yang, K.L. Tsui, Real-time detection of railway track component via one-stage deep learning networks, Sensors 20 (2020) 4325.

[20]

T. Wang, Z. Zhang, K.L. Tsui, A deep generative approach for rail foreign object detections via semisupervised learning, IEEE Trans. Ind. Inform. 19 (2022) 459-468.

[21]

X. Zheng, Y. He, Y. Luo, et al., Railway side slope hazard detection system based on generative models, IEEE Sens. J. 25 (9) (2025) 16281-16296.

[22]

M. Karakose, O. Yaman, M. Baygin, et al., A new computer vision based method for rail track detection and fault diagnosis in railways, Int. J. Mech. Eng. Robot. Res. 6 (2017) 17-22.

[23]

S. Faghih-Roohi, S. Hajizadeh, A. Núñez, et al., Deep Convolutional Neural Networks for Detection of Rail Surface Defects, 2016 International Joint Conference on Neural Networks (IJCNN), IEEE, 2016, pp. 2584-2589.

[24]

X.T. Jin, Y.N. Wang, H. Zhang, et al., Deeprail: Automatic visual detection system for railway surface defect using Bayesian cnn and attention network, Acta Autom. Sin. 45 (2019) 2312-2327.

[25]

X. Gibert, V.M. Patel, R. Chellappa, Deep multitask learning for railway track inspection, IEEE Trans. Intell. Transp. Syst. 18 (2016) 153-164.

[26]

F. Guo, J. Liu, Y. Qian, et al., Rail surface defect detection using a transformer-based network, J. Ind. Inf. Integr. 38 (2024) 100584.

[27]

W. Ye, J. Ren, C. Li, et al., Intelligent detection of surface defects in high-speed railway ballastless track based on self-attention and transfer learning, Struct. Control Health Monit. 2024 (2024) 2967927.

[28]

A. Creswell, T. White, V. Dumoulin, et al., Generative adversarial networks: An overview, IEEE Signal Process. Mag. 35 (2018) 53-65.

[29]

I. Goodfellow, J. Pouget-Abadie, M. Mirza, et al., Generative adversarial networks, Commun. ACM 63 (2020) 139-144.

[30]

D.P. Kingma, M. Welling, Auto-encoding variational bayes. arXiv:1312.6114, 2022.

[31]

Y. Luo, K. Chen, M. Zhu, Granp: A graph recurrent attentive neural process model for vehicle trajectory prediction, 2024 IEEE Intelligent Vehicles Symposium (IV), IEEE, 2024, pp. 370-375.

[32]

K. Chen, Y. Luo, M. Zhu, et al., Human-like interactive lane-change modeling based on reward-guided diffusive predictor and planner, IEEE Trans. Intell. Transp. Syst. 26 (3) (2024) 3903-3916.

[33]

J. Ho, A. Jain, P. Abbeel, Denoising diffusion probabilistic models, Adv. Neural Inf. Process. Syst. 33 (2020) 6840-6851.

[34]

C. Zhang, Y. Zhang, Q. Shao, et al., Chattraffic: Text-to-traffic generation via diffusion model, IEEE Trans. Intell. Transp. Syst. 26 (2) (2024) 2656-2668.

[35]

N. Sivaroopan, D. Bandara, C. Madarasingha, et al., Netdiffus: Network traffic generation by diffusion models through time-series imaging, Comput. Netw. 251 (2024) 110616.

[36]

X. Jiang, S. Liu, A. Gember-Jacobson, et al., Netdiffusion: Network data augmentation through protocol-constrained traffic generation, Proc. ACM Meas. Anal. Comput. Syst. 8 (2024) 1-32.

[37]

J. Lu, S. Azam, G. Alcan, et al., Data-driven diffusion models for enhancing safety in autonomous vehicle traffic simulations. arXiv:2410.04809, 2024.

[38]

C. Xu, A. Petiushko, D. Zhao, et al. Diffscene: iffusion-based safety-critical scenario generation for autonomous vehicles, Proceedings of the AAAI Conference on Artificial Intelligence (2025) 8797-8805.

[39]

C. Yang, Y. He, A.X. Tian, et al., Wcdt: World-centric diffusion transformer for traffic scene generation. arXiv:2404.02082, 2024.

[40]

L. Floridi, M. Chiriatti, GPT-3: Its nature, scope, limits, and consequences, Minds Mach. 30 (2020) 681-694.

[41]

M. Heusel, H. Ramsauer, T. Unterthiner, et al., Gans trained by a two time-scale update rule converge to a local nash equilibrium, Advances in Neural Information Processing Systems, 2017, 6626-6637.

[42]

T. Salimans, I. Goodfellow, W. Zaremba, et al., Improved techniques for training gans, Advances in Neural Information Processing Systems, 2016, 2234-2242.

[43]

R. Zhang, P. Isola, A.A. Efros, et al., The unreasonable effectiveness of deep features as a perceptual metric, Proceedings of the IEEE Conference On Computer Vision and Pattern Recognition, 2018, 586-595.

[44]

Z. Wang, A.C. Bovik, H.R. Sheikh, et al., Image quality assessment: From error visibility to structural similarity, IEEE Trans. Image Process. 13 (2004) 600-612.

PDF (3701KB)

15

Accesses

0

Citation

Detail

Sections
Recommended

/