Computational intelligence for road pavement condition assessment: a deep learning perspective
Yu Hu , Rakiba Rayhana , Ling Bai , Zheng Liu
Urban Lifeline ›› 2026, Vol. 4 ›› Issue (1) : 18
Pavement defects such as cracks and potholes compromise road safety and demand timely maintenance. Traditional manual inspection is slow and exposes workers to safety risks, whereas automated systems offer a promising alternative. This survey provides a comprehensive review of deep learning methods for road condition assessment. We first examine 2D image-based approaches, tracing their evolution from convolutional neural networks (CNNs) to Transformers. Although these methods are widely adopted, they remain sensitive to lighting conditions and cannot directly capture physical properties such as defect depth. To address these limitations, we review 3D sensing and subsurface diagnostic techniques, which provide essential geometric information for severity assessment. The primary focus of this paper is on evaluation: we summarize key public datasets and evaluation metrics and analyze the persistent gap between algorithmic performance and the practical needs of engineering, emphasizing the importance of assessing the actual utility of models in the field. Finally, we discuss several key challenges and promising research avenues, arguing that future work should prioritize model robustness, reliability, and the integration of these systems into real-world maintenance workflows.
Pavement defect detection / Road condition assessment / Deep learning / Transformers / Self-supervised learning / Vision-language models / LiDAR/3D
| [1] |
Lim RS, La HM, Shan Z, Sheng W (2011) Developing a crack inspection robot for bridge maintenance. In: 2011 IEEE International Conference on Robotics and Automation, IEEE, pp 6288–6293 |
| [2] |
|
| [3] |
|
| [4] |
|
| [5] |
Miller JS, Bellinger WY et al (2003) Distress identification manual for the long-term pavement performance program. Tech rep, United States. Department of Transportation. Federal Highway Administration |
| [6] |
|
| [7] |
Bureau of Infrastructure and Transport Research Economics (2024) Australian infrastructure and transport statistics yearbook 2024: transport safety. https://www.bitre.gov.au/publications/2024/australian-infrastructure-and-transport-statistics-yearbook-2024/transport-safety. Accessed 27 Apr 2026 |
| [8] |
|
| [9] |
|
| [10] |
|
| [11] |
|
| [12] |
|
| [13] |
Ohtsu M (2020) Acoustic emission and related non-destructive evaluation techniques in the fracture mechanics of concrete: fundamentals and applications, 2nd ed. Oxford, UK: Woodhead Publishing. https://shop.elsevier.com/books/acoustic-emission-and-related-non-destructive-evaluation-techniques-in-the-fracture-mechanics-of-concrete/ohtsu/978-0-12-822136-5 |
| [14] |
Graham-Jones J, Summerscales J (2015) Marine applications of advanced fibre-reinforced composites. Amsterdam, The Netherlands: Woodhead Publishing. https://shop.elsevier.com/books/marine-applications-of-advanced-fibre-reinforced-composites/graham-jones/978-1-78242-250-1 |
| [15] |
|
| [16] |
|
| [17] |
|
| [18] |
Du Y, Zhou Z, Wu Q, Huang H, Xu M, Cao J, Hu G (2020) A pothole detection method based on 3D point cloud segmentation. In: Twelfth International Conference on Digital Image Processing (ICDIP 2020), SPIE, vol 11519, pp 56–64 |
| [19] |
Yu Y, Li J, Guan H, Wang C (2014) 3D crack skeleton extraction from mobile lidar point clouds. In: 2014 IEEE geoscience and remote sensing symposium, IEEE, pp 914–917 |
| [20] |
|
| [21] |
|
| [22] |
|
| [23] |
|
| [24] |
Krizhevsky A, Sutskever I, Hinton GE (2012) Imagenet classification with deep convolutional neural networks. In: Advances in Neural Information Processing Systems, Curran Associates, vol 25, pp 1097–1105 |
| [25] |
Fan J, Bocus MJ, Wang L, Fan R (2021) Deep convolutional neural networks for road crack detection: qualitative and quantitative comparisons. In: 2021 IEEE International Conference on Imaging Systems and Techniques (IST), IEEE, pp 1–6 |
| [26] |
|
| [27] |
|
| [28] |
|
| [29] |
|
| [30] |
|
| [31] |
|
| [32] |
|
| [33] |
|
| [34] |
|
| [35] |
Ma N, Song Z, Hu Q, Liu CW, Han Y, Zhang Y, Fan R, Xie L (2025) Vehicular road crack detection with deep learning: a new online benchmark for comprehensive evaluation of existing algorithms. Preprint at https://arxiv.org/abs/2503.18082 |
| [36] |
|
| [37] |
Liu H, Miao X, Mertz C, Xu C, Kong H (2021) Crackformer: transformer network for fine-grained crack detection. In: Proceedings of the IEEE/CVF international conference on computer vision, pp 3783–3792 |
| [38] |
|
| [39] |
|
| [40] |
|
| [41] |
Liu X, Wu K, Cai X, Huang W (2024) Semi-supervised semantic segmentation using cross-consistency training for pavement crack detection. Road Mater Pavement Des 25(6):1368–1380 |
| [42] |
|
| [43] |
Jongwiriyanurak N, Zeng Z, Goo JM, Wang X, Ilyankou I, Sriroongvikrai K, Christie N, Wang M, Chen H, Haworth J (2024) V-roast: visual road assessment. Can vlm be a road safety assessor using the irap standard? Preprint at https://arxiv.org/abs/2408.10872 |
| [44] |
|
| [45] |
Zan C, Du S, Ikenaga T (2025) Clip-guided cross-modal feature fusion based few-shot learning for nighttime pavement defect detection. In: 2025 19th International Conference on Machine Vision and Applications (MVA), IEEE, pp 1–5 |
| [46] |
Federal Highway Administration (2023) Successful practices for quality management of pavement surface condition data collection and analysis. Tech Rep FHWA-RC-23-0002, Federal Highway Administration |
| [47] |
Federal Highway Administration (2023) Unmanned aircraft systems (UAS) for highway construction and maintenance: a field use guide. Tech Rep FHWA-HIF-23-012, Federal Highway Administration |
| [48] |
Chang GK, Sankaranarayanan S, Gilliland A (2024) NDT and ICT for asphalt pavement construction: Techbrief (PMTP/IC/DPS/VETA). Tech Rep FHWA-HIF-24-031, Federal Highway Administration. https://rosap.ntl.bts.gov/view/dot/78767. Accessed 20 Jan 2026 |
| [49] |
American Association of State Highway and Transportation Officials (2022) Standard practice for continuous thermal profile of asphalt mixture during construction. AASHTO R110-22, Washington, DC, USA. https://store.accuristech.com/standards/aashto-r-110-22. Accessed 20 Jan 2026 |
| [50] |
ASTM International (2022) Standard test method for measuring the p-wave speed and the thickness of concrete plates using the impact-echo method. ASTM C1383-15(2022). https://www.astm.org/c1383-15r22.html. Accessed 20 Jan 2026 |
| [51] |
Edmund Optics (2024) Successful light polarization techniques. Edmund Optics Knowledge Center. https://www.edmundoptics.com/knowledge-center/application-notes/illumination/successful-light-polarization-techniques/. Accessed 20 Jan 2026 |
| [52] |
|
| [53] |
|
| [54] |
Eriksson J, Girod L, Hull B, Newton R, Madden S, Balakrishnan H (2008) The pothole patrol: using a mobile sensor network for road surface monitoring. In: Proceedings of MobiSys 2008, pp 29–39. https://nms.csail.mit.edu/papers/p2-mobisys-2008.pdf. Accessed 12 Oct 2025 |
| [55] |
Olsen MJ (2013) Guidelines for the use of mobile LIDAR in transportation applications. NCHRP Rep 748. Washington, DC, USA: Transportation Research Board. https://highways.fhwa.dot.gov/safety/data-analysis-tools/rsdp/rsdp-tools/national-cooperative-highway-research-program-nchrp-5 |
| [56] |
US Geological Survey (2024) Lidar Base Specification 2024 Rev. A. https://www.usgs.gov/media/files/lidar-base-specification-2024-rev-a. Accessed 28 Jan 2026 |
| [57] |
American Society for Photogrammetry and Remote Sensing (2019) Las specification, version 1.4–R15. https://www.asprs.org/wp-content/uploads/2019/07/LAS_1_4_r15.pdf. Accessed 18 Dec 2025 |
| [58] |
Pavemetrics Systems Inc (2021) LCMS-2D / LCMS-3D whitepaper. White paper. https://www.pavemetrics.com/downloads/lcms-2d-3d-whitepaper/. Accessed 28 Jan 2026 |
| [59] |
International Organization for Standardization (2002) Characterization of pavement texture by use of surface profiles - Part 3: Specification and classification of profilometers, ISO 13473-3:2002. International Organization for Standardization, Geneva. https://www.iso.org/standard/29426.html |
| [60] |
ASTM International (2023) Standard practice for calculating pavement macrotexture mean profile depth. ASTM E1845-23. https://www.astm.org/e1845-23.html. Accessed 28 Jan 2026 |
| [61] |
Fan R, Liu Y, Yang X, Bocus MJ, Dahnoun N, Tancock S (2018) Real-time stereo vision for road surface 3-d reconstruction. In: 2018 IEEE International Conference on Imaging Systems and Techniques (IST), IEEE, pp 1–6 |
| [62] |
Basler AG (2023) Time-of-flight versus stereo vision – who scores where? Technical article. https://www.baslerweb.com/en/learning/time-of-flight-stereovision/. Accessed 10 Jan 2026 |
| [63] |
|
| [64] |
American Association of State Highway and Transportation Officials (2018) AASHTO R 37-04 (2018) Standard practice for application of ground penetrating radar (GPR) to highways. American Association of State Highway and Transportation Officials, Washington. https://store.transportation.org. Accessed 28 Jan 2026 |
| [65] |
ASTM International (2022) Standard test method for evaluating asphalt-covered concrete bridge decks using ground penetrating radar, ASTM D6087–22. Standard. https://doi.org/10.1520/D6087-22 |
| [66] |
Texas Department of Transportation (2024) Seismic evaluation tools (PSPA/DSPA) — TxDOT pavement evaluation manual. Technical manual. https://www.txdot.gov/. Accessed 20 Jan 2026 |
| [67] |
Federal Aviation Administration (2021) Portable seismic property analyzer (PSPA). https://www.airporttech.tc.faa.gov/Airport-Pavement/Evaluation-Management/Nondestructive-Testing-Technology/NDT-Technology/Portable-Seismic-Pavement-Analyzer. Accessed 20 Jan 2026 |
| [68] |
ASTM International (2018) Standard test method for measuring the longitudinal profile of traveled surfaces with an accelerometer-established inertial profiling reference, ASTM E950/E950M-09(2018). Standard. https://www.astm.org/e0950_e0950m-09r18.html. Accessed 20 Jan 2026 |
| [69] |
ASTM International (2021) Standard practice for computing international roughness index of roads from longitudinal profile measurements, ASTM E1926-08(2021). https://www.astm.org/e1926-08r21.html. Accessed 20 Jan 2026 |
| [70] |
Sayers MW, Gillespie TD, Queiroz CAV (1986) The international road roughness experiment: establishing correlation and a calibration standard for measurements. Technical Paper 45, World Bank, Washington, DC. https://documents.worldbank.org/en/publication/documents-reports/documentdetail/326081468740204115. Accessed 20 Jan 2026 |
| [71] |
Sayers MW, Gillespie TD, Queiroz CAV (1986) The international road roughness experiment: a basis for establishing a standard scale for road roughness measurements. Transportation Research Record (1084):76–85. https://onlinepubs.trb.org/Onlinepubs/trr/1986/1084/1084-010.pdf. Accessed 20 Jan 2026 |
| [72] |
American Association of State Highway and Transportation Officials (2020) AASHTO R 32-20: Standard recommended practice for calibrating the load cell and deflection sensors for a falling weight deflectometer. Standard, Washington, DC. https://store.accuristech.com/standards/aashto-r-32-20. Accessed 20 Jan 2026 |
| [73] |
ASTM International (2020) Standard test method for deflections with a falling-weight-type impulse load device, ASTM D4694-09(2020).https://www.astm.org/d4694-09r20.html. Accessed 20 Jan 2026 |
| [74] |
ASTM International (2020) Standard guide for calculating in situ equivalent elastic moduli of pavement materials using layered elastic theory, ASTM D5858-96(2020). Standard. https://www.astm.org/d5858-96r20.html. Accessed 20 Jan 2026 |
| [75] |
Federal Highway Administration (2017) 23 CFR Part 490 Subpart C: national performance management measures for assessing pavement condition. Office of the Federal Register, National Archives and Records Administration. https://www.ecfr.gov/current/title-23/chapter-I/subchapter-E/part-490/subpart-C. Accessed 20 Jan 2026 |
| [76] |
Federal Highway Administration (2024) Highway performance monitoring system (HPMS) field manual. Tech Rep No. FHWA-2023-0014-0003. https://downloads.regulations.gov/FHWA-2023-0014-0003/attachment_1.pdf. Accessed 20 Jan 2026 |
| [77] |
Grogg M, Van T, Rozycki R, Vaughn R, Roff T, Clarke J, Beatty W, Buck J, Christenson A, Chang C (2018) Computation procedure for the pavement condition measures. Tech Rep No. FHWA-HIF-18-022. https://www.fhwa.dot.gov/tpm/guidance/hif18022.pdf. Accessed 20 Jan 2026 |
| [78] |
ASTM International (2023) Standard practice for roads and parking lots pavement condition index surveys, ASTM D6433-23. https://www.astm.org/d6433-23.html. Accessed 20 Jan 2026 |
| [79] |
ASTM International (2023) Standard test method for measuring rut-depth of pavement surfaces using a straightedge, ASTM E1703/E1703M-10(2023). https://www.astm.org/e1703_e1703m-10r23.html. Accessed 20 Jan 2026 |
| [80] |
Federal Highway Administration (2014) Distress identification manual for the long-term pavement performance program (fifth revised edition). Tech Rep No. FHWA-HRT-13-092, Federal Highway Administration, Office of Infrastructure Research and Development. https://highways.dot.gov/sites/fhwa.dot.gov/files/docs/research/long-term-pavement-performance/products/1401/distress-identification-manual-13092.pdf. Accessed 20 Jan 2026 |
| [81] |
|
| [82] |
Swarna S, Tech M, Hossain K (2018) Effect of interface bonds on pavement performance. In: Proceedings of the 2018 TAC Conference, Saskatoon, Saskatchewan. Transportation Association of Canada (TAC), Ottawa, Ontario |
| [83] |
|
| [84] |
Simonyan K, Zisserman A (2015) Very deep convolutional networks for large-scale image recognition. Preprint at https://arxiv.org/abs/1409.1556 |
| [85] |
He K, Zhang X, Ren S, Sun J (2016) Deep residual learning for image recognition. In: CVPR, pp 770–778. https://doi.org/10.1109/CVPR.2016.90 |
| [86] |
Su QF, Zhou Z, Zhu M, Wang H, Ye J (2020) Concrete crack detection based on EfficientNet-B0 with transfer learning. IEEE Access |
| [87] |
Ren S, He K, Girshick R, Sun J (2015) Faster R-CNN: towards real-time object detection with region proposal networks. In: in Proc Adv Neural Inf Process Syst, pp 91–99 |
| [88] |
Redmon J, Divvala S, Girshick R, Farhadi A (2016) You only look once: unified, real-time object detection. In: in Proc IEEE Conf Comput Vis Pattern Recognit, pp 779–788 |
| [89] |
Lin TY, Goyal P, Girshick R, He K, Dollár P (2018) Focal loss for dense object detection. Preprint at https://arxiv.org/abs/1708.02002 |
| [90] |
Ronneberger O, Fischer P, Brox T (2015) U-net: convolutional networks for biomedical image segmentation. In: Proc MICCAI, Springer, pp 234–241. https://doi.org/10.1007/978-3-319-24574-4_28 |
| [91] |
Chen LC, Zhu Y, Papandreou G, Schroff F, Adam H (2018) Encoder-decoder with atrous separable convolution for semantic image segmentation. In: Proc ECCV, pp 801–818. Preprint at https://arxiv.org/abs/1802.02611 |
| [92] |
He K, Gkioxari G, Dollár P, Girshick R (2017) Mask R-CNN. In: Proceedings of the IEEE international conference on computer vision, pp 2961–2969 |
| [93] |
Tang W, Huang S, Zhang X, Huangfu L (2022) Pict: a slim weakly supervised vision transformer for pavement distress classification. Preprint at https://arxiv.org/abs/2209.10074 |
| [94] |
|
| [95] |
|
| [96] |
|
| [97] |
|
| [98] |
|
| [99] |
Zhao Y, Lv W, Xu S, Wei J, Wang G, Dang Q, Liu Y, Chen J (2024) DETRs beat YOLOs on real-time object detection. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp 16965–16974. https://doi.org/10.1109/CVPR52733.2024.01605 |
| [100] |
|
| [101] |
|
| [102] |
|
| [103] |
Yu M, Wu D, Rao W, Cheng L, Li R, Li Y (2022) Automated road crack detection method based on visual transformer with multi-head cross-attention. In: 2022 IEEE International Conference on Sensing, Diagnostics, Prognostics, and Control (SDPC), pp 328–332. https://doi.org/10.1109/SDPC55702.2022.9915808 |
| [104] |
|
| [105] |
|
| [106] |
|
| [107] |
|
| [108] |
Li Q, Arnab A, Yang Y, Dehghani M, Hassani A, Gritsenko A, Wang X, Zhai X, Lučić M, Houlsby N (2022) Language-driven semantic segmentation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), LSeg |
| [109] |
Liang F, Li Q, Yu H, Wang W et al (2023) Open-vocabulary semantic segmentation with mask-adapted clip. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, oVSeg |
| [110] |
|
| [111] |
|
| [112] |
Minderer M, Gritsenko A, Stone A, Neumann M, Weissenborn D, Dosovitskiy A, Mahendran A, Arnab A, Dehghani M, Shen Z et al (2022) Simple open-vocabulary object detection. In: European conference on computer vision, Springer, pp 728–755 |
| [113] |
Liu S, Zeng Z, Ren T, Li F, Zhang H, Yang J, Li C, Yang J, Su H, Zhu J, Zhang L (2023) Grounding dino: marrying dino with grounded pre-training for open-set object detection. Preprint at https://arxiv.org/abs/2303.05499 |
| [114] |
Kirillov A, Mintun E, Ravi N, Mao H, Rolland C, Gustafson L, Xiao T, Whitehead S, Berg AC, Lo WY et al (2023) Segment anything. In: Proceedings of the IEEE/CVF international conference on computer vision, pp 4015–4026 |
| [115] |
Xu S et al (2025) Zero-shot pavement monitoring with large language models. Preprint at https://arxiv.org/abs/2504.06785 |
| [116] |
|
| [117] |
Gao P, Geng S, Jiang R, Yuan N, Hsieh TY, Qiao Y (2022) Tip-adapter: training-free adaption of clip for few-shot classification. In: Computer Vision – ECCV 2022 |
| [118] |
|
| [119] |
Radford A, Kim JW, Hallacy C, Ramesh A, Goh G, Agarwal S, Sastry G, Askell A, Mishkin P, Clark J, Krueger G, Sutskever I (2021) Learning transferable visual models from natural language supervision. Preprint at https://arxiv.org/abs/2103.00020 |
| [120] |
|
| [121] |
|
| [122] |
|
| [123] |
|
| [124] |
Chollet F (2017) Xception: deep learning with depthwise separable convolutions. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, IEEE, pp 1251–1258 |
| [125] |
Hu J, Shen L, Sun G (2018) Squeeze-and-excitation networks. In: CVPR, pp 7132–7141. https://doi.org/10.1109/CVPR.2018.00745 |
| [126] |
|
| [127] |
|
| [128] |
|
| [129] |
|
| [130] |
|
| [131] |
|
| [132] |
|
| [133] |
|
| [134] |
|
| [135] |
|
| [136] |
|
| [137] |
|
| [138] |
Howard AG, Zhu M, Chen B, Kalenichenko D, Wang W, Weyand T, Andreetto M, Adam H (2017) Mobilenets: efficient convolutional neural networks for mobile vision applications. Preprint at https://arxiv.org/abs/1704.04861 |
| [139] |
Hou X, Zhang Z, Li W et al (2021) MobileCrack: an adaptive lightweight CNN model for pavement crack image classification. J Transp Eng B Pavements |
| [140] |
Liu Z, Mao H, Wu CY, Feichtenhofer C, Darrell T, Xie S (2022) A convnet for the 2020s. CVPR pp 11976–11986. https://doi.org/10.1109/CVPR52688.2022.01167 |
| [141] |
Song C, Zhang W, Li H et al (2025) Automatic crack defect detection via multiscale feature aggregation and adaptive fusion with multiple-dimension attention. Autom Constr |
| [142] |
Vishwakarma R, Vennelakanti R (2021) CNN model & tuning for global road damage detection. In: IEEE Big Data Cup 2020. Preprint at https://arxiv.org/abs/2103.09512 |
| [143] |
Cai Z, Vasconcelos N (2018) Cascade R-CNN: delving into high quality object detection. In: Proc IEEE Conf Comput Vis Pattern Recognit, pp 6154–6162 |
| [144] |
|
| [145] |
Shen T, Nie M (2020) Pavement damage detection based on cascade R-CNN. In: Proceedings of the 4th International Conference on Computer Science and Application Engineering (CSAE), ACM, pp 1–5 |
| [146] |
Fu R, Cao M, Novak D, Qian X, Alkayem NF (2023) Extended efficient convolutional neural network for concrete crack detection with illustrated merits. Autom Constr 156:105098. Elsevier |
| [147] |
Marin B, Brown KE, Erden MS (2021) Automated masonry crack detection with Faster R-CNN. In: 2021 IEEE 17th International Conference on Automation Science and Engineering (CASE). IEEE. https://doi.org/10.1109/CASE49439.2021.9551683 |
| [148] |
|
| [149] |
|
| [150] |
|
| [151] |
|
| [152] |
Wu P, Liu A, Fu J, Ye X, Zhao Y (2022) Autonomous surface crack identification of concrete structures based on an improved one-stage object detection algorithm. Eng Struct 272:114962. Elsevier |
| [153] |
|
| [154] |
|
| [155] |
Chu Y, Xiang X, Wang Y, Huang B (2022) Pavement disease detection through improved yolov5s neural network. Comput Intell Neurosci 2022 |
| [156] |
Guo K, He C, Yang M, Wang S (2022) A pavement distresses identification method optimized for yolov5s. Sci Rep 12(1) |
| [157] |
Xiang W, Wang H, Xu Y, Zhao Y, Zhang L, Duan Y (2023) Road disease detection algorithm based on yolov5s-dsg. J Real-Time Image Process 20(3) |
| [158] |
|
| [159] |
Liu W (2016) Ssd: single shot multibox detector. In: Proc 14th Eur Conf Comput Vis, pp 21–37 |
| [160] |
|
| [161] |
|
| [162] |
|
| [163] |
|
| [164] |
|
| [165] |
Long J, Shelhamer E, Darrell T (2015) Fully convolutional networks for semantic segmentation. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, pp 3431–3440. https://doi.org/10.1109/CVPR.2015.7298965 |
| [166] |
|
| [167] |
|
| [168] |
|
| [169] |
|
| [170] |
|
| [171] |
Zhang et al (2023) Asymmetric dual-decoder-U-net for pavement crack semantic segmentation. Adv Eng Inform. https://doi.org/10.1016/j.sc.2023.003989 |
| [172] |
|
| [173] |
Zim AH, Iqbal A, Al-Huda Z, Malik A, Kuribayashi M (2025) Efficientcracknet: a lightweight model for crack segmentation. In: 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), IEEE, pp 6279–6289 |
| [174] |
|
| [175] |
|
| [176] |
Xu Q et al (2022) Pixel-level pavement crack detection using enhanced high-resolution semantic dataset. Int J Pavement Eng 23(14). https://doi.org/10.1080/10298436.2021.1985491 |
| [177] |
|
| [178] |
|
| [179] |
Badrinarayanan V, Kendall A, Cipolla R (2016) Segnet: a deep convolutional encoder-decoder architecture for image segmentation. Preprint at https://arxiv.org/abs/1511.00561 |
| [180] |
|
| [181] |
Chen X et al (2024) A two-stage framework for pixel-level pavement surface crack detection and segmentation. Eng Appl Artif Intell. https://doi.org/10.1016/j.engappai.2024.110867 |
| [182] |
|
| [183] |
|
| [184] |
|
| [185] |
|
| [186] |
|
| [187] |
|
| [188] |
|
| [189] |
Chen J, Lu Y, Yu Q, Luo X, Adeli E, Wang Y, Lu L, Yuille A, Zhou Y (2021) Transunet: Transformers make strong encoders for medical image segmentation. Preprint at https://arxiv.org/abs/2102.04306 |
| [190] |
|
| [191] |
|
| [192] |
|
| [193] |
|
| [194] |
|
| [195] |
|
| [196] |
|
| [197] |
|
| [198] |
|
| [199] |
|
| [200] |
|
| [201] |
Chen Z, Zou Y, González VA, Ingham J, Wotherspoon LM (2025) Bridge inspection using a multi-modal vision language model. In: Proceedings of the 6th International Conference on Civil and Building Engineering Informatics, vol 8, p 11 |
| [202] |
Hu EJ, Shen Y, Wallis P, Allen-Zhu Z, Li Y, Wang S, Wang L, Chen W (2021) Lora: low-rank adaptation of large language models. Preprint at https://arxiv.org/abs/2106.09685 |
| [203] |
|
| [204] |
Yu N, Meng X, Zhang D, Hu S, Lang C (2020) Unsupervised pixel-level road defect detection via adversarial image-to-frequency transform. In: Proc. IEEE Intelligent Vehicles Symposium (IV), pp 972–979. https://doi.org/10.1109/IV47402.2020.9304587 |
| [205] |
|
| [206] |
Duan Y, Gu Y, Li X (2023) Unsupervised crack image binarization via image-to-image translation with GANs. Sci Iran |
| [207] |
Yu D, Chen X, Liu H et al (2023) Multi-source domain adaptation and domain alignment with outlier relocation for road defect segmentation. In: Proc IEEE Int Conf Robot Autom (ICRA) Workshops |
| [208] |
|
| [209] |
Mubashshira S, Azam MM, Ahsan SMM (2020) An unsupervised approach for road surface crack detection. In: 2020 IEEE Region 10 Symposium (TENSYMP), pp 1596–1599. IEEE |
| [210] |
Shamsolmoali P, K MHR et al (2019) Adversarial spatial pyramid networks for road detection. IEEE Geosci Remote Sens Lett 16(11):1815–1819. https://doi.org/10.1109/LGRS.2019.2900539 |
| [211] |
|
| [212] |
Ouali Y, Hudelot C, Tami M (2020) Semi-supervised semantic segmentation with cross-consistency training. In: Proc IEEE/CVF Conf on Computer Vision and Pattern Recognition (CVPR), pp 12674–12684. https://doi.org/10.1109/CVPR42600.2020.01269 |
| [213] |
Liu X, Wu K, Cai X, Huang W (2024) Semi-supervised semantic segmentation using cross-consistency training for pavement crack detection. Road Mater Pavement Des 25(6):1368–1380. Taylor & Francis |
| [214] |
Zhou B, Khosla A, Lapedriza À, Oliva A, Torralba A (2016) Learning deep features for discriminative localization. In: Proc IEEE/CVF Conf on Computer Vision and Pattern Recognition (CVPR), pp 2921–2929. https://doi.org/10.1109/CVPR.2016.319 |
| [215] |
Krähenbühl P, Koltun V (2011) Efficient inference in fully connected CRFs with Gaussian edge potentials. In: Advances in Neural Information Processing Systems (NeurIPS), pp 109–117 |
| [216] |
|
| [217] |
|
| [218] |
|
| [219] |
|
| [220] |
Pizer SM, Amburn EP, Austin JD, Cromartie R, Geselowitz A, Greer T, ter Haar Romeny B, Zimmerman JB, Zuiderveld K (1990) Adaptive histogram equalization and its variations (CLAHE). In: Proc SPIE, Visualization in Biomedical Computing, pp 337–345. https://doi.org/10.1109/VBC.1990.109340 |
| [221] |
|
| [222] |
Zhang X, Rajan D, Story B (2019) Concrete crack detection using context-aware deep semantic segmentation network. Comput-Aided Civ Infrastruct Eng 34(11):951–971. https://doi.org/10.1111/mice.12477 |
| [223] |
|
| [224] |
|
| [225] |
|
| [226] |
Laurent J, Talbot M, Doucet M (2008) A road surface crack detection and measurement system based on 3D laser profiling. In: 7th symposium on pavement surface characteristics: SURF 2008, pp 1–11 |
| [227] |
|
| [228] |
Wang KC, Li JQ (2021) A review of 3D point cloud applied to pavement engineering. J Traffic Transp Eng 8(2):177–193 |
| [229] |
|
| [230] |
Graham B, Engelcke M, van der Maaten L (2018) 3D semantic segmentation with submanifold sparse convolutional networks. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp 4555–4564. https://doi.org/10.1109/CVPR.2018.00479 |
| [231] |
Choy C, Gwak J, Savarese S (2019) 4D spatio-temporal ConvNets: Minkowski convolutional neural networks. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp 3075–3084. https://doi.org/10.1109/CVPR.2019.00319 |
| [232] |
Zhu X, Zhou H, Wang T, Hong F, Ma Y, Li W, Li H, Lin D (2021) Cylindrical and asymmetrical 3D convolution networks for LiDAR segmentation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp 9939–9948. https://doi.org/10.1109/CVPR46437.2021.00981 |
| [233] |
Wu B, Wang P, Bänzigar MD, Shen XS, Keutzer K (2018) Squeezeseg: convolutional neural nets with recurrent CRF for real-time road-object segmentation from 3D LiDAR point cloud. In: 2018 IEEE International Conference on Robotics and Automation (ICRA), pp 1887–1894. https://doi.org/10.1109/ICRA.2018.8462926 |
| [234] |
Milioto A, Vizzo I, Behley J, Stachniss C (2019) Rangenet++: fast and accurate LiDAR semantic segmentation. In: 2019 IEEE/RSJ international conference on intelligent robots and systems (IROS), IEEE, pp 4213–4220 |
| [235] |
|
| [236] |
|
| [237] |
|
| [238] |
|
| [239] |
Qi CR, Su H, Mo K, Guibas LJ (2017) Pointnet: Deep learning on point sets for 3D classification and segmentation. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp 652–660 |
| [240] |
Qi CR, Yi L, Su H, Guibas LJ (2017) Pointnet++: deep hierarchical feature learning on point sets in a metric space. Adv Neural Inform Process Syst 30 |
| [241] |
|
| [242] |
|
| [243] |
Zhao H, Jiang L, Jia J, Torr PHS, Koltun V (2021) Point transformer. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp 16259–16268 |
| [244] |
|
| [245] |
|
| [246] |
|
| [247] |
Tseng TY, Lyu H, Li J, Berrio JS, Shan M, Worrall S (2025) M2s-road: multi-modal semantic segmentation for road damage using camera and LiDAR data. Preprint at https://arxiv.org/abs/2504.10123 |
| [248] |
|
| [249] |
|
| [250] |
Gupta S, Girshick R, Arbelaez P, Malik J (2014) Learning rich features from RGB-D images for object detection and semantic segmentation. In: European Conference on Computer Vision (ECCV), Springer, pp 345–360. https://doi.org/10.1007/978-3-319-10584-0_23 |
| [251] |
Vora S, Lang AH, Helou B, Beijbom O (2020) Pointpainting: sequential fusion for 3D object detection. In: IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp 4603–4611. https://doi.org/10.1109/CVPR42600.2020.00466 |
| [252] |
Wang Z, Li B, Fan T, He M, Wu W, Ouyang W, Feng X, Qiao Y (2021) Pointaugmenting: cross-modal augmentation for 3D object detection. In: IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp 11794–11803. https://doi.org/10.1109/CVPR46437.2021.01163 |
| [253] |
Xu D, Anguelov D, Jain A (2018) Pointfusion: deep sensor fusion for 3D bounding box estimation. In: IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp 244–253. https://doi.org/10.1109/CVPR.2018.00033 |
| [254] |
Wang C, Xu D, Zhu Y, Martín-Martín R, Lu C, Savarese S, Fei-Fei L (2019) Densefusion: 6D object pose estimation by iterative dense fusion. In: IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp 3343–3352. https://doi.org/10.1109/CVPR.2019.00346 |
| [255] |
|
| [256] |
|
| [257] |
|
| [258] |
|
| [259] |
Liu Y, Wang Z et al (2024) Road surface defect detection—from image-based to non-image-based: a survey. Preprint at https://arxiv.org/abs/2409.20118 |
| [260] |
|
| [261] |
|
| [262] |
Federal Highway Administration. Highway performance monitoring system (hpms) field manual, 2023, Federal Highway Administration, Technical report |
| [263] |
Sharma M, Al-Hammadi M, Mork H, Klein-Paste A (2025) Combined dataset for subjective panel rating, international roughness index and images for road damage detection of low volume road in norway. Dataset. https://doi.org/10.18710/EHMQU7 |
| [264] |
Federal Highway Administration (2018) LTPP InfoPave brochure. Technical Report No. FHWA-HRT-18-011, Federal Highway Administration, Washington, DC. https://www.fhwa.dot.gov/publications/research/infrastructure/pavements/ltpp/18011/18011.pdf |
| [265] |
Federal Highway Administration (2006) LTPP directive D-44: distress survey photographs. Technical Report, Federal Highway Administration, Washington, DC. https://infopave.fhwa.dot.gov/InfoPave_Repository/Reports/B10/B10_20/D-44.pdf |
| [266] |
Owor NJ, Du H, Daud A, Aboah A, Adu-Gyamfi Y (2023) Image2pci – a multitask learning framework for estimating pavement condition indices directly from images. Preprint at https://arxiv.org/abs/2310.08538 |
| [267] |
Government of Alberta (2015) International roughness index and rut data. Open Government dataset. https://open.canada.ca/data/en/dataset/5161dfca-bca9-4e11-bfe4-683ef3c8aad7. Accessed 28 Apr 2026 |
| [268] |
Ontario Ministry of Transportation (2023) Pavement condition for provincial highways. Ontario Data Catalogue dataset. https://data.ontario.ca/dataset/pavement-condition-for-provincial-highways. Accessed 28 Apr 2026 |
| [269] |
San Francisco Public Works (2026) Streets data–pavement condition index (PCI) scores. DataSF dataset. https://data.sfgov.org/City-Infrastructure/Streets-Data-Pavement-Condition-Index-PCI-Scores/5aye-4rtt. Accessed 28 Apr 2026 |
| [270] |
Data.Gov (2025) Pavement condition index. Data.gov dataset catalog entry. Montgomery County, Maryland. https://catalog.data.gov/dataset/pavement-condition-index-2019. Accessed 28 Apr 2026 |
| [271] |
Cheng B, Girshick R, Dollár P, Berg AC, Kirillov A (2021) Boundary IOU: Improving object-centric image segmentation evaluation. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp 15334–15342 |
| [272] |
Shit S, Paetzold JC, Ezhov I, Sekuboyina A, Unger A, Zhylka A, Pluim JPW, Bauer U, Menze BH (2021) CLDICE – a novel topology-preserving loss function for tubular structure segmentation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp 16555–16564.https://doi.org/10.1109/CVPR46437.2021.01629 |
| [273] |
|
| [274] |
Yuan Y, Chen X, Wang J (2020) Object-contextual representations for semantic segmentation. In: Proc ECCV, pp 173–190. https://doi.org/10.1007/978-3-030-58548-8_11 |
| [275] |
Oktay O, Schlemper J, Folgoc LL et al (2018) Attention U-Net: learning where to look for the pancreas. Preprint at https://arxiv.org/abs/1804.03999 |
| [276] |
Goo JR et al (2025) Hybrid-segmentor: high-resolution and efficient pavement crack segmentation. Automation in Construction (In press). TechRxiv. https://doi.org/10.36227/techrxiv.170775714.48894818. arXiv:2409.02866 |
| [277] |
Tao H, Liu B, Cui J, Zhang H (2023) A convolutional-transformer network for crack segmentation with boundary awareness. In: Proc IEEE ICIP, pp 86–90. https://doi.org/10.1109/ICIP49359.2023.10349671 |
| [278] |
|
| [279] |
|
| [280] |
Xie E, Wang W, Yu Z, Anandkumar A, Alvarez JM, Luo P (2021) SegFormer: simple and efficient design for semantic segmentation with transformers. Preprint at https://arxiv.org/abs/2105.15203 |
| [281] |
Yeom SU, Klitzing J (2025) U-MixFormer: UNet-like transformer with mix-attention for efficient semantic segmentation. In: Proc IEEE/CVF Winter Conf Appl Comput Vis (WACV), pp 7710–7719. https://doi.org/10.1109/WACV61041.2025.00750 |
| [282] |
Guo M, Lu C, Hou Q, Liu Z, Cheng M, Hu S (2022) SegNeXt: rethinking convolutional attention design for semantic segmentation. In: 36th Conference on Neural Information Processing Systems (NeurIPS 2022). https://proceedings.neurips.cc/paper_files/paper/2022/file/08050f40fff41616ccfc3080e60a301a-Paper-Conference.pdf |
| [283] |
Cheng B, Misra I, Girdhar R, Kirillov A (2022) Masked-attention mask transformer for universal image segmentation. In: Proc CVPR. https://doi.org/10.1109/CVPR52688.2022.01964 |
| [284] |
Shan J, Huang Y, Jiang W (2024) DCUFormer: enhancing pavement crack segmentation in complex scenarios by dual-branch connected U-Former. Expert Syst Appl 264:125891. https://doi.org/10.1016/j.eswa.2024.125891 |
| [285] |
Yan H, Wu M, Zhang C (2024) Multi-scale representations by varying window attention for semantic segmentation (vwformer). Preprint at https://arxiv.org/abs/2404.16573 |
| [286] |
Yu C, Wang J, Peng C, Gao C, Yu G, Sang N (2018) Bisenet: bilateral segmentation network for real-time semantic segmentation. In: ECCV. https://openaccess.thecvf.com/content_ECCV_2018/papers/Changqian_Yu_BiSeNet_Bilateral_Segmentation_ECCV_2018_paper.pdf |
| [287] |
Zhao H, Shi J, Qi X, Wang X, Jia J (2017) Pyramid scene parsing network. In: Proc CVPR, pp 2881–2890. https://doi.org/10.1109/CVPR.2017.660 |
| [288] |
Wang Y, Zhou Q, Liu J, Xiong J, Gao G, Wu X, Latecki LJ (2019) Lednet: a lightweight encoder-decoder network for real-time semantic segmentation. In: 2019 IEEE International Conference on Image Processing (ICIP), IEEE, pp 1860–1864. https://doi.org/10.1109/ICIP.2019.8803154. Accessed 20 Apr 2026 |
| [289] |
|
| [290] |
|
| [291] |
|
| [292] |
|
| [293] |
|
| [294] |
|
| [295] |
Liu Z, Lin Y, Cao Y, Hu H, Wei Y, Zhang Z, Lin S, Guo B (2021) Swin transformer: hierarchical vision transformer using shifted windows. In: Proc ICCV, pp 10012–10022. https://doi.org/10.1109/ICCV48922.2021.00986 |
| [296] |
|
| [297] |
|
| [298] |
Chen J et al (2022) Refined crack detection via LECSFormer for autonomous road inspection vehicles. IEEE Trans Intell Veh 8(3):2049–2061. https://doi.org/10.1109/TIV.2022.3204583 |
The Author(s)
/
| 〈 |
|
〉 |