Artificial intelligence for piping and instrumentation diagrams: a review and perspective
Jifei Ma , Wenjie Peng , Yaoyu Pan , Xiaoshuai Yuan , Jibin Zhou , Tao Zhang , Mao Ye , Zhongmin Liu
ENG. Chem. Eng. ›› 2026, Vol. 20 ›› Issue (9) : 70
Driven by the urgent demands for process efficiency, operational safety, and industrial intelligence, piping and instrumentation diagrams are evolving from static design documents into dynamic knowledge-intensive carriers. However, the intelligent digitization of piping and instrumentation diagrams faces systemic challenges due to their highly unstructured data format, dense symbolic information, and heterogeneous drafting styles, which hinder automated information extraction and result in a fragmented data ecosystem. Recent advancements in artificial intelligence, particularly in pattern recognition and nonlinear feature modeling, have enabled the automated extraction of core elements such as symbols, annotations, and connecting lines from piping and instrumentation diagrams, the reconstruction of process topologies, and the establishment of semantically enriched knowledge models. These developments provide a foundational framework for high-level applications including automated compliance checking, intelligent piping and instrumentation diagram generation, hazard and operability analysis, and digital twin development. This paper provides a systematic review of the state-of-the-art artificial intelligence-driven methodologies across the ‘perception-cognition-application’ pipeline, analyzes current technical bottlenecks, and outlines future directions.
artificial intelligence / piping and instrumentation diagrams / image processing / information extraction
| [1] |
Elyan E, Garcia C M, Jayne C. Symbols classification in engineering drawings. In: Proceedings of the 2018 International Joint Conference on Neural Networks (IJCNN). Rio de Janeiro: IEEE, 2018: 1–8 |
| [2] |
|
| [3] |
|
| [4] |
|
| [5] |
|
| [6] |
Stinner F, Wiecek M, Baranski M, Kümpel A, Müller D. Automatic digital twin generation of building energy systems using piping and instrumentation diagrams. In: Proceedings of the 34th International Conference on Efficiency, Cost, Optimization, Simulation, and Environmental Impact of Energy Systems (ECOS 2021). Tokyo: JapanECOS 2021 Program Organizers, 2022: 1854–1865 |
| [7] |
Alpúsig S, Pruna E, Escobar I, Guano A. Extended Reality. Cham: Springer, 2024: 101–113 |
| [8] |
|
| [9] |
|
| [10] |
|
| [11] |
|
| [12] |
|
| [13] |
|
| [14] |
|
| [15] |
Wolf T, Debut L, Sanh V, Chaumond J, Delangue C, Moi A, Cistac P, Rault T, Louf R, Funtowicz M, et al. Transformers: state-of-the-art natural language processing. In: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations. Stroudsburg: ACL, 2020: 38–45 |
| [16] |
|
| [17] |
Peng J, Bu X, Sun M, Zhang Z, Tan T, Yan J. Large-scale object detection in the wild from imbalanced multi-labels. In: Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Seattle: IEEE, 2020: 9706–9715 |
| [18] |
|
| [19] |
|
| [20] |
Tian Z, Huang W, He T, He P, Qiao Y. Computer Vision-ECCV 2016. Cham: Springer, 2016: 56–72 |
| [21] |
Zhou X, Yao C, Wen H, Wang Y, Zhou S, He W, Liang J. EAST: an efficient and accurate scene text detector. In: Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Honolulu: IEEE, 2017: 2642–2651 |
| [22] |
|
| [23] |
Veličković P, Cucurull G, Casanova A, Romero A, Liò P, Bengio Y. Graph attention networks. In: Proceedings of the International Conference on Learning Representations, Vancouve: ICLR, 2018 |
| [24] |
Kipf T N, Welling M. Semi-supervised classification with graph convolutional networks. In: Proceedings of the International Conference on Learning Representations, Toulon: ICLR, 2017 |
| [25] |
|
| [26] |
|
| [27] |
Purohit S, Van N, Chin G. Semantic property graph for scalable knowledge graph analytics. In: Proceedings of the 2021 IEEE International Conference on Big Data (Big Data). Orlando: IEEE, 2021: 2672–2677 |
| [28] |
|
| [29] |
|
| [30] |
Srinivas S S, Gupta S, Runkana V. AutoChemSchematic AI: agentic physics-aware automation for chemical manufacturing scale-up. Neural Information Processing Systems 2025 Workshop on AI for Science, 2025 |
| [31] |
Radford A, Kim J W, Hallacy C, Ramesh A, Goh G, Agarwal S, Sastry G, Askell A, Mishkin P, Clark J, et al. Learning transferable visual models from natural language supervision. International Conference on Machine Learning, 2021, 8748–8763 |
| [32] |
|
| [33] |
|
| [34] |
|
| [35] |
|
| [36] |
|
| [37] |
Tan W C, Chen I-M, Tan H K. Automated identification of components in raster piping and instrumentation diagram with minimal pre-processing. In: Proceedings of 2016 IEEE International Conference on Automation Science and Engineering (CASE). Fort Worth: IEEE, 2016: 1301–1306 |
| [38] |
|
| [39] |
|
| [40] |
|
| [41] |
|
| [42] |
Liu S. Design and implementation of pipeline reconstruction system for petrochemical P&ID drawings. Dissertation for the Doctoral Degree. Jinan: University of Jinan, 2024 |
| [43] |
|
| [44] |
Toghraei M. Piping and Instrumentation Diagram Development. New York: John Wiley & Sons, 2019 |
| [45] |
|
| [46] |
Mani S, Haddad M A, Constantini D, Douhard W, Li Q W, Poirier L. Automatic digitization of engineering diagrams using deep learning and graph search. In: Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops. Seattle: IEEE, 2020: 673–679 |
| [47] |
Mesli-Kesraoui S, Kesraoui D, Oquendo F, Bignon A, Toguyeni A, Berruet P. Software Architecture. Cham: Springer, 2016: 210–226 |
| [48] |
|
| [49] |
|
| [50] |
|
| [51] |
|
| [52] |
Datta R, De Sekhar Mandal P, Chanda B. Detection and identification of logic gates from document images using mathematical morphology. In: Proceedings of the 2015 Fifth National Conference on Computer Vision, Pattern Recognition, and Image Processing and Graphics (NCVPRIPG). Patna: IEEE, 2015: 1–4 |
| [53] |
|
| [54] |
|
| [55] |
|
| [56] |
|
| [57] |
Hain A, Gölzhäuser S, Réhault N, Brox T, Demant M. Machine Learning and Knowledge Discovery in Databases: Research Track and Applied Data Science Track. Berlin, Heidelberg: Springer, 2026: 403–421 |
| [58] |
Bailey D, Norman A, Moretti G, North P. Electronic Schematic Recognition. Palmerston North: Massey University, 1995 |
| [59] |
Singh S, Singh R. Comparison of various edge detection techniques. In: Proceedings of the 2015 2nd International Conference on Computing for Sustainable Global Development (INDIACom). New Delhi: IEEE, 2015: 393–396 |
| [60] |
|
| [61] |
|
| [62] |
|
| [63] |
Viola P, Jones M. Rapid object detection using a boosted cascade of simple features. In: Proceedings of the 2001 IEEE Computer Society Conference on Computer Vision and Pattern Recognition. Kauai: IEEE, 2001: I |
| [64] |
Suhailam P, Yerolla R, Besta C S. Enhancing efficiency in piping & instrumentation diagrams (P&IDs) through AI-driven digitization. In: Proceedings of the 2025 6th International Conference on Control, Communication, and Computing (ICCC). Thiruvanathapuram: IEEE, 2025: 1–6 |
| [65] |
|
| [66] |
|
| [67] |
|
| [68] |
Girshick R, Donahue J, Darrell T, Malik J. Rich feature hierarchies for accurate object detection and semantic segmentation. In: Proceedings of the 2014 IEEE Conference on Computer Vision and Pattern Recognition. Columbus: IEEE, 2014: 580–587 |
| [69] |
Girshick R. Fast R-CNN. 2015 IEEE International Conference on Computer Vision (ICCV). Santiago: IEEE, 2015: 1440–1448 |
| [70] |
Ren S, He K, Girshick R, Sun J. Faster R-CNN: towards real-time object detection with region proposal networks. In: Advances in Neural Information Processing Systems. Montreal: NeurIPS, 2015, 91–99 |
| [71] |
He K, Zhang X, Ren S, Sun J. Deep residual learning for image recognition. In: Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Las Vegas: IEEE, 2016: 770–778 |
| [72] |
|
| [73] |
|
| [74] |
|
| [75] |
Szegedy C, Liu W, Jia Y, Sermanet P, Reed S, Anguelov D, Erhan D, Vanhoucke V, Rabinovich A. Going deeper with convolutions. In: Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Boston: IEEE, 2015: 1–9 |
| [76] |
Gada M. Object detection for P&ID images using various deep learning techniques. In: Proceedings of the 2021 International Conference on Computer Communication and Informatics (ICCCI). Coimbatore: IEEE, 2021: 1–5 |
| [77] |
Howard A, Sandler M, Chen B, Wang W, Chen L-C, Tan M, Chu G, Vasudevan V, Zhu Y, Pang R, et al. Searching for MobileNetV3. In: Proceedings of the 2019 IEEE/CVF International Conference on Computer Vision (ICCV). Seoul: IEEE, 2019: 1314–1324 |
| [78] |
Dai J, Li Y, He K, Sun J. R-FCN: object detection via region-based fully convolutional networks. In: Advances in Neural Information Processing Systems. Barcelona : NeurIPS, 2016, 379–387 |
| [79] |
Redmon J, Divvala S, Girshick R, Farhadi A. You only look once: unified, real-time object detection. In: Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Las Vegas: IEEE, 2016: 779–788 |
| [80] |
Redmon J, Farhadi A. YOLO9000: better, faster, stronger. In: Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Honolulu: IEEE, 2017: 6517–6525 |
| [81] |
Redmon J, Farhadi A. YOLOv3: an incremental improvement. 2018, arXiv: 1804.02767 |
| [82] |
Bochkovskiy A, Wang C-Y, Liao H-Y M. YOLOv4: optimal speed and accuracy of object detection. 2020, arXiv: 2004.10934 |
| [83] |
Li C, Li L, Jiang H, Weng K, Geng Y, Li L, Ke Z, Li Q, Cheng M, Nie W, et al. YOLOv6: a single-stage object detection framework for industrial applications. 2022, arXiv: 2209.02976 |
| [84] |
Wang C-Y, Bochkovskiy A, Liao H-Y M. YOLOv7: trainable bag-of-freebies sets new state-of-the-art for real-time object detectors. In: Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Vancouver: IEEE, 2023: 7464–7475 |
| [85] |
|
| [86] |
Bertinetto L, Valmadre J, Henriques J F, Vedaldi A, Torr P H S. Computer Vision-ECCV 2016 Workshops. Cham: Springer, 2016: 850–865 |
| [87] |
Hantach R, Lechuga G, Calvez P. Document Analysis and Recognition-ICDAR 2021 Workshops. Cham: Springer, 2021: 504–508 |
| [88] |
Bhanbhro H, Hooi Y K, Hassan Z, Sohu N. Modern deep learning approaches for symbol detection in complex engineering drawings. In: Proceedings of the 2022 International Conference on Digital Transformation and Intelligence (ICDI). Kuching: IEEE, 2022: 121–126 |
| [89] |
Haar C, Kim H, Koberg L. Flexible Automation and Intelligent Manufacturing: the Human-Data-Technology Nexus. Cham: Springer, 2023: 374–382 |
| [90] |
Lin T-Y, Goyal P, Girshick R, He K, Dollár P. Focal loss for dense object detection. In: Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV). Venice: IEEE, 2017: 2999–3007 |
| [91] |
Liu W, Anguelov D, Erhan D, Szegedy C, Reed S, Fu C-Y, Berg A C. Computer Vision-ECCV 2016. Cham: Springer, 2016: 21–37 |
| [92] |
Howard A G, Zhu M, Chen B, Kalenichenko D, Wang W, Weyand T, Andreetto M, Adam H. MobileNets: efficient convolutional neural networks for mobile vision applications. 2017, arXiv: 1704.04861 |
| [93] |
Long J, Shelhamer E, Darrell T. Fully convolutional networks for semantic segmentation. In: Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Boston: IEEE, 2015: 3431–3440 |
| [94] |
Paliwal S, Jain A, Sharma M, Vig L. Trends and Applications in Knowledge Discovery and Data Mining. Cham: Springer, 2021: 168–180 |
| [95] |
Zhang F, Li M, Zhai G, Liu Y. MultiMedia Modeling. Cham: Springer, 2021: 136–147 |
| [96] |
Dosovitskiy A, Beyer L, Kolesnikov A, Weissenborn D, Zhai X, Unterthiner T, Dehghani M, Minderer M, Heigold G, Gelly S, et al. An image is worth 16 × 16 words: transformers for image recognition at scale. In: International Conference on Learning Representations.Vienna: ICLR, 2021 |
| [97] |
Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez A N, Kaiser L, Polosukhin I. Attention is all you need. In: Advances in Neural Information Processing Systems. Long Beach: NeurIPS, 2017, 5998–6008 |
| [98] |
|
| [99] |
|
| [100] |
|
| [101] |
|
| [102] |
|
| [103] |
Jamieson L, Moreno-Garcia C F, Elyan E. Deep learning for text detection and recognition in complex engineering diagrams. In: Proceedings of the 2020 International Joint Conference on Neural Networks (IJCNN). Glasgow: IEEE, 2020: 1–7 |
| [104] |
Shi B, Bai X, Belongie S. Detecting oriented text in natural images by linking segments. In: Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Honolulu: IEEE, 2017: 3482–3490 |
| [105] |
|
| [106] |
Sun P, Zhang R, Jiang Y, Kong T, Xu C, Zhan W, Tomizuka M, Li L, Yuan Z, Wang C, et al. Sparse R-CNN: end-to-end object detection with learnable proposals. In: Proceedings of the 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Nashville: IEEE, 2021: 14449–14458 |
| [107] |
Smith R. An overview of the tesseract OCR engine. In: Ninth International Conference on Document Analysis and Recognition. Kyoto: ICDAR, 2007: 629–633 |
| [108] |
Francois M, Eglin V, Biou M. Document Analysis Systems. Cham: Springer, 2022: 726–740 |
| [109] |
Du Y, Li C, Guo R, Yin X, Liu W, Zhou J, Bai Y, Yu Z, Yang Y, Dang Q, et al. PP-OCR: a practical ultra lightweight OCR system. 2020, arXiv: 2009.09941 |
| [110] |
|
| [111] |
|
| [112] |
|
| [113] |
|
| [114] |
Bresson X, Laurent T. Residual gated graph ConvNets. 2017, arXiv: 1711.07553 |
| [115] |
Sinha A, Bayer J, Bukhari S S. Table localization and field value extraction in piping and instrumentation diagram images. In: Proceedings of the 2019 International Conference on Document Analysis and Recognition Workshops (ICDARW). Sydney: IEEE, 2019: 26–31 |
| [116] |
|
| [117] |
|
| [118] |
Shit S, Koner R, Wittmann B, Paetzold J, Ezhov I, Li H W, Pan J Z, Sharifzadeh S, Kaissis G, Tresp V, et al. Relationformer: a unified framework for image-to-graph generation. 2022, arXiv: 2203.10202 |
| [119] |
Stürmer J M, Graumann M, Koch T. From engineering diagrams to graphs: digitizing P&IDs with transformers. In: Proceedings of the 2025 IEEE 12th International Conference on Data Science and Advanced Analytics (DSAA). Birmingham: IEEE, 2025: 1–11 |
| [120] |
Carion N, Massa F, Synnaeve G, Usunier N, Kirillov A, Zagoruyko S. Computer Vision-ECCV 2020. Cham: Springer, 2020: 213–229 |
| [121] |
Brown T B, Mann B, Ryder N, Subbiah M, Kaplan J, Dhariwal P, Neelakantan A, Shyam P, Sastry G, Askell A, et al. Language models are few-shot learners. In: Advances in Neural Information Processing Systems. Curran Associates, 2020, 1877–1901. |
| [122] |
|
| [123] |
|
| [124] |
Gowaikar S, Iyengar S, Segal S, Kalyanaraman S. An agentic approach to automatic creation of P&ID diagrams from natural language descriptions. 2024, arXiv: 2412.12898 |
| [125] |
|
| [126] |
|
| [127] |
Procko T T, Ochoa O. Graph retrieval-augmented generation for large language models: a survey. 2024 Conference on AI, Science, Engineering, and Technology (AIxSET). Laguna Hills: IEEE, 2024: 166–169 |
| [128] |
|
| [129] |
|
| [130] |
Sanh V, Debut L, Chaumond J, Wolf T. DistilBERT, a distilled version of BERT: smaller, faster, cheaper, and lighter. 2019, arXiv: 1910.01108 |
| [131] |
|
| [132] |
|
| [133] |
Elhosary E, Moselhi O. Utilization of artificial intelligence in HAZOP studies and reports. Proceedings of the International Symposium on Automation and Robotics in Construction (IAARC), 2025: 617–624 |
| [134] |
Grootendorst M. BERTopic: neural topic modeling with a class-based TF-IDF procedure. 2022, arXiv: 2203.05794 |
| [135] |
|
| [136] |
|
| [137] |
Koltun G, Kolter M, Vogel-Heuser B. Automated generation of modular PLC control software from P&ID diagrams in process industry. In: Proceedings of the 2018 IEEE International Systems Engineering Symposium (ISSE). Rome: IEEE, 2018: 1–8 |
| [138] |
|
| [139] |
d’Anterroches L. Process flow sheet generation and design through a group contribution approach. Dissertation for the Doctoral Degree. Kgs Lyngby: Technical University of Denmark, 2006 |
| [140] |
|
| [141] |
|
| [142] |
|
| [143] |
Gill M S, Vyas J, Markaj A, Gehlhoff F, Mercangöz M. Leveraging LLM agents and digital twins for fault handling in process plants. In: Proceedings of the 2025 IEEE 30th International Conference on Emerging Technologies and Factory Automation (ETFA). Porto: IEEE, 2025: 1–8 |
| [144] |
|
| [145] |
|
| [146] |
|
| [147] |
|
| [148] |
|
| [149] |
|
| [150] |
|
| [151] |
Moreno-Garcia C F, Elyan E. Digitisation of assets from the oil & gas industry: challenges and opportunities. 2019 International Conference on Document Analysis and Recognition Workshops (ICDARW). Sydney: IEEE, 2019: 2–5 |
| [152] |
|
| [153] |
|
| [154] |
Dzhusupova R, Bosch J, Olsson H H. Challenges in developing and deploying AI in the engineering, procurement, and construction industry. In: Proceedings of the 2022 IEEE 46th Annual Computers, Software, and Applications Conference (COMPSAC). Los Alamitos: IEEE, 2022: 1070–1075 |
| [155] |
Vogt L, Urbas L. Automation of automation: mapping LLM capabilities to the modular plant engineering workflow. In: Proceedings of the 2025 IEEE 30th International Conference on Emerging Technologies and Factory Automation (ETFA). Porto: IEEE, 2025: 1–8 |
| [156] |
|
Higher Education Press
/
| 〈 |
|
〉 |