Dual-Model Fusion for Ultra-Accurate Embedded Object Detection

Leendert Remmelzwaal

Smart Wearable Technology ›› 2025, Vol. 1 ›› Issue (1) : 52026032

PDF (1199KB)
Smart Wearable Technology ›› 2025, Vol. 1 ›› Issue (1) :52026032 DOI: 10.47852/bonviewSWT52026032
RESEARCH ARTICLE
research-article
Dual-Model Fusion for Ultra-Accurate Embedded Object Detection
Author information +
History +
PDF (1199KB)

Abstract

Numerous industries consider their needs ultra-high accuracy with regard to an artificial intelligence (AI)-detected object and pose a challenge with real-world variations. This study specifically focuses on industrial manufacturing lines, where quality control is critical. The methods we discuss here should be transferable to other industry domains with similar constraints, such as logistics or packaging. The primary objective is to achieve greater than 99.9% accuracy with object detection in real-time industrial environments, without significantly impacting the latency. For that, we worked on an SSD_MobileNet model that was refined to the utmost precision and implemented alongside a dual-model system that used a generalist surrogate trained on blurred synthetic images. To achieve blur efficacy, the second model had to be blur trained, blending contextual depth and resilience. Both models’ outputs are fused through a low-computational-cost, high-confidence detection using Intersection over Union metrics selection (>=0.8) to strike a balance between efficiency and detection reliability. Model fusion has better results compared to model stacking or score-based thresholds because it decides on the best detection by considering the spatial overlap of detections and the agreement of class IDs. On the Nvidia Jetson Orin NX platform, deploying this ensemble achieved 99.8% accuracy and further boosted the system to 99.97% without expanding inference passes. Smart dual-model implementation helps increase precision and fault-tolerant parameters while maintaining streamlined recalibrated embedded systems thresholds, proving non-breach. This work supports the shift toward AI-powered advanced industrial surveillance, and research focuses on multidisciplinary approaches toward precise, reliable object detection.

Keywords

artificial intelligence / deep learning / ensemble AI

Cite this article

Download citation ▾
Leendert Remmelzwaal. Dual-Model Fusion for Ultra-Accurate Embedded Object Detection. Smart Wearable Technology, 2025, 1 (1) : 52026032 DOI:10.47852/bonviewSWT52026032

登录浏览全文

4963

注册一个新账户 忘记密码

References

[1]

Sinha, D., & El—Sharkawy, M. (2019). Thin MobileNet: An enhanced MobileNet architecture. In 2019 IEEE 10th Annual Ubiquitous Computing, Electronics & Mobile Communication Conference, 0280-0285. https://doi.org/10.1109/UEMCON47517.2019.8993089

[2]

Choi, S.—Y., Choi, J.—H., & Lim, S.—H. (2023). SSD—Mobilenet—V2 model—eul sayonghan Edge Device eseoui gaegchegeomchul seongneung bigyo mich bunseog [Comparative analysis of object detection performance on edge devices using SSD—Mobilenet—V2 model]. In Annual Conference of KIPS, 79-80. https://doi.org/10.3745/PKIPS.Y2023M05A.79

[3]

Oleiwi, B. K., & Kadhim, M. R. (2022). Real time embedded system for object detection using deep learning. AIP Conference Proceedings, 2415(1), 070003. https://doi.org/10.1063/5.0093469

[4]

Nijkamp, N., Sallou, J., van der Heijden, N., & Cruz, L. (2024). Green AI in action: Strategic model selection for ensembles in production. In Proceedings of the 1st ACM International Conference on AI—Powered Software, 50-58. https://doi.org/10.1145/3664646.3664763

[5]

Chen, S., He, W., Ren, J., & Jiang, X. (2022). Attention—based dual—stream vision transformer for radar gait recognition. In 2022 IEEE International Conference on Acoustics, Speech and Signal Processing, 3668-3672. https://doi.org/10.1109/ICASSP43922.2022.9746565

[6]

Hussain, F., Hussain, R., Hassan, S. A., & Hossain, E. (2020). Machine learning in IoT security: Current solutions and future challenges. IEEE Communications Surveys & Tutorials, 22(3), 1686-1721. https://doi.org/10.1109/COMST.2020.2986444

[7]

Ansari, M. F., Dash, B., Sharma, P., & Yathiraju, N. (2022). The impact and limitations of artificial intelligence in cybersecurity: A literature review. International Journal of Advanced Research in Computer and Communication Engineering, 11(9), 81-90. https://doi.org/10.17148/IJARCCE.2022.11912

[8]

Djenouri, Y., Belhadi, A., Yazidi, A., Srivastava, G., & Lin, J. C.—W. (2024). Artificial intelligence of medical things for disease detection using ensemble deep learning and attention mechanism. Expert Systems, 41(6), e13093. https://doi.org/10.1111/exsy.13093

[9]

Kwasniewska, A., MacAllister, A., Nicolas, R., & Garza, J. (2023). Multi—sensor ensemble—guided attention network for aerial vehicle perception beyond visible spectrum. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 345-353. https://doi.org/10.1109/CVPRW59228.2023.00040

[10]

Wang, Y., & Chung, S. H. (2022). Artificial intelligence in safety—critical systems: A systematic review. Industrial Management & Data Systems, 122(2), 442-470. https://doi.org/10.1108/IMDS-07-2021-0419

[11]

Katkoria, D., Sreevalsan—Nair, J., Sati, M., & Karunakaran, S. (2024). WBF—ODAL: Weighted boxes fusion for 3D object detection from automotive LiDAR point clouds. In 2024 International Conference on Vehicular Technology and Transportation Systems, 1-6. https://doi.org/10.1109/ICVTTS62812.2024.10763933

[12]

Hong, J., He, X., Deng, Z., & Yang, C. (2024). IoU—aware feature fusion R—CNN for dense object detection. Machine Vision and Applications, 35(1), 3. https://doi.org/10.1007/s00138-023-01483-2

[13]

Duong, V. H., Nguyen, D. Q., van Luong, T., Vu, H., & Nguyen, T. C. (2024). Robust data augmentation and ensemble method for object detection in fisheye camera images. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 7017-7026. https://doi.org/10.1109/CVPRW63382.2024.00695

[14]

Alomar, K., Aysel, H. I., & Cai, X. (2023). Data augmentation in classification and segmentation: A survey and new strategies. Journal of Imaging, 9(2), 46. https://doi.org/10.3390/jimaging9020046

[15]

Vasiljevic, I., Chakrabarti, A., & Shakhnarovich, G. (2016). Examining the impact of blur on recognition by convolutional networks. arXiv. https://doi.org/10.48550/arXiv.1611.05760

[16]

Lébl, M., Šroubek, F., & Flusser, J. (2023). Impact of image blur on classification and augmentation of deep convolutional networks. In Image Analysis: 22nd Scandinavian Conference, 108-117. https://doi.org/10.1007/978-3-031-31438-4_8

[17]

Zhou, Y., Song, S., & Cheung, N.—M. (2017). On classification of distorted images with deep convolutional neural networks. In 2017 IEEE International Conference on Acoustics, Speech and Signal Processing, 1213-1217. https://doi.org/10.1109/ICASSP.2017.7952349

[18]

Jia, Z., Li, X., Ling, Z., Liu, S., Wu, Y., & Su, H. (2022). Improving policy optimization with generalist—specialist learning. In Proceedings of the 39th International Conference on Machine Learning, 162, 10104-10119.

[19]

Pasupuleti, S., Ramalakshmi, K., Gunasekaran, H., Arokiaraj, R. M., Debnath, S., & Jebaseeli, T. J. (2025). An enhancement of object detection using YOLO V8 and mobile net in challenging conditions. SN Computer Science, 6(4), 321. https://doi.org/10.1007/s42979-025-03856-y

[20]

Chompookham, T., & Surinta, O. (2021). Ensemble methods with deep convolutional neural networks for plant leaf recognition. ICIC Express Letters, 15(6), 553-565. https://doi.org/10.24507/icicel.15.06.553

[21]

Remmelzwaal, L. (2023). An AI—based early fire detection system utilizing HD cameras and real—time image analysis. Artificial Intelligence and Applications. Advance online publication. https://doi.org/10.47852/bonviewAIA3202975

[22]

Forecr. (2025). NVIDIA® Jetson Orin NX Industrial Fanless PC — DSBOX—ORNX . https://www.forecr.io/products/jetson-orin-nx-industrial-fanless-pc-dsbox-ornx

[23]

Shah, A., Ali, B., Habib, M., Frnda, J., Ullah, I., & Anwar, M. S. (2023). An ensemble face recognition mechanism based on three—way decisions. Journal of King Saud University — Computer and Information Sciences, 35(4), 196-208. https://doi.org/10.1016/j.jksuci.2023.03.016

[24]

Li, M., Wu, J., Wang, X., Chen, C., Qin, J., Xiao, X., :::, & Pan, X. (2023). AlignDet: Aligning pre—training and fine—tuning in object detection. In 2023 IEEE/CVF International Conference on Computer Vision, 6843-6853. https://doi.org/10.1109/ICCV51070.2023.00632

[25]

Ouyang, W., Wang, X., Zhang, C., & Yang, X. (2016). Factors in finetuning deep model for object detection with long—tail distribution. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, 864-873. https://doi.org/10.1109/CVPR.2016.100

[26]

Sun, L. (2024). Global vision object detection using an improved Gaussian mixture model based on contour. PeerJ Computer Science, 10, e1812. https://doi.org/10.7717/peerj-cs.1812

[27]

Yoshihara, S., Fukiage, T., & Nishida, S. (2023). Does training with blurred images bring convolutional neural networks closer to humans with respect to robust object recognition and internal representations? Frontiers in Psychology, 14, 1047694. https://doi.org/10.3389/fpsyg.2023.1047694

[28]

Kaur, R., & Singh, S. (2023). A comprehensive review of object detection with deep learning. Digital Signal Processing, 132, 103812. https://doi.org/10.1016/j.dsp.2022.103812

[29]

Kanadath, A., Jothi, J. A. A., & Urolagin, S. (2024). CViTS—Net: A CNN—ViT network with skip connections for histopathology image classification. IEEE Access, 12, 117627-117649. https://doi.org/10.1109/ACCESS.2024.3448302

[30]

Ounoughi, C., & Ben Yahia, S. (2023). Data fusion for ITS: A systematic literature review. Information Fusion, 89, 267-291. https://doi.org/10.1016/j.inffus.2022.08.016

[31]

Chen, Z., Liu, J., Shen, Y., Simsek, M., Kantarci, B., Mouftah, H. T., & Djukic, P. (2023). Machine learning—enabled IoT security: Open issues and challenges under advanced persistent threats. ACM Computing Surveys, 55(5), 105. https://doi.org/10.1145/3530812

[32]

Chen, Y., Li, Y., Kong, T., Qi, L., Chu, R., Li, L., & Jia, J. (2021). Scale—aware automatic augmentation for object detection. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 9558-9567. https://doi.org/10.1109/CVPR46437.2021.00944

PDF (1199KB)

0

Accesses

0

Citation

Detail

Sections
Recommended

/