Review of Research on Black-Box Model Inversion Attack

Shifei HE , Chen WANG , Xin LUO , Guangwei XU

Journal of Donghua University(English Edition) ›› 2026, Vol. 43 ›› Issue (4) : 49 -58.

PDF (2512KB)
Journal of Donghua University(English Edition) ›› 2026, Vol. 43 ›› Issue (4) :49 -58. DOI: 10.19884/j.1672-5220.202505004
Information Technology and Artificial Intelligence
research-article
Review of Research on Black-Box Model Inversion Attack
Author information +
History +
PDF (2512KB)

Abstract

The development of artificial intelligence has given rise to the paradigm of machine learning as a service, enabling users to either train the requisite models or employ pre-existing models to make predictions based on the training data that users provide. Despite various measures taken to prevent privacy leakage in machine learning, such as data deletion and anonymization, model inversion attack (MIA) remains capable of inferring sensitive information from user data to a certain extent. This paper reviews the current state of research on black-box MIA, with a particular focus on confidence-based MIA and label-based MIA. It further analyzes the MIA methodologies applied to the emerging modes such as text and audio in recent years, filling the gap in the current review of such modes. Finally, the paper discusses the future challenges and research directions for black-box MIA.

Keywords

machine learning / model inversion attack / black-box model / privacy protection

Cite this article

Download citation ▾
Shifei HE, Chen WANG, Xin LUO, Guangwei XU. Review of Research on Black-Box Model Inversion Attack. Journal of Donghua University(English Edition), 2026, 43 (4) : 49-58 DOI:10.19884/j.1672-5220.202505004

登录浏览全文

4963

注册一个新账户 忘记密码

References

[1]

Raju K, Rao B C, Saikumar K, et al. An optimal hybrid solution to local and global facial recognition through machine learning[M]// A Fusion of Artificial Intelligence and Internet of Things for Emerging Cyber Systems. Cham: Springer International Publishing, 2021: 203-226.

[2]

Wood A, Najarian K, Kahrobaei D. Homomorphic encryption for machine learning in medicine and bioinformatics[J]. ACM Computing Surveys, 2020, 53(4): 1-35.

[3]

Bello O A, Folorunso A, Ejiofor O E, et al. Machine learning approaches for enhancing fraud prevention in financial transactions[J]. International Journal of Management Technology, 2023, 10(1): 85-108.

[4]

Weiss J C, Natarajan S, Peissig P L, et al. Machine learning for personalized medicine: predicting primary myocardial infarction from electronic health records[J]. AI Magazine, 2012, 33(4): 33-45.

[5]

Schroff F, Kalenichenko D, Philbin J. FaceNet: a unified embedding for face recognition and clustering[C]// 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway, NJ: IEEE, 2015: 815-823.

[6]

Fredrikson M, Lantz E, Jha S, et al. Privacy in pharmacogenetics: an end—to—end case study of personalized warfarin dosing[J]. Proceedings of the USENIX Security Symposium, 2014, 2014: 17-32.

[7]

Shokri R, Stronati M, Song C Z, et al. Membership inference attacks against machine learning models[C]// 2017 IEEE Symposium on Security and Privacy (SP). Piscataway, NJ: IEEE, 2017: 3-18.

[8]

Melis L, Song C Z, De Cristofaro E, et al. Exploiting unintended feature leakage in collaborative learning[C]// 2019 IEEE Symposium on Security and Privacy (SP). Piscataway, NJ: IEEE, 2019: 691-706.

[9]

Zhu L G, Liu Z J, Han S. Deep Leakage from Gradients[C]// Proceedings of the 32nd International Conference on Advances in Neural Information Processing Systems. Vancouver: NeurIPS, 2019: 14747-14756.

[10]

Song C Z, Ristenpart T, Shmatikov V. Machine learning models that remember too much[C]// Proceedings of the 2017 ACM SIGSAC Conference on Computer and Communications Security. New York: ACM, 2017: 587-601.

[11]

Dwork C. Differential privacy[C]// Automata, Languages and Programming. Berlin, Heidelberg: Springer, 2006: 1-12.

[12]

Abadi M, Chu A, Goodfellow I, et al. Deep learning with differential privacy[C]// Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security. New York: ACM, 2016: 308-318.

[13]

Murphy K, Schölkopf B, Srivastava N, et al. Dropout: a simple way to prevent neural networks from overfitting[J]. Journal of Machine Learning Research, 2014, 15(1): 1929-1958.

[14]

Wang T H, Zhang Y H, Jia R X. Improving robustness to model inversion attacks via mutual information regularization[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2021, 35(13): 11666-11673.

[15]

Wang X B, Hou R, Zhu Y F, et al. NPUFort: a secure architecture of DNN accelerator against model inversion attack[C]// Proceedings of the 16th ACM International Conference on Computing Frontiers. New York: ACM, 2019: 190-196.

[16]

Mehnaz S, Dibbo S V, De Viti R, et al. Are your sensitive attributes private? Novel model inversion attribute inference attacks on classification models[C]// 31st USENIX Security Symposium. Berkeley: USENIX Association, 2022: 4579-4596.

[17]

Liu R X, Chen H, Guo R Y, et al. Survey on privacy attacks and defenses in machine learning[J]. Journal of Software, 2020, 31(3): 866-892. (in Chinese)

[18]

Fredrikson M, Jha S, Ristenpart T. Model inversion attacks that exploit confidence information and basic countermeasures[C]// Proceedings of the 22nd ACM SIGSAC Conference on Computer and Communications Security. New York: ACM, 2015: 1322-1333.

[19]

Kahla M, Chen S, Just H A, et al. Label—only model inversion attacks via boundary repulsion[C]// 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway, NJ: IEEE, 2022: 15025-15033.

[20]

Ma W L, Wang D R, Song Y L, et al. TrapNet: model inversion defense via trapdoor[J]. IEEE Transactions on Information Forensics and Security, 2025, 20: 4469-4483.

[21]

Yang W C, Wang S, Wu D, et al. Deep learning model inversion attacks and defenses: a comprehensive survey[J]. Artificial Intelligence Review, 2025, 58(8): 242.

[22]

Bao H, Wei K M, Wu Y D, et al. Distributional black—box model inversion attack with multi—agent reinforcement learning[J]. IEEE Transactions on Information Forensics and Security, 2025, 20: 5425-5437.

[23]

Yang Z Q, Zhang J Y, Chang E C, et al. Neural network inversion in adversarial setting via background knowledge alignment[C]// Proceedings of the 2019 ACM SIGSAC Conference on Computer and Communications Security. New York: ACM, 2019: 225-240.

[24]

Yin H X, Molchanov P, Alvarez J M, et al. Dreaming to distill: data—free knowledge transfer via DeepInversion[C]// 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway, NJ: IEEE, 2020: 8712-8721.

[25]

Zhang Y H, Jia R X, Pei H Z, et al. The secret revealer: generative model—inversion attacks against deep neural networks[C]// 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway, NJ: IEEE, 2020: 250-258.

[26]

Liu R K, Wang D, Ren Y Z, et al. Unstoppable attack: label—only model inversion via conditional diffusion model[J]. IEEE Transactions on Information Forensics and Security, 2024, 19: 3958-3973.

[27]

Zhang R S, Hidano S, Koushanfar F. Text revealer: private text reconstruction via model inversion attacks against transformers[PP/OL]. arXiv (2022—09—21)[2025—04—21]. https://arxiv.org/abs/2209.10505.

[28]

Yin H X, Mallya A, Vahdat A, et al. See through gradients: image batch recovery via GradInversion[C]// 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway, NJ: IEEE, 2021: 16332-16341.

[29]

Liu Y F, Zhang W Q, Wu D Y, et al. Prediction exposes your face: black—box model inversion via prediction alignment[C]// Computer Vision — ECCV 2024. Cham: Springer, 2025: 288-306.

[30]

Pizzi K, Boenisch F, Sahin U, et al. Introducing model inversion attacks on automatic speaker recognition[PP/OL]. arXiv (2023—01—09)[2025—04—21]. https://arxiv.org/abs/2301.03206.

[31]

Chen G K, Zhao Z, Song F, et al. Towards understanding and mitigating audio adversarial examples for speaker recognition[J]. IEEE Transactions on Dependable and Secure Computing, 2023, 20(5): 3970-3987.

[32]

Li O X, Hao Y B, Wang Z C, et al. Model inversion attacks through target—specific conditional diffusion models[PP/OL]. arXiv (2024—07—16)[2025—04—21]. https://arxiv.org/abs/2407.11424.

[33]

Li Z A, Zhang H G, Wang J, et al. From head to tail: efficient black—box model inversion attack via long—tailed learning[C]// 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway, NJ: IEEE, 2025: 29288-29298.

[34]

Xu Q, Arafin M T, Qu G. An approximate memory based defense against model inversion attacks to neural networks[J]. IEEE Transactions on Emerging Topics in Computing, 2022, 10(4): 1733-1745.

[35]

Zhu T Q, Ye D Y, Zhou S, et al. Label—only model inversion attacks: attack with the least information[J]. IEEE Transactions on Information Forensics and Security, 2023, 18: 991-1005.

[36]

Liu Z T, Chen S T. Trap—mid: trapdoor—based defense against model inversion attacks[J]. Advances in Neural Information Processing Systems, 2024, 37: 88486-88526.

[37]

Ye Z P, Luo W J, Naseem M L, et al. C2FMI: corse—to—fine black—box model inversion attack[J]. IEEE Transactions on Dependable and Secure Computing, 2024, 21(3): 1437-1450.

[38]

Nguyen N B, Chandrasegaran K, Abdollahzadeh M, et al. Label—only model inversion attacks via knowledge transfer[J]. Advances in Neural Information Processing Systems, 2023, 36: 68895-68907.

[39]

Zhou S, Zhu T Q, Ye D Y, et al. Boosting model inversion attacks with adversarial examples[J]. IEEE Transactions on Dependable and Secure Computing, 2024, 21(3): 1451-1468.

[40]

Han G, Choi J, Lee H, et al. Reinforcement learning—based black—box model inversion attacks[C]// 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway, NJ: IEEE, 2023: 20504-20513.

[41]

Guo P X, Zeng S, Chen W H, et al. A new federated learning framework against gradient inversion attacks[J]. Proceedings of the AAAI Conference on Artificial Intelligence, 2025, 39(16): 16969-16977.

[42]

Zhang Q C, Ma J, Xiao Y H, et al. Broadening differential privacy for deep learning against model inversion attacks[C]// 2020 IEEE International Conference on Big Data. Piscataway, NJ: IEEE, 2021: 1061-1070.

[43]

Ho S T, Hao K J, Chandrasegaran K, et al. Model inversion robustness: can transfer learning help[C]// 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Piscataway, NJ: IEEE, 2024: 12183-12193.

[44]

Zhuang T Q, Yu H Y, Qiu Y X, et al. Stealthy shield defense: a conditional mutual information—based approach against black—box model inversion attacks[C]// The Thirteenth International Conference on Learning Representations (ICLR). Cham: Springer Nature Switzerland, 2025: p0DjhjPXl3.

[45]

Petrov I, Dimitrov D I, Baader M, et al. Dager: exact gradient inversion for large language models[J]. Advances in Neural Information Processing Systems, 2024, 37: 87801-87830.

[46]

Chen Y Y, Lent H, Bjerva J. Text embedding inversion security for multilingual language models[PP/OL]. arXiv (2024—01—22)[2025—04—21]. https://arxiv.org/abs/2401.12192.

[47]

Shi W Y, Shea R, Chen S, et al. Just fine—tune twice: selective differential privacy for large language models[PP/OL]. arXiv (2022—04—15)[2025—04—21]. https://arxiv.org/abs/2204.07667.

[48]

Gharib S, Tran M, Luong D, et al. Adversarial representation learning for robust privacy preservation in audio[J]. IEEE Open Journal of Signal Processing, 2024, 5: 294-302.

[49]

Yang C H, Chen I F, Stolcke A, et al. An experimental study on private aggregation of teacher ensemble learning for end—to—end speech recognition[C]// 2022 IEEE Spoken Language Technology Workshop (SLT). Piscataway, NJ: IEEE, 2023: 1074-1080.

[50]

Lecun Y, Bottou L, Bengio Y, et al. Gradient—based learning applied to document recognition[J]. Proceedings of the IEEE, 1998, 86(11): 2278-2324.

[51]

Norc at the University of Chicago. The general social survey[EB/OL]. (2025—05—22)[2026—06—25]. https://gss.norc.org/.

[52]

Walt H. Fivethirtyeight.com datalab: how Americans like their steak. https://fivethirtyeight.com/features/how—americans—like—their—steak/.

[53]

Liu Z W, Luo P, Wang X G, et al. Deep learning face attributes in the wild[C]// Proceedings of IEEE International Conference on Computer Vision (ICCV). Piscataway, NJ: IEEE, 2016: 3730-3738.

[54]

Deng J, Dong W, Socher R, et al. ImageNet: a large—scale hierarchical image database[C]// 2009 IEEE Conference on Computer Vision and Pattern Recognition. Piscataway, NJ: IEEE, 2009: 248-255.

[55]

Socher R, Perelygin A, Wu J, et al. Recursive deep models for semantic compositionality over a sentiment treebank[C]// Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing. Seattle: Association for Computational Linguistics, 2013: 1631-1642.

[56]

Carlini N, Tramer F, Wallace E, et al. Extracting training data from large language models[C]// 30th USENIX Security Symposium (USENIX Security 21). Berkeley, CA: USENIX Association, 2021: 2633-2650.

[57]

Panayotov V, Chen G G, Povey D, et al. Librispeech: an ASR corpus based on public domain audio books[C]// 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Piscataway, NJ: IEEE, 2015: 5206-5210.

[58]

Ardila R, Branson M, Davis K, et al. Common voice: a massively—multilingual speech corpus[PP/OL]. arXiv (2019—12—13)[2025—04—21]. https://arxiv.org/abs/1912.06670.

[59]

Jaber A, Fritsch L. Towards AI—powered cybersecurity attack modeling with simulation tools: review of attack simulators[C]// Advances on P2P, Parallel, Grid, Cloud and Internet Computing. Cham: Springer, 2023: 249-257.

[60]

He Z C, Zhang T W, Lee R B. Model inversion attacks against collaborative inference[C]// Proceedings of the 35th Annual Computer Security Applications Conference. New York: ACM, 2019: 148-162.

[61]

Zhang S J, Yuan W, Yin H Z. Comprehensive privacy analysis on federated recommender system against attribute inference attacks[J]. IEEE Transactions on Knowledge and Data Engineering, 2024, 36(3): 987-999.

PDF (2512KB)

0

Accesses

0

Citation

Detail

Sections
Recommended

/