Data Augmentation for Fake Review Detection Based on Large Language Model with Multi-stage Prompting

Qingxu Li , Jindong Chen , Wen Zhang

Journal of Systems Science and Systems Engineering ›› : 1 -22.

PDF
Journal of Systems Science and Systems Engineering ›› :1 -22. DOI: 10.1007/s11518-026-5731-y
Article
review-article
Data Augmentation for Fake Review Detection Based on Large Language Model with Multi-stage Prompting
Author information +
History +
PDF

Abstract

Fake reviews on e-commerce platforms mislead consumer decisions, damage merchant reputations, and undermine market order. Detecting fake review has become vital for protecting customer rights and maintaining platform fairness. Existing fake review datasets often suffer from severe class imbalance, which degrades the performance of detection model. To address this issue, we propose a data augmentation method based on Large Language Models with Multi-Stage Prompting (LLM-MSP) for fake review detection. First, the original reviews are parsed by LLMs to extract key elements such as review targets, detailed descriptions, and sentiment tendencies, and converted into structured data. Second, prompts are then constructed based on the extracted structures to guide LLMs to generate natural and semantically coherent reviews, and the high-quality generated reviews are selected to augment the original dataset. Finally, different classification models based on BERT are adopted to test the effectiveness of LLM-MSP for fake review detection. Experimental results show that, compared with the single-stage prompting, LLM-MSP improves the novelty and diversity of the generated fake reviews. Furthermore, among various data augmentation methods, LLM-MSP achieved the best classification performance. Specifically, compared with the original imbalanced dataset, the classification accuracy increased about 10% by incorporating the data generated by LLM-MSP. These findings validate the effectiveness and practical value of LLM-MSP for data augmentation task of fake review detection.

Keywords

Fake review detection / data augmentation / LLMs / multi-stage prompting

Cite this article

Download citation ▾
Qingxu Li, Jindong Chen, Wen Zhang. Data Augmentation for Fake Review Detection Based on Large Language Model with Multi-stage Prompting. Journal of Systems Science and Systems Engineering 1-22 DOI:10.1007/s11518-026-5731-y

登录浏览全文

4963

注册一个新账户 忘记密码

References

[1]

Budhi G S, Chiong R, Wang Z. Resampling imbalanced data to detect fake reviews using machine learning classifiers and textual based features. Multimedia Tools and Applications, 2021, 80: 13079-13097

[2]

Cohen Y, Aperstein Y. A review of generative pretrained multi-step prompting schemes – And a new multi-step prompting framework, 2024: 2024050720 Preprints 2024

[3]

Dai W, Jin G, Lee J, et al. . Aggregation of consumer ratings: An application to Yelp.com. Quantitative Marketing and Economics, 2018, 16: 289-339

[4]

Devlin J, Chang M W, Lee K, et al. . BERT: Pretraining of deep bidirectional transformers for language understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), 2019: 4171-4186

[5]

Dieng A B, Kim Y, Rush A M, et al. . Avoiding latent variable collapse with generative skip models. Proceedings of the 22nd International Conference on Artificial Intelligence and Statistics, 2019: 2397-2405

[6]

Dharma E M, Gaol F L, Warnars H, et al. . The accuracy comparison among Word2Vec, GloVe, and FastText towards convolution neural network (CNN) text classification. Journal of Theoretical and Applied Information Technology, 2022, 100(2): 31

[7]

Duth S, Vedavathi S, Roshan S. Herbal leaf classification using RCNN, Fast RCNN, Faster RCNN. 2023 7th International Conference on Computing, Communication, Control and Automation (ICCUBEA), 2023: 1-8

[8]

Gao M, Li T, Huang P. Text classification research based on improved Word2Vec and CNN. Service-Oriented Computing–ICSOC 2018 Workshops, 2018, Cham, Springer: 126-135

[9]

Grattafiori A, Dubey A, Jauhri A, et al. (2024). The Llama 3 herd of models. arXiv Preprint arXiv: 2407.21783.

[10]

Graves A, Fernández S, Schmidhuber J. Bidirectional LSTM networks for improved phoneme classification and recognition. International Conference on Artificial Neural Networks, 2005, Berlin, Heidelberg, Springer: 799-804

[11]

Goodfellow I, Pouget-Abadie J, Mirza M, et al. . Generative adversarial nets. Advances in Neural Information Processing Systems, 2014, 27: 2672-2680

[12]

Guo J, Lu S, Cai H, et al. . Long text generation via adversarial training with leaked information. Proceedings of the AAAI Conference on Artificial Intelligence, 2018, 32(1): 5141-5148

[13]

Guo D, Yang D, Zhang H, et al. (2025). DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning. arXiv Preprint arXiv: 2501.12948.

[14]

Ji Z, Lee N, Frieske R, et al. . Survey of hallucination in natural language generation. ACM Computing Surveys, 2023, 55(12): 1-38

[15]

Khondaker M T I, Naeem N, Khan F, et al. . Benchmarking Llama-3 on Arabic language generation tasks. Proceedings of the Second Arabic Natural Language Processing Conference, 2024: 283-297

[16]

Kingma D P, Welling M. An introduction to variational autoencoders. Foundations and Trends in Machine Learning, 2019, 12(4): 307-392

[17]

Li J, Luong M T, Jurafsky D (2015). A hierarchical neural autoencoder for paragraphs and documents. arXiv Preprint arXiv: 1506.01057.

[18]

Li Z, Liu F, Yang W, et al. . A survey of convolutional neural networks: Analysis, applications, and prospects. IEEE Transactions on Neural Networks and Learning Systems, 2021, 33(12): 6999-7019

[19]

Liu Y, Wang L, Shi T, et al. . Detection of spam reviews through a hierarchical attention architecture with N-gram CNN and Bi-LSTM. Information Systems, 2022, 103: 101865

[20]

Lotter W, Sorensen G, Cox D. A multi-scale CNN and curriculum learning strategy for mammogram classification. International Workshop on Deep Learning in Medical Image Analysis, 2017, Cham, Springer: 169-177

[21]

Maity S, Deroy A, Sarkar S. A novel multi-stage prompting approach for language agnostic MCQ generation using GPT. European Conference on Information Retrieval, 2024, Cham, Springer: 268-277

[22]

Maharana K, Mondal S, Nemade B. A review: Data pre-processing and data augmentation techniques. Global Transitions Proceedings, 2022, 3(1): 91-99

[23]

Madaan N, Saha D, Bedathur S. Counterfactual sentence generation with plug-and-play perturbation. 2023 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML), 2023: 306-315

[24]

Mir A Q, Khan F Y, Chishti M A (2023). Online fake review detection using supervised machine learning and BERT model. arXiv Preprint arXiv: 2301.03225.

[25]

Mohawesh R, Xu S, Tran S N, et al. . Fake reviews detection: A survey. IEEE Access, 2021, 9: 65771-65802

[26]

Mo Y, Qin H, Dong Y, et al. (2024). Large language model (LLM) AI text generation detection based on transformer deep learning algorithm. arXiv Preprint arXiv: 2405.06652.

[27]

Nagano T, Kurata G, Thomas S, et al. . LLM based text generation for improved low-resource speech recognition models. ICASSP 2025–2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2025: 1-5

[28]

Nan Q, Sheng Q, Cao J, et al. . Let silence speak: Enhancing fake news detection with generated comments from large language models. Proceedings of the 33rd ACM International Conference on Information and Knowledge Management, 2024: 1732-1742

[29]

Paruchuri V L, Rajesh P. CyberNet: A hybrid deep CNN with N-gram feature selection for cyberbullying detection in online social networks. Evolutionary Intelligence, 2023, 16(6): 1935-1949

[30]

Pearton S J, Zolper J C, Shul R J, et al. . GaN: Processing, defects, and devices. Journal of Applied Physics, 1999, 86(1): 1-78

[31]

Qiao S, Ou Y, Zhang N, et al. (2022). Reasoning with language model prompting: A survey. arXiv Preprint arXiv: 2212.09597.

[32]

Salehi P, Chalechale A, Taghizadeh M (2020). Generative adversarial networks (GANs): An overview of theoretical model, evaluation metrics, and recent developments. arXiv Preprint arXiv: 2005.13178.

[33]

Sathyanandani S, Sreedharan D. An e-commerce feedback review mining for a trusted seller’s profile by classifying fake and authentic feedback comments. 2017 International Conference on Circuit, Power and Computing Technologies (ICCPCT), 2017: 1-6

[34]

Shan L, Liu Y, Tang M, et al. . CNN-BiLSTM hybrid neural networks with attention mechanism for well log prediction. Journal of Petroleum Science and Engineering, 2021, 205: 108838

[35]

Suhaeni C, Yong H S. Mitigating class imbalance in sentiment analysis through GPT-3-generated synthetic sentences. Applied Sciences, 2023, 13(17): 9766

[36]

Team G L M, Zeng A, Xu B, et al. (2024). ChatGLM: A family of large language models from GLM-130B to GLM-4 all tools. arXiv e-prints: arXiv: 2406.12793.

[37]

Van Den Oord A, Vinyals O. Neural discrete representation learning. Advances in Neural Information Processing Systems, 2017, 30: 6309-6318

[38]

Vaswani A, Shazeer N, Parmar N, et al. . Attention is all you need. Advances in Neural Information Processing Systems, 2017, 30: 5998-6008

[39]

Wang H. Multi-label text classification using GloVe and neural network models. 2023 3rd International Conference on Electronic Information Engineering and Computer (EIECT), 2023: 512-515

[40]

Wang P, Fan E, Wang P. Comparative analysis of image classification algorithms based on traditional machine learning and deep learning. Pattern Recognition Letters, 2021, 141: 61-67

[41]

Wang P, Bai S, Tan S, et al. (2024). Qwen2-VL: Enhancing vision-language model’s perception of the world at any resolution. arXiv Preprint arXiv: 2409.12191.

[42]

Wei J, Wang X, Schuurmans D, et al. (2022). Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems: 24824–24837.

[43]

Wu J, Yang S, Zhan R, et al. (2025). A survey on LLM-generated text detection: Necessity, methods, and future directions. Computational Linguistics: 1–66.

[44]

Yang A, Li Z, Li J (2024). Advancing GenAI assisted programming – A comparative study on prompt efficiency and code quality between GPT-4 and GLM-4. arXiv Preprint arXiv: 2402.12782.

[45]

Yu L, Zhang W, Wang J, et al. . SeqGAN: Sequence generative adversarial nets with policy gradient. Proceedings of the AAAI Conference on Artificial Intelligence, 2017, 31(1): 2852-2858

[46]

Zhang W, Yoshida T, Tang X. Text classification based on multi-word with support vector machine. Knowledge-Based Systems, 2008, 21(8): 879-886

[47]

Zhang W, Zhao J, Quan P, et al. . Prediction of influent wastewater quality based on wavelet transform and residual LSTM. Applied Soft Computing, 2023, 148: 110858

[48]

Zhou M Z, Chen J D, Zhang W, et al. . Fraud detection based on GNNs with local augmentation and adaptive relation aggregation. Expert Systems with Applications, 2026, 299: 130110

RIGHTS & PERMISSIONS

Systems Engineering Society of China and Springer-Verlag GmbH Germany

PDF

0

Accesses

0

Citation

Detail

Sections
Recommended

/