PDF
Abstract
Faced with the various and massive information resources, it is prominent to provide users with accurate, personalized service content efficiently and comprehensively, addressing their diverse retrieval needs. To this end, for University Digital Libraries (UDLs), we propose a Personalized Information Retrieval method for UDLs based on Probabilistic Graphical Model (PIRPGM). This method integrates both keyword retrieval and semantic retrieval to enhance the relevance and precision of search results. We introduce a “Spike and Slab” prior and design a novel Collapsed Variational Bayesian (CVB) inference algorithm to estimate model parameter. The PIRPGM not only offers deep insights into topic but also alleviates data sparsity. Its effectiveness is verified across different lengths (e.g., short texts and long texts) and scenarios, particularly benefiting cold start users.
Keywords
Probabilistic graphical model
/
semantic retrieval
/
personalized information retrieval
/
sparse data
Cite this article
Download citation ▾
Xuan Tang, Zhiyan Ye, Tingting Zhu.
Personalized Information Retrieval based on Probabilistic Graphical Model for Sparse Data.
Journal of Systems Science and Systems Engineering 1-19 DOI:10.1007/s11518-026-5753-5
| [1] |
Abbasiantaeb Z, Momtazi S. Text-based question answering from information retrieval and deep neural network perspectives: A survey. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 2021, 11(6): e1412
|
| [2] |
Alagarsamy R, Sahaaya Arul Mary S. Intelligent rule-based approach for effective information retrieval and dynamic storage in local repositories. The Journal of Supercomputing, 2020, 76(6): 3984-3998
|
| [3] |
Alhabashneh O, Iqbal R, Doctor F, James A. Fuzzy rule based profiling approach for enterprise information seeking and retrieval. Information Sciences, 2017, 394: 18-37
|
| [4] |
Ali J (2025). Probabilistic hesitant fuzzy group decision analysis using partitioned Maclaurin symmetric mean operators. Journal of Applied Mathematics and Computing: 1–29.
|
| [5] |
Ali J, Pamucar D (2025). Normal wiggly probabilistic hesitant fuzzy-based TODIM approach for optimal solid waste disposal method selection. Heliyon 11(2).
|
| [6] |
Arms W Y. Digital Libraries, 2001
|
| [7] |
Asuncion A, Welling M, Smyth P, Teh Y W (2012). On smoothing and inference for topic models. arXiv Preprint arXiv: 1205.2662.
|
| [8] |
Bashir Z, Ali J, Rashid T. Consensus-based robust decision making methods under a novel study of probabilistic uncertain linguistic information and their application in Forex investment. Artificial Intelligence Review, 2021, 54(3): 2091-2132
|
| [9] |
Benahal A R (2024). Evaluation of AI-generated keywords for information retrieval in library catalogues. Journal of Information and Knowledge: 197–203.
|
| [10] |
Blei D M, Ng A Y, Jordan M I. Latent dirichlet allocation. Journal of Machine Learning Research, 2003, 3(Jan): 993-1022
|
| [11] |
Bonifacio L, Abonizio H, Fadaee M, Nogueira R (2022). InPars: Data augmentation for information retrieval using large language models. arXiv Preprint arXiv: 2202.05144.
|
| [12] |
Borgman C L. The science library catalog project: Comparison of children’s searching behavior in hypertext and a keyword search system. Proceedings of the ASIST Annual Meeting, 1991, 28: 162-169
|
| [13] |
Brants T, Popat A C, Xu P, Och F J, Dean J. Large language models in machine translation. Proceedings of the 2007 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning, 2007858-867
|
| [14] |
Cai F, Chen H. A probabilistic model for information retrieval by mining user behaviors. Cognitive Computation, 2016, 8: 494-504
|
| [15] |
Carlini N, Tramer F, Wallace E, Jagielski M, Herbert-Voss A, Lee K, et al.. Extracting training data from large language models. 30th USENIX Security Symposium (USENIX Security 21), 20212633-2650
|
| [16] |
Chao H. Assessing the quality of academic libraries on the web: The development and testing of criteria. Library & Information Science Research, 2002, 24(2): 169-194
|
| [17] |
Chen G, Xiao L. Selecting publication keywords for domain analysis in bibliometrics: A comparison of three methods. Journal of Informetrics, 2016, 10(1): 212-223
|
| [18] |
Chen L, Jose J M, Yu H, Yuan F, Zhang D. A semantic graph based topic model for question retrieval in community question answering. Proceedings of the Ninth ACM International Conference on Web Search and Data Mining, 2016287-296
|
| [19] |
Chen X, et al.. Information fusion and artificial intelligence for smart healthcare: A bibliometric study. Information Processing & Management, 2023, 60(1): 103113
|
| [20] |
Djenouri Y, Belhadi A, Fournier-Viger P, Lin J C W. Fast and effective cluster-based information retrieval using frequent closed itemsets. Information Sciences, 2018, 453: 154-167
|
| [21] |
Gao J, Toutanova K, Yih W T. Clickthrough-based latent semantic models for web search. Proceedings of the 34th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2011675-684
|
| [22] |
Gao J, Peng P, Lu F, Claramunt C, Xu Y. Towards travel recommendation interpretability: Disentangling tourist decision-making process via knowledge graph. Information Processing & Management, 2023, 60(4): 103369
|
| [23] |
Goker A, Davies J. Information Retrieval: Searching in the 21st Century, 2009
|
| [24] |
Griffiths T L, Steyvers M. Finding scientific topics. Proceedings of the National Academy of Sciences, 2004, 101: 5228-5235
|
| [25] |
Guo J, Cai Y, Fan Y, Sun F, Zhang R, Cheng X. Semantic models for the first-stage retrieval: A comprehensive review. ACM Transactions on Information Systems (TOIS), 2022, 40(4): 1-42
|
| [26] |
Huang M, Peng W, Wang D (2021). TPRM: A topic-based personalized ranking model for web search. arXiv Preprint arXiv: 2108.06014.
|
| [27] |
Ishwaran H, Rao J S. Spike and slab variable selection: Frequentist and Bayesian strategies. The Annals of Statistics, 2005, 33(2): 730-773
|
| [28] |
Jiang Y. Semantically-enhanced information retrieval using multiple knowledge sources. Cluster Computing, 2020, 23(4): 2925-2944
|
| [29] |
Khademizadeh S, Nematollahi Z, Danesh F. Analysis of book circulation data and a book recommendation system in academic libraries using data mining techniques. Library & Information Science Research, 2022, 44(4): 101191
|
| [30] |
Lashkari A H, Mahdavi F, Ghomi V. A boolean model in information retrieval for search engines. 2009 International Conference on Information Management and Engineering, 2009385-389
|
| [31] |
Li L, Xu Q, Gan T, Tan C, Lim J H. A probabilistic model of social working memory for information retrieval in social interactions. IEEE Transactions on Cybernetics, 2017, 48(5): 1540-1552
|
| [32] |
Lian D, Zheng K, Ge Y, Cao L, Chen E, Xie X. GeoMF++ scalable location recommendation via joint geographical modeling and matrix factorization. ACM Transactions on Information Systems (TOIS), 2018, 36(3): 1-29
|
| [33] |
Lin T, Tian W, Mei Q, Cheng H. The dual-sparse topic model: Mining focused topics and focused terms in short text. Proceedings of the 23rd International Conference on World Wide Web, 2014539-550
|
| [34] |
Liu X, Croft W B. Statistical language modeling for information retrieval. Annual Review of Information Science and Technology, 2005, 39(1): 1-31
|
| [35] |
Liu J, Toubia O. A semantic approach for estimating consumer content preferences from online search queries. Marketing Science, 2018, 37(6): 930-952
|
| [36] |
Malik M A, Bashir Z, Rashid T, Ali J. Probabilistic hesitant intuitionistic linguistic term sets in multi-attribute group decision making. Symmetry, 2018, 10(9): 392
|
| [37] |
Nguyen G, Dlugolinsky S, Bobák M, Tran V, López García Á, Heredia I, et al.. Machine learning and deep learning frameworks and libraries for large-scale data mining: A survey. Artificial Intelligence Review, 2019, 52: 77-124
|
| [38] |
Peinelt N, Nguyen D, Liakata M. tBERT: Topic models and BERT joining forces for semantic similarity detection. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 20207047-7055
|
| [39] |
Porcel C, Herrera-Viedma E. Dealing with incomplete information in a fuzzy linguistic recommender system to disseminate information in university digital libraries. Knowledge-Based Systems, 2010, 23(1): 32-39
|
| [40] |
Qian Y, Liu Y, Jiang Y, Liu X. Detecting topic-level influencers in large-scale scientific networks. World Wide Web, 2020, 23(2): 831-851
|
| [41] |
Qian Y, Jiang Y, Shang J, Chai Y, Liu Y. Why some products compete and others don’t: A competitive attribution model from customer perspective. Decision Support Systems, 2023, 169: 113956
|
| [42] |
Sato I, Nakagawa H (2012). Rethinking collapsed variational Bayes inference for LDA. arXiv Preprint arXiv: 1206.6435.
|
| [43] |
Schatz B R. Information retrieval in digital libraries: Bringing search to the net. Science, 1997, 275(5298): 327-334
|
| [44] |
Teh Y, Newman D, Welling M (2006). A collapsed variational Bayesian inference algorithm for latent Dirichlet allocation. Advances in Neural Information Processing Systems: 19.
|
| [45] |
Tian Y, Zheng B, Wang Y, Zhang Y, Wu Q. College library personalized recommendation system based on hybrid recommendation algorithm. Procedia CIRP, 2019, 83: 490-494
|
| [46] |
Varelas G, Voutsakis E, Raftopoulou P, Petrakis E G, Milios E E. Semantic similarity methods in wordnet and their application to information retrieval on the web. Proceedings of the 7th Annual ACM International Workshop on Web Information and Data Management, 200510-16
|
| [47] |
Voorhees E M. Using WordNet to disambiguate word senses for text retrieval. Proceedings of the 16th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, 1993171-180
|
| [48] |
Wang W, Barnaghi P M, Bargiela A. Probabilistic topic models for learning terminological ontologies. IEEE Transactions on Knowledge and Data Engineering, 2009, 22(7): 1028-1040
|
| [49] |
Wang W, Li Q, Xie J, Hu N, Wang Z, Zhang N. Research on emotional semantic retrieval of attention mechanism oriented to audio-visual synesthesia. Neurocomputing, 2023, 519: 194-204
|
| [50] |
Weikum G, Kasneci G, Ramanath M, Suchanek F. Database and information-retrieval methods for knowledge discovery. Communications of the ACM, 2009, 52(4): 56-64
|
| [51] |
Wu Y, Wang X, Zhao W, Lv X. A novel topic clustering algorithm based on graph neural network for question topic diversity. Information Sciences, 2023, 629: 685-702
|
| [52] |
Xianghua F, Guo L, Yanyan G, Zhiqiang W. Multi-aspect sentiment analysis for Chinese online social reviews based on topic modeling and HowNet lexicon. Knowledge-Based Systems, 2013, 37: 186-195
|
| [53] |
Zamani H, Dumais S, Craswell N, Bennett P, Lueck G. Generating clarifying questions for information retrieval. Proceedings of the Web Conference, 2020418-428
|
| [54] |
Zhu S, Li Y, Shao Y. Research on construction and automatic expansion of multi-source lexical semantic knowledge base. China Conference on Knowledge Graph and Semantic Computing, 201974-85
|
| [55] |
Zuo Y, Li C, Lin H, Wu J. Topic modeling of short texts: A pseudo-document view with word embedding enhancement. IEEE Transactions on Knowledge and Data Engineering, 2021, 35(1): 972-985
|
Rights & permissions
Systems Engineering Society of China and Springer-Verlag GmbH Germany
Just Accepted
This article has successfully passed peer review and final editorial review, and will soon enter typesetting, proofreading and other publishing processes. The currently displayed version is the accepted final manuscript. The officially published version will be updated with format, DOI and citation information upon launch. We recommend that you pay attention to subsequent journal notifications and preferentially cite the officially published version. Thank you for your support and cooperation.