About the journal
Browse
Collections
Multimedia collections
Authors & reviewers
ColdChat: benchmarking large language model personalization using limited real-user interaction history
Wenbin ZHANG , Zheni ZENG , Zhiyuan LIU , Wanxiang CHE
Front. Comput. Sci. ›› 2026, Vol. 20 ›› Issue (9) : 2009373
Existing methods for Large Language Models (LLMs) personalization typically rely on extensive pre-collected user data. However, practical personalization for LLMs-based chatbots often starts from a “cold-start” scenario with limited interaction history, where forming an accurate initial impression is crucial for user acquisition and retention. This critical challenge of “cold-start” personalization is further compounded by the absence of dedicated benchmarks for its evaluation. To address this gap, we introduce ColdChat, the first benchmark designed to assess LLM personalization using brief interaction histories. The collection process for ColdChat involved tasking human annotators with (1) engage in multi-session open-domain dialogues with LLM, (2) annotate personalized user profiles based on their dialogue history, and (3) label a user-specific test set for the evaluation of personalization. Our experiment on ColdChat reveals that state-of-the-art LLMs struggle to be personalized in this “cold-start” setup. To enhance LLM personalization in this scenario, we propose EPIC, a novel framework for Extracting user Profile from Interaction Context. Experimental results on ColdChat reveal that the profiles extracted by EPIC yield substantial improvements for personalization, yielding a +27.2% gain in Spearman’s for the personalized ranking task and +12.7% in LLM-as-a-judge on the personalized generation task compared to the baseline method that relies solely on dialogue history. Together, our work establishes a foundational benchmark and a robust framework to advance LLM personalization from early interaction, representing a critical first step towards effective “cold-start” LLM personalization.
large language models / personalized LLM / dialogue system / human-AI interaction
| [1] |
|
| [2] |
|
| [3] |
|
| [4] |
|
| [5] |
|
| [6] |
|
| [7] |
|
| [8] |
|
| [9] |
|
| [10] |
|
| [11] |
|
| [12] |
|
| [13] |
|
| [14] |
|
| [15] |
Ni J, Li J, McAuley J. Justifying recommendations using distantly-labeled reviews and fine-grained aspects. In: Proceedings of 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019, 188−197 |
| [16] |
|
| [17] |
|
| [18] |
|
| [19] |
|
| [20] |
|
| [21] |
Zhang W N, Li L, Cao D, Liu T. Exploring implicit feedback for open domain conversation generation. In: Proceedings of the 32nd AAAI Conference on Artificial Intelligence. 2018, 68 |
| [22] |
|
| [23] |
|
| [24] |
|
| [25] |
|
| [26] |
|
| [27] |
|
| [28] |
|
| [29] |
|
| [30] |
|
| [31] |
|
| [32] |
|
| [33] |
Gu J C, Ling Z H, Zhu X, Liu Q. Dually interactive matching network for personalized response selection in retrieval-based chatbots. In: Proceedings of 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019, 1845−1854 |
| [34] |
|
| [35] |
|
| [36] |
Li H, Yang C, Zhang A, Deng Y, Wang X, Chua T S. Hello again! LLM-powered personalized agent for long-term dialogue. In: Proceedings of 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025, 5259−5276 |
| [37] |
|
| [38] |
|
| [39] |
|
| [40] |
|
| [41] |
|
| [42] |
|
| [43] |
Zhong H, Dou Z, Zhu Y, Qian H, Wen J R. Less is more: learning to refine dialogue history for personalized dialogue generation. In: Proceedings of 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022, 5808−5820 |
| [44] |
|
| [45] |
|
| [46] |
|
| [47] |
|
| [48] |
|
| [49] |
|
| [50] |
|
| [51] |
|
| [52] |
|
| [53] |
|
| [54] |
OpenAI , Hurst A, Lerer A, Goucher A P, Perelman A, Ramesh A, Clark A, Ostrow A, Welihinda A, Hayes A, et al . Gpt-4o system card. ArXiv, 2024, abs/2410.21276 |
| [55] |
|
| [56] |
|
| [57] |
|
| [58] |
Ding N, Chen Y, Xu B, Qin Y, Hu S, Liu Z, Sun M, Zhou B. Enhancing chat language models by scaling high-quality instructional conversations. In: Proceedings of 2023 Conference on Empirical Methods in Natural Language Processing. 2023, 3029−3051 |
| [59] |
|
| [60] |
|
| [61] |
|
| [62] |
Chen L, Li J, Dong X, Zhang P, He C, Wang J, Zhao F, Lin D. ShareGPT4V: improving large multi-modal models with better captions. In: Proceedings of the 18th European Conference on Computer Vision. 2024, 370−387 |
| [63] |
|
| [64] |
|
| [65] |
|
| [66] |
|
| [67] |
|
Higher Education Press
/
| 〈 |
|
〉 |