ColdChat: benchmarking large language model personalization using limited real-user interaction history

Wenbin ZHANG , Zheni ZENG , Zhiyuan LIU , Wanxiang CHE

Front. Comput. Sci. ›› 2026, Vol. 20 ›› Issue (9) : 2009373

PDF (8260KB)
Front. Comput. Sci. ›› 2026, Vol. 20 ›› Issue (9) :2009373 DOI: 10.1007/s11704-026-51716-6
Artificial Intelligence
RESEARCH ARTICLE
ColdChat: benchmarking large language model personalization using limited real-user interaction history
Author information +
History +
PDF (8260KB)

Abstract

Existing methods for Large Language Models (LLMs) personalization typically rely on extensive pre-collected user data. However, practical personalization for LLMs-based chatbots often starts from a “cold-start” scenario with limited interaction history, where forming an accurate initial impression is crucial for user acquisition and retention. This critical challenge of “cold-start” personalization is further compounded by the absence of dedicated benchmarks for its evaluation. To address this gap, we introduce ColdChat, the first benchmark designed to assess LLM personalization using brief interaction histories. The collection process for ColdChat involved tasking human annotators with (1) engage in multi-session open-domain dialogues with LLM, (2) annotate personalized user profiles based on their dialogue history, and (3) label a user-specific test set for the evaluation of personalization. Our experiment on ColdChat reveals that state-of-the-art LLMs struggle to be personalized in this “cold-start” setup. To enhance LLM personalization in this scenario, we propose EPIC, a novel framework for Extracting user Profile from Interaction Context. Experimental results on ColdChat reveal that the profiles extracted by EPIC yield substantial improvements for personalization, yielding a +27.2% gain in Spearman’s ρ for the personalized ranking task and +12.7% in LLM-as-a-judge on the personalized generation task compared to the baseline method that relies solely on dialogue history. Together, our work establishes a foundational benchmark and a robust framework to advance LLM personalization from early interaction, representing a critical first step towards effective “cold-start” LLM personalization.

Graphical abstract

Keywords

large language models / personalized LLM / dialogue system / human-AI interaction

Cite this article

Download citation ▾
Wenbin ZHANG, Zheni ZENG, Zhiyuan LIU, Wanxiang CHE. ColdChat: benchmarking large language model personalization using limited real-user interaction history. Front. Comput. Sci., 2026, 20 (9) : 2009373 DOI:10.1007/s11704-026-51716-6

登录浏览全文

4963

注册一个新账户 忘记密码

References

[1]

Gemma Team. Gemma 2: improving open language models at a practical size. 2024, arXiv preprint arXiv: 2408.00118

[2]

Llama Team. The Llama 3 herd of models. 2024, arXiv preprint arXiv: 2407.21783

[3]

Yang A, Yang B, Zhang B, Hui B, Zheng B, , et al. Qwen2.5 technical report. 2024, arXiv preprint arXiv: 2412.15115

[4]

Ouyang L, Wu J, Jiang X, Almeida D, Wainwright C L, , et al. Training language models to follow instructions with human feedback. In: Proceedings of the 36th International Conference on Neural Information Processing Systems. 2022, 2011

[5]

Rafailov R, Sharma A, Mitchell E, Ermon S, Manning C D, Finn C. Direct preference optimization: your language model is secretly a reward model. In: Proceedings of the 37th International Conference on Neural Information Processing Systems. 2023, 2338

[6]

Askell A, Bai Y, Chen A, Drain D, Ganguli D, , et al. A general language assistant as a laboratory for alignment. 2021, arXiv preprint arXiv: 2112.00861

[7]

Bai Y, Jones A, Ndousse K, Askell A, Chen A, , et al. Training a helpful and harmless assistant with reinforcement learning from human feedback. 2022, arXiv preprint arXiv: 2204.05862

[8]

Kirk H R, Whitefield A, Röttger P, Bean A, Margatina K, Ciro J, Mosquera R, Bartolo M, Williams A, He H, Vidgen B, Hale S A. The PRISM alignment dataset: what participatory, representative and individualised human feedback reveals about the subjective and multicultural alignment of large language models. In: Proceedings of the 38th International Conference on Neural Information Processing Systems. 2024, 3342

[9]

Aroyo L, Taylor A S, Díaz M, Homan C M, Parrish A, Serapio-García G, Prabhakaran V, Wang D. DICES dataset: diversity in conversational AI evaluation for safety. In: Proceedings of the 37th International Conference on Neural Information Processing Systems. 2023, 2321

[10]

Siththaranjan A, Laidlaw C, Hadfield-Menell D. Distributional preference learning: understanding and accounting for hidden context in RLHF. In: Proceedings of the 12th International Conference on Learning Representations. 2024

[11]

Sorensen T, Moore J, Fisher J, Gordon M, Mireshghallah N, Rytting C M, Ye A, Jiang L, Lu X, Dziri N, Althoff T, Choi Y. Position: a roadmap to pluralistic alignment. In: Proceedings of the 41st International Conference on Machine Learning. 2024, 1882

[12]

Zhang Z, Rossi R A, Kveton B, Shao Y, Yang D, Zamani H, Dernoncourt F, Barrow J, Yu T, Kim S, Zhang R, Gu J, Derr T, Chen H, Wu J, Chen X, Wang Z, Mitra S, Lipka N, Ahmed N K, Wang Y. Personalization of large language models: a survey. Transactions on Machine Learning Research, 2025

[13]

Tseng Y M, Huang Y C, Hsiao T Y, Chen W L, Huang C W, Meng Y, Chen Y N. Two tales of persona in LLMs: a survey of role-playing and personalization. In: Proceedings of Findings of the Association for Computational Linguistics. 2024, 16612−16631

[14]

Castricato L, Lile N, Rafailov R, Fränken J P, Finn C. PERSONA: a reproducible testbed for pluralistic alignment. In: Proceedings of the 31st International Conference on Computational Linguistics. 2025, 11348−11368

[15]

Ni J, Li J, McAuley J. Justifying recommendations using distantly-labeled reviews and fine-grained aspects. In: Proceedings of 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019, 188−197

[16]

Li J N, Guan J, Wu S, Wu W, Yan R. From 1,000,000 users to every user: scaling up personalized preference for user-level alignment. 2025, arXiv preprint arXiv: 2503.15463

[17]

Zhao S, Dang J, Grover A. Group preference optimization: few-shot alignment of large language models. In: Proceedings of the 12th International Conference on Learning Representations. 2024

[18]

Singh A, Hsu S, Hsu K, Mitchell E, Ermon S, Hashimoto T, Sharma A, Finn C. FSPO: few-shot preference optimization of synthetic preference data elicits LLM personalization to real users. In: Proceedings of the 2nd Workshop on Models of Human Feedback for AI Alignment. 2025

[19]

Purificato E, Boratto L, de Luca E W. User modeling and user profiling: a comprehensive survey. 2024, arXiv preprint arXiv: 2402.09660

[20]

Peintner B, Viappiani P, Yorke-Smith N . Preferences in interactive systems: technical challenges and case studies. AI Magazine, 2008, 29( 4): 13–24

[21]

Zhang W N, Li L, Cao D, Liu T. Exploring implicit feedback for open domain conversation generation. In: Proceedings of the 32nd AAAI Conference on Artificial Intelligence. 2018, 68

[22]

Wu S, Fung Y R, Qian C, Kim J, Hakkani-Tur D, Ji H. Aligning LLMs with individual preferences via interaction. In: Proceedings of the 31st International Conference on Computational Linguistics. 2025, 7648−7662

[23]

Pang R Y, Roller S, Cho K, He H, Weston J. Leveraging implicit feedback from deployment data in dialogue. In: Proceedings of the 18th Conference of the European Chapter of the Association for Computational. 2024, 60−75

[24]

Chen J, Liu Z, Huang X, Wu C, Liu Q, Jiang G, Pu Y, Lei Y, Chen X, Wang X, Zheng K, Lian D, Chen E . When large language models meet personalization: perspectives of challenges and opportunities. World Wide Web, 2024, 27( 4): 42

[25]

Hwang E, Majumder B, Tandon N. Aligning language models to user opinions. In: Proceedings of Findings of the Association for Computational Linguistics. 2023, 5906−5919

[26]

Kirk H R, Vidgen B, Röttger P, Hale S A. Personalisation within bounds: a risk taxonomy and policy framework for the alignment of large language models with personalised feedback. 2023, arXiv preprint arXiv: 2303.05453

[27]

Jang J, Kim S, Lin B Y, Wang Y, Hessel J, Zettlemoyer L, Hajishirzi H, Choi Y, Ammanabrolu P. Personalized soups: personalized large language model alignment via post-hoc parameter merging. In: Proceedings of Adaptive Foundation Models: Evolving AI for Personalized and Efficient Learning. 2024

[28]

Salemi A, Mysore S, Bendersky M, Zamani H. LaMP: when large language models meet personalization. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics. 2024, 7370−7392

[29]

Li C, Zhang M, Mei Q, Wang Y, Hombaiah S A, Liang Y, Bendersky M. Teach LLMs to personalize - an approach inspired by writing education. 2023, arXiv preprint arXiv: 2308.07968

[30]

Mysore S, Lu Z, Wan M, Yang L, Sarrafzadeh B, Menezes S, Baghaee T, Gonzalez E B, Neville J, Safavi T. Pearl: personalizing large language model writing assistants with generation-calibrated retrievers. In: Proceedings of the 1st Workshop on Customizable NLP: Progress and Challenges in Customizing NLP for a Domain, Application, Group, or Individual (CustomNLP4U). 2024, 198−219

[31]

Richardson C, Zhang Y, Gillespie K, Kar S, Singh A, Raeesy Z, Khan O Z, Sethy A. Integrating summarization and retrieval for enhanced personalization via large language models. 2023, arXiv preprint arXiv: 2310.20081

[32]

Liu Q, Chen N, Sakai T, Wu X M. ONCE: boosting content-based recommendation with both open- and closed-source large language models. In: Proceedings of the 17th ACM International Conference on Web Search and Data Mining. 2024, 452−461

[33]

Gu J C, Ling Z H, Zhu X, Liu Q. Dually interactive matching network for personalized response selection in retrieval-based chatbots. In: Proceedings of 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019, 1845−1854

[34]

Zhang J. Guided profile generation improves personalization with large language models. In: Proceedings of Findings of the Association for Computational Linguistics. 2024, 4005−4016

[35]

Ji K, Lian Y, Li L, Gao J, Li W, Dai B. Enhancing persona consistency for LLMs’ role-playing using persona-aware contrastive learning. In: Proceedings of Findings of the Association for Computational Linguistics. 2025, 26221−26238

[36]

Li H, Yang C, Zhang A, Deng Y, Wang X, Chua T S. Hello again! LLM-powered personalized agent for long-term dialogue. In: Proceedings of 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025, 5259−5276

[37]

Santurkar S, Durmus E, Ladhak F, Lee C, Liang P, Hashimoto T. Whose opinions do language models reflect?. In: Proceedings of the 40th International Conference on Machine Learning. 2023, 1244

[38]

Durmus E, Nguyen K, Liao T, Schiefer N, Askell A, Bakhtin A, Chen C, Hatfield-Dodds Z, Hernandez D, Joseph N, Lovitt L, McCandlish S, Sikder O, Tamkin A, Thamkul J, Kaplan J, Clark J, Ganguli D. Towards measuring the representation of subjective global opinions in language models. In: Proceedings of the 1st Conference on Language Modeling. 2024

[39]

Kanoje S, Girase S, Mukhopadhyay D . User profiling trends, techniques and applications. International Journal of Advance Foundation and Research in Computer, 2014, 1( 11): 119–124

[40]

Sutcliffe R. A survey of personality, persona, and profile in conversational agents and chatbots. 2023, arXiv preprint arXiv: 2401.00609

[41]

Qian H, Dou Z, Zhu Y, Ma Y, Wen J R. Learning implicit user profile for personalized retrieval-based chatbot. In: Proceedings of the 30th ACM International Conference on Information & Knowledge Management. 2021, 1467−1477

[42]

Ma Z, Dou Z, Zhu Y, Zhong H, Wen J R. One chatbot per person: Creating personalized chatbots based on implicit user profiles. In: Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval. 2021, 555−564

[43]

Zhong H, Dou Z, Zhu Y, Qian H, Wen J R. Less is more: learning to refine dialogue history for personalized dialogue generation. In: Proceedings of 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022, 5808−5820

[44]

Wang K, Li X, Yang S, Zhou L, Jiang F, Li H. Know you first and be you better: modeling human-like user simulators via implicit profiles. In: Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025, 21082−21107

[45]

Cheng Y, Liu W, Xu K, Hou W, Ouyang Y, Leong C T, Li W, Wu X, Zheng Y. AutoPal: autonomous adaptation to users for personal AI companionship. 2025, arXiv preprint arXiv: 2406.13960

[46]

OpenAI . Memory and new controls for ChatGPT. See simonwillison.net/2024/Feb/14/memory-and-new-controls-for-chatgpt/website, 2024

[47]

Zhong W, Guo L, Gao Q, Ye H, Wang Y. MemoryBank: enhancing large language models with long-term memory. In: Proceedings of the 38th AAAI Conference on Artificial Intelligence. 2024, 2198

[48]

Modarressi A, Imani A, Fayyaz M, Schütze H. RET-LLM: towards a general read-write memory for large language models. 2023, arXiv preprint arXiv: 2305.14322

[49]

Yuan R, Sun S, Li Y, Wang Z, Cao Z, Li W. Personalized large language model assistant with evolving conditional memory. In: Proceedings of the 31st International Conference on Computational Linguistics. 2025, 3764−3777

[50]

Tan Z, Yan J, Hsu I H, Han R, Wang Z, Le L T, Song Y, Chen Y, Palangi H, Lee G, Iyer A R, Chen T, Liu H, Lee C Y, Pfister T. In prospect and retrospect: reflective memory management for long-term personalized dialogue agents. In: Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics. 2025, 8416−8439

[51]

Wang Y, Jiang Z, Chen Z, Yang F, Zhou Y, Cho E, Fan X, Lu Y, Huang X, Yang Y. RecMind: large language model powered agent for recommendation. In: Proceedings of Findings of the Association for Computational Linguistics. 2024, 4351−4364

[52]

Li Y, Yu Y, Li H, Chen Z, Khashanah K. TradingGPT: multi-agent system with layered memory and distinct characters for enhanced financial trading performance. 2023, arXiv preprint arXiv: 2309.03736

[53]

Qian C, Liu W, Liu H, Chen N, Dang Y, Li J, Yang C, Chen W, Su Y, Cong X, Xu J, Li D, Liu Z, Sun M. ChatDev: communicative agents for software development. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics. 2024, 15174−15186

[54]

OpenAI , Hurst A, Lerer A, Goucher A P, Perelman A, Ramesh A, Clark A, Ostrow A, Welihinda A, Hayes A, et al . Gpt-4o system card. ArXiv, 2024, abs/2410.21276

[55]

Chaves A P, Doerry E, Egbert J, Gerosa M. It’s how you say it: identifying appropriate register for chatbot language design. In: Proceedings of the 7th International Conference on Human-Agent Interaction. 2019, 102−109

[56]

Lønvik A T. The highly sensitive person and chatbots: what personality traits are more important to have in a chatbot, based on a person’s HSP value. Norwegian University of Science and Technology, Dissertation, 2022

[57]

Wei J W, Huang D, Lu Y, Zhou D, Le Q V. Simple synthetic data reduces sycophancy in large language models. 2023, arXiv preprint arXiv: 2308.03958

[58]

Ding N, Chen Y, Xu B, Qin Y, Hu S, Liu Z, Sun M, Zhou B. Enhancing chat language models by scaling high-quality instructional conversations. In: Proceedings of 2023 Conference on Empirical Methods in Natural Language Processing. 2023, 3029−3051

[59]

Cui G, Yuan L, Ding N, Yao G, He B, Zhu W, Ni Y, Xie G, Xie R, Lin Y, Liu Z, Sun M. ULTRAFEEDBACK: boosting language models with scaled AI feedback. In: Proceedings of the 41st International Conference on Machine Learning. 2024, 384

[60]

OpenAI . New embedding models and API updates. See openai.com/index/new-embedding-models-and-api-updates/website, 2024

[61]

Anthropic . Claude 3.5 sonnet model card addendum. See theresanaiforthat.com/paper/claude-3-5-sonnet-model-card-addendum/website, 2024

[62]

Chen L, Li J, Dong X, Zhang P, He C, Wang J, Zhao F, Lin D. ShareGPT4V: improving large multi-modal models with better captions. In: Proceedings of the 18th European Conference on Computer Vision. 2024, 370−387

[63]

Zheng L, Chiang W L, Sheng Y, Li T, Zhuang S, Wu Z, Zhuang Y, Li Z, Lin Z, Xing E, Gonzalez J E, Stoica I, Zhang H. LMSYS-Chat-1M: a large-scale real-world LLM conversation dataset. In: Proceedings of the 12th International Conference on Learning Representations. 2024

[64]

Chen D, Chen Y, Rege A, Wang Z, Vinayak R K. PAL: sample-efficient personalized reward modeling for pluralistic alignment. In: Proceedings of the 13th International Conference on Learning Representations. 2025

[65]

Bradley R A, Terry M E . Rank analysis of incomplete block designs: I. the method of paired comparisons. Biometrika, 1952, 39( 3-4): 324–345

[66]

Zhao S, Hong M, Liu Y, Hazarika D, Lin K. Do LLMs recognize your preferences? Evaluating personalized preference following in LLMs. In: Proceedings of the 13th International Conference on Learning Representations. 2025

[67]

DeepSeek-AI . DeepSeek-V3.2: pushing the frontier of open large language models. 2025, arXiv preprint arXiv: 2512.02556

RIGHTS & PERMISSIONS

Higher Education Press

PDF (8260KB)

Supplementary files

highlights

1530

Accesses

0

Citation

Detail

Sections
Recommended

/