KaLM: knowledge-aligned autoregressive language modeling via dual-view knowledge graph contrastive learning

Peng YU , Cheng DENG , Beiya DAI , Luoyi FU , Xinbing WANG , Guihai CHEN , Ying WEN

Front. Comput. Sci. ›› 2027, Vol. 21 ›› Issue (2) : 2102348

PDF (2787KB)
Front. Comput. Sci. ›› 2027, Vol. 21 ›› Issue (2) :2102348 DOI: 10.1007/s11704-026-50906-6
Artificial Intelligence
RESEARCH ARTICLE
KaLM: knowledge-aligned autoregressive language modeling via dual-view knowledge graph contrastive learning
Author information +
History +
PDF (2787KB)

Abstract

Autoregressive large language models (LLMs) pre-trained by next token prediction are inherently proficient in generative tasks. However, their performance on knowledge-driven tasks such as factual knowledge reasoning remains unsatisfactory. Knowledge graphs (KGs), as high-quality structured knowledge bases, can provide reliable knowledge for LLMs, potentially compensating for their knowledge deficiencies. Aligning LLMs with explicit, structured knowledge from KGs has been a challenge; previous attempts either failed to effectively align knowledge representations or compromised the generative capabilities of LLMs, leading to less-than-optimal outcomes. This paper proposes KaLM, a Knowledge-aligned Language Modeling approach, which fine-tunes autoregressive LLMs to align with KG knowledge via the joint objective of explicit knowledge alignment and implicit knowledge alignment. The explicit knowledge alignment objective aims to directly optimize the knowledge representation of LLMs through dual-view knowledge graph contrastive learning. The implicit knowledge alignment objective focuses on incorporating textual patterns of knowledge into LLMs through triple completion language modeling. The proposed KaLM is a unified framework designed to leverage domain knowledge from KGs for post-training of LLMs, aiming to enhance their knowledge reasoning capabilities and achieve generalization of knowledge representations. Our method achieves a significant performance boost in evaluations of knowledge-driven tasks, particularly in embedding-based knowledge graph completion and generation-based knowledge graph question answering, and also demonstrates strong generalization in out-of-domain (OOD) knowledge representation experiments.

Graphical abstract

Keywords

large language models / knowledge alignment / post training

Cite this article

Download citation ▾
Peng YU, Cheng DENG, Beiya DAI, Luoyi FU, Xinbing WANG, Guihai CHEN, Ying WEN. KaLM: knowledge-aligned autoregressive language modeling via dual-view knowledge graph contrastive learning. Front. Comput. Sci., 2027, 21 (2) : 2102348 DOI:10.1007/s11704-026-50906-6

登录浏览全文

4963

注册一个新账户 忘记密码

References

[1]

Anil R, Dai A M, Firat O, Johnson M, Lepikhin D, , et al. PaLM 2 technical report. 2023, arXiv preprint arXiv: 2305.10403

[2]

Achiam J, Adler S, Agarwal S, Ahmad L, Akkaya I, , et al. GPT-4 technical report. 2023, arXiv preprint arXiv: 2303.08774

[3]

Li J, Tang T, Zhao W X, Wen J R. Pretrained language models for text generation: A survey. 2021, arXiv preprint arXiv: 2201.05273

[4]

Su D, Xu Y, Winata G I, Xu P, Kim H, Liu Z, Fung P. Generalizing question answering system with pre-trained language model fine-tuning. In: Proceedings of the 2nd Workshop on Machine Reading for Question Answering. 2019, 203–211

[5]

Muennighoff N. SGPT: GPT sentence embeddings for semantic search. 2022, arXiv preprint arXiv: 2202.08904

[6]

Ma X, Wang L, Yang N, Wei F, Lin J. Fine-tuning llama for multi-stage text retrieval. In: Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval. 2023, 2421–2425

[7]

Touvron H, Martin L, Stone K, Albert P, Almahairi A, , et al. Llama 2: Open foundation and fine-tuned chat models. 2023, arXiv preprint arXiv: 2307.09288

[8]

Ethayarajh K. How contextual are contextualized word representations? Comparing the geometry of BERT, ELMo, and GPT-2 embeddings. In: Proceedings of 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019, 55–65

[9]

Li B, Zhou H, He J, Wang M, Yang Y, Li L. On the sentence embeddings from pre-trained language models. In: Proceedings of 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 2020, 9119–9130

[10]

Su J, Cao J, Liu W, Ou Y. Whitening sentence representations for better semantics and faster retrieval. 2021, arXiv preprint arXiv: 2103.15316

[11]

Su Y, Lan T, Wang Y, Yogatama D, Kong L, Collier N. A contrastive framework for neural text generation. In: Proceedings of the 36th International Conference on Neural Information Processing Systems. 2022, 1566

[12]

Shen J, Wang C, Gong L, Song D. Joint language semantic and structure embedding for knowledge graph completion. In: Proceedings of the 29th International Conference on Computational Linguistics. 2022, 1965–1978

[13]

Wang X, He Q, Liang J, Xiao Y. Language models as knowledge embeddings. In: Proceedings of the 31st International Joint Conference on Artificial Intelligence. 2022, 2291–2297

[14]

Chen C, Wang Y, Li B, Lam K Y. Knowledge is flat: a Seq2Seq generative framework for various knowledge graph completion. In: Proceedings of the 29th International Conference on Computational Linguistics. 2022, 4005–4017

[15]

Yao L, Peng J, Mao C, Luo Y. Exploring large language models for knowledge graph completion. In: Proceedings of 2025 IEEE International Conference on Acoustics, Speech and Signal Processing. 2025, 1–5

[16]

Wang L, Zhao W, Wei Z, Liu J. SimKGC: Simple contrastive knowledge graph completion with pre-trained language models. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022, 4281–4294

[17]

Fu P, Zhang Y, Wang H, Qiu W, Zhao J. Revisiting the knowledge injection frameworks. In: Proceedings of 2023 Conference on Empirical Methods in Natural Language Processing. 2023, 10983–10997

[18]

Sun J, Xu C, Tang L, Wang S, Lin C, Gong Y, Ni L M, Shum H Y, Guo J. Think-on-graph: Deep and responsible reasoning of large language model on knowledge graph. 2023, arXiv preprint arXiv: 2307.07697

[19]

Jiang J, Zhou K, Dong Z, Ye K, Zhao X, Wen J R. StructGPT: A general framework for large language model to reason over structured data. In: Proceedings of 2023 Conference on Empirical Methods in Natural Language Processing. 2023, 9237–9251

[20]

Feng Z, Ma W, Yu W, Huang L, Wang H, Chen Q, Peng W, Feng X, Qin B, Liu T. Trends in integration of knowledge and large language models: A survey and taxonomy of methods, benchmarks, and applications. 2023, arXiv preprint arXiv: 2311.05876

[21]

Wang X, Gao T, Zhu Z, Zhang Z, Liu Z, Li J, Tang J . Kepler: A unified model for knowledge embedding and pre-trained language representation. Transactions of the Association for Computational Linguistics, 2021, 9: 176–194

[22]

Yasunaga M, Bosselut A, Ren H, Zhang X, Manning C D, Liang P, Leskovec J. Deep bidirectional language-knowledge graph pretraining. In: Proceedings of the 36th International Conference on Neural Information Processing Systems. 2022, 2704

[23]

Zhang Y, Chen Z, Guo L, Xu Y, Zhang W, Chen H. Making large language models perform better in knowledge graph completion. In: Proceedings of the 32nd ACM International Conference on Multimedia. 2024, 233–242

[24]

Rafailov R, Sharma A, Mitchell E, Ermon S, Manning C D, Finn C. Direct preference optimization: Your language model is secretly a reward model. In: Proceedings of the 37th International Conference on Neural Information Processing Systems. 2023, 2338

[25]

Taori R, Gulrajani I, Zhang T, Dubois Y, Li X, Guestrin C, Liang P, Hashimoto T B. Stanford alpaca: An instruction-following llama model, 2023

[26]

Chen T, Kornblith S, Norouzi M, Hinton G E. A simple framework for contrastive learning of visual representations. In: Proceedings of the 37th International Conference on Machine Learning. 2020, 1597–1607

[27]

Wang T, Isola P. Understanding contrastive representation learning through alignment and uniformity on the hypersphere. In: Proceedings of the 37th International Conference on Machine Learning. 2020, 9929–9939

[28]

Dettmers T, Minervini P, Stenetorp P, Riedel S. Convolutional 2D knowledge graph embeddings. In: Proceedings of the 32th AAAI Conference on Artificial Intelligence. 2018, 221

[29]

Toutanova K, Chen D. Observed versus latent features for knowledge base and text inference. In: Proceedings of the 3rd Workshop on Continuous Vector Space Models and Their Compositionality. 2015, 57–66

[30]

Bordes A, Usunier N, Garcia-Duran A, Weston J, Yakhnenko O. Translating embeddings for modeling multi-relational data. In: Proceedings of the 27th International Conference on Neural Information Processing Systems. 2013, 2787–2795

[31]

Yao L, Mao C, Luo Y. KG-BERT: BERT for knowledge graph completion. 2019, arXiv preprint arXiv: 1909.03193

[32]

Hu E J, Shen Y, Wallis P, Allen-Zhu Z, Li Y, Wang S, Wang L, Chen W. LoRA: low-rank adaptation of large language models. In: Proceedings of the 10th International Conference on Learning Representations. 2022

[33]

Yang B, Yih W T, He X, Gao J, Deng L. Embedding entities and relations for learning and inference in knowledge bases. In: Proceedings of the 3rd International Conference on Learning Representations. 2015

[34]

Sun Z, Deng Z H, Nie J Y, Tang J. RotatE: Knowledge graph embedding by relational rotation in complex space. In: Proceedings of the 7th International Conference on Learning Representations. 2019

[35]

Wang B, Shen T, Long G, Zhou T, Wang Y, Chang Y. Structure-augmented text representation learning for efficient knowledge graph completion. In: Proceedings of the Web Conference 2021. 2021, 1737–1748

[36]

Hendrycks D, Burns C, Basart S, Zou A, Mazeika M, Song D, Steinhardt J. Measuring massive multitask language understanding. In: Proceedings of the Web Conference 2021. 2021

[37]

Cobbe K, Kosaraju V, Bavarian M, Chen M, Jun H, Kaiser L, Plappert M, Tworek J, Hilton J, Nakano R, Hesse C, Schulman J. Training verifiers to solve math word problems. 2021, arXiv preprint arXiv: 2110.14168

[38]

Lightman H, Kosaraju V, Burda Y, Edwards H, Baker B, Lee T, Leike J, Schulman J, Sutskever I, Cobbe K. Let’s verify step by step. In: Proceedings of the 12th International Conference on Learning Representations. 2024

[39]

Lu P, Qiu L, Chang K W, Wu Y N, Zhu S C, Rajpurohit T, Clark P, Kalyan A. Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning. In: Proceedings of the 11th International Conference on Learning Representations. 2023

[40]

Merity S, Xiong C, Bradbury J, Socher R. Pointer sentinel mixture models. In: Proceedings of the 5th International Conference on Learning Representations. 2017

[41]

Yih W T, Richardson M, Meek C, Chang M W, Suh J. The value of semantic parse labeling for knowledge base question answering. In: Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). 2016, 201–206

[42]

Talmor A, Berant J. The web as a knowledge-base for answering complex questions. In: Proceedings of 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018, 641–651

[43]

He J, Zhou C, Ma X, Berg-Kirkpatrick T, Neubig G. Towards a unified view of parameter-efficient transfer learning. In: Proceedings of the 10th International Conference on Learning Representations. 2022

[44]

Liu S, Fan H, Qian S, Chen Y, Ding W, Wang Z. HiT: Hierarchical transformer with momentum contrast for video-text retrieval. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. 2021, 11895–11905

[45]

Gunel B, Du J, Conneau A, Stoyanov V. Supervised contrastive learning for pre-trained language model fine-tuning. In: Proceedings of the 9th International Conference on Learning Representations. 2021

[46]

Wang F, Liu H. Understanding the behaviour of contrastive loss. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2021, 2495–2504

[47]

Dubey A, Jauhri A, Pandey A, Kadian A, Al-Dahle A, , et al. The llama 3 herd of models. 2024, arXiv preprint arXiv: 2407.21783

[48]

Yang A, Yang B, Zhang B, Hui B, Zheng B, , et al. Qwen2.5 technical report. 2025, arXiv preprint arXiv: 2412.12115

[49]

Jiang A Q, Sablayrolles A, Mensch A, Bamford C, Chaplot D S, de las Casas D, Bressand F, Lengyel G, Lample G, Saulnier L, Lavaud L R, Lachaux M A, Stock P, Le Scao T, Lavril T, Wang T, Lacroix T, El Sayed W. Mistral 7B. 2023, arXiv preprint arXiv: 2310.06825

[50]

Zhang S, Roller S, Goyal N, Artetxe M, Chen M, Chen S, Dewan C, Diab M, Li X, Lin X V, Mihaylov T, Ott M, Shleifer S, Shuster K, Simig D, Koura P S, Sridhar A, Wang T, Zettlemoyer L. OPT: Open pre-trained transformer language models. 2022, arXiv preprint arXiv: 2205.01068

[51]

Biderman S, Schoelkopf H, Anthony Q G, Bradley H, O’Brien K, Hallahan E, Khan M A, Purohit S, USVSN Sai Prashanth, Raff E, Skowron A, Sutawika L, van der Wal O. Pythia: A suite for analyzing large language models across training and scaling. In: Proceedings of the 40th International Conference on Machine Learning. 2023, 2397–2430

RIGHTS & PERMISSIONS

Higher Education Press

PDF (2787KB)

299

Accesses

0

Citation

Detail

Sections
Recommended

/