Language models as QSPR predictors: Unleashing potential through in-context learning and instruction tuning
Chenyue Tao , Chengcheng Liu , Chenxuan Li , Bin Yang
ENG.Energy ›› 2026, Vol. 20 ›› Issue (5) : 10760
Quantitative structure–property relationship (QSPR) modeling is a fundamental approach for property-oriented fuel design because it establishes quantitative relationships between molecular structures and physicochemical properties. Conventional QSPR models generally rely on manually engineered molecular descriptors, are usually developed to predict a single property, and often cannot readily incorporate newly available data, limiting their scalability and long-term applicability. Recent advances in large language models (LLMs) have demonstrated remarkable capabilities in text understanding and parsing knowledge in specialized fields such as chemistry. This study proposes FuelProp-LM, a QSPR framework based on instruction tuning and in-context learning that predicts multiple fuel properties with only the SMILES notation of a fuel molecule as input. Four open-source language models were fine-tuned using 100000 data entries obtained from the PubChem database. To enhance contextual reasoning, high-information prompts were constructed through MACCS molecular fingerprint similarity retrieval. The proposed framework was evaluated on 10 fuel physicochemical property datasets and compared with the widely-adopted LLM DeepSeek-V3.2 and existing QSPR approaches. FuelProp-LM consistently achieved competitive or superior predictive performance across downstream tasks. Quantitative analysis further revealed how instruction tuning and in-context learning improve molecular property prediction by enhancing chemically relevant attention patterns. The accuracy, generalizability, data scalability, and ease of use of this work demonstrate that FuelProp-LM can serve as a practical and flexible tool for fuel property prediction and play a significant role in future AI-assisted fuel design.
quantitative structure–property relationship (QSPR) / domain-specific language model / instruction tuning / in-context learning / fuel design
| [1] |
|
| [2] |
|
| [3] |
|
| [4] |
|
| [5] |
|
| [6] |
|
| [7] |
|
| [8] |
|
| [9] |
|
| [10] |
|
| [11] |
Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need. In: Proceedings of the 31st International Conference on Neural Information Processing Systems, Long Beach, USA, 2017, 6000–6010 |
| [12] |
|
| [13] |
|
| [14] |
|
| [15] |
|
| [16] |
|
| [17] |
|
| [18] |
Yu B T, Baker F N, Chen Z Q, et al. LlaSMol: Advancing large language models for chemistry with a large-scale, comprehensive, high-quality instruction tuning dataset. In: 1st Conference on Language Modeling, Philadelphia, USA, 2024 |
| [19] |
Edwards C, Lai T, Ros K, et al. Translation between molecules and natural language. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Abu Dhabi, United Arab Emirates, 2022, 375–413 |
| [20] |
|
| [21] |
|
| [22] |
Dai D M, Sun Y T, Dong L, et al. Why can GPT learn in-context? Language models secretly perform gradient descent as meta-optimizers. In: Proceedings of the Findings of the Association for Computational Linguistics, Toronto, Canada, 2023, 4005–34019 |
| [23] |
|
| [24] |
|
| [25] |
|
| [26] |
|
| [27] |
|
| [28] |
|
| [29] |
Xian Z T, Gu J W, Li L B, et al. MolRAG: Unlocking the power of large language models for molecular property prediction. In: Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, Vienna, Austria, 2025, 15513–15531 |
| [30] |
|
| [31] |
Rajbhandari S, Rasley J, Ruwase O, et al. ZeRO: Memory optimizations toward training trillion parameter models. In: Proceedings of the SC20: International Conference for High Performance Computing, Networking, Storage and Analysis, Atlanta, USA, 2020, 1–16 |
| [32] |
|
| [33] |
|
| [34] |
|
| [35] |
|
| [36] |
|
| [37] |
|
| [38] |
|
| [39] |
|
| [40] |
|
| [41] |
Fang Y, Liang X Z, Zhang N Y, et al. Mol-instructions: A large-scale biomolecular instruction dataset for large language models. In: Proceedings of the 12th International Conference on Learning Representations, Vienna, Austria, 2024 |
| [42] |
|
| [43] |
|
| [44] |
|
| [45] |
Zhou G M, Gao Z F, Ding Q K, et al. Uni-Mol: A universal 3D molecular representation learning framework. In: the11th International Conference on Learning Representationss, Kigali, Rwanda, 2023 |
| [46] |
|
Higher Education Press
/
| 〈 |
|
〉 |