Language models as QSPR predictors: Unleashing potential through in-context learning and instruction tuning

Chenyue Tao , Chengcheng Liu , Chenxuan Li , Bin Yang

ENG.Energy ›› 2026, Vol. 20 ›› Issue (5) : 10760

PDF (3480KB)
ENG.Energy ›› 2026, Vol. 20 ›› Issue (5) :10760 DOI: 10.1007/s11708-026-1076-y
RESEARCH ARTICLE
Language models as QSPR predictors: Unleashing potential through in-context learning and instruction tuning
Author information +
History +
PDF (3480KB)

Abstract

Quantitative structure–property relationship (QSPR) modeling is a fundamental approach for property-oriented fuel design because it establishes quantitative relationships between molecular structures and physicochemical properties. Conventional QSPR models generally rely on manually engineered molecular descriptors, are usually developed to predict a single property, and often cannot readily incorporate newly available data, limiting their scalability and long-term applicability. Recent advances in large language models (LLMs) have demonstrated remarkable capabilities in text understanding and parsing knowledge in specialized fields such as chemistry. This study proposes FuelProp-LM, a QSPR framework based on instruction tuning and in-context learning that predicts multiple fuel properties with only the SMILES notation of a fuel molecule as input. Four open-source language models were fine-tuned using 100000 data entries obtained from the PubChem database. To enhance contextual reasoning, high-information prompts were constructed through MACCS molecular fingerprint similarity retrieval. The proposed framework was evaluated on 10 fuel physicochemical property datasets and compared with the widely-adopted LLM DeepSeek-V3.2 and existing QSPR approaches. FuelProp-LM consistently achieved competitive or superior predictive performance across downstream tasks. Quantitative analysis further revealed how instruction tuning and in-context learning improve molecular property prediction by enhancing chemically relevant attention patterns. The accuracy, generalizability, data scalability, and ease of use of this work demonstrate that FuelProp-LM can serve as a practical and flexible tool for fuel property prediction and play a significant role in future AI-assisted fuel design.

Graphical abstract

Keywords

quantitative structure–property relationship (QSPR) / domain-specific language model / instruction tuning / in-context learning / fuel design

Cite this article

Download citation ▾
Chenyue Tao, Chengcheng Liu, Chenxuan Li, Bin Yang. Language models as QSPR predictors: Unleashing potential through in-context learning and instruction tuning. ENG.Energy, 2026, 20 (5) : 10760 DOI:10.1007/s11708-026-1076-y

登录浏览全文

4963

注册一个新账户 忘记密码

References

[1]

Sarathy S M , Eraqi B A . Artificial intelligence for novel fuel design. Proceedings of the Combustion Institute, 2024, 40(1–4): 105630

[2]

Kuzhagaliyeva N , Horváth S , Williams J . et al. Artificial intelligence-driven design of fuel mixtures. Communications Chemistry, 2022, 5(1): 111

[3]

Joback K G , Reid R C . Estimation of pure-component properties from group-contributions. Chemical Engineering Communications, 1987, 57(1–6): 233–243

[4]

Pepiot-Desjardins P , Pitsch H , Malhotra R . et al. Structural group analysis for soot reduction tendency of oxygenated fuels. Combustion and Flame, 2008, 154(1–2): 191–205

[5]

Chinta S , Rengaswamy R . Machine learning derived quantitative structure property relationship (QSPR) to predict drug solubility in binary solvent systems. Industrial & Engineering Chemistry Research, 2019, 58(8): 3082–3092

[6]

Yalamanchi K K , Pal P , Mohan B . et al. A Variational Autoencoder model toward molecular structure representation learning of fuels. Journal of Energy Resources Technology, Part A: Sustainable and Renewable Energy, 2025, 1(5): 052301

[7]

Üstün C E , Da Silva Machado De Freitas R , Okafor E C . et al. Machine learning applications for predicting fuel ignition and flame properties: Current status and future perspectives. Energy & Fuels, 2025, 39(28): 13281–13314

[8]

Han X , Jia M , Chang Y C . et al. Directed message passing neural network (D-MPNN) with graph edge attention (GEA) for property prediction of biofuel-relevant species. Energy and AI, 2022, 10: 100201

[9]

Scarselli F , Gori M , Tsoi A C . et al. The graph neural network model. IEEE Transactions on Neural Networks, 2009, 20(1): 61–80

[10]

Schweidtmann A M , Rittig J G , Weber J M . et al. Physical pooling functions in graph neural networks for molecular property prediction. Computers & Chemical Engineering, 2023, 172: 108202

[11]

Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need. In: Proceedings of the 31st International Conference on Neural Information Processing Systems, Long Beach, USA, 2017, 6000–6010

[12]

Gao S H , Fang A D , Huang Y P . et al. Empowering biomedical discovery with AI agents. Cell, 2024, 187(22): 6125–6151

[13]

Jablonka K M , Schwaller P , Ortega-Guerrero A . et al. Leveraging large language models for predictive chemistry. Nature Machine Intelligence, 2024, 6(2): 161–169

[14]

Yüksel A , Ulusoy E , Ünlü A . et al. SELFormer: Molecular representation learning via SELFIES language models. Machine Learning: Science and Technology, 2023, 4(2): 025035

[15]

Ahmad W , Simon E , Chithrananda S . et al. ChemBERTa-2: Towards chemical foundation models. arXiv preprint. arXiv: 2209.01712, 2022,

[16]

Devlin J , Chang M W , Lee K . et al. BERT: Pre-training of deep bidirectional transformers for language understanding. In: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Minneapolis, Minnesota, 2019, 4171–4186

[17]

Yang M X , Song G L , Cheng L H . et al. SMILES token additivity model with interpretability and generalizability for fuel property predictions. Journal of Chemical Information and Modeling, 2025, 65(15): 8022–8032

[18]

Yu B T, Baker F N, Chen Z Q, et al. LlaSMol: Advancing large language models for chemistry with a large-scale, comprehensive, high-quality instruction tuning dataset. In: 1st Conference on Language Modeling, Philadelphia, USA, 2024

[19]

Edwards C, Lai T, Ros K, et al. Translation between molecules and natural language. In: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Abu Dhabi, United Arab Emirates, 2022, 375–413

[20]

Ramos M C , Collison C J , White A D . A review of large language models and autonomous agents in chemistry. Chemical Science, 2025, 16(6): 2514–2572

[21]

Castro Nascimento C M , Pimentel A S . Do large language models understand chemistry? A conversation with ChatGPT. Journal of Chemical Information and Modeling, 2023, 63(6): 1649–1655

[22]

Dai D M, Sun Y T, Dong L, et al. Why can GPT learn in-context? Language models secretly perform gradient descent as meta-optimizers. In: Proceedings of the Findings of the Association for Computational Linguistics, Toronto, Canada, 2023, 4005–34019

[23]

Ramos M C , Michtavy S S , White A D . et al. Bayesian optimization of catalysis with in-context learning. ACS Central Science, 2026, 12(5): 599–615

[24]

Liu Y Y , Ding S R , Zhou S . et al. MolecularGPT: Open large language model (LLM) for few-shot molecular property prediction. arXiv preprint. arXiv: 2406.12950, 2024,

[25]

Weininger D . SMILES, a chemical language and information system. 1. Introduction to methodology and encoding rules. Journal of Chemical Information and Computer Sciences, 1988, 28(1): 31–36

[26]

Kim S , Chen J , Cheng T J . et al. PubChem 2025 update. Nucleic Acids Research, 2025, 53(D1): D1516–D1525

[27]

Gaulton A , Bellis L J , Bento A P . et al. ChEMBL: A large-scale bioactivity database for drug discovery. Nucleic Acids Research, 2012, 40(D1): D1100–D1107

[28]

Durant J L , Leland B A , Henry D R . et al. Reoptimization of MDL keys for use in drug discovery. Journal of Chemical Information and Computer Sciences, 2002, 42(6): 1273–1280

[29]

Xian Z T, Gu J W, Li L B, et al. MolRAG: Unlocking the power of large language models for molecular property prediction. In: Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, Vienna, Austria, 2025, 15513–15531

[30]

Taori R , Gulrajani I , Zhang T Y . et al. Alpaca: A strong, replicable instruction-following model. Stanford Center for Research on Foundation Models. Available online, ,

[31]

Rajbhandari S, Rasley J, Ruwase O, et al. ZeRO: Memory optimizations toward training trillion parameter models. In: Proceedings of the SC20: International Conference for High Performance Computing, Networking, Storage and Analysis, Atlanta, USA, 2020, 1–16

[32]

Tanimoto T T . An Elementary Mathematical Theory of Classification and Prediction. IBM Internal Report, 1958,

[33]

Malkov Y A , Yashunin D A . Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020, 42(4): 824–836

[34]

Das D D , John P C S , McEnally C S . et al. Measuring and predicting sooting tendencies of oxygenates, alkanes, alkenes, cycloalkanes, and aromatics on a unified scale. Combustion and Flame, 2018, 190: 349–364

[35]

Schweidtmann A M , Rittig J G , König A . et al. Graph neural networks for prediction of fuel ignition quality. Energy & Fuels, 2020, 34(9): 11395–11407

[36]

Kim Y , Cho J , Naser N . et al. Physics-informed graph neural networks for predicting cetane number with systematic data quality analysis. Proceedings of the Combustion Institute, 2023, 39(4): 4969–4978

[37]

Freitas R S M , Jiang X . Descriptors-based machine-learning prediction of cetane number using quantitative structure–property relationship. Energy and AI, 2024, 17: 100385

[38]

Wu Z Q , Ramsundar B , Feinberg E N . et al. MoleculeNet: A benchmark for molecular machine learning. Chemical Science, 2018, 9(2): 513–530

[39]

Chen T Q , Guestrin C . XGBoost: A scalable tree boosting system. In: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, USA, 2016, 785–794

[40]

Rogers D , Hahn M . Extended-connectivity fingerprints. Journal of Chemical Information and Modeling, 2010, 50(5): 742–754

[41]

Fang Y, Liang X Z, Zhang N Y, et al. Mol-instructions: A large-scale biomolecular instruction dataset for large language models. In: Proceedings of the 12th International Conference on Learning Representations, Vienna, Austria, 2024

[42]

Taylor R , Kardas M , Cucurull G . et al. Galactica: A large language model for science. arXiv preprint: arXiv: 2211.09085, 2022,

[43]

Su B , Du D Z , Yang Z . et al. A molecular multimodal foundation model associating molecule graphs with natural language. arXiv preprint: arXiv: 2209.05481, 2022,

[44]

Rollins Z A , Cheng A C , Metwally E . MolPROP: Molecular property prediction with multimodal language and graph fusion. Journal of Cheminformatics, 2024, 16(1): 56

[45]

Zhou G M, Gao Z F, Ding Q K, et al. Uni-Mol: A universal 3D molecular representation learning framework. In: the11th International Conference on Learning Representationss, Kigali, Rwanda, 2023

[46]

Heid E , Greenman K P , Chung Y . et al. Chemprop: A machine learning package for chemical property prediction. Journal of Chemical Information and Modeling, 2024, 64(1): 9–17

Rights & permissions

Higher Education Press

PDF (3480KB)

10

Accesses

0

Citation

Detail

Sections
Recommended

/