FMM-Agent: Evolving feature meta-models for industrial imbalanced scenarios via LLMs

Yu ZHOU , Guanghua LYU , Hainan GUO , Sam KWONG , Qingfu ZHANG

Eng. Manag ››

PDF (5299KB)
Eng. Manag ›› DOI: 10.1007/s42524-026-6014-5
RESEARCH ARTICLE
FMM-Agent: Evolving feature meta-models for industrial imbalanced scenarios via LLMs
Author information +
History +
PDF (5299KB)

Abstract

Industrial classification tasks often face challenges such as class imbalance, noise and non-stationary data distributions. Most feature engineering methods aided by evolutionary algorithms and large language models (LLMs) of-ten rely on the predictive performance of downstream classification metrics, while neglecting the feature distribution structure and the relationship between features and labels under distribution shifts. To address these issues, we propose the Feature Meta-Model Agent (FMM-Agent), a framework that evolves features within a meta-model space defined by statistical information, rather than operating directly on raw data. FMM-Agent enables LLMs to perform operator restructure and refine the chain-of-thought to obtain better feature shaping. We further introduce a unified scoring mechanism to jointly evaluate label relevance and distribution stability, allowing the feature pool to gradually move toward better feature distribution shapes. Experiments conducted on 11 data sets show that FMM-Agent consistently improves the recognition ability of minority classes in terms of balanced accuracy, g-mean, and recall, and its performance is superior to other comparative methods. Ablation studies confirm the necessity of evolutionary restructure and generation mechanism strategies. In addition, experimental results with different LLMs show that although stronger models can produce more stable evolutionary process, the overall performance improvement of FMM-Agent does not depend on a specific model. It is worth noting that although FMM-Agent incurs additional inference time costs due to evolutionary feature generation, it achieves a good balance be-tween computational overhead and performance improvement.

Graphical abstract

Keywords

agent / evolutionary search / feature meta-model / feature engineering / Large Language Models

Cite this article

Download citation ▾
Yu ZHOU, Guanghua LYU, Hainan GUO, Sam KWONG, Qingfu ZHANG. FMM-Agent: Evolving feature meta-models for industrial imbalanced scenarios via LLMs. Eng. Manag DOI:10.1007/s42524-026-6014-5

登录浏览全文

4963

注册一个新账户 忘记密码

References

[1]

Abhyankar NShojaee PReddy C K (2025). LLM-FE: Automated feature engineering for tabular data with LLMs as evolutionary optimizers. Preprint at arXiv. arXiv:2503.14434

[2]

Chawla N V, Bowyer K W, Hall L O, Kegelmeyer W P, (2002). SMOTE: Synthetic minority over-sampling technique. Journal of Artificial Intelligence Research, 16: 321–357

[3]

Cheng H, Yi J, Xia W, Pu H, Luo J, (2024). Adaptive memetic algorithm with dual-level local search for cooperative route planning of multi-robot surveillance systems. Complex System Modeling and Simulation, 4( 2): 210–221

[4]

Dablain D, Krawczyk B, Chawla N V, (2023). DeepSMOTE: Fusing deep learning and SMOTE for imbalanced data. IEEE Transactions on Neural Networks and Learning Systems, 34( 9): 6390–6404

[5]

Das I, (1999). On characterizing the “knee” of the Pareto curve based on normal-boundary intersection. Structural Optimization, 18( 2): 107–115

[6]

Ditzler G, Roveri M, Alippi C, Polikar R, (2015). Learning in nonstationary environments: A survey. IEEE Computational Intelligence Magazine, 10( 4): 12–25

[7]

Elkan C (2001). The foundations of cost-sensitive learning. In: Proceedings of the 17th International Joint Conference on Artificial Intelligence. Seattle: Morgan Kaufmann, 973–978

[8]

Guo SDeng CWen YChen HChang YWang J (2024). DS-Agent: Automated data science by empowering large language models with case-based reasoning. Preprint at arXiv. arXiv:2402.17453

[9]

Hollmann N, Müller S, Eggensperger K, Hutter F, (2022). .

[10]

Hollmann N, Müller S, Hutter F, (2023). Large language models for automated data science: Introducing CAAFE for context-aware automated feature engineering. In: Advances in Neural Information Processing Systems, 36: 44753–44775

[11]

Hong S, Lin Y, Liu B, Liu B, Wu B, Zhang C, Wu C (2025). Data interpreter: An LLM agent for data science. In: Findings of the Association for Computational Linguistics: ACL 2025. Stroudsburg: Association for Computational Linguistics, 19796–19821

[12]

Jiang M, Wang Z, Hong H, Yen G G, (2021). Knee point-based imbalanced transfer learning for dynamic multiobjective optimization. IEEE Transactions on Evolutionary Computation, 25( 1): 117–129

[13]

Kanter J M, Veeramachaneni K (2015). Deep feature synthesis: Towards automating data science endeavors. In: Proceedings of the IEEE International Conference on Data Science and Advanced Analytics. Paris: IEEE, 1–10

[14]

Katz G, Shin E C R, Song D (2016). ExploreKit: Automatic feature generation and selection. In: Proceedings of the IEEE International Conference on Data Mining. Barcelona: IEEE, 979–984

[15]

Ko J, Park G, Lee D, Lee K (2025). FeRG-LLM: Feature engineering by reason generation large language models. In: Findings of the Association for Computational Linguistics: NAACL 2025. Albuquerque: Association for Computational Linguistics, 4211–4228

[16]

Khan S H, Hayat M, Bennamoun M, Sohel F A, Togneri R, (2018). Cost-sensitive learning of deep feature representations from imbalanced data. IEEE Transactions on Neural Networks and Learning Systems, 29( 8): 3573–3587

[17]

Khurana UNargesian FSamulowitz HKhalil ETuraga D (2016). Automating feature engineering. In: NIPS Workshop on Artificial Intelligence for Data Science

[18]

Krawczyk B, (2016). Learning from imbalanced data: Open challenges and future directions. Progress in Artificial Intelligence, 5( 4): 221–232

[19]

Li W, Ye X, Huang Y, Mahmoodi S, (2022). Adaptive dimensional learning with a tolerance framework for the differential evolution algorithm. Complex System Modeling and Simulation, 2( 1): 59–77

[20]

Lin JGuo YHan YHu SNi ZWang LWang H (2025). SE-Agent: Self-evolution trajectory optimization in multi-step reasoning with LLM-based agents. Preprint at arXiv. arXiv:2508.02085

[21]

Liu FTong XYuan MLin XLuo FWang ZZhang Q (2024). Evolution of heuristics: Towards efficient automatic algorithm design using large language models. Preprint at arXiv. arXiv:2401.02051

[22]

Liu X Y, Wu J, Zhou Z H, (2008). Exploratory undersampling for class-imbalance learning. IEEE Transactions on Systems, Man, and Cybernetics. Part B, Cybernetics, 39( 2): 539–550

[23]

Ma Y J, Liang W, Wang G, Huang D A, Bastani O, Jayaraman D, Anandkumar A, (2023). Eureka: Human-level reward design via coding large language models. In: International conference on learning. Representations. 2024: 26516–26560

[24]

Maldonado S, Weber R, Famili F, (2014). Feature selection for high-dimensional class-imbalanced data sets using support vector machines. Information Sciences, 286: 228–246

[25]

Mundhenk T NLandajuela MGlatt RSantiago C PFaissol D MPetersen B K (2021). Symbolic regression via neural-guided genetic programming population seeding. Preprint at arXiv. arXiv:2111.00053

[26]

Nejjar IAhmed FFink O (2024). IM-Context: In-context learning for imbalanced regression tasks. Preprint at arXiv. arXiv:2405.18202

[27]

Novikov AVu NEisenberger MDupont EHuang P SWagner A ZShirobokov SKozlovskii BRuiz F J RMehrabian AKumar M PSee AChaudhuri SHolland GDavies ANowozin SKohli PBalog M (2025). AlphaEvolve: A coding agent for scientific and algorithmic discovery. Preprint at arXiv. arXiv:2506.13131

[28]

Romera-Paredes B, Barekatain M, Novikov A, Balog M, Kumar M P, Dupont E, Ruiz F J R, Ellenberg J S, Wang P, Fawzi O, Kohli P, Fawzi A, (2024). Mathematical discoveries from program search with large language models. Nature, 625( 7995): 468–475

[29]

Saadatmand H, Akbarzadeh-T M R, (2024). Many-objective Jaccard-based evolutionary feature selection for high-dimensional imbalanced data classification. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46( 12): 8820–8835

[30]

Shen Y, Song K, Tan X, Li D, Lu W, Zhuang Y, (2023). HuggingGPT: Solving AI tasks with ChatGPT and its friends in Hugging Face. In: Advances in Neural Information Processing Systems, 36: 38154–38180

[31]

Shu Z, Song A, Wu G, Pedrycz W, (2023). Variable reduction strategy integrated variable neighborhood search and NSGA-II hybrid algorithm for emergency material scheduling. Complex System Modeling and Simulation, 3( 2): 83–101

[32]

Tan M, Zhang Z, Ren Y, Richard I, Zhang Y, (2023). Multi-agent system for electric vehicle charging scheduling in parking lots. Complex System Modeling and Simulation, 3( 2): 129–142

[33]

Wang L, Pan Z, Wang J, (2021). A review of reinforcement learning-based intelligent optimization for manufacturing scheduling. Complex System Modeling and Simulation, 1( 4): 257–270

[34]

Wei J, Wang X, Schuurmans D, Bosma M, Xia F, Chi E, Zhou D, (2022). Chain-of-thought prompting elicits reasoning in large language models. In: Advances in Neural Information Processing Systems, 35: 24824–24837

[35]

Xue B, Zhang M, Browne W N, Yao X, (2016). A survey on evolutionary computation approaches to feature selection. IEEE Transactions on Evolutionary Computation, 20( 4): 606–626

[36]

Yang HYue SHe Y (2023). Auto-GPT for online decision making: Benchmarks and additional opinions. Preprint at arXiv. arXiv:2306.02224

[37]

Yang Q T, Xu X X, Zhan Z H, Zhong J, Kwong S, Zhang J, (2025). Evolutionary multitask optimization for multiform feature selection in classification. IEEE Transactions on Cybernetics, 55( 4): 1673–1686

[38]

Zhang Q, Xu C, Li J, Sun Y, Bao J, Zhang D, (2025). LLM-TSFD: An industrial time series human-in-the-loop fault diagnosis method based on a large language model. Expert Systems with Applications, 264: 125861

[39]

Zhan Z H, Li J Y, Kwong S, Zhang J, (2023). Learning-aided evolution for optimization. IEEE Transactions on Evolutionary Computation, 27( 6): 1794–1808

[40]

Zhou Y, Gao L, Wang D, Wu W, Zhou Z, Ye T, (2023). Imbalanced multifault diagnosis via improved localized feature selection. IEEE Transactions on Instrumentation and Measurement, 72: 1–11

[41]

Zhou Y, Zhang W, Kang J, Zhang X, Wang X, (2021). A problem-specific non-dominated sorting genetic algorithm for supervised feature selection. Information Sciences, 547: 841–859

[42]

Zhou Y, Zhang X, Kwong S (2025). Computational Intelligence for High-Dimensional Machine Learning: A Feature Selection Perspective and Its Real-World Applications. Cham: Springer Nature

RIGHTS & PERMISSIONS

Higher Education Press

PDF (5299KB)

Supplementary files

Supplementary materials

153

Accesses

0

Citation

Detail

Sections
Recommended

/