Benchmarking large language models for binary function name prediction
Yunxiang GE , Bing XIA , Yong DING , Wenbo LIU
Eng Inform Technol Electron Eng ›› 2026, Vol. 27 ›› Issue (8) : 260017
In practical reverse engineering, stripped and optimized binaries lack high-level semantics, hindering automated function understanding. This paper adopts binary function name prediction as a benchmark to evaluate open-source large language models (LLMs) for function-level semantic inference under realistic conditions. We systematically examine key factors affecting performance, including the pretraining domain, model scale, architecture and optimization settings, prompting, and contextual information. Experiments on a large-scale multi-project binary dataset reveal clear limitations in recovering function semantics from stripped binaries. Code-oriented LLMs consistently outperform general-purpose models, while increasing the model size alone does not yield monotonic gains. We further show that realistically recoverable contextual signals, especially structurally recoverable cross-function contexts, substantially mitigate semantic sparsity and improve prediction quality. These results delineate current capability boundaries of LLM-based binary semantic understanding and suggest future directions in contextual modeling and domain-adaptive techniques for practical reverse engineering.
Binary function name prediction / Large language models (LLMs) / Reverse engineering (RE) / Binary code analysis
| [1] |
|
| [2] |
|
| [3] |
|
| [4] |
|
| [5] |
|
| [6] |
|
| [7] |
|
| [8] |
|
| [9] |
|
| [10] |
|
| [11] |
|
| [12] |
|
| [13] |
|
| [14] |
|
| [15] |
|
| [16] |
|
| [17] |
|
| [18] |
|
| [19] |
|
| [20] |
|
| [21] |
|
| [22] |
|
| [23] |
|
| [24] |
|
| [25] |
|
| [26] |
|
| [27] |
|
| [28] |
|
| [29] |
|
| [30] |
|
| [31] |
|
| [32] |
|
| [33] |
|
| [34] |
|
| [35] |
|
| [36] |
|
| [37] |
|
| [38] |
|
| [39] |
|
The Authors. Published by Zhejiang University Press Co., Ltd.
/
| 〈 |
|
〉 |