LineShine: online acceleration for HPC-AI converged scientific computing

Yutong LU , Guangnan FENG , Haohuan FU

Front. Comput. Sci. ›› 2026, Vol. 20 ›› Issue (10) : 2010111

PDF (5493KB)
Front. Comput. Sci. ›› 2026, Vol. 20 ›› Issue (10) :2010111 DOI: 10.1007/s11704-026-61591-w
Architecture
RESEARCH ARTICLE
LineShine: online acceleration for HPC-AI converged scientific computing
Author information +
History +
PDF (5493KB)

Abstract

Modern scientific workflows increasingly interleave high-precision numerical simulation, low-precision AI computation, irregular computation, complex communication, and data-intensive processing within continuous execution pipelines. When tightly coupled stages repeatedly cross separate host and accelerator domains, data transfers, explicit synchronization, and distinct software paths can introduce substantial coordination overhead. We present LineShine, a newly deployed exascale supercomputer whose online acceleration architecture provides a new design point for converged HPC and AI. LineShine adopts a many-core online acceleration architecture in which matrix engines are embedded within CPU cores and share a unified execution environment, memory hierarchy, and software stack with scalar and vector execution. This capability is extended to system scale through balanced full-stack co-design, including HBM–DDR hierarchical memory, the Lingqi interconnect, symmetric distributed tiered storage, and an architecture-aware software stack. LineShine achieves leading results on major HPC benchmarks. In the June 2026 rankings, it sustains 2.1984 EFlops on HPL with 80.37% computation efficiency and 52.07 GFlops/W system energy efficiency, and delivers 22.0049 PFlops on HPCG, ranking first on both benchmarks. Beyond benchmark performance, LineShine has supported numerous full-system-scale production applications across a broad range of scientific domains, including computational fluid dynamics, neuroscience, biomedicine, and Earth observation. These applications demonstrate that its balanced architecture and software stack can translate system capability into high throughput, robust scalability, and sustained scientific productivity at exascale.

Graphical abstract

Keywords

LineShine Supercomputer / exascale supercomputing / online acceleration / hardware and software co-design / HPC-AI converged system

Cite this article

Download citation ▾
Yutong LU, Guangnan FENG, Haohuan FU. LineShine: online acceleration for HPC-AI converged scientific computing. Front. Comput. Sci., 2026, 20 (10) : 2010111 DOI:10.1007/s11704-026-61591-w

登录浏览全文

4963

注册一个新账户 忘记密码

References

[1]

Deelman E, Dongarra J, Hendrickson B, Randles A, Reed D, Seidel E, Yelick K . High-performance computing at a crossroads. Science, 2025, 387( 6736): 829–831

[2]

Kochkov D, Yuval J, Langmore I, Norgaard P, Smith J, Mooers G, Klöwer M, Lottes J, Rasp S, Düben P, Hatfield S, Battaglia P, Sanchez-Gonzalez A, Willson M, Brenner M P, Hoyer S . Neural general circulation models for weather and climate. Nature, 2024, 632( 8027): 1060–1066

[3]

Bi K, Xie L, Zhang H, Chen X, Gu X, Tian Q . Accurate medium-range global weather forecasting with 3D neural networks. Nature, 2023, 619( 7970): 533–538

[4]

Das S, Kanungo B, Subramanian V, Panigrahi G, Motamarri P, Rogers D, Zimmerman P, Gavini V. Large-scale materials modeling at quantum accuracy: ab initio simulations of quasicrystals and interacting extended defects in metallic alloys. In: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, SC ’23. 2023, 1−12

[5]

Roccon A, Amati G, Brandt L, Calhoun D, Costa P, Lu W, Pirozzoli S, Richter D, Umair M, You D, Zahtila T, Marchioli C . GPU-accelerated simulations of turbulence: review of current applications and future perspectives. Physical Review Fluids, 2026, 11( 3): 034905

[6]

Merzari E, Hamilton S, Evans T, Min M, Fischer P, Kerkemeier S, Fang J, Romano P, Lan Y H, Phillips M, Biondo E, Royston K, Warburton T, Chalmers N, Rathnayake T. Exascale multiphysics nuclear reactor simulations for advanced designs. In: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, SC ’23. 2023, 3

[7]

Kuriyama R, Akira K, Green L, Herrera B, Dai K, Iura M, Gouaillardet G, Terasawa A, Kobayashi T, Igarashi J, Arkhipov A, Yamazaki T. Microscopic-level mouse whole cortex simulation composed of 9 million biophysical neurons and 26 billion synapses on the supercomputer Fugaku. In: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, SC ’25. 2025, 2158−2171

[8]

Wilfong B, Radhakrishnan A, Le Berre H, Vickers D, Prathi T, Tselepidis N, Dorschner B, Budiardja R, Cornille B, Abbott S, Schäfer F, Bryngelson S. Simulating many-engine spacecraft: exceeding 1 quadrillion degrees of freedom via information geometric regularization. In: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, SC ’25. 2025, 14−24

[9]

Alsaadi A, Hategan-Marandiuc M, Maheshwari K, Merzky A, Titov M, Turilli M, Wilke A, Wozniak J M, Chard K, da Silva R F, Jha S, Laney D . Exascale workflow applications and middleware: an ExaWorks retrospective. The International Journal of High Performance Computing Applications, 2025, 39( 4): 579–593

[10]

Atchley S, Badia R M, de Supinski B R, Fryman J, Kranzlmüller D, Manne S, Manninen P, Matsuoka S, Milojicic D, Shipman G, Van Hensbergen E, Wisniewski R W . Predicting the future of supercomputing. Computer, 2025, 58( 7): 110–120

[11]

Lu Y. “Tackling Fragmentation in Exascale Supercomputing and Beyond,” keynote presented at ISC High Performance 2025, Berlin, Germany, Jun. 12, 2025. See isc.app.swapcard.com/event/isc-high-performance-2025 website

[12]

Salpekar O, Varma R, Yu K, Ivanov V, Wang Y, Sharif A, Si M, Xu S, Tian F, Zheng S, Rice T, Garg A, Peng S, Siravara S, Fu W, Castro R D, Gangidi A, Obraztsov A, Narang S, Edunov S, Naumov M, Tang C, Oldham M. Training LLMs with fault tolerant HSDP on 100,000 GPUs. 2026, arXiv preprint arXiv: 2602.00277

[13]

xAI. Grok 4. See Xai/news/grok-4 website, 2025

[14]

Solera-Rico A, Sanmiguel Vila C, Gómez-López M, Wang Y, Almashjary A, Dawson S T M, Vinuesa R . β-Variational autoencoders and transformers for reduced-order modelling of fluid flows. Nature Communications, 2024, 15( 1): 1361

[15]

Lam R, Sanchez-Gonzalez A, Willson M, Wirnsberger P, Fortunato M, Alet F, Ravuri S, Ewalds T, Eaton-Rosen Z, Hu W, Merose A, Hoyer S, Holland G, Vinyals O, Stott J, Pritzel A, Mohamed S, Battaglia P . Learning skillful medium-range global weather forecasting. Science, 2023, 382( 6677): 1416–1421

[16]

Li Y, Han W, Li H, Duan W, Chen L, Zhong X, Wang J, Liu Y, Sun X . FuXi-En4DVar: an assimilation system based on machine learning weather forecasting model ensuring physical constraints. Geophysical Research Letters, 2024, 51( 22): e2024GL111136

[17]

Raissi M, Perdikaris P, Karniadakis G E . Physics-informed neural networks: a deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations. Journal of Computational Physics, 2019, 378: 686–707

[18]

Boiko D A, MacKnight R, Kline B, Gomes G . Autonomous chemical research with large language models. Nature, 2023, 624( 7992): 570–578

[19]

Lu Y, Huang D, Chen P . HPC-AI coupling methodology for scientific applications. The International Journal of High Performance Computing Applications, 2026, 40( 3): 334–351

[20]

Goto K, Ibeid H, Kumaran K, Muralidharan S, Nguyen A T, Nishtala A. Sustaining exascale performance: lessons from HPL and HPL-MxP on aurora. 2026, arXiv preprint arXiv: 2604.09517

[21]

Kashi A, Koukpaizan N, Lu H, Matheson M, Oral S, Wang F. Scaling the memory wall using mixed-precision - HPG-MxP on an exascale machine. In: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, SC ’25. 2025, 281−297

[22]

Arm Limited. Arm architecture reference manual for a-profile architecture. Manual DDI 0487. Cambridge: Arm Limited, 2026

[23]

Dongarra J J, Luszczek P, Petitet A . The LINPACK benchmark: past, present and future. Concurrency and Computation: Practice and Experience, 2003, 15( 9): 803–820

[24]

TOP500. June 2026. See Top500.org/lists/top500/2026/06/ website, 2026

[25]

Dongarra J, Heroux M A, Luszczek P . High-performance conjugate-gradient benchmark. International Journal of High Performance Computing Applications, 2016, 30( 1): 3–10

[26]

TOP500. HPCG - June 2026. See Top500.org/lists/hpcg/2026/06/ website, 2026

[27]

Dongarra J, Luszczek P . HPL-MxP benchmark: mixed-precision algorithms, iterative refinement, and scalable data generation. The International Journal of High Performance Computing Applications, 2026, 40( 1): 52–62

[28]

Huang H, Xie J, Feng G, Zhang X, Huang D, Chen Z, Lu Y. HStencil: matrix-vector stencil computation with interleaved outer product and MLA. In: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, SC ’25. 2025, 1816−1829

[29]

Huang L, Huang H, Chen Z, Lu Y. KirbyMM: outer-product based matrix multiplication on ARMv9 processor. In: Proceedings of 2026 Design, Automation & Test in Europe Conference (DATE). 2026, 1−7

[30]

Jiang J, Yao X, Chen J, Wei J, Huang D, Lu Y. ASM-SpMM: unleashing the potential of arm SME for sparse matrix multiplication acceleration. In: Proceedings of the 31st ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming, PPoPP ’26. 2026, 232−244

[31]

Feng G, Xie J, Dong D, Lu Y. UNR: unified notifiable RMA library for HPC. In: Proceedings of the International Conference for High Performance Computing, Networking, Storage, and Analysis, SC ’24. 2024, 1−15

[32]

Yang X, Li S, Yuan F, Dong D. DBSR: an efficient storage format for vectorizing sparse triangular solvers on structured grids. In: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. 2024, 59

[33]

Smits A J, McKeon B J, Marusic I . High–Reynolds number wall turbulence. Annual Review of Fluid Mechanics, 2011, 43( 1): 353–375

[34]

Chen X, Sreenivasan K R . Law of bounded dissipation and its consequences in turbulent wall flows. Journal of Fluid Mechanics, 2022, 933: A20

[35]

Lee M, Malaya N, Moser R D. Petascale direct numerical simulation of turbulent channel flow on up to 786K cores. In: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis, SC ’13. 2013, 1−11

[36]

Lee M, Moser R D . Direct numerical simulation of turbulent channel flow up to Reτ≈5200. Journal of Fluid Mechanics, 2015, 774: 395–415

[37]

Xie J, He J, Bao Y, Chen X . A low-communication-overhead parallel DNS method for the 3D incompressible wall turbulence. International Journal of Computational Fluid Dynamics, 2021, 35( 6): 413–432

[38]

Xie J, Feng G, Huang H, Feng J, Chen Z, Lu Y. Extreme-scale direct numerical simulation of incompressible turbulence on the heterogeneous many-core system. In: Proceedings of the 29th ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming, PPoPP ’24. 2024, 120−132

[39]

Tsuji Y. Asymptotic behavior theory of turbulent dissipation rate for high-Reynolds number limits in wall-turbulence. HPCI User Report hp230138. Tokai-mura: Research Organization for Information Science and Technology, 2024

[40]

Fan X, Markram H . A brief history of simulation neuroscience. Frontiers in Neuroinformatics, 2019, 13: 32

[41]

Lu W, Du X, Wang J, Zeng L, Ye L, Xiang S, Zheng Q, Zhang J, Xu N, Feng J . Simulation and assimilation of the digital human brain. Nature Computational Science, 2024, 4( 12): 890–898

[42]

Kunkel S, Schmidt M, Eppler J M, Plesser H E, Masumoto G, Igarashi J, Ishii S, Fukai T, Morrison A, Diesmann M, Helias M . Spiking network simulation code for petascale computers. Frontiers in Neuroinformatics, 2014, 8: 78

[43]

Morrison A, Mehring C, Geisel T, Aertsen A, Diesmann M . Advancing the boundaries of high-connectivity network simulation with distributed computing. Neural Computation, 2005, 17( 8): 1776–1801

[44]

Knight J C, Nowotny T . Larger GPU-accelerated brain simulations with procedural connectivity. Nature Computational Science, 2021, 1( 2): 136–142

[45]

Sadybekov A A, Sadybekov A V, Liu Y, Iliopoulos-Tsoutsouvas C, Huang X P, Pickett J, Houser B, Patel N, Tran N K, Tong F, Zvonok N, Jain M K, Savych O, Radchenko D S, Nikas S P, Petasis N A, Moroz Y S, Roth B L, Makriyannis A, Katritch V . Synthon-based ligand discovery in virtual libraries of over 11 billion compounds. Nature, 2022, 601( 7893): 452–459

[46]

Gao W, Coley C W . The synthesizability of molecules proposed by generative models. Journal of Chemical Information and Modeling, 2020, 60( 12): 5714–5723

[47]

Lyu J, Wang S, Balius T E, Singh I, Levit A, Moroz Y S, O’Meara M J, Che T, Algaa E, Tolmachova K, Tolmachev A A, Shoichet B K, Roth B L, Irwin J J . Ultra-large library docking for discovering new chemotypes. Nature, 2019, 566( 7743): 224–229

[48]

Mobley D L, Gilson M K . Predicting binding free energies: frontiers and benchmarks. Annual Review of Biophysics, 2017, 46( 1): 531–558

[49]

Shao Z, Wang P, Zhu Q, Xu R, Song J, Bi X, Zhang H, Zhang M, Li Y K, Wu Y, Guo D. DeepSeekMath: pushing the limits of mathematical reasoning in open language models. 2024, arXiv preprint arXiv: 2402.03300

[50]

Tran-Nguyen V K, Jacquemard C, Rognan D . LIT-PCBA: an unbiased data set for machine learning and virtual screening. Journal of Chemical Information and Modeling, 2020, 60( 9): 4263–4273

[51]

Earth Science Data Systems NASA. ESDS data metrics. See Earthdata.nasa.gov/about/data-metrics website, 2026

[52]

Chi M, Plaza A, Benediktsson J A, Sun Z, Shen J, Zhu Y . Big data for remote sensing: challenges and opportunities. Proceedings of the IEEE, 2016, 104( 11): 2207–2219

[53]

Lipman Y, Chen R T Q, Ben-Hamu H, Nickel M, Le M. Flow matching for generative modeling. 2023, arXiv preprint arXiv: 2210.02747

[54]

Zhang J, Dong R, Wu X, Huang X, Cheng S, Yang Y, Zhou Z, Xu Y, Luo Z, Yang M, Wei F, Chen M, You Y, Zheng J, Li W, Lu Y, Fu H. Transforming the use of earth observation data: exascale training of a generative compression model with historical priors for up to 10,000x data reduction. 2026, arXiv preprint arXiv: 2605.08633

Rights & permissions

Higher Education Press

PDF (5493KB)

Supplementary files

Highlights

920

Accesses

0

Citation

Detail

Sections
Recommended

/

〈 〉