Area-efficient layer normalization via hybrid stochastic-binary computing

Xiaoyu ZHANG , Liang WANG , Limin XIAO

Front. Comput. Sci. ››

PDF (695KB)
Front. Comput. Sci. ›› DOI: 10.1007/s11704-026-61022-w
RESEARCH ARTICLE
Area-efficient layer normalization via hybrid stochastic-binary computing
Author information +
History +
PDF (695KB)

Abstract

Layer normalization is an essential component of Transformer networks, yet its hardware implementation poses significant challenges in terms of area and power consumption owing to the non-linear square-root and reciprocal operations it entails. This paper presents a hybrid stochastic- binary computing architecture for layer normalization acceleration, in which variance estimation is performed in the stochastic computing domain using AND-gate squarers in place of the conventional parallel multiplier array, while mean computation and final normalization are retained in the binary domain. The proposed architecture integrates several key stochastic computing techniques, including bitstream decorrelation via paired linear feedback shift registers employing distinct primitive polynomials, stochastic-to-binary conversion through an accumulative parallel counter, signal-to-noise ratio enhancement via dynamic prescaling, and high-precision inverse square-root computation through time-multiplexed Newton-Raphson iteration on a single shared multiplier. Synthesized in SMIC 55 nm technology, the proposed design achieves approximately 86% reduction in cell area, over 80% reduction in dynamic power, and nearly 90% reduction in leakage power compared with two baseline accelerators based on piecewise linear approximation. Across a wide range of input distributions, and at a fixed stochastic bitstream length of 32 samples per element, the proposed design achieves a 3.7–7.5× improvement in mean squared error over the piecewise linear baselines, attributable to the substantially higher effective precision of Newton-Raphson iteration relative to cascaded piecewise linear approximation. The composite power-mean-squared-error figure of merit demonstrates that the proposed design outperforms the baselines across all tested ranges. These area and power gains are obtained at the expense of throughput and energy efficiency: throughput-, energy-, and energy-delay-product-normalized comparisons show that the design trades time for silicon area and peak power. These results indicate that selectively applying stochastic computing to the computational bottleneck of layer normalization constitutes an effective design strategy for area- and peak-power-constrained, low-throughput edge artificial intelligence accelerators.

Keywords

layer normalization / stochastic computing / Transformer accelerator / fixed-point arithmetic / Newton-Raphson / ASIC synthesis / edge AI

Cite this article

Download citation ▾
Xiaoyu ZHANG, Liang WANG, Limin XIAO. Area-efficient layer normalization via hybrid stochastic-binary computing. Front. Comput. Sci. DOI:10.1007/s11704-026-61022-w

登录浏览全文

4963

注册一个新账户 忘记密码

References

RIGHTS & PERMISSIONS

Higher Education Press 2026

PDF (695KB)

0

Accesses

0

Citation

Detail

Sections
Recommended

/