About the journal
Browse
Collections
Multimedia collections
Authors & reviewers
Area-efficient layer normalization via hybrid stochastic-binary computing
Xiaoyu ZHANG , Liang WANG , Limin XIAO
Layer normalization is an essential component of Transformer networks, yet its hardware implementation poses significant challenges in terms of area and power consumption owing to the non-linear square-root and reciprocal operations it entails. This paper presents a hybrid stochastic- binary computing architecture for layer normalization acceleration, in which variance estimation is performed in the stochastic computing domain using AND-gate squarers in place of the conventional parallel multiplier array, while mean computation and final normalization are retained in the binary domain. The proposed architecture integrates several key stochastic computing techniques, including bitstream decorrelation via paired linear feedback shift registers employing distinct primitive polynomials, stochastic-to-binary conversion through an accumulative parallel counter, signal-to-noise ratio enhancement via dynamic prescaling, and high-precision inverse square-root computation through time-multiplexed Newton-Raphson iteration on a single shared multiplier. Synthesized in SMIC 55 nm technology, the proposed design achieves approximately 86% reduction in cell area, over 80% reduction in dynamic power, and nearly 90% reduction in leakage power compared with two baseline accelerators based on piecewise linear approximation. Across a wide range of input distributions, and at a fixed stochastic bitstream length of 32 samples per element, the proposed design achieves a 3.7–7.5× improvement in mean squared error over the piecewise linear baselines, attributable to the substantially higher effective precision of Newton-Raphson iteration relative to cascaded piecewise linear approximation. The composite power-mean-squared-error figure of merit demonstrates that the proposed design outperforms the baselines across all tested ranges. These area and power gains are obtained at the expense of throughput and energy efficiency: throughput-, energy-, and energy-delay-product-normalized comparisons show that the design trades time for silicon area and peak power. These results indicate that selectively applying stochastic computing to the computational bottleneck of layer normalization constitutes an effective design strategy for area- and peak-power-constrained, low-throughput edge artificial intelligence accelerators.
layer normalization / stochastic computing / Transformer accelerator / fixed-point arithmetic / Newton-Raphson / ASIC synthesis / edge AI
Higher Education Press 2026
/
| 〈 |
|
〉 |