TP-ViT: truncated uniform-log2 quantizer and progressive bit-decline reconstruction for vision Transformer quantization

Xichuan ZHOU , Sihuan ZHAO , Rui DING , Jiayu SHI , Jing NIE , Lihui CHEN , Haijun LIU

Eng Inform Technol Electron Eng ›› 2026, Vol. 27 ›› Issue (1) : 250081

PDF (1356KB)
Eng Inform Technol Electron Eng ›› 2026, Vol. 27 ›› Issue (1) :250081 DOI: 10.1631/ENG.ITEE.2025.0081
Research Article
TP-ViT: truncated uniform-log2 quantizer and progressive bit-decline reconstruction for vision Transformer quantization
Author information +
History +
PDF (1356KB)

Abstract

Vision Transformers (ViTs) have achieved remarkable success across various artificial intelligence-based computer vision applications. However, their demanding computational and memory requirements pose significant challenges for deployment on resource-constrained edge devices. Although post-training quantization (PTQ) provides a promising solution by reducing model precision with minimal calibration data, aggressive low-bit quantization typically leads to substantial performance degradation. To address this challenge, we present the truncated uniform-log2 quantizer and progressive bit-decline reconstruction method for vision Transformer quantization (TP-ViT). It is an innovative PTQ framework specifically designed for ViTs, featuring two key technical contributions: (1) truncated uniform-log2 quantizer, a novel quantization approach which effectively handles outlier values in post-Softmax activations, significantly reducing quantization errors; (2) bit-decline optimization strategy, which employs transition weights to gradually reduce bit precision while maintaining model performance under extreme quantization conditions. Comprehensive experiments on image classification, object detection, and instance segmentation tasks demonstrate TP-ViT's superior performance compared to state-of-the-art PTQ methods, particularly in challenging 3-bit quantization scenarios. Our framework achieves a notable 6.18 percentage points improvement in top-1 accuracy for ViT-small under 3-bit quantization. These results validate TP-ViT's robustness and general applicability, paving the way for more efficient deployment of ViT models in computer vision applications on edge hardware.

Keywords

Vision Transformers / Post-training quantization / Block reconstruction / Image classification / Object detection / Instance segmentation

Cite this article

Download citation ▾
Xichuan ZHOU, Sihuan ZHAO, Rui DING, Jiayu SHI, Jing NIE, Lihui CHEN, Haijun LIU. TP-ViT: truncated uniform-log2 quantizer and progressive bit-decline reconstruction for vision Transformer quantization. Eng Inform Technol Electron Eng, 2026, 27 (1) : 250081 DOI:10.1631/ENG.ITEE.2025.0081

登录浏览全文

4963

注册一个新账户 忘记密码

References

Rights & permissions

The Authors. Published by Zhejiang University Press Co., Ltd.

PDF (1356KB)

Supplementary files

EITEE20250081-04-XCZ-suppl 2

267

Accesses

0

Citation

Detail

Sections
Recommended

/