A

GUPrecision: Group-Wise Uniform Precision Accelerator for Depthwise Separable Convolution using Hardware-Algorithm Co-Design

ACM Great Lakes Symposium on VLSI (GLSVLSI), pp. 876–882

Abstract

Quantization and pruning are effective techniques for reducing neural network size and improving energy efficiency. Although fixed word length networks are well-suited for hardware acceleration, mixed-precision and pruned networks still suffer from efficient hardware support. Depthwise separable convolution (DSC) has become a key building block for resource-constrained devices; however, applying quantization and pruning to DSC models remains challenging. To address these challenges, we propose GUPrecision, a group-wise mixed-precision uniform quantization framework that inherently supports network pruning while enabling efficient hardware realization. GUPrecision achieves hardware compatibility by dividing channels into subgroups with a fixed total bit budget. Within each group, the mixed-precision multipliers in the PE array can dynamically adapt to varying word lengths using simple shifting and multiplexing operations. We evaluated GUPrecision on the MobileNetV1 model and implemented it using GlobalFoundries 22 nm FDSOI technology. The DSC accelerator operates at 1 GHz and 0.8 V after signoff, occupying an area of 0.71 mm2 . At 75% effective sparsity, it achieves a peak energy efficiency of 17.4 TOPS/W, with a corresponding throughput of 8136 GOPS and an area efficiency of 11459 GOPS/mm2.

Authors 3

  1. Jie Lou Aachen

    RWTH Aachen University

    Affiliation as printed

    RWTH Aachen University, Aachen, Germany

  2. Malte Wabnitz Aachen

    RWTH Aachen University

    Affiliation as printed

    RWTH Aachen University, Aachen, Germany

  3. RWTH Aachen University

    Affiliation as printed

    RWTH Aachen University, Aachen, Germany

Cited by 1 stored of 1

1 result

No patents citing this paper on Lens.org (checked 2026-10-06).

References 24