A

Mixed-Precision Training and Compilation for RRAM-based Computing-in-Memory Accelerators

Abstract

Computing-in-Memory (CIM) accelerators are a promising solution for accelerating Machine Learning (ML) workloads, as they perform Matrix-Vector Multiplications (MVMs) on crossbar arrays directly in memory. Although the bit widths of the crossbar inputs and cells are very limited, most CIM compilers do not support quantization below 8 bit. As a result, a single MVM requires many compute cycles, and weights cannot be efficiently stored in a single crossbar cell.To address this problem, we propose a mixed-precision training and compilation framework for CIM architectures. The biggest challenge is the massive search space, that makes it difficult to find good quantization parameters. This is why we introduce a reinforcement learning-based strategy to find suitable quantization configurations that balance latency and accuracy. In the best case, our approach achieves up to a 2.48× speedup over existing state-of-the-art solutions, with an accuracy loss of only 0.086 %.

Authors 6

  1. Rebecca Pelke Aachen

    RWTH Aachen University

    Affiliation as printed

    RWTH Aachen University,Germany

  2. Joel Klein Aachen

    RWTH Aachen University

    Affiliation as printed

    RWTH Aachen University,Germany

  3. RWTH Aachen University

    Affiliation as printed

    RWTH Aachen University,Germany

  4. Nils Bosbach Aachen

    RWTH Aachen University

    Affiliation as printed

    RWTH Aachen University,Germany

  5. RWTH Aachen University

    Affiliation as printed

    RWTH Aachen University,Germany

  6. RWTH Aachen University

    Affiliation as printed

    RWTH Aachen University,Germany

Cited by 0 stored of 0

Cited by patents worldwide 1 (Lens.org)

References 44