Mixed-Precision Training and Compilation for RRAM-based Computing-in-Memory Accelerators
Abstract
Computing-in-Memory (CIM) accelerators are a promising solution for accelerating Machine Learning (ML) workloads, as they perform Matrix-Vector Multiplications (MVMs) on crossbar arrays directly in memory. Although the bit widths of the crossbar inputs and cells are very limited, most CIM compilers do not support quantization below 8 bit. As a result, a single MVM requires many compute cycles, and weights cannot be efficiently stored in a single crossbar cell.To address this problem, we propose a mixed-precision training and compilation framework for CIM architectures. The biggest challenge is the massive search space, that makes it difficult to find good quantization parameters. This is why we introduce a reinforcement learning-based strategy to find suitable quantization configurations that balance latency and accuracy. In the best case, our approach achieves up to a 2.48× speedup over existing state-of-the-art solutions, with an accuracy loss of only 0.086 %.
Authors 6
-
Rebecca Pelke Aachen
Affiliation as printed
RWTH Aachen University,Germany
-
Joel Klein Aachen
Affiliation as printed
RWTH Aachen University,Germany
-
José Cubero-Cascante Aachen
Affiliation as printed
RWTH Aachen University,Germany
-
Nils Bosbach Aachen
Affiliation as printed
RWTH Aachen University,Germany
-
Jan Moritz Joseph Aachen
Affiliation as printed
RWTH Aachen University,Germany
-
Rainer Leupers Aachen
Affiliation as printed
RWTH Aachen University,Germany
Cited by 0 stored of 0
Cited by patents worldwide 1 (Lens.org)
-
In-memory calculation-oriented mixed precision inverse quantization system and methodCN122154796A 2026-06-05 Active
References 44
-
W2963122961details pending0citations
-
W4220958508details pending0citations
-
W3212430008details pending0citations
-
W3018945530details pending0citations
-
W4213436331details pending0citations
-
W2979365412details pending0citations
-
W3198899759details pending0citations