A

On Effectiveness of Exploration Strategies in Deep Reinforcement Learning for Power Allocation in Multi-Carrier Wireless Systems

Abstract

This paper presents a comprehensive study on the efficiency and effectiveness of exploration policies for deep reinforcement (DRL) algorithms with applications to the power allocation problem in multi-carrier wireless systems. We propose three distinct exploration functions, i.e., linear, fast and slow, to balance exploration and exploitation in the dynamic wireless environment. We analyze the effect of exploration on the initial training length as well as learning models' sum-rate performance and power violation probabilities. Our results indicate that the DRL algorithms with the proposed exploration functions reach close-to-optimal sum-rate performance within only 1000 training episodes (i.e., equivalent to 8.01 s) while satisfying the predefined power constraint of the base station.

Authors 3

  1. RWTH Aachen University

    Affiliation as printed

    Institute for Communication Technologies and Embedded Systems, RWTH Aachen University,Aachen,Germany

    Institute for Communication Technologies and Embedded Systems, RWTH Aachen University, Aachen, Germany

  2. RWTH Aachen University

    Affiliation as printed

    Institute for Communication Technologies and Embedded Systems, RWTH Aachen University,Aachen,Germany

    Institute for Communication Technologies and Embedded Systems, RWTH Aachen University, Aachen, Germany

  3. RWTH Aachen University

    Affiliation as printed

    Institute for Communication Technologies and Embedded Systems, RWTH Aachen University,Aachen,Germany

    Institute for Communication Technologies and Embedded Systems, RWTH Aachen University, Aachen, Germany

Cited by 3 stored of 3

3 results

No patents citing this paper on Lens.org (checked 2026-10-06).

References 18

18 results