A

Multi-Agent Proximal Policy Optimization for a Deadlock Capable Transport System in a Simulation-Based Learning Environment

Abstract

In this paper, we explore the potential of multi-agent reinforcement learning (MARL) for managing the driving behavior of autonomous guided vehicles (AGVs) in production logistics environments with single-lane tracks, where deadlocks pose a significant challenge. We build upon previous work and adopt a MARL approach using the Proximal Policy Optimization (PPO) algorithm. We conduct a thorough hyperparameter search and investigate the impact of varying numbers of agents on the performance of the AGVs. Our results demonstrate the effectiveness of the MARL approach in addressing deadlocks and coordinating AGV behavior, as well as the scalability of the learned policy to different numbers of agents. The Bayesian optimization process and increased iteration count contribute to improved performance and more stable learning curves.

Authors 4

  1. Otto-von-Guericke-Universität Magdeburg

    Affiliation as printed

    Otto von Guericke University Magdeburg,Institute of Logistics and Material Handling Systems,Magdeburg,Germany,39106

  2. RWTH Aachen University

    Affiliation as printed

    RWTH Aachen University,Data & Business Analytics,Aachen,Germany,52072

  3. Otto-von-Guericke-Universität Magdeburg

    Affiliation as printed

    Otto von Guericke University Magdeburg,Institute of Logistics and Material Handling Systems,Magdeburg,Germany,39106

  4. Otto-von-Guericke-Universität Magdeburg

    Affiliation as printed

    Otto von Guericke University Magdeburg,Institute of Logistics and Material Handling Systems,Magdeburg,Germany,39106

Cited by 2 stored of 2

2 results

No patents citing this paper on Lens.org (checked 2026-10-06).

References 47