Xingyu Liu, Esa Apriaskar, Lyudmila Mihaylova
The introduction of autonomous vehicles (AVs) presents a novel approach to regulating and optimising traffic flow through the automated control of AVs. In this context, the AV is defined as the actuator and an optimal control policy is desired to make control decisions. Deep Reinforcement Learning (DRL) is a novel method which aims to maximize the cumulative rewards given by the predefined reward function by making sequential decisions in a stochastic environment. In light of the above, we propose a DRL-based vehicular control method to train an optimal policy for the control of AV in a model-free fashion, and consequently improve the traffic efficiency with the obtained control policy. A single-lane circular road environment with both AV and human-driven vehicles is selected to serve as the mixed autonomy traffic system in the Simulation of Urban MObility (SUMO) [1] traffic simulator, and the Proximal Policy Optimization (PPO) algorithm is applied for the policy improvement. Simulation results demonstrate that our strategy is effective in mitigating the unstable stop-and-go waves, increasing 67.7% of the average driving speed and reducing 19.3% of the average energy consumption in a closed-ring road environment. © 2024 IEEE.
University of Sheffield, Department of Automatic Control and Systems Engineering, Sheffield, S1 3JD, United Kingdom; Universitas Negeri Semarang, Faculty of Engineering, Department of Electrical Engineering, Semarang, 50229, Indonesia