TY - GEN
T1 - A Graph-based Reinforcement Learning Framework for Urban Air Mobility Fleet Scheduling
AU - Paul, Steve
AU - Chowdhury, Souma
N1 - Publisher Copyright:
© 2022, American Institute of Aeronautics and Astronautics Inc, AIAA. All rights reserved.
PY - 2022
Y1 - 2022
N2 - Optimal scheduling of the fleet of aircraft comprising an urban air mobility (UAM) network is key to economically viable and sustainable integration of UAM networks within our existing urban and suburban transportation ecosystems. To this end, this paper firstly formulates the UAM fleet scheduling problem as a Markov Decision Process (MDP) over a graph space, with the graph representing the network of vertiports (and their dynamic properties, e.g., demand) being served by these aircraft. A simulation environment that incorporates real-world constraints associated with aircraft characteristics (e.g., max speed and battery capacity), passenger transport demand and electricity pricing is developed and used to evaluate schedules modeled by this MDP. The event-triggered action of each aircraft is determined in a decentralized manner using a novel policy model embodied by a neural network comprising a Graph Neural Network (GNN) based encoder and a Multi-head attention (MHA) based decoder. A policy gradient based reinforcement learning (RL) method is used to train this model. Motivated by the emerging work in learning to solve combinatorial optimization problems, this GNN-based policy model is expected to capture the local and global structural information of the UAM network, allowing the trained policies to generalize across demand and aircraft initialization scenarios. Compared to a simple feasible randomized baseline and a typical multi-layer neural network based policy, our method demonstrates a remarkable 25% better performance in terms of the estimated average daily profit.
AB - Optimal scheduling of the fleet of aircraft comprising an urban air mobility (UAM) network is key to economically viable and sustainable integration of UAM networks within our existing urban and suburban transportation ecosystems. To this end, this paper firstly formulates the UAM fleet scheduling problem as a Markov Decision Process (MDP) over a graph space, with the graph representing the network of vertiports (and their dynamic properties, e.g., demand) being served by these aircraft. A simulation environment that incorporates real-world constraints associated with aircraft characteristics (e.g., max speed and battery capacity), passenger transport demand and electricity pricing is developed and used to evaluate schedules modeled by this MDP. The event-triggered action of each aircraft is determined in a decentralized manner using a novel policy model embodied by a neural network comprising a Graph Neural Network (GNN) based encoder and a Multi-head attention (MHA) based decoder. A policy gradient based reinforcement learning (RL) method is used to train this model. Motivated by the emerging work in learning to solve combinatorial optimization problems, this GNN-based policy model is expected to capture the local and global structural information of the UAM network, allowing the trained policies to generalize across demand and aircraft initialization scenarios. Compared to a simple feasible randomized baseline and a typical multi-layer neural network based policy, our method demonstrates a remarkable 25% better performance in terms of the estimated average daily profit.
KW - Fleet scheduling
KW - Graph neural network
KW - Reinforcement learning
KW - Urban Air Mobility (UAM)
UR - https://www.scopus.com/pages/publications/85135247129
U2 - 10.2514/6.2022-3911
DO - 10.2514/6.2022-3911
M3 - Conference contribution
AN - SCOPUS:85135247129
SN - 9781624106354
T3 - AIAA AVIATION 2022 Forum
BT - AIAA AVIATION 2022 Forum
PB - American Institute of Aeronautics and Astronautics Inc, AIAA
T2 - AIAA AVIATION 2022 Forum
Y2 - 27 June 2022 through 1 July 2022
ER -