Heterogeneous Multi-Agent Proximal Policy Optimization for Power Distribution System Restoration
By: Parya Dolatyabi, Mahdi Khodayar
Potential Business Impact:
Fixes power grids faster after blackouts.
Restoring power distribution systems (PDS) after large-scale outages requires sequential switching operations that reconfigure feeder topology and coordinate distributed energy resources (DERs) under nonlinear constraints such as power balance, voltage limits, and thermal ratings. These challenges make conventional optimization and value-based RL approaches computationally inefficient and difficult to scale. This paper applies a Heterogeneous-Agent Reinforcement Learning (HARL) framework, instantiated through Heterogeneous-Agent Proximal Policy Optimization (HAPPO), to enable coordinated restoration across interconnected microgrids. Each agent controls a distinct microgrid with different loads, DER capacities, and switch counts, introducing practical structural heterogeneity. Decentralized actor policies are trained with a centralized critic to compute advantage values for stable on-policy updates. A physics-informed OpenDSS environment provides full power flow feedback and enforces operational limits via differentiable penalty signals rather than invalid action masking. The total DER generation is capped at 2400 kW, and each microgrid must satisfy local supply-demand feasibility. Experiments on the IEEE 123-bus and IEEE 8500-node systems show that HAPPO achieves faster convergence, higher restored power, and smoother multi-seed training than DQN, PPO, MAES, MAGDPG, MADQN, Mean-Field RL, and QMIX. Results demonstrate that incorporating microgrid-level heterogeneity within the HARL framework yields a scalable, stable, and constraint-aware solution for complex PDS restoration.
Similar Papers
Agentic-AI based Mathematical Framework for Commercialization of Energy Resilience in Electrical Distribution System Planning and Operation
Systems and Control
Makes power grids stronger against storms and attacks.
Optimisation of Resource Allocation in Heterogeneous Wireless Networks Using Deep Reinforcement Learning
Machine Learning (CS)
Makes phone signals faster and better for everyone.
Situationally Aware Rolling Horizon Multi-Tier Load Restoration Considering Behind-The-Meter DER
Systems and Control
Fixes power outages faster by connecting distant parts.