Score: 0

Heterogeneous Multi-Agent Proximal Policy Optimization for Power Distribution System Restoration

Published: November 18, 2025 | arXiv ID: 2511.14730v1

By: Parya Dolatyabi, Mahdi Khodayar

Potential Business Impact:

Fixes power grids faster after blackouts.

Business Areas:

Power Grid Energy

Restoring power distribution systems (PDS) after large-scale outages requires sequential switching operations that reconfigure feeder topology and coordinate distributed energy resources (DERs) under nonlinear constraints such as power balance, voltage limits, and thermal ratings. These challenges make conventional optimization and value-based RL approaches computationally inefficient and difficult to scale. This paper applies a Heterogeneous-Agent Reinforcement Learning (HARL) framework, instantiated through Heterogeneous-Agent Proximal Policy Optimization (HAPPO), to enable coordinated restoration across interconnected microgrids. Each agent controls a distinct microgrid with different loads, DER capacities, and switch counts, introducing practical structural heterogeneity. Decentralized actor policies are trained with a centralized critic to compute advantage values for stable on-policy updates. A physics-informed OpenDSS environment provides full power flow feedback and enforces operational limits via differentiable penalty signals rather than invalid action masking. The total DER generation is capped at 2400 kW, and each microgrid must satisfy local supply-demand feasibility. Experiments on the IEEE 123-bus and IEEE 8500-node systems show that HAPPO achieves faster convergence, higher restored power, and smoother multi-seed training than DQN, PPO, MAES, MAGDPG, MADQN, Mean-Field RL, and QMIX. Results demonstrate that incorporating microgrid-level heterogeneity within the HARL framework yields a scalable, stable, and constraint-aware solution for complex PDS restoration.

Agentic-AI based Mathematical Framework for Commercialization of Energy Resilience in Electrical Distribution System Planning and Operation

Systems and Control

Makes power grids stronger against storms and attacks.

6 Aug 2025 0

89%

Optimisation of Resource Allocation in Heterogeneous Wireless Networks Using Deep Reinforcement Learning

Machine Learning (CS)

Makes phone signals faster and better for everyone.

29 Sep 2025 0

89%

Situationally Aware Rolling Horizon Multi-Tier Load Restoration Considering Behind-The-Meter DER

Systems and Control

Fixes power outages faster by connecting distant parts.

2 Oct 2025 0

View PDF Login to Bookmark

Country of Origin

🇺🇸 United States

Page Count

6 pages

Heterogeneous Multi-Agent Proximal Policy Optimization for Power Distribution System Restoration

Fixes power grids faster after blackouts.

Technical Abstract

Agentic-AI based Mathematical Framework for Commercialization of Energy Resilience in Electrical Distribution System Planning and Operation

Optimisation of Resource Allocation in Heterogeneous Wireless Networks Using Deep Reinforcement Learning

Situationally Aware Rolling Horizon Multi-Tier Load Restoration Considering Behind-The-Meter DER