Date of Award
8-2026
Document Type
Thesis
Degree Name
Master of Science
Department
Electrical Engineering
Abstract
The increasing integration of renewable energies into modern power systems has created significant operational challenges, including transmission congestion, renewable energy curtailment, and increased real-time supply-demand variability, which makes effective coordination and optimization of multiple distributed Battery Energy Storage Systems (BESS) across power systems a critical problem, a useful solution, yet an unresolved challenge. Existing centralized optimization methods require perfect future information and are incompatible with the confidentiality constraints of deregulated electricity markets, while rule-based approaches fail to capture the multi-objective complexity of real-world power grid operations. This thesis proposed a Multi-Agent Proximal Policy Optimization (MAPPO) framework with Centralized Training and Decentralized Execution (CTDE) to coordinate multiple utility-scale BESS agents placed at strategic locations across large-scale power systems. The proposed framework was trained using an enhanced 18-dimensional observation space and an eight-component reward function that simultaneously captured price arbitrage, congestion relief, curtailment mitigation, and battery degradation objectives, while a Gurobi-based Mixed-Integer Quadratic Program (MIQP) solver with perfect foresight about the future electricity market and operating information served as a benchmark to establish a theoretical performance upper bound. During training, the CTDE approach employed a centralized critic that observed the complete global state of all agents, enabling accurate advantage estimation and stable multi-agent learning, while each actor operated fully independently at deployment using only local information with no inter-agent communication.
The synthetic Texas 2000-bus grid was studied as a representative of the Independent System Operator (ISO)-level power systems to evaluate the effectiveness and efficiency of the proposed methods. Experimental results showed that the proposed MAPPO framework achieved 78.4% of the Gurobi-based optimal benchmark performance with a Sharpe ratio of 5.59 compared to 2.89 for the best rule-based baseline, with the remaining 21.6% optimality gap representing the fundamental value attributable to the perfect foresight of future LMP trajectories and system-wide operating states that the conventional optimization solver exploits, but no real-time policy can really achieve. The numerical studies demonstrated that the proposed MAPPO framework can achieve high-quality and computationally efficient BESS coordination while maintaining full operational feasibility, which outperforms the conventional rule-based baseline methods.
Index Terms- Battery energy storage systems, power system operations, distributed optimization, multi-agent reinforcement learning
Committee Chair/Advisor
Lin Gong
Committee Member
Lijun Qian
Committee Member
Xiangfang Li
Publisher
Prairie View A&M University
Rights
© 2021 Prairie View A & M University
This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
Date of Digitization
8/27/2026
Contributing Institution
J. B Coleman Library
City of Publication
Prairie View
MIME Type
Application/PDF
Recommended Citation
Debnath, P. (2026). Distributed Coordination And Optimization Of Utility-Scale Battery Energy Storage Systems Via Multi-Agent Reinforcement Learning. Retrieved from https://digitalcommons.pvamu.edu/pvamu-theses/1683