Abstract
Many real-world tasks require high-level coordination among interdependent subtasks, often represented by intricate relationships in a directed acyclic graph (DAG). Existing multi-agent reinforcement learning (MARL) methods, however, are not well-suited to such scenarios. To address this challenge, we formalize the problem as a Markov decision process with DAG constraints and propose a novel MARL framework designed to enable multiple agents to learn and coordinate effectively within these constraints. Theoretically, we demonstrate that each agent can focus on maximizing its own synthetic rewards, which are carefully designed to align with the overall team rewards, thereby ensuring high team rewards. Computationally, we propose a practical algorithm that exploits new notion of a leader agent and reward generator and distributor (RGD) agent to guide the decomposed follower agents for better coordination. Empirically, we evaluate our method across four DAG environments, including a real-world semiconductor scheduling task for Intel's high-volume packaging and test factory. Our results demonstrate that the proposed method outperforms non-DAG approaches, with the leader agent and RGD agent playing a key role in driving this performance improvement.
| Original language | English |
|---|---|
| Article number | 127744 |
| Journal | Expert Systems with Applications |
| Volume | 283 |
| DOIs | |
| State | Published - 15 Jul 2025 |
Bibliographical note
Publisher Copyright:© 2025 Elsevier Ltd
Keywords
- Multi-agent systems
- Reinforcement learning
- Reward shaping
- Scalable model
- Synthetic reward
Fingerprint
Dive into the research topics of 'Learning multiple coordinated agents under directed acyclic graph constraints'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver