Skip to main navigation Skip to search Skip to main content

Learning multiple coordinated agents under directed acyclic graph constraints

  • Jaeyeon Jang
  • , Diego Klabjan
  • , Han Liu
  • , Nital S. Patel
  • , Xiuqi Li
  • , Balakrishnan Ananthanarayanan
  • , Husam Dauod
  • , Tzung Han Juang
  • Northwestern University
  • Intel

Research output: Contribution to journalArticlepeer-review

1 Scopus citations

Abstract

Many real-world tasks require high-level coordination among interdependent subtasks, often represented by intricate relationships in a directed acyclic graph (DAG). Existing multi-agent reinforcement learning (MARL) methods, however, are not well-suited to such scenarios. To address this challenge, we formalize the problem as a Markov decision process with DAG constraints and propose a novel MARL framework designed to enable multiple agents to learn and coordinate effectively within these constraints. Theoretically, we demonstrate that each agent can focus on maximizing its own synthetic rewards, which are carefully designed to align with the overall team rewards, thereby ensuring high team rewards. Computationally, we propose a practical algorithm that exploits new notion of a leader agent and reward generator and distributor (RGD) agent to guide the decomposed follower agents for better coordination. Empirically, we evaluate our method across four DAG environments, including a real-world semiconductor scheduling task for Intel's high-volume packaging and test factory. Our results demonstrate that the proposed method outperforms non-DAG approaches, with the leader agent and RGD agent playing a key role in driving this performance improvement.

Original languageEnglish
Article number127744
JournalExpert Systems with Applications
Volume283
DOIs
StatePublished - 15 Jul 2025

Bibliographical note

Publisher Copyright:
© 2025 Elsevier Ltd

Keywords

  • Multi-agent systems
  • Reinforcement learning
  • Reward shaping
  • Scalable model
  • Synthetic reward

Fingerprint

Dive into the research topics of 'Learning multiple coordinated agents under directed acyclic graph constraints'. Together they form a unique fingerprint.

Cite this