The research proposes Hierarchical Dynamic Role-Graph Multi-Agent Proximal Policy Optimization (DRG-MAPPO) to address limitations in multi-agent reinforcement learning for air combat. The framework incorporates graph-based relational modeling to capture complex interactions between battlefield entities. A dynamic role assignment mechanism is used to determine tactical responsibilities, such as ‘leader’ and ‘supporter’. The system employs graph attention mechanisms to extract critical relational features, conditioning a low-level policy to execute maneuver actions. Experimental results show a win rate of 87%.
Source: https://arxiv.org/abs/2609.11155