Safe-CTDE: Safety-Aware Joint Mobility and Power Control for Multi-UAV Relay Networks with Dynamic User Assignment
编号:60 访问权限:仅限参会人 更新:2026-10-04 23:29:08 浏览:13次 In-person

报告开始:2026年10月12日 16:00(Asia/Ho_Chi_Minh)

报告时间:15min

所在会场:[S1] Track 1: Mobile computing, communications, 5G and beyond [S1-2] Track 1: Mobile computing, communications, 5G and beyond

演示文件

提示:该报告下的文件权限为仅限参会人,您尚未登录,暂时无法查看。

摘要
Coordinating multiple unmanned aerial vehicle (UAV) relays requires joint control of UAV mobility and transmission power, dynamic association between UAVs and device-to-device (D2D) pairs, and management of inter-UAV proximity risk. This paper presents Safe-CTDE, a heuristic-augmented, value-based centralized training with decentralized execution (CTDE) framework for multi-UAV relay networks. A parameter-shared Q-network is evaluated separately for each UAV using its local observation, while a centralized state-value network provides system-level temporal-difference (TD) feedback during training. UAV–device-pair association is handled separately by a capacity-aware Hungarian scheduler in which each D2D pair is assigned to at most one UAV, while each UAV may serve multiple pairs subject to a training-dependent capacity. The local Q-target combines an agent-specific shaping reward with a clipped global TD signal. During Q-guided action selection, deterministic motion/power priors and a one-step proximity-aware safety filter bias the selected action away from locally predicted separation violations; uniform ε-random exploration itself is not filtered. Reward weights follow a curriculum during the first part of training and a feedback-based adaptation mechanism thereafter. In the reported single-run experiment with five UAVs and five D2D pairs, the final Safe-CTDE training trajectory reaches approximately 69 kbps aggregate throughput with a proximity surrogate of approximately 0.30. The aligned non-learning reference curves span approximately 57–81 kbps in throughput and 0.19–0.49 in the adopted proximity metric. These results are interpreted as a proof-of-concept intermediate operating point among the tested behaviors rather than as statistical superiority or a hard collision-avoidance guarantee.
关键词
Multi-UAV relay network, multi-agent reinforcement learning, CTDE, trajectory control, user association, action masking, safety-aware learning
报告人
Thai-Duong Nguyen
Lecturer VNU University of Engineering and Technology (VNU-UET)

稿件作者
Thu Hoàng Thị PTIT
Thai-Duong Nguyen VNU University of Engineering and Technology (VNU-UET)
Tan Nguyen Ngoc VNU - University of Engineering and Technology
发表评论
验证码 看不清楚,更换一张
全部评论
重要日期
  • 会议日期

    10月11日

    2026

    至

    10月14日

    2026

  • 12月30日 2025

    报告提交截止日期

  • 09月28日 2026

    提前注册日期

  • 10月10日 2026

    初稿截稿日期

  • 10月14日 2026

    注册截止日期

主办单位
United Societies of Science
承办单位
Posts and Telecommunications Institute of Technology
协办单位
IEEE Section
IEEE Vietnam Section
移动端
在手机上打开
小程序
打开微信小程序
客服
扫码或点此咨询