Abstract—Unmanned aerial vehicle (UAV) relay swarms can
maintain connectivity when terrestrial infrastructure is unavail
able, but their open wireless links expose them to both external
jamming and internal packet-forwarding attacks. This paper
develops a trust-aware, hybrid centralized-training decentralized
execution (CTDE) architecture for a 10-UAV swarm serving 10
source–destination (SU–DU) pairs over two-hop links within a
120 m × 120 m tactical area in the presence of two external jam
mers and an identity-hidden gray-hole UAV. The compromised
UAV follows the legitimate mobility and power-control policy,
but after a dormant period selectively drops forwarded packets
with probability 0.2, 0.5, or 0.8. Its identity is hidden from the
learned policy and the network controller. Each UAV applies a
parameter-shared Double Deep Q-Network (Double-DQN) over a
53-dimensional local observation to select one of 21 joint motion–
power actions, while a centralized critic uses a 90-dimensional
global state for advantage shaping. A channel-calibrated trust
controller uses trusted packet-observation points, rather than
relay-reported counters, to isolate suspicious relays and reassign
traffic. We report scheduled Shannon throughput (in Mbit/s) and
ACK-confirmed packet goodput (in kbit/s) as distinct metrics.
The reported results are preliminary point estimates from one
1.2-million-step training run (seed 42, initialized from compatible
legacy weights) and one 500-slot rollout per attack severity; they
do not establish statistical significance. In these runs, scheduled
throughput is approximately 0.6–1.0 Mbit/s across the swarm.
At 50% and 80% packet dropping, the trust controller detects
the attacker within 4 and 3 slots, respectively, and the observed
packet-delivery ratios exceed the no-defense point estimates by
7.2 and 7.1 percentage points.
Index Terms—UAV swarm, anti-jamming, gray-hole attack,
selective forwarding, trust management, centralized training,
Double-DQN, relay reassignment.
发表评论