Application of Reinforcement Learning and Large Language Models for Energy Optimization in Wireless Networks
编号:117
访问权限:仅限参会人
更新:2026-10-04 23:45:24 浏览:20次
In-person
摘要
This paper investigates the combination of Proximal Policy Optimization (PPO) with DeepSeek-R1-Distill-Qwen-7B to improve network performance. Simulations show that the proposed PPO-LLM framework achieves the lowest energy consumption (2038 Wh) across all evaluated methods, including the Enhanced Heuristic Switching Algorithm (EHSA), Deep Deterministic Policy Gradient (DDPG), Soft Actor-Critic (SAC), Twin Delayed DDPG (TD3), and PPO. This corresponds to a 47% reduction relative to the unoptimized baseline while maintaining a competitive downlink throughput of 28,885 Mbps. Compared to standard Deep Reinforcement Learning (DRL) baselines, PPO-LLM outperforms SAC in energy efficiency by 23% and surpasses TD3 and DDPG in throughput, demonstrating a superior throughput–energy trade-off. These results suggest that LLM-guided reward engineering is a promising approach to automating reward design, enabling efficient and adaptive energy management in 5G heterogeneous networks
关键词
Large Language Models (LLM),5G Heterogeneous Cellular Networks,energy-efficiency,reward function
稿件作者
Bui Duong
VNU
Le Giang
VNU
Le Hoang
VNU
Hoang Hoc
VNU
Nguyen Tuan
VNU
Pham Thinh
Viettel
Thai-Mai Dinh Thi
VNU
发表评论