Deep Reinforcement Learning-Based Target Wake Time Scheduling for Tactical Edge Wi-Fi Networks

M. Menezes de Carvalho
Texas State University, Texas, United States

Keywords: Deep Reinforcement Learning; Wi-Fi; Target Wake Time; Tactical Edge; Internet of Battlefield Things

Tactical Wi-Fi networks are facing ever-growing mix of heterogeneous devices with diverse QoS requirements. They may host reconnaissance sensors, real-time video feeds, and push-to-talk voice in one command-and-control fabric. Legacy contention-based MAC can no longer guarantee reliable QoS to each device. Target Wake Time (TWT), introduced in IEEE 802.11ax, opens new control plane for AP where it can unilaterally `time-slice' the channel into `service-periods' (SPs) assigned to individual/grouped stations, allowing them to enter `sleep-mode' outside their SP. A well-designed TWT schedule improves throughput, latency, and battery life dramatically. However, optimal scheduling is NP-hard, analytical models degrade under realistic traffic, and heuristics cannot guarantee optimality. We present our reinforcement-learning (RL) based TWT scheduler. Its core features are: (1) MLP/LSTM-PPO agents driving the scheduler that adapt TWT schedules to shifting traffic in real time; (2) protocol-compliant observation space limited to signals AP can realistically measure; (3) end-to-end training and evaluation on full-fidelity ns-3 simulation platform with realistic MAC/PHY dynamics; and (4) substantial gains over strong analytical M/D/1 baseline across competing objectives. Because our framework utilizes nothing beyond the standard 802.11 protocol, it can be easily integrated into existing Wi-Fi infrastructure, bringing smarter, QoS-aware Wi-Fi to the connected home and the contested edge.