Deep Reinforcement Learning for Adaptive Bitrate Control in Rural Telemedicine Networks

In rural Bangladesh, weather-induced bandwidth volatility, monsoon induced fading, and peak hour congestion are all huge Quality-of-Service (QoS) issues with telemedicine video streaming over LTE networks. Static adaptive bitrate (ABR) heuristics fail under these non-stationary conditions, causing rebuffering events that disrupt clinical consultations. This paper proposes a deep reinforcement learning (DRL) framework for ABR control tailored to Bangladesh LTE, formulating the problem as a Markov Decision Process (MDP) with a sevendimensional state and a clinically calibrated Pensieve reward that heavily penalises rebuffering. A Proximal Policy Optimization (PPO) agent is trained on a 200-trace Gilbert-Elliott Markov channel dataset anchored to BTRC 2024 field measurements across five weather conditions and evaluated against DQN and four classical baselines. PPO achieves a mean per-step reward of +0.604 +/- 0.119, outperforming the best classical baseline by 37.3% and MPC-Heuristic by 61.5% (Wilcoxon p < 0.001, Cohen’s d = 1.10), while maintaining a stall ratio of 0.36%, well below the 2% clinical target, with consistent positive rewards across all weather conditions including adversarial Rain/Monsoon and Peak Hours scenarios where competing agents fail.