Vision-Based Traffic Signal Control with an Enhanced Deep Q-Network: Evaluation and Reward Decomposition on a Camera-Derived Benchmark

An end-to-end vision-based controller. We present a dueling, double-target EDQN with 9.54 million parameters that maps four raw 128×128 RGB camera views directly to a signal phase, removing the detection and state-estimation stages that intervene in conventional pipelines.
A reproducible offline protocol. We define a complete ten-stage pipeline – preprocessing, MDP construction, training, greedy evaluation, metric estimation and baseline comparison – over a camera-derived corpus of 728 synchronised four-view frames, with every reported figure traceable to a stated closed-form estimator.
Quantified performance against baselines. The EDQN attains a cumulative return of -4358.4 against -9748.4 for a random policy and -9868.4 for a fixed-time policy, a 2.24× improvement, and identifies the highest-demand approach in 99.0% of steps compared with 25.0% and 23.3% respectively.
A reward decomposition diagnostic. We derive a closed-form decomposition of the reward and of the delay estimator that attributes the entire performance gap to the approach-selection term, proves that queue, waiting time, throughput and speed are policy-invariant under open-loop replay, and shows that an apparent 0.47 s delay regression is an estimator artefact rather than a behavioural effect.