Our contributions are:
• A walk-forward (rolling-origin) evaluation of six forecasting models spanning classical econometrics, ensemble ML, and quantile-regression deep learning, with bootstrap confidence intervals and formal Diebold-Mariano
significance testing between every model pair.
• Aregime-conditional SHAP analysis that quantifies how each model’s reliance on macroeconomic features shifts between crisis and calm periods—an interpretability angle largely absent from existing GDP forecasting litera
ture.
• A learning-curve analysis quantifying the training-data threshold at which each model class stabilizes, directly addressing the small-sample concern rather than treating it as an unexamined limitation.
