Depression-severity classifiers built on the Center for Epidemiologic Studies Depression Scale (CES-D) are usually evaluated on accuracy alone, without checking whether the 20item instrument behaves coherently on the population studied. We address both questions on 896 Bangladeshi university students. First, a bank of 12 classifiers is benchmarked on mapping the 20 CES-D items to four severity categories; leave-one-out ablation and forward selection identify a two-member ensemble of Logistic Regression and TabPFN, a pre-trained tabular foundation model applied zero-shot. On a locked test set, this ensemble reaches ROC-AUC 1.000, balanced accuracy 0.986, and F1 0.986 for binary elevated-symptomatology classification, and balanced accuracy 0.983 for the four-class task. Because the severity label is a deterministic function of the same 20 items, this is reported as classification-rule recovery rather than independent prediction. Second, the CES-D’s internal structure is validated using itemnetwork predictability– each item modeled from the other 19, with no circular dependency on the target– cross-checked against exploratory factor analysis (KMO = 0.929, Cronbach’s α = 0.852). Both methods converge on the same core symptom cluster (felt sad, felt depressed, loneliness), consistent with the standard four-factor CES-D structure. Calibration, subgroup fairness, and SHAP and permutation-importance explainability are reported for the ensemble as a unit.
