This paper demonstrates that standard supervised machine learning models hit a strict predictive ceiling when using tabular self-reported AI usage data to forecast student outcomes across 50,000 records. Across six distinct model families—spanning linear, kernel, ensemble, and neural network architectures—performance converged to a narrow band without statistically significant separation, indicating that predictive bottlenecks stem from limitations in the feature space rather than algorithm capacity. Furthermore, a feature group ablation study revealed that AI usage data alone is highly informative for psychological targets, capturing 94% of full-model performance for burnout risk, but is virtually uninformative on its own for academic outcomes like grade change, where predictive power depends almost entirely on complex interactions between study habits, prior GPA, and AI usage patterns.
