We introduce RoadAccidentScenes, a large-scale annotated dataset containing 68,642 accident and non-accident road-scene images collected from heterogeneous sources and covering diverse traffic environments, viewpoints, and lighting conditions.
We develop a unified image-classification framework that investigates two complementary visual representations: a handcrafted feature representation combining color, texture, statistical, shape, and edge characteristics, and a direct pixel representation without handcrafted feature extraction.
We systematically evaluate multiple conventional machine-learning classifiers under a consistent five-fold cross-validation framework and analyze their classification performance and computational requirements across the two visual representations.
We demonstrate that the effectiveness of visual representation is classifier-dependent, with handcrafted features substantially improving the performance of several classifiers, including Logistic Regression and Linear SVM, while direct pixel representations provide better performance for selected models.
