An Explainable Speech-Based Conversational AI System for Sustainable, Automated English-Speaking Assessment and Personalized Feedback for EFL Learners

Automated spoken-language assessment is essential to intelligent educational systems, particularly for scaling sustainable, high-quality feedback in resource-constrained EFL settings. Yet most conversational AI tutors provide dialogue without explainable, learner-specific speaking evaluation. We propose an Explainable Speech-Based Conversational AI Framework that combines Whisper-based automatic speech recognition (ASR), multilevel speech and linguistic feature extraction, and large language model (LLM) feedback in a transparent pipeline. The system transcribes learner speech, measures fluency, lexical, grammatical, and pronunciation features, and maps detected problems to evidence-grounded explanations and personalized recommendations through structured LLM prompting. We evaluate the framework with 60 university-level EFL learners in a six-week mixed-method study: 30 used the AI system and 30 followed conventional speaking practice. Compared with ratings from three certified instructors, AI feedback achieved Cohen’s kappa = 0.74 for weakness identification. The experimental group showed significantly greater gains in fluency (Cohen’s d = 0.86) and grammatical accuracy (d = 0.79). By automating high-quality, evidence-grounded feedback, the framework offers a scalable, sustainability-oriented approach to reducing reliance on limited human-instructor resources.