A Comparative Analysis of Multimodal Frameworks for Multilingual Human vs. AI Caption Detection

This paper presents a comprehensive comparative analysis of multimodal frameworks for detecting human vs. AI-generated captions in a multilingual context. It identifies the most robust framework architectures and provides critical insights into cross-lingual detection accuracy, helping improve the reliability of AI-generated content auditing.