A Human-Centered Evaluation of AI-Generated Responses for Cultural Sensitivity in Large Language Models

The use of Large Language Models (LLMs) in culturally diverse settings has grown, but their capacity to produce culturally relevant output is not fully understood at the human level. There is a need to develop an AI framework that can evaluate responses for cultural sensitivity to empower non-Human Resource professionals to perform this task. Two dimensions of cultural sensitivity in AI-generated responses are proposed and evaluated in this study, cultural respect and contextual appropriateness, in an AI framework. 121 participants were recruited for the quantitative evaluation, which involved rating three different response blocks generated by AI on a five-point Likert scale. The two indicators were combined to create an overall measure of perceived cultural sensitivity, the proposed Cultural Sensitivity Score (CSS). Overall, the perception of the AI responses was positive with a mean CSS of 3.652 (SD = 0.869), well above the neutral value of 5 (t(120) = 8.247, p < .001). There were no significant differences between the three response blocks, suggesting equivalent perceived cultural sensitivity. Additionally, cultural respect and contextual appropriateness had a significant and positive correlation (r = 0.883, p < .001). The results demonstrate the value of human-centered evaluation to better understand culturally appropriate AI behavior and the need for cross-cultural studies and more comprehensive evaluation frameworks.