Large Language Models (LLMs) are increasingly used in contexts where interaction quality depends not only on technical performance but also on human-centered qualities. This study presents a user-centered comparative evaluation of four LLMs—ChatGPT, Gemini, Claude, and Grok—across empathy, helpfulness, cultural sensitivity, safety, and overall human-centeredness. A total of 112 participants evaluated 16 selected responses (four responses per human-centered dimension) drawn from a common set of 24 prompts; the full response pool contained 96 LLM-generated responses. Quantitative analysis was conducted using descriptive statistics and repeated-measures analysis of variance, supplemented by post-hoc comparisons and categorical preference frequencies. The results indicate that the models were generally perceived similarly across most dimensions. Before correction for multiple comparisons, empathy and overall human-centeredness showed the strongest unadjusted differences, with Gemini receiving higher ratings than ChatGPT; however, after Holm correction across the five outcome tests, no dimension reached statistical significance. Preference patterns nevertheless varied by dimension, with Gemini most frequently selected for empathy and overall response preference, while Claude was most frequently selected for helpfulness and trust. These findings suggest that user judgments may vary by dimension, supporting the value of multidimensional user-centered evaluation.
