A challenge exists in evaluating the quality of explanations produced by explainable AI (XAI) methods. Current approaches frequently rely on subjective human judgment, which limits reproducibility and comparability. This research investigates whether LLMs can provide a scalable and reproducible mechanism for assessing XAI explanation quality. The XAI-Arena framework was introduced to allow for multidimensional comparison of explanations, considering dimensions such as simplicity, clarity, task adequacy, trust calibration, actionability, transparency, faithfulness, and overall interpretability. Benchmarking was conducted across various datasets, machine learning models, and stakeholder personas. Human validation demonstrated a strong positive association between LLM-generated and human ratings (Spearman's rho=.693, p<.001). This framework captures systematic differences in XAI explanation quality and offers a scalable and reproducible approach for comparative assessment.
Source: https://arxiv.org/abs/2609.09428