ORCID

Abstract

Explainable AI (XAI) is increasingly used in clinical machine learning, yet quantitative evaluation of explanation quality is often reported inconsistently across methods and datasets. We present a reproducible, metric-driven framework for evaluating XAI methods on healthcare tabular data. The framework consolidates six established, family-specific metrics, fidelity, simplicity, consistency, robustness, precision, and coverage, into explicit equations; pairs them with a pre-specified focal-model protocol; and releases open-source code with a method-metric applicability map. We evaluate LIME, SHAP, Anchors, EBM, and TabNet across four public healthcare tabular datasets. Post-hoc explainers are applied to a single selected Random Forest focal predictor to control model-induced variability, whereas EBM and TabNet are assessed through their native interpretability mechanisms. Global explanation summaries are reported descriptively only. The results show that SHAP/TreeSHAP provides exact score reconstruction for the Random Forest setting, while LIME produces simpler but lower-fidelity explanations with greater instance-level variability. LIME and SHAP show the strongest rank agreement among the evaluated pairs, although agreement varies across datasets. TabNet often yields compact native explanations, but these must be interpreted alongside its dataset-specific predictive performance. EBM and TabNet show low sensitivity under the fixed Gaussian-jitter robustness protocol, while Anchors produces high-precision rules with reduced coverage at stricter thresholds. Overall, the framework enables controlled comparison under explicit method-metric and focal-model assumptions, supporting more transparent XAI selection for tabular machine learning. Although demonstrated in healthcare, the framework is transferable to other high-stakes tabular domains. Source code: https://github.com/matifq/XAI_Tab_Health.

Publication Date

2026-01-01

Publication Title

PLoS ONE

Volume

21

Issue

7

First Page

351473

Last Page

351473

Deposit Date

2026-09-24

Funding

This publication has emanated from research conducted with the support of Research Ireland under award nos. GOIPG/2022/660 and GOIPG/2021/1354 and grant number 18/CRT/6183 from Research Ireland Centre for Research Training in Machine Learning (ML-Labs) at Technological University Dublin, with the financial support of Research Ireland under grant no. 13/RC/2106_P2 at the ADAPT Research Centre at Technological University Dublin. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.

Creative Commons License

Creative Commons Attribution 4.0 International License
This work is licensed under a Creative Commons Attribution 4.0 International License.


Share

COinS