Implementation of SHAP in Ensemble Learning for Interpreting Audit Risk Classification: A Preliminary Study Using a Public Audit Dataset

Authors

  • Yanuangga Galahartlambang Institut Teknologi dan Bisnis Ahmad Dahlan Lamongan Author
  • Titik Khotiah Institut Teknologi dan Bisnis Ahmad Dahlan Lamongan Author
  • Ilham Basri K Institut Teknologi dan Bisnis Ahmad Dahlan Lamongan Author
  • Masrur Anwar Institut Teknologi dan Bisnis Ahmad Dahlan Lamongan Author
  • Mohamad Akhsan Rofiqi Institut Teknologi dan Bisnis Ahmad Dahlan Lamongan Author

Keywords:

Audit Risk, Ensemble Learning, Explainable AI, SHAP, XAI.

Abstract

Ensemble learning models can achieve strong predictive performance on structured audit-risk data, yet their complex decision logic can limit transparency and technical accountability. This study develops a proof-of-concept Explainable Artificial Intelligence pipeline using SHapley Additive exPlanations (SHAP) to interpret Random Forest and XGBoost classifiers on the public UCI Audit Data. The experiment processed 776 observations and 27 columns, with one missing value in Money_Value imputed using the median of 0.09 and the non-numeric LOCATION_ID removed, resulting in 25 predictors and one binary target. A stratified 80:20 split with random_state = 42 produced 620 training observations and 156 test observations, with Random Forest achieving 1.0000 across accuracy, precision, recall, F1-score, and ROC-AUC, while XGBoost achieved 0.9936 accuracy, 1.0000 precision, 0.9836 recall, 0.9917 F1-score, and 1.0000 ROC-AUC. SHAP analysis identified Audit_Risk as the dominant global predictor, while its contribution to a representative at-risk prediction reached +5.82, shifting the raw model margin from −0.506 to 5.107 and yielding an estimated probability of approximately 0.994. However, because the dataset's target construction is closely associated with the Audit Risk Score, retaining Audit_Risk as a predictor introduces potential structural target leakage. The findings position the near-perfect performance as evidence of successful pipeline implementation and model interpretability rather than external generalization, supporting future leakage-controlled ablation, cross-validation, and external-data validation.

Downloads

Download data is not yet available.

References

Arrieta, A. B., Díaz-Rodríguez, N., Del Ser, J., Bennetot, A., Tabik, S., Barbado, A., ... & Herrera, F. (2020). Explainable Artificial Intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI. Information fusion, 58, 82-115. https://doi.org/10.1016/j.inffus.2019.12.012

Ashtari, S., et al. (2023). Improving audit opinion prediction accuracy using metaheuristics-tuned XGBoost algorithm with interpretable results through SHAP value analysis. Applied Soft Computing, 149, 110955. https://doi.org/10.1016/j.asoc.2023.110955

Burkart, N., & Huber, M. F. (2021). A survey on the explainability of supervised machine learning. Journal of Artificial Intelligence Research, 70, 245-317. https://doi.org/10.1613/jair.1.12228

Chen, T., & Guestrin, C. (2016). XGBoost: A Scalable Tree Boosting System. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 785–794. https://doi.org/10.1145/2939672.2939785

Erasmus, N. C., & Paepae, T. (2026). An Explainability-Driven SHAP-Weighted Ensemble Framework for Fraud Detection: Insights into Model Contribution Dynamics. Information, 17(6), 607. https://doi.org/10.3390/info17060607

Hooda, N., Bawa, S., & Rana, P. S. (2018). Fraudulent Firm Classification: A Case Study of an External Audit. Applied Artificial Intelligence, 32(1), 48–64. https://doi.org/10.1080/08839514.2018.1451032

Hooda, N. (2018). Audit Data [Dataset]. UCI Machine Learning Repository. https://doi.org/10.24432/C5930Q.

Lin, K., & Gao, Y. (2022). Model interpretability of financial fraud detection by group SHAP. Expert Systems with Applications, 210, 118354. https://doi.org/10.1016/j.eswa.2022.118354

Lundberg, S. M., & Lee, S. I. (2017). A unified approach to interpreting model predictions. Advances in neural information processing systems, 30. https://dl.acm.org/doi/10.5555/3295222.3295230

Lundberg, S. M., Erion, G., Chen, H., DeGrave, A., Prutkin, J. M., Nair, B., ... & Lee, S. I. (2020). From local explanations to global understanding with explainable AI for trees. Nature machine intelligence, 2(1), 56-67. https://doi.org/10.1038/s42256-019-0138-9

Minh, D., Wang, H. X., Li, Y. F., & Nguyen, T. N. (2022). Explainable artificial intelligence: a comprehensive review. Artificial Intelligence Review, 55(5), 3503-3568. https://doi.org/10.1007/s10462-021-10088-y

Mupenzi, J. (2026). An Explainable Hybrid Machine Learning Framework for Financial and Tax Fraud Analytics in Emerging Economies. Ultimatics: Jurnal Teknik Informatika, 17(2), 194–202. https://doi.org/10.31937/ti.v17i2.4483

Nguyen, H. H., Viviani, J. L., & Ben Jabeur, S. (2025). Bankruptcy prediction using machine learning and Shapley additive explanations: H.-H. Nguyen et al. Review of Quantitative Finance and Accounting, 65(1), 107-148. https://doi.org/10.1007/s11156-023-01192-x

Qazi, A., & Jadoon, A. A. K. (2026). Explainable ensemble learning using SHAP for ERP anomaly detection. Scientific Reports. https://doi.org/10.1038/s41598-026-57913-4

Resul, A. P. A. K., & GANJI, F. (2025). Using decision tree algorithms and artificial intelligence to increase audit quality: A data-based approach to predicting financial risks. International Journal of Business Management and Entrepreneurship, 4(1), 87-99. https://mbajournal.ir/index.php/IJBME/article/view/65

Rudin, C. (2019). Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nature machine intelligence, 1(5), 206-215. https://doi.org/10.1038/s42256-019-0048-x

Salih, A. M., Raisi‐Estabragh, Z., Galazzo, I. B., Radeva, P., Petersen, S. E., Lekadir, K., & Menegaz, G. (2025). A perspective on explainable artificial intelligence methods: SHAP and LIME. Advanced Intelligent Systems, 7(1), 2400304. https://doi.org/10.1002/aisy.202400304

Samek, W., Wiegand, T., & Müller, K. R. (2017). Explainable artificial intelligence: Understanding, visualizing and interpreting deep learning models. arXiv preprint arXiv:1708.08296. http://arxiv.org/abs/1708.08296

Selvam, S., & Sughasiny, M. (2025). Smart and Explainable Credit Card Fraud Detection Using XGBoost and SHAP. Journal of ISMAC, 7(2), 155–169. https://doi.org/10.36548/jismac.2025.2.004

Sibagariang, S. (2025). Interpretable Machine Learning for Job Placement Prediction: A SHAP-Based Feature Analysis. Jurnal Nasional Teknik Elektro dan Teknologi Informasi, 14(3), 190-198. https://doi.org/10.22146/jnteti.v14i3.20516

Thanathamathee, P., Sawangarreerak, S., Chantamunee, S., & Mohd Nizam, D. N. (2024). SHAP-Instance Weighted and Anchor Explainable AI: Enhancing XGBoost for Financial Fraud Detection. Emerging Science Journal, 8(6), 2404–2430. https://doi.org/10.28991/ESJ-2024-08-06-016

Ulambayar, T., Pagjii, O., Luvsandash, O., Garamdorj, G., & Sambuu, U. (2026). Audit Detection Risk Reduction Using Machine Learning: Evidence from 3.3 million Transactions. Research Square https://doi.org/10.21203/rs.3.rs-9968348/v1

Yanuangga G. (2025). Explainable AI (XAI) Analysis Using SHAP for Credit Card Fraud. Scripta-Technica, 1(2). https://doi.org/10.65310/scxk4755

Downloads

Published

2026-06-28

How to Cite

Implementation of SHAP in Ensemble Learning for Interpreting Audit Risk Classification: A Preliminary Study Using a Public Audit Dataset. (2026). Technema: Journal of Intelligent Engineering and Computing, 1(2), 308-323. https://sovereignresearch.org/technema/article/view/239