Can LLMs Explain AI? A Study of Explainable AI in Cybersecurity Machine Learning Models

  • Elton D. Freitas UFC
  • Beloin S. F. Rodrigues UFC
  • Emanuel B. Rodrigues UFC
  • Rossana M. C. Andrade UFC
  • Dário V. Conceição Ecole d’Ingénieur des Technologies de l’Information et de la Communication
  • Larisse C. Lucas UFC
  • George P. S. Fernandes UFC
  • Oracio C. Melo UFC
  • Miguel F. Castro UFC

Resumo


This work evaluates whether Large Language Models (LLMs) can support Explainable Artificial Intelligence (XAI) in cybersecurity, as standalone tools or combined with SHAP and LIME. Experiments with Random Forests on three intrusion detection datasets compared GPT-5 and GPT-OSS-20B explanations to SHAP and LIME. Standalone LLMs showed semantic bias and hallucinated feature importance, while SHAP/LIME-grounded prompts improved faithfulness and coherence. In a qualitative study on explainability in cybersecurity with 38 participants, LLM+SHAP/LIME explanations were preferred over SHAP/LIME-only. GPT-5 achieved the highest preference, while GPT-OSS-20B demonstrated the potential of lightweight local LLMs for XAI.

Referências

Ahmed, M., Afreen, N., Ahmed, M., Sameer, M., and Ahamed, J. (2023). An inception v3 approach for malware classification using machine learning and transfer learning. International Journal of Intelligent Networks, 4:11–18.

Ali, T. and Kostakos, P. (2023). Huntgpt: Integrating machine learning-based anomaly detection and explainable ai with large language models (llms).

Alnahdi, A. and Narain, S. (2024). Towards transparent intrusion detection: A coherence-based framework in explainable ai integrating large language models. pages 87–96.

Baral, S., Saha, S., and Haque, A. (2024). An adaptive end-to-end iot security framework using explainable ai and llms.

Barredo Arrieta, A., Díaz-Rodríguez, N., Del Ser, J., Bennetot, A., Tabik, S., Barbado, A., Garcia, S., Gil-Lopez, S., Molina, D., Benjamins, R., Chatila, R., and Herrera, F. (2020). Explainable artificial intelligence (xai): Concepts, taxonomies, opportunities and challenges toward responsible ai. Information Fusion, 58:82–115.

Basheer, N., Islam, S., Alwaheidi, M., Mouratidis, H., and Papastergiou, S. (2025). Large language model based hybrid framework for automatic vulnerability detection with explainable ai for cybersecurity enhancement. Integrated Computer-Aided Engineering, 33.

Bilal, A., Ebert, D., and Lin, B. (2025). Llms for explainable ai: A comprehensive survey.

Boonstra, L. (2025). Prompt engineering. Google, [link].

Chatzimiltis, S., Shojafar, M., Mashhadi, M. B., and Tafazolli, R. (2026). Ai-on-ran for cyber defense: An xai-llm framework for interpretable anomaly detection. IEEE Transactions on Network Science and Engineering, 13:3301–3319.

Choubisa, M., Doshi, R., Khatri, N., and Kant Hiran, K. (2022). A simple and robust approach of random forest for intrusion detection system in cyber security. In 2022 International Conference on IoT and Blockchain Technology (ICIBT), pages 1–5.

Elmaghraby, R. T., Abdel Aziem, N. M., Sobh, M. A., and Bahaa-Eldin, A. M. (2024). Encrypted network traffic classification based on machine learning. Ain Shams Engineering Journal, 15(2):102361.

Ghazal, T. M., Janjua, J. I., Abushiba, W., Ahmad, M., Ihsan, A., and Al-Dmour, N. A. (2024). Cybersecurity revolution via large language models and explainable ai. In 2024 17th International Conference on Security of Information and Networks (SIN), pages 1–6.

Giarimpampa, D., Meier, R., Bissyande, T. F., Lenders, V., and Klein, J. (2025). Exploring the role of artificial intelligence in enhancing security operations: A systematic review. ACM Comput. Surv., 58(3).

Gravereaux, S. C. and Islam, S. R. (2025). Accuracy and efficiency trade-offs in llm-based malware detection and explanation: A comparative study of parameter tuning vs. full fine-tuning.

Hozouri, A., Mirzaei, A., and Zhu, Effatparvar, M. (2025). A comprehensive survey on intrusion detection systems with advances in machine learning, deep learning and emerging cybersecurity challenges. Discover Artificial Intelligence.

Lim, B., Huerta, R., Sotelo, A., Quintela, A., and Kumar, P. (2025). Explicate: Enhancing phishing detection through explainable ai and llm-powered interpretability.

Lundberg, S. and Lee, S.-I. (2017). A unified approach to interpreting model predictions.

Mao, Q., Li, Z., Hu, X., Liu, K., Xia, X., and Sun, J. (2025). Towards explainable vulnerability detection with large language models.

Papastergiou, S., Basheer, N., Lampropoulos, K., Verrios, P., and Islam, S. (2026). Explainable ai based dynamic cybersecurity risk management for cyber insurability. Int. J. Inf. Secur., 25(1).

Pinto, A., Herrera, L.-C., Donoso, Y., and Gutierrez, J. A. (2023). Survey on intrusion detection systems based on machine learning techniques for the protection of critical infrastructure. Sensors, 23(5).

Ribeiro, M. T., Singh, S., and Guestrin, C. (2016). ”why should i trust you?”: Explaining the predictions of any classifier.

Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q., and Zhou, D. (2023). Chain-of-thought prompting elicits reasoning in large language models.

Wu, T., Fan, H., Zhu, H., You, C., Zhou, H., and Huang, X. (2022). Intrusion detection system combined enhanced random forest with smote algorithm. EURASIP Journal on Advances in Signal Processing.
Publicado
01/09/2026
FREITAS, Elton D. et al. Can LLMs Explain AI? A Study of Explainable AI in Cybersecurity Machine Learning Models. In: WORKSHOP DE CIBERSEGURANÇA EM IA - SIMPÓSIO BRASILEIRO DE CIBERSEGURANÇA (SBSEG), 26. , 2026, Armação dos Búzios/RJ. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 1024-1031. DOI: https://doi.org/10.5753/sbseg_estendido.2026.33818.

Artigos mais lidos do(s) mesmo(s) autor(es)

<< < 1 2 3 4 5 6 7 > >>