Towards SLMs Usage in Contextualized Question Generation for Security Information Audits: A Study Using the G-Eval Framework
Resumo
Conducting information security audits today is no trivial task. One key aspect is the quantity and quality of controls to be selected for the assessment, especially considering different types of organizations. This paper evaluates the ability of Small Language Models (SLMs) with Chain-of-Thought to adapt generic questions from the ISO/IEC 27001:2022 and 27002:2022 control lists to specific organizational contexts (hospital, bank, retail, development, and government). Five SLMs were tested using an automatic evaluator based on the G-Eval framework (1–5 scale) and human validation (n=25), revealing that intermediate models (3B–8B) achieve acceptable quality (mean ≥ 4.0). Agreement between the automatic evaluator and human evaluators varied by up to 1.2 points in some dimensions, and limitations were also observed in Portuguese. The results suggest feasibility for supporting security audits with human supervision.Referências
Ayyaz, S. and Malik, S. M. (2024). A comprehensive study of generative adversarial networks (gan) and generative pre-trained transformers (gpt) in cybersecurity. In 2024 Sixth International Conference on Intelligent Computing in Data Sciences (ICDS), pages 1–8. IEEE.
Brezavšček, A. and Baggia, A. (2025). Recent trends in information and cyber security maturity assessment: A systematic literature review. Systems, 13(1).
Chrissis, M., Konrad, M., and Shrum, S. (2011). CMMI for Development: Guidelines for Process Integration and Product Improvement. SEI Series in Software Engineering. Pearson Education.
Ferrag, M. A., Alwahedi, F., Battah, A., Cherif, B., Mechri, A., Tihanyi, N., Bisztray, T., and Debbah, M. (2025). Generative ai in cybersecurity: A comprehensive review of llm applications and vulnerabilities. Internet of Things and Cyber-Physical Systems, 5:1–46.
Fotoh, L. E. and Mugwira, T. (2025). Exploring large language models in external audits: Implications and ethical considerations. International Journal of Accounting Information Systems, 56:100748.
Gong, C., Li, Z., and Li, X. (2026). Information security based on llm approaches: A review.
ISO (2022a). Iso/iec 27001: 2022, information security, cybersecurity and privacy protection–information security management systems, requirements.
ISO (2022b). Iso/iec 27002: 2022, information security, cybersecurity and privacy protection information security controls.
Liu, Y., Iter, D., Xu, Y., Wang, S., Xu, R., and Zhu, C. (2023). G-eval: NLG evaluation using gpt-4 with better human alignment. In Bouamor, H., Pino, J., and Bali, K., editors, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 2511–2522, Singapore. Association for Computational Linguistics.
Lu, Z., Li, X., Cai, D., Yi, R., Liu, F., Zhang, X., Lane, N. D., and Xu, M. (2025). Small language models: Survey, measurements, and insights.
Magister, L. C., Mallinson, J., Adamek, J., Malmi, E., and Severyn, A. (2023). Teaching small language models to reason. In Rogers, A., Boyd-Graber, J., and Okazaki, N., editors, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 1773–1781, Toronto, Canada. Association for Computational Linguistics.
Nguyen, C. V., Shen, X., Aponte, R., Xia, Y., Basu, S., Hu, Z., Chen, J., Parmar, M., Kunapuli, S., Barrow3, J., Wu, J., Singh, A., Wang, Y., Gu, J., K. Ahmed, N., Lipka, N., Zhang, R., Chen, X., Yu, T., Kim, S., Deilamsalehy, H., Park, N., Rimer, M., Zhang, Z., Yang, H., Mathur, P., Wu, G., Dernoncourt, F., Rossi, R., and Nguyen, T. H. (2025). A survey on small language models. In Angelova, G., Kunilovskaya, M., Escribe, M., and Mitkov, R., editors, Proceedings of the 15th International Conference on Recent Advances in Natural Language Processing - Natural Language Processing in the Generative AI Era, pages 807–821, Varna, Bulgaria. INCOMA Ltd., Shoumen, Bulgaria.
Rabii, A., Saliha, A., Khadija, O. T., and Roudies, O. (2020). Information and cyber security maturity models: a systematic literature review. Information & Computer Security, ahead-of-print.
Ranaldi, L. and Freitas, A. (2024). Aligning large and small language models via chain-of-thought reasoning. In Graham, Y. and Purver, M., editors, Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1812–1827, St. Julian’s, Malta. Association for Computational Linguistics.
Santos, M. (2024). Sectum: O chatbot de segurança da informação. In Anais Estendidos do XXIV Simpósio Brasileiro de Segurança da Informação e de Sistemas Computacionais, pages 161–168, Porto Alegre, RS, Brasil. SBC.
Schmitz, C., Schmid, M., Harborth, D., and Pape, S. (2021). Maturity level assessments of information security controls: An empirical analysis of practitioners assessment capabilities. Computers & Security, 108:102306.
Spanos, G. and Angelis, L. (2016). The impact of information security events to the stock market: A systematic literature review. Computers & Security, 58:216–229.
Yigit, Y., Buchanan, W. J., Tehrani, M. G., and Maglaras, L. (2024). Review of generative ai methods in cybersecurity.
Zhang, J., Bu, H., Wen, H., Liu, Y., Fei, H., Xi, R., Li, L., Yang, Y., Zhu, H., and Meng, D. (2025). When llms meet cybersecurity: A systematic literature review. Cybersecurity, 8(1):55.
Brezavšček, A. and Baggia, A. (2025). Recent trends in information and cyber security maturity assessment: A systematic literature review. Systems, 13(1).
Chrissis, M., Konrad, M., and Shrum, S. (2011). CMMI for Development: Guidelines for Process Integration and Product Improvement. SEI Series in Software Engineering. Pearson Education.
Ferrag, M. A., Alwahedi, F., Battah, A., Cherif, B., Mechri, A., Tihanyi, N., Bisztray, T., and Debbah, M. (2025). Generative ai in cybersecurity: A comprehensive review of llm applications and vulnerabilities. Internet of Things and Cyber-Physical Systems, 5:1–46.
Fotoh, L. E. and Mugwira, T. (2025). Exploring large language models in external audits: Implications and ethical considerations. International Journal of Accounting Information Systems, 56:100748.
Gong, C., Li, Z., and Li, X. (2026). Information security based on llm approaches: A review.
ISO (2022a). Iso/iec 27001: 2022, information security, cybersecurity and privacy protection–information security management systems, requirements.
ISO (2022b). Iso/iec 27002: 2022, information security, cybersecurity and privacy protection information security controls.
Liu, Y., Iter, D., Xu, Y., Wang, S., Xu, R., and Zhu, C. (2023). G-eval: NLG evaluation using gpt-4 with better human alignment. In Bouamor, H., Pino, J., and Bali, K., editors, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 2511–2522, Singapore. Association for Computational Linguistics.
Lu, Z., Li, X., Cai, D., Yi, R., Liu, F., Zhang, X., Lane, N. D., and Xu, M. (2025). Small language models: Survey, measurements, and insights.
Magister, L. C., Mallinson, J., Adamek, J., Malmi, E., and Severyn, A. (2023). Teaching small language models to reason. In Rogers, A., Boyd-Graber, J., and Okazaki, N., editors, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 1773–1781, Toronto, Canada. Association for Computational Linguistics.
Nguyen, C. V., Shen, X., Aponte, R., Xia, Y., Basu, S., Hu, Z., Chen, J., Parmar, M., Kunapuli, S., Barrow3, J., Wu, J., Singh, A., Wang, Y., Gu, J., K. Ahmed, N., Lipka, N., Zhang, R., Chen, X., Yu, T., Kim, S., Deilamsalehy, H., Park, N., Rimer, M., Zhang, Z., Yang, H., Mathur, P., Wu, G., Dernoncourt, F., Rossi, R., and Nguyen, T. H. (2025). A survey on small language models. In Angelova, G., Kunilovskaya, M., Escribe, M., and Mitkov, R., editors, Proceedings of the 15th International Conference on Recent Advances in Natural Language Processing - Natural Language Processing in the Generative AI Era, pages 807–821, Varna, Bulgaria. INCOMA Ltd., Shoumen, Bulgaria.
Rabii, A., Saliha, A., Khadija, O. T., and Roudies, O. (2020). Information and cyber security maturity models: a systematic literature review. Information & Computer Security, ahead-of-print.
Ranaldi, L. and Freitas, A. (2024). Aligning large and small language models via chain-of-thought reasoning. In Graham, Y. and Purver, M., editors, Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1812–1827, St. Julian’s, Malta. Association for Computational Linguistics.
Santos, M. (2024). Sectum: O chatbot de segurança da informação. In Anais Estendidos do XXIV Simpósio Brasileiro de Segurança da Informação e de Sistemas Computacionais, pages 161–168, Porto Alegre, RS, Brasil. SBC.
Schmitz, C., Schmid, M., Harborth, D., and Pape, S. (2021). Maturity level assessments of information security controls: An empirical analysis of practitioners assessment capabilities. Computers & Security, 108:102306.
Spanos, G. and Angelis, L. (2016). The impact of information security events to the stock market: A systematic literature review. Computers & Security, 58:216–229.
Yigit, Y., Buchanan, W. J., Tehrani, M. G., and Maglaras, L. (2024). Review of generative ai methods in cybersecurity.
Zhang, J., Bu, H., Wen, H., Liu, Y., Fei, H., Xi, R., Li, L., Yang, Y., Zhu, H., and Meng, D. (2025). When llms meet cybersecurity: A systematic literature review. Cybersecurity, 8(1):55.
Publicado
01/09/2026
Como Citar
SOUZA, Vinícius F.; GOMES, Diego R.; MOTA, Rafael P. B.; AIRES, Fernando.
Towards SLMs Usage in Contextualized Question Generation for Security Information Audits: A Study Using the G-Eval Framework. In: WORKSHOP DE TRABALHOS DE INICIAÇÃO CIENTÍFICA E DE GRADUAÇÃO - SIMPÓSIO BRASILEIRO DE CIBERSEGURANÇA (SBSEG), 26. , 2026, Armação dos Búzios/RJ.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2026
.
p. 738-749.
DOI: https://doi.org/10.5753/sbseg_estendido.2026.29353.
