Avaliação de Risco usando Engenharia de Prompt em LLMs

  • Arthur P. Labaki UFU
  • Gustavo R. Pinto UFU
  • Rodrigo S. Miani UFU

Resumo


Modelos de Linguagem de Grande Escala (LLMs) têm sido explorados como ferramentas de apoio à avaliação de risco em cibersegurança, mas evidências anteriores indicam que tendem a subestimar riscos em comparação com profissionais da área. Este artigo avalia se estratégias de engenharia de prompt melhoram o alinhamento entre avaliações de risco produzidas por LLMs e por humanos, utilizando um cenário baseado nos CIS Controls com 18 perguntas, avaliações de 50 profissionais, cinco famílias de LLMs e 16 condições de prompt, totalizando 7.200 avaliações individuais. Os resultados mostram que prompts focados em análise de evidências, lacunas, cobertura parcial e auditabilidade reduziram o gap entre LLMs e humanos, com a melhor estratégia reduzindo o MAE em 38,9% em relação a profissionais seniores e especialistas. Ainda assim, as LLMs continuaram atribuindo notas inferiores às humanas, reforçando seu uso como ferramentas de apoio, e não como avaliadores autônomos de risco.

Referências

Benz, M. and Chatterjee, D. (2020). Calculated risk? a cybersecurity evaluation tool for smes. Business horizons, 63(4):531–540.

Bhusal, D., Alam, M. T., Nguyen, L., Mahara, A., Lightcap, Z., Frazier, R., Fieblinger, R., Torales, G. L., Blakely, B. A., and Rastogi, N. (2024). Secure: Benchmarking large language models for cybersecurity. In 2024 Annual Computer Security Applications Conference (ACSAC), pages 15–30. IEEE.

Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020). Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901.

CIS (2024). CIS Critical Security Controls v8.1. White paper. Publicado em 24 jun. 2024. Acesso em: 3 maio 2026.

Ekstedt, M., Afzal, Z., Mukherjee, P., Hacks, S., and Lagerström, R. (2023). Yet another cybersecurity risk assessment framework. International Journal of Information Security, 22(6):1713–1729.

ISC2 (2024). Global cybersecurity workforce prepares for an ai-driven world. Technical report, ISC2.

ISO/IEC (2022). ISO/IEC 27005:2022 information security, cybersecurity and privacy protection — guidance on managing information security risks. [link].

Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., Alon, U., Dziri, N., Prabhumoye, S., Yang, Y., et al. (2023). Self-refine: Iterative refinement with self-feedback. Advances in neural information processing systems, 36:46534–46594.

NIST (2012). Guide for conducting risk assessments. Technical Report NIST Special Publication 800-30 Revision 1, National Institute of Standards and Technology.

NIST (2024). The nist cybersecurity framework (csf) 2.0. Technical Report NIST CSWP 29, National Institute of Standards and Technology.

Nong, Y., Aldeen, M., Cheng, L., Hu, H., Chen, F., and Cai, H. (2024). Chain-of-thought prompting of large language models for discovering and fixing software vulnerabilities. arXiv preprint arXiv:2402.17230.

Perez, E., Huang, S., Song, F., Cai, T., Ring, R., Aslanides, J., Glaese, A., McAleese, N., and Irving, G. (2022). Red teaming language models with language models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 3419–3448.

Pinto, G. R., do Prado Labaki, A., and Miani, R. S. (2026). Evaluating the reliability of multiple large language models in risk assessment: A cis controls based approach.

Sahoo, P., Singh, A. K., Saha, S., Jain, V., Mondal, S., and Chadha, A. (2024). A systematic survey of prompt engineering in large language models: Techniques and applications. arXiv preprint arXiv:2402.07927, 1.

Tihanyi, N., Ferrag, M. A., Jain, R., Bisztray, T., and Debbah, M. (2024). Cybermetric: A benchmark dataset based on retrieval-augmented generation for evaluating llms in cybersecurity knowledge. In 2024 IEEE International Conference on Cyber Security and Resilience (CSR), pages 296–302. IEEE.

Van Haastrecht, M., Sarhan, I., Shojaifar, A., Baumgartner, L., Mallouli, W., and Spruit, M. (2021). A threat-based cybersecurity risk assessment approach addressing sme needs. In Proceedings of the 16th International Conference on Availability, Reliability and Security, pages 1–12.

Veuthey, J. R., Majid, Z. A., Hariharan, S., and Haimes, J. (2025). Meqa: A meta-evaluation framework for question & answer llm benchmarks. arXiv preprint arXiv:2504.14039.

Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al. (2022). Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824–24837.

White, J., Fu, Q., Hays, S., Sandborn, M., Olea, C., Gilbert, H., Elnashar, A., Spencer-Smith, J., and Schmidt, D. C. (2023). A prompt pattern catalog to enhance prompt engineering with chatgpt. arXiv preprint arXiv:2302.11382.

Xu, H., Wang, S., Li, N., Wang, K., Zhao, Y., Chen, K., Yu, T., Liu, Y., and Wang, H. (2024). Large language models for cyber security: A systematic literature review. ACM Transactions on Software Engineering and Methodology.

Zheng, H. S., Mishra, S., Chen, X., Cheng, H.-T., Chi, E. H., Le, Q. V., and Zhou, D. (2024). Take a step back: Evoking reasoning via abstraction in large language models. In International Conference on Learning Representations, volume 2024, pages 20279–20316.

Zhou, D., Schärli, N., Hou, L., Wei, J., Scales, N., Wang, X., Schuurmans, D., Cui, C., Bousquet, O., Le, Q., et al. (2022). Least-to-most prompting enables complex reasoning in large language models. arXiv preprint arXiv:2205.10625.
Publicado
01/09/2026
LABAKI, Arthur P.; PINTO, Gustavo R.; MIANI, Rodrigo S.. Avaliação de Risco usando Engenharia de Prompt em LLMs. In: SIMPÓSIO BRASILEIRO DE CIBERSEGURANÇA (SBSEG), 26. , 2026, Armação dos Búzios/RJ. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 238-253. DOI: https://doi.org/10.5753/sbseg.2026.26985.

Artigos mais lidos do(s) mesmo(s) autor(es)