Um Framework Metodológico para Avaliação de Segurança em Código Gerado por LLMs

  • Augusto César Graeml UTFPR
  • Thiago O. Bispo de Jesus UTFPR
  • Daniel F. Pigatto UTFPR
  • Juliana de Santi UTFPR
  • Ana Cristina B. Kochem Vendramin UTFPR

Resumo


Os Grandes Modelos de Linguagem têm sido amplamente utilizados no desenvolvimento de software. Entretanto, a adoção de modelos compactos, motivada pela redução de custos e latência, levanta preocupações de segurança. Este trabalho propõe um framework para auditoria de código gerado por LLMs compactos, utilizando prompts baseados no MITRE Top 25 CWE e análise SAST com CodeQL. Os resultados mostram vulnerabilidades em mais de 40% dos cenários, principalmente relacionadas à validação inadequada de entradas e à injeção de comandos. Os achados demonstram a eficácia do framework na validação contínua da segurança de código gerado por inteligência artificial.

Referências

Asare, O., Nagappan, M., and Asokan, N. (2022). Is GitHub’s Copilot as Bad as Humans at Introducing Vulnerabilities in Code? DOI: 10.48550/arXiv.2204.04741.

Ayyamperumal, S. G. and Ge, L. (2024). Current state of llm risks and ai guardrails. DOI: 10.48550/arXiv.2406.12934.

DeepSeek (2025). DeepSeek. [link].

Gasiba, T. E., Oguzhan, K., Kessba, I., Lechner, U., and Pinto-Albuquerque, M. (2023). I’m Sorry Dave, I’m Afraid I Can’t Fix Your Code: On ChatGPT, CyberSecurity, and Secure Coding. In OpenAccess Series in Informatics, volume 112. Schloss Dagstuhl-Leibniz-Zentrum fur Informatik GmbH, Dagstuhl Publishing.

GitHub Inc. (2024). CodeQL: The static analysis engine powering GitHub Advanced Security. [link]. Documentação técnica oficial.

Hamer, S., D’Amorim, M., and Williams, L. (2024). Just another copy and paste? comparing the security vulnerabilities of ChatGPT generated code and stackoverflow answers. In IEEE Symposium on Security and Privacy Workshops, pages 87–94.

Han, T., Wang, Z., Fang, C., Zhao, S., Ma, S., and Chen, Z. (2024). Token-budget-aware llm reasoning. DOI: 10.48550/arXiv.2412.18547.

Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., and Narasimhan, K. R. (2024). SWE-bench: Can language models resolve real-world GitHub issues? In International Conference on Learning Representations. [link].

Khoury, R., Avila, A. R., Brunelle, J., and Camara, B. M. (2023). How Secure is Code Generated by ChatGPT? DOI: 10.48550/arXiv.2304.09655.

Liu, Z., Tang, Y., Luo, X., Zhou, Y., and Zhang, L. F. (2024). No Need to Lift a Finger Anymore? Assessing the Quality of Code Generation by ChatGPT. IEEE Transactions on Software Engineering, 50:1548–1584.

Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., Alon, U., Dziri, N., Prabhumoye, S., Yang, Y., et al. (2023). Self-refine: Iterative refinement with self-feedback. DOI: 10.48550/arXiv.2303.17651.

Manik, M. M. H. (2025). Chatgpt vs. deepseek: A comparative study on ai-based code generation. DOI: 10.48550/arXiv.2502.18467.

MITRE (2024). CWE - CWE Top 25 Most Dangerous Software Weaknesses — cwe.mitre.org. [link].

Mohsin, A., Janicke, H., Wood, A., Sarker, I. H., Maglaras, L., and Janjua, N. (2024). Can We Trust Large Language Models Generated Code? A Framework for In-Context Learning, Security Patterns, and Code Evaluations Across Diverse LLMs. DOI: 10.48550/arXiv.2406.12513.

Negri-Ribalta, C., Geraud-Stewart, R., Sergeeva, A., and Lenzini, G. (2024). A systematic literature review on the impact of ai models on the security of code generation. Frontiers in Big Data, Volume 7. [link].

OASIS Standard (2019). Static Analysis Results Interchange Format (SARIF) Version 2.1.0. [link]. OASIS Committee Specification.

OpenAI (2025). OpenAI. [link].

Pearce, H., Ahmad, B., Tan, B., Dolan-Gavitt, B., and Karri, R. (2021). Asleep at the keyboard? assessing the security of github copilot’s code contributions. DOI: 10.48550/arXiv.2108.09293.

Perry, N., Srivastava, M., Kumar, D., and Boneh, D. (2023). Do users write more insecure code with AI assistants? In ACM SIGSAC Conference on Computer and Communications Security, pages 2785–2799. Association for Computing Machinery, Inc.

Silva, E., Quinaia, E., Silva, D., and Braga, A. (2024). Análise comparativa de ias generativas como ferramentas de apoio à programação segura. In Anais do XXIV Simpósio Brasileiro de Segurança da Informação e de Sistemas Computacionais, pages 732–738, Porto Alegre, RS, Brasil. SBC.

White, J., Fu, Q., Hays, S., Sandborn, M., Olea, C., Gilbert, H., Elnashar, A., Spencer-Smith, J., and Schmidt, D. C. (2023). A prompt pattern catalog to enhance prompt engineering with chatgpt. DOI: 10.48550/arXiv.2302.11382.
Publicado
01/09/2026
GRAEML, Augusto César; JESUS, Thiago O. Bispo de; PIGATTO, Daniel F.; SANTI, Juliana de; VENDRAMIN, Ana Cristina B. Kochem. Um Framework Metodológico para Avaliação de Segurança em Código Gerado por LLMs. In: WORKSHOP DE CIBERSEGURANÇA EM IA - SIMPÓSIO BRASILEIRO DE CIBERSEGURANÇA (SBSEG), 26. , 2026, Armação dos Búzios/RJ. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 1073-1080. DOI: https://doi.org/10.5753/sbseg_estendido.2026.33740.