Security by Default? Injection Risks in Routine LLM-Generated Web Components
Resumo
This study evaluates the security posture of LLM-generated web components, assessing injection vulnerability prevalence and the effectiveness of the Recursive Criticism and Improvement (RCI) mitigation. A cross-analysis of frontier models (Claude, ChatGPT, Gemini), three languages (Python, JavaScript, PHP), and three SAST tools (SonarQube, CodeQL, Semgrep) indicates that “security by design” in LLMs is still evolving. Baseline vulnerability incidence reached 85.71%. RCI self-correction proved inconsistent, mitigating flaws by 87.5% in some scenarios while increasing detections by 197.89% in others. Developers are therefore encouraged to integrate multi-tool SAST workflows with deep taint analysis rather than relying on model reflection.
Referências
Aydin, D. and Bahtiyar, S. (2025). Security vulnerabilities in ai-generated javascript: A comparative study of large language models. In 2025 IEEE International Conference on Cyber Security and Resilience (CSR), pages 200–205.
FIRST (2023). Common vulnerability scoring system version 4.0: Specification document. Available at: [link]. Access date: 2026-05-16.
Fu, Y., Liang, P., Tahir, A., Li, Z., Shahin, M., Yu, J., and Chen, J. (2025). Security weaknesses of copilot-generated code in github projects: An empirical study. ACM Trans. Softw. Eng. Methodol., 34(8).
Hamer, S., d’Amorim, M., and Williams, L. (2024). Just another copy and paste? comparing the security vulnerabilities of chatgpt generated code and stackoverflow answers. In 2024 IEEE Security and Privacy Workshops (SPW), pages 87–94.
He, J. and Vechev, M. (2023). Large language models for code: Security hardening and adversarial testing. In Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security, CCS ’23, page 1865–1879, New York, NY, USA. Association for Computing Machinery.
Jamdade, M. and Liu, Y. (2024). A pilot study on secure code generation with chatgpt for web applications. In Proceedings of the 2024 ACM Southeast Conference, ACMSE ’24, page 229–234, New York, NY, USA. Association for Computing Machinery.
Melo, R. (2025). Securing Language Models Against Vulnerability Encoding, page 1277–1278. Association for Computing Machinery, New York, NY, USA.
MITRE (2025). 2025 cwe top 25 most dangerous software weaknesses. Available at: [link]. Access date: 2026-05-19.
Nazzal, M., Khalil, I., Khreishah, A., and Phan, N. (2024). Promsec: Prompt optimization for secure generation of functional source code with large language models (llms). In Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security, CCS ’24, page 2266–2280, New York, NY, USA. Association for Computing Machinery.
OWASP (2025). A05:2025 - injection. Available at: [link]. Access date: 2026-05-19.
Pandey, M. and Kumar, S. (2026). Secure code generation with open source generative ai models. IEEE Software, pages 1–9.
Pearce, H., Ahmad, B., Tan, B., Dolan-Gavitt, B., and Karri, R. (2022). Asleep at the keyboard? assessing the security of github copilot’s code contributions. In 2022 IEEE Symposium on Security and Privacy (SP), pages 754–768, San Francisco, CA, USA. IEEE.
Peng, J., Cui, L., Huang, K., Yang, J., and Ray, B. (2025). Cweval: Outcome-driven evaluation on functionality and security of llm code generation. In 2025 IEEE/ACM International Workshop on Large Language Models for Code (LLM4Code), pages 33–40.
Pimentel, R. and Progetti, C. (2025). Avaliação comparativa do desempenho de inteligências artificiais generativas e ferramentas tradicionais na análise de código-fonte javascript. In Anais Estendidos do XXV Simpósio Brasileiro de Cibersegurança, pages 170–179, Porto Alegre, RS, Brasil. SBC.
Sousa, F., Braga, J., Sousa, A., Dantas, V., Andrade, R., and Santos, I. (2025). Mapeamento sistemático de comparativos de ferramentas de análise estática de código com foco na segurança. In Anais Estendidos do XXV Simpósio Brasileiro de Cibersegurança, pages 238–249, Porto Alegre, RS, Brasil. SBC.
Stack Overflow (2025). 2025 developer survey - most popular technologies. Available at: [link]. Access date: 2026-05-13.
Tony, C., Díaz Ferreyra, N. E., Mutas, M., Dhif, S., and Scandariato, R. (2025a). Prompting techniques for secure code generation: A systematic investigation. ACM Trans. Softw. Eng. Methodol., 34(8).
Tony, C., Iannone, E., and Scandariato, R. (2025b). Retrieve, refine, or both? using task-specific guidelines for secure python code generation. In 2025 IEEE International Conference on Software Maintenance and Evolution (ICSME), pages 368–379.
Tony, C., Mutas, M., Ferreyra, N. E. D., and Scandariato, R. (2023). Llmseceval: A dataset of natural language prompts for security evaluations. In 2023 IEEE/ACM 20th International Conference on Mining Software Repositories (MSR), pages 588–592.
Veracode (2025). October 2025 update: Genai code security report. Technical report, Veracode Inc. Available at: [link]. Access date: 2026-04-27.
