Auditing Technical Security Debt through Retrieval-Augmented Generation: A CWE, CAPEC, and STRIDE-Based Approach

  • Kleiton Ewerton de Oliveira UFJF
  • Gleiph Ghiotto Lima de Menezes UFJF
  • André Luiz de Oliveira UFJF

Resumo


Technical Security Debt (TSD) refers to latent vulnerabilities that may not immediately compromise software functionality but progressively weaken its defenses. Conventional Static Application Security Testing (SAST) tools provide the basis for vulnerability detection and remain essential in secure development workflows. However, complementary semantic layers can support contextualizing the findings in terms of data flows, attack patterns, and architectural impact. In this paper, we propose a neuro-symbolic approach for auditing TSD by integrating Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), and security ontologies. Our approach uses the CWE → CAPEC → STRIDE chain as the supporting structure to link implementationlevel weaknesses to attack patterns, and preliminary architectural threat hypotheses. We use the Open Worldwide Application Security Project (OWASP) Benchmark v1.2, we built a pipeline that enriches code fragments with ontological metadata and retrieves semantically similar examples during inference. The RAG-assisted configuration achieved 90.88% of accuracy, and a weighted F1-score of 0.9142 in CWE classification, outperforming an LLM-only baseline by 20.44% points in accuracy. A paired McNemar test confirmed that this improvement was statistically significant (p = 3.85 × 10−34). The results indicate that retrieval and ontological grounding can reduce semantic drift in vulnerability auditing, while STRIDE-based threat mapping should be interpreted as a preliminary contextualization layer rather than as a definitive threat-modeling oracle.

Referências

AbdulGhaffar, A. and Matrawy, A. (2025). Llms’ suitability for network security: A case study of stride threat modeling. arXiv preprint arXiv:2505.04101.

Costa, L. A. M., Fontão, A., Dos Santos, R. P., and Serebrenik, A. (2025). Applying generative artificial intelligence for vulnerability fixing in a proprietary software ecosystem. Journal of Systems and Software, page 112723.

Fu, Y., Wang, T., Li, S., Ding, J., Zhou, S., Jia, Z., Li, W., Jiang, Y., and Liao, X. (2024). Missconf: Llm-enhanced reproduction of configuration-triggered bugs. In Proceedings of the 2024 IEEE/ACM 46th International Conference on Software Engineering: Companion Proceedings, pages 484–495.

Izurieta, C., Rice, D., Kimball, K., and Valentien, T. (2018). A position study to investigate technical debt associated with security weaknesses. In Proceedings of the 2018 International Conference on technical debt, pages 138–142.

Khan, J. Y. and Uddin, G. (2022). Automatic detection and analysis of technical debts in peer-review documentation of r packages. In 2022 IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER), pages 765–776. IEEE.

Lenarduzzi, V., Besker, T., Taibi, D., Martini, A., and Fontana, F. (2021). A systematic literature review on technical debt prioritization: Strategies, processes, factors, and tools. J. Syst. Softw., 171:110827.

Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., et al. (2020). Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in neural information processing systems, 33:9459–9474.

Martinez, J., Quintano, N., Ruiz, A., Santamaria, I., de Soria, I. M., and Arias, J. (2021). Security debt: characteristics, product life-cycle integration and items. In 2021 IEEE/ACM International Conference on Technical Debt (TechDebt), pages 1–5. IEEE.

MITRE (2026a). Common attack pattern enumeration and classification (capec). [link]. Accessed em: 29 jan. 2026.

MITRE (2026b). Common vulnerabilities and exposures (cve). [link]. Accessed em: 29 jan. 2026.

MITRE (2026c). Common weakness enumeration (cwe). [link]. Accessed em: 29 jan. 2026.

MITRE Corporation (2026). CVE and NVD Relationship. [link] relationship.html. Accessed: 2026-05-22.

Nam, D., Macvean, A., Hellendoorn, V., Vasilescu, B., and Myers, B. (2024). Using an llm to help with code understanding. In Proceedings of the IEEE/ACM 46th International Conference on Software Engineering, pages 1–13.

Njeru, V. K., Kuria, J., and Kituku, B. (2025). Hybrid bert, cnn-bilstm model for detecting self-admitted technical debt in software development. In 2025 IEEE Conference on Computer Applications (ICCA), pages 1–6. IEEE.

OWASP Benchmark Project (2026). OWASP Benchmark for Java. [link]. Accessed: 2026-05-20.

OWASP Foundation (2026). OWASP Benchmark Project. [link]. Accessed: 2026-05-20.

Paul, D. G., Zhu, H., and Bayley, I. (2025a). Investigating the smells of llm generated code. arXiv preprint arXiv:2510.03029.

Paul, S., Alemi, F., and Macwan, R. (2025b). Llm-assisted proactive threat intelligence for automated reasoning. arXiv preprint arXiv:2504.00428.

Ren, Z., Ju, X., Chen, X., and Shen, H. (2024). Prorlearn: boosting prompt tuning-based vulnerability detection by reinforcement learning. Automated Software Engineering, 31(2):38.

Russo, B., Melegati, J., and Mock, M. (2025). Leveraging multi-task learning to improve the detection of satd and vulnerability. arXiv preprint arXiv:2501.15934.

Sharma, R., Shahbazi, R., Fard, F. H., Codabux, Z., and Vidoni, M. (2022). Self-admitted technical debt in r: detection and causes. Automated Software Engineering, 29(2):53.

Xue, Z., Zhang, X., Gao, Z., Hu, X., Gao, S., Xia, X., and Li, S. (2025). Clean code, better models: Enhancing llm performance with smell-cleaned dataset. arXiv preprint arXiv:2508.11958.

Yao, Y., Duan, J., Xu, K., Cai, Y., Sun, Z., and Zhang, Y. (2024). A survey on large language model (llm) security and privacy: The good, the bad, and the ugly. High-Confidence Computing, 4(2):100211.
Publicado
01/09/2026
OLIVEIRA, Kleiton Ewerton de; MENEZES, Gleiph Ghiotto Lima de; OLIVEIRA, André Luiz de. Auditing Technical Security Debt through Retrieval-Augmented Generation: A CWE, CAPEC, and STRIDE-Based Approach. In: SIMPÓSIO BRASILEIRO DE CIBERSEGURANÇA (SBSEG), 26. , 2026, Armação dos Búzios/RJ. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 207-221. DOI: https://doi.org/10.5753/sbseg.2026.27795.