APEX: Agentic Pentesting Execution

Resumo


Penetration testing validates whether software systems resist realistic attacks, but it remains costly, expertise-intensive, and difficult to scale. Recent advances in Large Language Models (LLMs), Retrieval-Augmented Generation (RAG), and tool-using agents enable pentesting to be reframed as a controlled agentic workflow. This paper proposes APEX, an Agentic Pentesting Execution architecture that combines policy-bounded planning, specialized agents, constrained tool execution, evidence tracking, human approval, risk triage, and remediation-oriented reporting. APEX aims to support authorized assessments with greater traceability, reproducibility, safety, and educational value.

Referências

Adam, H. M., Widyawan, and Putra, G. D. (2023). A review of penetration testing frameworks, tools, and application areas. In 2023 IEEE 7th International Conference on Information Technology, Information Systems and Electrical Engineering (ICITISEE), pages 319–324. IEEE.

Deng, G., Liu, Y., Mayoral-Vilches, V., Liu, P., Li, Y., Xu, Y., Zhang, T., Liu, Y., Pinzger, M., and Rass, S. (2024). PentestGPT: Evaluating and harnessing large language models for automated penetration testing. In Proceedings of the 33rd USENIX Security Symposium. USENIX Association.

FIRST (2023). Common Vulnerability Scoring System Version 4.0. [link]. Accessed: 2026-06-05.

Garg, D. and Bansal, N. (2021). A systematic review on penetration testing. In 2021 2nd Global Conference for Advancement in Technology (GCAT). IEEE.

Greshake, K., Abdelnabi, S., Mishra, S., Endres, C., Holz, T., and Fritz, M. (2023). Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection. In Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security. ACM.

Guan, J., Blanchard, T., Foerster, H., Jia, H., Huang, G., and Papernot, N. (2026). AI Agents Enable Adaptive Computer Worms.

Herzog, P. (2010). OSSTMM 3: The Open Source Security Testing Methodology Manual. ISECOM.

Huang, J. and Zhu, Q. (2024). PenHeal: A two-stage llm framework for automated pentesting and optimal remediation. In Proceedings of the 2nd ACM Workshop on Secure and Trustworthy Large Language Models. ACM.

International Organization for Standardization (2022). ISO/IEC 15408: Information security, cybersecurity and privacy protection – evaluation criteria for it security. [link]. Accessed: 2026-06-05.

Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., Riedel, S., and Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive nlp tasks. In Advances in Neural Information Processing Systems, volume 33, pages 9459–9474.

Li, W. G., Abuadbba, A., Moore, K., and Kim, D. D. (2026). APT-Agent: Automated penetration testing using large language models.

MITRE Corporation (2024a). Common Vulnerabilities and Exposures (CVE). [link]. Accessed: 2026-06-05.

MITRE Corporation (2024b). Common Weakness Enumeration (CWE). [link]. Accessed: 2026-06-05.

MITRE Corporation (2024c). MITRE ATT&CK. [link]. Accessed: 2026-06-05.

Muzsai, L., Imolai, D., and Lukács, A. (2024). HackSynth: Llm agent and evaluation framework for autonomous penetration testing.

OWASP Foundation (2024). Web Security Testing Guide. [link]. Accessed: 2026-06-05.

OWASP Foundation (2025). OWASP Top 10 for Large Language Model Applications. [link]. Accessed: 2026-06-05.

Penetration Testing Execution Standard (2014). The Penetration Testing Execution Standard. [link]. Accessed: 2026-06-05.

Scarfone, K., Souppaya, M., Cody, A., and Orebaugh, A. (2008). Technical guide to information security testing and assessment. Technical Report Special Publication 800-115, National Institute of Standards and Technology.

Schick, T., Dwivedi-Yu, J., Dessı̀, R., Raileanu, R., Lomeli, M., Zettlemoyer, L., Cancedda, N., and Scialom, T. (2023). Toolformer: Language models can teach themselves to use tools. In Advances in Neural Information Processing Systems, volume 36.

Shen, X., Wang, L., Li, Z., Chen, Y., Zhao, W., Sun, D., Wang, J., and Ruan, W. (2025). PentestAgent: Incorporating llm agents to automated penetration testing. In Proceedings of the 20th ACM Asia Conference on Computer and Communications Security. ACM.

Shinn, N., Cassano, F., Gopinath, A., Narasimhan, K., and Yao, S. (2023). Reflexion: Language agents with verbal reinforcement learning. In Advances in Neural Information Processing Systems, volume 36.

Spínola, R. O. (2004). Conhecendo a iso 15408: Trabalhando com segurança da informação. Engenharia de Software Magazine, 45:31–36.

Wang, L., Shi, X., Li, Z., Jiang, Y., Tan, S., Jiang, Y., Cheng, J., Chen, W., Shen, X., Li, Z., and Chen, Y. (2025). Automated penetration testing with llm agents and classical planning.

Wu, Q., Bansal, G., Zhang, J., Wu, Y., Li, B., Zhu, E., Jiang, L., Zhang, X., Zhang, S., Liu, J., Awadallah, A. H., White, R. W., Burger, D., and Wang, C. (2023). AutoGen: Enabling next-gen llm applications via multi-agent conversation.

Yang, R., Cheng, M., Deng, G., Zhang, T., Wang, J., and Xie, X. (2025). PentestEval: Benchmarking llm-based penetration testing with modular and stage-level design.

Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y. (2023). ReAct: Synergizing reasoning and acting in language models. In Proceedings of the 11th International Conference on Learning Representations.
Publicado
01/09/2026
BARCELOS, Tuigg R.; QUINCOZES, Camilla B.; BELLAGAMBA, Gabriel Pereira; SOUZA, Paulo; QUINCOZES, Silvio E.. APEX: Agentic Pentesting Execution. In: WORKSHOP DE TRABALHOS DE INICIAÇÃO CIENTÍFICA E DE GRADUAÇÃO EM ANDAMENTO - SIMPÓSIO BRASILEIRO DE CIBERSEGURANÇA (SBSEG), 26. , 2026, Armação dos Búzios/RJ. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 785-791. DOI: https://doi.org/10.5753/sbseg_estendido.2026.29810.

Artigos mais lidos do(s) mesmo(s) autor(es)

1 2 > >>