Método Reprodutível para Avaliar a Resiliência de Agentes LLM a Ataques de Injeção Indireta de Prompt
Resumo
Com agentes LLM operando sobre conteúdo não confiável, a injeção indireta de prompt (IPI) sustenta riscos agênticos como o Agent Goal Hijack. Este artigo propõe um método reprodutível de ataque por IPI: instrumento de avaliação de resiliência com desenho adaptativo in-the-loop, oráculo canary, invariantes e experimento estatístico (N=30 por célula; 900 ensaios) com três defensores sob condições idênticas. As taxas foram 0% para o Opus 4.8 (0/300; IC Wilson 99% ≤2,2%), 42,0% para o Haiku 4.5 e 61,7% para o Sonnet 4.5, todas as diferenças significativas; ressalvado que atacante e defensor mais resistente são o mesmo modelo (Opus 4.8). No desenho avaliado, o vazamento exigiu receita funcional de localização; o Sonnet vazou mais que o Haiku.
Palavras-chave:
injeção indireta de prompt, agentes LLM, avaliação de resiliência, método de ataque adversarial, segurança de IA agêntica, benchmark reprodutível
Referências
Brasil (2018). Lei n. 13.709, de 14 de agosto de 2018 — lei geral de proteção de dados pessoais (LGPD). Diário Oficial da União. Acesso em: 25 jun. 2026.
Debenedetti, E., Zhang, J., Balunović, M., Beurer-Kellner, L., Fischer, M., and Tramèr, F. (2024). AgentDojo: A dynamic environment to evaluate prompt injection attacks and defenses for LLM agents. In Advances in Neural Information Processing Systems 37 (NeurIPS), pages 82895–82920. Curran Associates, Inc. Datasets and Benchmarks Track. arXiv:2406.13352. Acesso em: 25 jun. 2026.
European Union (2024). Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (AI act). Official Journal of the European Union, L series, 12 jul. 2024. Acesso em: 25 jun. 2026.
Gartner (2025). Gartner predicts 40% of enterprise apps will feature task-specific AI agents by 2026. Press release, 26 ago. 2025. Acesso em: 13 jun. 2026.
Greshake, K., Abdelnabi, S., Mishra, S., Endres, C., Holz, T., and Fritz, M. (2023). Not what you’ve signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection. In Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security (AISec), pages 79–90. ACM. arXiv:2302.12173. Acesso em: 13 maio 2026.
Guan, A. (2026). Comment and control: Prompt injection to credential theft in claude code, gemini cli, and github copilot agent. 15 abr. 2026. Acesso em: 13 maio 2026.
He, F., Zhu, T., Ye, D., Liu, B., Zhou, W., and Yu, P. S. (2025). The emerged security and privacy of LLM agent: A survey with case studies. ACM Computing Surveys, 58(6):1–36. arXiv:2407.19354. Acesso em: 25 jun. 2026.
Nasr, M., Carlini, N., Sitawarin, C., Schulhoff, S. V., Hayes, J., Ilie, M., Pluto, J., Song, S., Chaudhari, H., Shumailov, I., Thakurta, A., Xiao, K. Y., Terzis, A., and Tramèr, F. (2025). The attacker moves second: Stronger adaptive attacks bypass defenses against LLM jailbreaks and prompt injections. arXiv preprint arXiv:2510.09023. Acesso em: 26 jun. 2026.
OWASP (2024). OWASP top 10 for large language model applications: 2025 edition. OWASP GenAI Security Project. Publicado em 18 nov. 2024. Acesso em: 13 maio 2026.
OWASP (2025). OWASP top 10 for agentic applications for 2026. OWASP GenAI Security Project. Publicado em 10 dez. 2025. Acesso em: 13 maio 2026.
Wang, P., Li, X., Xiang, C., Zhang, J., Li, Y., Zhang, L., Wang, X., and Tian, Y. (2026). The landscape of prompt injection threats in LLM agents: From taxonomy to analysis. arXiv preprint arXiv:2602.10453. Acesso em: 25 jun. 2026.
Wohlin, C., Runeson, P., Höst, M., Ohlsson, M. C., Regnell, B., and Wesslén, A. (2024). Experimentation in Software Engineering. Springer, Berlin, Heidelberg, 2nd edition. Acesso em: 25 jun. 2026.
Zhan, Q., Fang, R., Panchal, H. S., and Kang, D. (2025). Adaptive attacks break defenses against indirect prompt injection attacks on LLM agents. In Findings of the Association for Computational Linguistics: NAACL 2025, pages 7116–7132, Albuquerque, New Mexico. Association for Computational Linguistics. arXiv:2503.00061. Acesso em: 26 jun. 2026.
Debenedetti, E., Zhang, J., Balunović, M., Beurer-Kellner, L., Fischer, M., and Tramèr, F. (2024). AgentDojo: A dynamic environment to evaluate prompt injection attacks and defenses for LLM agents. In Advances in Neural Information Processing Systems 37 (NeurIPS), pages 82895–82920. Curran Associates, Inc. Datasets and Benchmarks Track. arXiv:2406.13352. Acesso em: 25 jun. 2026.
European Union (2024). Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (AI act). Official Journal of the European Union, L series, 12 jul. 2024. Acesso em: 25 jun. 2026.
Gartner (2025). Gartner predicts 40% of enterprise apps will feature task-specific AI agents by 2026. Press release, 26 ago. 2025. Acesso em: 13 jun. 2026.
Greshake, K., Abdelnabi, S., Mishra, S., Endres, C., Holz, T., and Fritz, M. (2023). Not what you’ve signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection. In Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security (AISec), pages 79–90. ACM. arXiv:2302.12173. Acesso em: 13 maio 2026.
Guan, A. (2026). Comment and control: Prompt injection to credential theft in claude code, gemini cli, and github copilot agent. 15 abr. 2026. Acesso em: 13 maio 2026.
He, F., Zhu, T., Ye, D., Liu, B., Zhou, W., and Yu, P. S. (2025). The emerged security and privacy of LLM agent: A survey with case studies. ACM Computing Surveys, 58(6):1–36. arXiv:2407.19354. Acesso em: 25 jun. 2026.
Nasr, M., Carlini, N., Sitawarin, C., Schulhoff, S. V., Hayes, J., Ilie, M., Pluto, J., Song, S., Chaudhari, H., Shumailov, I., Thakurta, A., Xiao, K. Y., Terzis, A., and Tramèr, F. (2025). The attacker moves second: Stronger adaptive attacks bypass defenses against LLM jailbreaks and prompt injections. arXiv preprint arXiv:2510.09023. Acesso em: 26 jun. 2026.
OWASP (2024). OWASP top 10 for large language model applications: 2025 edition. OWASP GenAI Security Project. Publicado em 18 nov. 2024. Acesso em: 13 maio 2026.
OWASP (2025). OWASP top 10 for agentic applications for 2026. OWASP GenAI Security Project. Publicado em 10 dez. 2025. Acesso em: 13 maio 2026.
Wang, P., Li, X., Xiang, C., Zhang, J., Li, Y., Zhang, L., Wang, X., and Tian, Y. (2026). The landscape of prompt injection threats in LLM agents: From taxonomy to analysis. arXiv preprint arXiv:2602.10453. Acesso em: 25 jun. 2026.
Wohlin, C., Runeson, P., Höst, M., Ohlsson, M. C., Regnell, B., and Wesslén, A. (2024). Experimentation in Software Engineering. Springer, Berlin, Heidelberg, 2nd edition. Acesso em: 25 jun. 2026.
Zhan, Q., Fang, R., Panchal, H. S., and Kang, D. (2025). Adaptive attacks break defenses against indirect prompt injection attacks on LLM agents. In Findings of the Association for Computational Linguistics: NAACL 2025, pages 7116–7132, Albuquerque, New Mexico. Association for Computational Linguistics. arXiv:2503.00061. Acesso em: 26 jun. 2026.
Publicado
01/09/2026
Como Citar
SILVA, Marcos Paulo Pereira da; MESSIAS, Eric Hans; GONDIM, João José Costa; MELO, Laerte Peotta de.
Método Reprodutível para Avaliar a Resiliência de Agentes LLM a Ataques de Injeção Indireta de Prompt. In: WORKSHOP DE CIBERSEGURANÇA EM IA - SIMPÓSIO BRASILEIRO DE CIBERSEGURANÇA (SBSEG), 26. , 2026, Armação dos Búzios/RJ.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2026
.
p. 1057-1064.
DOI: https://doi.org/10.5753/sbseg_estendido.2026.33718.
