Comparative Analysis of Automated Red Teaming in LLMs: AttnGCG, GPTFuzzer, and PAPILLON

  • Giovanni Mandel Martignago UDESC
  • Milton Pedro Pagliuso Neto UDESC
  • Charles Christian Miers UDESC

Resumo


Automated red-teaming methods enable systematic adversarial testing of Large Language Models, yet direct comparisons across different attack paradigms remain scarce, particularly regarding the trade-off between attack effectiveness and operational cost. We present a controlled comparison of three frameworks spanning distinct strategies: (i) AttnGCG; (ii) GPTFuzzer; and (iii) PAPILLON. All three are evaluated against Llama 3.1 8B Instruct on 18 adversarial goals drawn from HarmBench, using LlamaGuard-3 as a unified judge and an equal budget of 500 iterations per goal. Our results suggest seed-pool-based mutation offers the most favorable cost-effectiveness ratio against the evaluated target model, while gradient-based suffix optimization alone proves largely ineffective against recent safety-aligned architectures within practical iteration budgets. Although PAPILLON is designed to overcome the dependence on pre-existing templates, this advantage comes at a significantly higher operational cost compared to GPTFuzzer, since each mutation is tailored to a specific goal rather than shared across all goals simultaneously.

Referências

Responsibilities Working Group 2025] AI Organizational Responsibilities Working Group (2025). Agentic ai red teaming guide. Technical report, Cloud Security Alliance (CSA).

Brokman, J., Hofman, O., Rachmil, O., Singh, I., Pahuja, V., Sabapathy, R., Priya, A., Giloni, A., Vainshtein, R., and Kojima, H. (2025). Insights and current gaps in open-source LLM vulnerability scanners: A comparative analysis. In 2025 IEEE/ACM International Workshop on Responsible AI Engineering (RAIE), pages 1–8.

Chao, P., Debenedetti, E., Robey, A., Andriushchenko, M., Croce, F., Sehwag, V., Dobriban, E., Flammarion, N., Pappas, G. J., Tramèr, F., Hassani, H., and Wong, E. (2024a). JailbreakBench: An open robustness benchmark for jailbreaking large language models. In Advances in Neural Information Processing Systems 37 (NeurIPS).

Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G. J., and Wong, E. (2024b). Jailbreaking black box large language models in twenty queries.

Deng, G., Liu, Y., Li, Y., Wang, K., Zhang, Y., Li, Z., Wang, H., Zhang, T., and Liu, Y. (2024). Masterkey: Automated jailbreaking of large language model chatbots. In Proceedings 2024 Network and Distributed System Security Symposium, NDSS 2024. Internet Society.

Gong, X., Li, M., Zhang, Y., Ran, F., Chen, C., Chen, Y., Wang, Q., and Lam, K.-Y. (2025). PAPILLON: Efficient and stealthy fuzz Testing-Powered jailbreaks for LLMs. In 34th USENIX Security Symposium (USENIX Security 25), pages 2401–2420, Seattle, WA. USENIX Association.

Jiang, F., Xu, Z., Niu, L., Xiang, Z., Ramasubramanian, B., Li, B., and Poovendran, R. (2024). Artprompt: Ascii art-based jailbreak attacks against aligned llms.

Liu, X., Xu, N., Chen, M., and Xiao, C. (2024). Autodan: Generating stealthy jailbreak prompts on aligned large language models.

Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. (2019). Roberta: A robustly optimized bert pretraining approach.

Liu, Y., Zhou, S., Lu, Y., Zhu, H., Wang, W., Lin, H., He, B., Han, X., and Sun, L. (2025). Auto-rt: Automatic jailbreak strategy exploration for red-teaming large language models.

Mazeika, M., Phan, L., Yin, X., Zou, A., Wang, Z., Mu, N., Sakhaee, E., Li, N., Basart, S., Li, B., Forsyth, D., and Hendrycks, D. (2024). Harmbench: A standardized evaluation framework for automated red teaming and robust refusal.

Mehrotra, A., Zampetakis, M., Kassianik, P., Nelson, B., Anderson, H., Singer, Y., and Karbasi, A. (2024). Tree of attacks: Jailbreaking black-box llms automatically.

OWASP (2025). OWASP Top 10 for Large Language Model Applications 2025. [link]. Accessed: May 2026.

Russinovich, M., Salem, A., and Eldan, R. (2025). Great, now write an article about that: The crescendo multi-turn llm jailbreak attack.

Vassilev, A., Oprea, A., Fordyce, A., Anderson, H., Davies, X., and Hamin, M. (2025). Adversarial machine learning: A taxonomy and terminology of attacks and mitigations. NIST Trustworthy and Responsible AI NIST AI 100-2e2025, National Institute of Standards and Technology, Gaithersburg, MD.

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. (2023). Attention is all you need.

Wang, Z., Tu, H., Mei, J., Zhao, B., Wang, Y., and Xie, C. (2024). Attngcg: Enhancing jailbreaking attacks on llms with attention manipulation.

Wei, Z., Wang, Y., Li, A., Mo, Y., and Wang, Y. (2024). Jailbreak and guard aligned language models with only few in-context demonstrations.

Yu, J., Lin, X., Yu, Z., and Xing, X. (2024). LLM-Fuzzer: Scaling assessment of large language model jailbreaks. In 33rd USENIX Security Symposium (USENIX Security 24), pages 4657–4674, Philadelphia, PA. USENIX Association.

Yuan, Y., Jiao, W., Wang, W., tse Huang, J., He, P., Shi, S., and Tu, Z. (2024). Gpt-4 is too smart to be safe: Stealthy chat with llms via cipher.

Zou, A., Wang, Z., Carlini, N., Nasr, M., Kolter, J. Z., and Fredrikson, M. (2023). Universal and transferable adversarial attacks on aligned language models.
Publicado
01/09/2026
MARTIGNAGO, Giovanni Mandel; PAGLIUSO NETO, Milton Pedro; MIERS, Charles Christian. Comparative Analysis of Automated Red Teaming in LLMs: AttnGCG, GPTFuzzer, and PAPILLON. In: WORKSHOP DE TRABALHOS DE INICIAÇÃO CIENTÍFICA E DE GRADUAÇÃO - SIMPÓSIO BRASILEIRO DE CIBERSEGURANÇA (SBSEG), 26. , 2026, Armação dos Búzios/RJ. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 431-442. DOI: https://doi.org/10.5753/sbseg_estendido.2026.29282.

Artigos mais lidos do(s) mesmo(s) autor(es)