Autonomous Red Teaming for Large Language Models: A Minimally Supervised Approach Using Specialized Agents

Resumo


This paper investigates the feasibility of using autonomous AI agents to conduct Red Team evaluations of Large Language Models (LLMs) with minimal human supervision. We propose a framework of specialized agents targeting sensitive domains, including homicide, human trafficking, and zoophilia, which interact with aligned and less aligned LLMs to elicit and identify harmful outputs. The method aligns encouragingly with expert majority judgments, though the low inter-rater agreement highlights the subjectivity of harmfulness assessment, which we discuss as a central limitation. Overall, the framework shows potential as a scalable, cost-efficient support tool for LLM safety evaluation.

Referências

Amirizaniani, M. et al. (2024). Auditllm: A tool for auditing large language models using multiprobe approach. In Proceedings of the 33rd ACM International Conference on Information and Knowledge Management, pages 5174–5179.

Aykut, A. and Sezenoz, A. S. (2024). Exploring the potential of code-free custom gpts in ophthalmology: An early analysis of gpt store and user-creator guidance. Ophthalmology and Therapy, 13(10):2697–2713.

Bombieri, M., Ponzetto, S. P., and Rospocher, M. (2025). The dangerous effects of a frustratingly easy llms jailbreak attack. IEEE Access, 13:126418–126431.

Chao, P., Debenedetti, E., Robey, A., Andriushchenko, M., Croce, F., Sehwag, V., Dobriban, E., Flammarion, N., Pappas, G. J., Tramèr, F., Hassani, H., and Wong, E. (2024). JailbreakBench: An open robustness benchmark for jailbreaking large language models. In Advances in Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track, volume 37.

Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G. J., and Wong, E. (2025). Jailbreaking black box large language models in twenty queries. In 2025 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML), pages 23–42. IEEE.

Cui, Z., Li, N., and Zhou, H. (2025). A large-scale replication of scenario-based experiments in psychology and management using large language models. Nature Computational Science, pages 1–8.

Da Costa Júnior, J. F. et al. (2024). Um estudo sobre o uso da escala de likert na coleta de dados qualitativos e sua correlação com as ferramentas estatísticas. Contribuciones a las Ciencias Sociales, 17(1):360–376.

Ganguli, D., Lovitt, L., Kern, J., et al. (2022). Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned. arXiv preprint arXiv:2209.07858.

Gillespie, N., Lockey, S., Ward, T., MacDade, A., and Hassed, G. (2025). Trust, attitudes and use of artificial intelligence: A global study 2025. Technical report, The University of Melbourne and KPMG.

GitHub (2025). Redteam calvin: Repositório do código-fonte. [link].

Glukhov, D., Shumailov, I., Gal, Y., Papernot, N., and Papyan, V. (2024). Position: Fundamental limitations of llm censorship necessitate new approaches. In Proceedings of the 41st International Conference on Machine Learning (ICML).

Jing, H., Wei, J., Wei, W., Tan, Y., Zheng, B., and Yao, Q. (2025). Multi-step adaptive attack agent: A dynamic approach for jailbreaking large language models. In 2025 IEEE 37th International Conference on Tools with Artificial Intelligence (ICTAI), pages 139–148. IEEE.

Kumar, P. et al. (2024). Refusal-trained llms are easily jailbroken as browser agents. arXiv preprint arXiv:2410.13886.

Landauer, M. et al. (2024). Red team redemption: A structured comparison of open-source tools for adversary emulation. In Proceedings of the IEEE International Conference on Trust, Security and Privacy in Computing and Communications (TrustCom), Chengdu, China.

Liu, X. et al. (2024). Autodan-turbo: A lifelong agent for strategy self-exploration to jailbreak llms. arXiv preprint arXiv:2410.05295.

Mazeika, M., Phan, L., Yin, X., Zou, A., Wang, Z., Mu, N., Sakhaee, E., Li, N., Basart, S., Li, B., Forsyth, D., and Hendrycks, D. (2024). HarmBench: A standardized evaluation framework for automated red teaming and robust refusal. In Proceedings of the 41st International Conference on Machine Learning (ICML).

Mehrotra, A., Zampetakis, M., Kassianik, P., Nelson, B., Anderson, H., Singer, Y., and Karbasi, A. (2024). Tree of attacks: Jailbreaking black-box LLMs automatically. In Advances in Neural Information Processing Systems (NeurIPS), volume 37, pages 61065–61105.

Namiot, D. and Zubareva, E. (2023). About ai red team. International Journal of Open Information Technologies, 11(10):130–139.

Park, L. H. and Kwon, T. (2025). Red-teaming llms with token control score: Efficient, universal, and transferable jailbreaks. In 2025 28th International Symposium on Research in Attacks, Intrusions and Defenses (RAID), pages 629–647. IEEE.

Perez, E., Huang, S., Song, F., Cai, T., Ring, R., Aslanides, J., Glaese, A., McAleese, N., and Irving, G. (2022). Red teaming language models with language models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 3419–3448.

Russinovich, M., Salem, A., and Eldan, R. (2025). Great, now write an article about that: The crescendo multi-turn llm jailbreak attack. In Proceedings of the 34th USENIX Security Symposium.

Shen, X., Chen, Z., Backes, M., Shen, Y., and Zhang, Y. (2024). “do anything now”: Characterizing and evaluating in-the-wild jailbreak prompts on large language models. In Proceedings of the 2024 ACM SIGSAC Conference on Computer and Communications Security (CCS).

Stanford Institute for Human-Centered AI (2025). The ai index 2025 annual report. AI Index Steering Committee, Stanford University.

Sun, X. et al. (2024). Multi-turn context jailbreak attack on large language models from first principles. arXiv preprint arXiv:2408.04686.

Wang, J. et al. (2024). Software testing with large language models: Survey, landscape, and vision. IEEE Transactions on Software Engineering, 50(4):911–936.

Zeng, Y. et al. (2024). Autodefense: Multi-agent llm defense against jailbreak attacks. arXiv preprint arXiv:2403.04783.

Zhou, A. et al. (2025). Autoredteamer: Autonomous red teaming with lifelong attack integration. arXiv preprint arXiv:2503.15754.

Zou, A., Wang, Z., Carlini, N., Nasr, M., Kolter, J. Z., and Fredrikson, M. (2023). Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043.
Publicado
01/09/2026
MEDEIROS, Débora et al. Autonomous Red Teaming for Large Language Models: A Minimally Supervised Approach Using Specialized Agents. In: SIMPÓSIO BRASILEIRO DE CIBERSEGURANÇA (SBSEG), 26. , 2026, Armação dos Búzios/RJ. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 222-237. DOI: https://doi.org/10.5753/sbseg.2026.22972.