Red Teaming an Industrial Maintenance Chatbot using PAIR and nanoGCG

  • Victor Takashi Hayashi USP
  • Milton Pedro Pagliuso Neto UDESC
  • Giovanni Mandel Martignago UDESC
  • Charles Christian Miers UDESC
  • Marcos Antonio Simplicio Junior USP

Resumo


Os Large Language Models (LLMs) têm sido cada vez mais adotados em ambientes industriais para suporte a atividades de manutenção, diagnóstico e tomada de decisão operacional. No entanto, sua integração a sistemas corporativos introduz novos riscos de segurança, especialmente relacionados a ataques adversariais baseados em prompts. Este trabalho apresenta uma avaliação de um chatbot de manutenção baseado em LLM, desenvolvido para uma empresa parceira do setor industrial (anonimizada). Foram avaliadas técnicas de red teaming black-box e white-box, incluindo geração automática de jailbreaks (PAIR) e otimização de prompts baseada em gradiente (nanoGCG). Os resultados revelam entradas adversariais geradas com PAIR que conseguem contornar mecanismos de segurança e induzir respostas inseguras ou violação de políticas. No cenário experimental avaliado, o PAIR demonstrou maior aplicabilidade prática e melhor relação custo-benefício em comparação ao nanoGCG, ainda que as diferenças metodológicas entre os dois experimentos (modelo vítima e LLM juiz distintos) impeçam uma comparação direta dos valores absolutos de ASR. Por fim, são discutidas implicações de segurança para a adoção de LLMs em ambientes industriais.

Referências

Brasil (2018). Lei no 13.709, de 14 de agosto de 2018. lei geral de proteção de dados pessoais (LGPD). Diário Oficial da União. Disponível em: [link]. Acesso em: 7 maio 2026.

Chao, P., Robey, A., Dobriban, E., Hassani, H., Pappas, G. J., and Wong, E. (2025). Jailbreaking black box large language models in twenty queries.

De Oliveira, D. E. G. C., Miers, C. C., Simplicio, M. A., and Hayashi, V. T. (2025). Between generation and judgment: A cloud-native framework for adversarial evaluation of llm alignment. In 2025 IEEE International Conference on Cloud Computing Technology and Science (CloudCom), pages 1–8.

Ganguli, D., Lovitt, L., Kernion, J., Askell, A., Bai, Y., Kadavath, S., Mann, B., Perez, E., Schiefer, N., Ndousse, K., et al. (2022). Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned. arXiv preprint arXiv:2209.07858.

GraySwanAI (2025). nanogcg: A fast + lightweight implementation of the gcg algorithm in pytorch. [link]. GitHub repository. Latest release: v0.3.0 (commit 33d8418).

Munoz, G. D. L., Minnich, A. J., Lutz, R., Lundeen, R., Dheekonda, R. S. R., Chikanov, N., Jagdagdorj, B.-E., Pouliot, M., Chawla, S., Maxwell, W., et al. (2024). Pyrit: A framework for security risk identification and red teaming in generative ai system. arXiv preprint arXiv:2410.02828.

Vassilev, A., Oprea, A., Fordyce, A., Anderson, H., Davies, X., and Hamin, M. (2025). Adversarial machine learning: A taxonomy and terminology of attacks and mitigations. NIST Technical Series Publication NIST AI 100-2e2025, National Institute of Standards and Technology, Gaithersburg, MD.

Yu, G., Wang, Y., Wang, Z., Chen, L., Zheng, Y., and Liu, Y. (2026). Llms in industrial domains: A systematic review of adaptation techniques and applications from the product lifecycle perspective. Advanced Engineering Informatics, 74:104655.

Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., and Stoica, I. (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. In Advances in Neural Information Processing Systems 36 (NeurIPS 2023) Datasets and Benchmarks Track.

Zou, A., Wang, Z., Carlini, N., Nasr, M., Kolter, J. Z., and Fredrikson, M. (2023). Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043.
Publicado
01/09/2026
HAYASHI, Victor Takashi; PAGLIUSO NETO, Milton Pedro; MARTIGNAGO, Giovanni Mandel; MIERS, Charles Christian; SIMPLICIO JUNIOR, Marcos Antonio. Red Teaming an Industrial Maintenance Chatbot using PAIR and nanoGCG. In: TRILHA DE INTERAÇÃO COM A INDÚSTRIA E DE INOVAÇÃO - SIMPÓSIO BRASILEIRO DE CIBERSEGURANÇA (SBSEG), 26. , 2026, Armação dos Búzios/RJ. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 940-946. DOI: https://doi.org/10.5753/sbseg_estendido.2026.26898.

Artigos mais lidos do(s) mesmo(s) autor(es)

<< < 1 2