Evaluating the Robustness of Small Language Models in Portuguese under Red Teaming with Linguistic Perturbations

  • Carlos Daniel S. Bunn UDESC
  • Matheus A. Sá UDESC
  • Milton P. Pagliuso Neto UDESC
  • Charles C. Miers UDESC
  • Marcos A. Simplicio Jr. USP

Resumo


Small Language Models are increasingly deployed in resource-constrained settings, yet their safety behavior under linguistic perturbations remains underexplored for Brazilian Portuguese. This work evaluates three open-box Small Language Models (SLMs), Llama 3.2 3B Instruct, Qwen 2.5 3B Instruct, and TucanoBR Tucano-2b4, against 480 adversarial prompts spanning six risk categories, each subjected to five character-level perturbations at three intensity levels, totaling 1,440 model runs. An automated pipeline, using Phi-4-mini-instruct as judge model, measures Objective Satisfaction Rate (OSR), Unsafe Compliance Rate (UCR), Intent Understanding Rate (IUR), and Noise Failure Rate (NFR) to separate genuine safety from refusals caused by incomprehension. Overall unsafe compliance was low across all models (0.42% – 2.71%), with TucanoBR exhibiting the highest UCR despite strong intent understanding, while heavy corruption degraded comprehension rather than eliciting unsafe outputs. Our findings indicate that literal pattern-matching filters are insufficient safeguards against surface-level adversarial evasion.

Referências

Alibaba (2024). Qwen2.5-3b-instruct. [link]. Acesso em: jun. 2026.

Boucher, N., Shumailov, I., Anderson, R., and Papernot, N. (2021). Bad characters: Imperceptible nlp attacks. arXiv preprint arXiv:2106.09898. Submitted 18 June 2021; revised 11 December 2021.

Ebrahimi, J., Rao, A., Lowd, D., and Dou, D. (2018). Hotflip: White-box adversarial examples for text classification. In Proceedings of ACL.

Hackett, W., Birch, L., Trawicki, S., Suri, N., and Garraghan, P. (2025). Bypassing llm guardrails: An empirical analysis of evasion attacks against prompt injection and jailbreak detection systems.

Huertas-García, Á., Martín, A., Huertas-Tato, J., and Camacho, D. (2024). Camouflage is all you need: Evaluating and enhancing language model robustness against camouflage adversarial attacks. arXiv preprint arXiv:2402.09874.

Li, J., Ji, S., Du, T., Li, B., and Wang, T. (2018). Textbugger: Generating adversarial text against real-world applications. arXiv preprint arXiv:1812.05271.

Meta (2024). Llama 3.2 3b instruct. [link]. Acesso em: jun. 2026.

OWASP (2025a). OWASP foundation - top 10 for large language model applications. [link]. Acesso em: jul. 2026.

OWASP (2025b). OWASP foundation AI - security and privacy guide. [link]. Acessed on: 12 jul. 2026.

Yi, S., Cong, T., He, X., Li, Q., and Song, J. (2025). Beyond the tip of efficiency: Uncovering the submerged threats of jailbreak attacks in small language models. In Findings of the Association for Computational Linguistics: ACL 2025, pages 17221–17234.

Zhang, W., Xu, H., Wang, Z., He, Z., Zhu, Z., and Ren, K. (2025). Can small language models reliably resist jailbreak attacks? a comprehensive evaluation.

Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., and Stoica, I. (2023). Judging LLM-as-a-judge with MT-bench and chatbot arena. In Advances in Neural Information Processing Systems, volume 36. Curran Associates, Inc.
Publicado
01/09/2026
BUNN, Carlos Daniel S.; SÁ, Matheus A.; PAGLIUSO NETO, Milton P.; MIERS, Charles C.; SIMPLICIO JR., Marcos A.. Evaluating the Robustness of Small Language Models in Portuguese under Red Teaming with Linguistic Perturbations. In: WORKSHOP DE CIBERSEGURANÇA EM IA - SIMPÓSIO BRASILEIRO DE CIBERSEGURANÇA (SBSEG), 26. , 2026, Armação dos Búzios/RJ. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 1036-1043. DOI: https://doi.org/10.5753/sbseg_estendido.2026.33816.

Artigos mais lidos do(s) mesmo(s) autor(es)

1 2 3 > >>