Além da Distância Linguística: Jailbreaks Multilíngues em LLMs Especializados por Idioma

Resumo


Este artigo avalia se o sucesso de jailbreaks multilíngues contra LLMs especializados por idioma é melhor explicado por distância tipológica ou por especialização pós-treinamento. Avaliamos oito assistentes instrucionais organizados em quatro pares fraco/forte alinhados a português, italiano, sueco e búlgaro, em treze idiomas de ataque, pontuando respostas com StrongREJECT e combinando distâncias URIEL+ e métricas de especialização do BELEBELE em modelos logísticos de efeitos mistos. Os resultados não sustentam a hipótese de distância: o risco é predominantemente específico do modelo, e modelos pareados mais fortes são, em geral, mais seguros que seus pares fracos.

Referências

Abonizio, H., Pires, R., Almeida, T., and Nogueira, R. (2024). Sabiá-3 technical report. arXiv:2410.12049.

AI Sweden NLU Team (2024). Llama-3-8b-instruct. Hugging Face model card.

Alexandrov, A., Nikolov, N. I., Koychev, I., et al. (2024). BgGPT 1.0: Extending english-centric LLMs to other languages. arXiv:2412.09833.

Bandarkar, L., Liang, D., Muller, B., Artetxe, M., Shukla, S. N., Husa, D., Goyal, N., Krishnan, A., Zettlemoyer, L., and Khabsa, M. (2024). The BELEBELE benchmark: A parallel reading comprehension dataset in 122 language variants. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics.

Basile, P., Semeraro, G., Ferilli, S., et al. (2023). LLaMAntino: LLaMA 2 models for effective text generation in italian language. arXiv:2312.09993.

Deng, Y., Zhang, W., Pan, S. J., and Bing, L. (2023). Multilingual jailbreak challenges in large language models. arXiv preprint arXiv:2310.06474.

Ekgren, A., Gyllensten, A. C., Gogoulou, E., Heiman, M., Gogoulou, E., and Sahlgren, M. (2024). GPT-SW3: An autoregressive language model for the scandinavian languages. arXiv preprint arXiv:2405.12987.

Ganguli, D., Lovitt, L., Kernion, J., Askell, A., Bai, Y., Kadavath, S., Mann, B., Perez, E., Schiefer, N., Ndousse, K., Jones, A., Bowman, S. R., Chen, A., Conerly, T., DasSarma, N., Drain, D., Elhage, N., El-Showk, S., Fort, S., Hatfield-Dodds, Z., Henighan, T., Hernandez, D., Johnston, S., Kravec, S., Nanda, N., Olsson, C., Ringer, S., TranJohnson, E., Amodei, D., Brown, T., Clark, J., Joseph, N., McCandlish, S., Olah, C., and Kaplan, J. (2022). Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned. arXiv:2209.07858.

Huang, L., Jin, H., Bi, Z., Yang, P., Zhao, P., Chen, T., Wu, X., Ma, L., and Chen, H. (2025). The tower of babel revisited: Multilingual jailbreak prompts on closed-source large language models.

INSAIT Institute (2024). Bggpt-7b-instruct-v0.2. Hugging Face model card.

Khan, N., Muradoglu, S., Saleva, J., and Bjerva, J. (2025). URIEL+ : Enhancing linguistic inclusion and usability in a typological and multilingual knowledge base. In Proceedings of the 31st International Conference on Computational Linguistics.

NLLB Team (2024). Scaling neural machine translation to 200 languages. Nature, 630(8018):841–846.

Oliveira, J. L. T. (2024). Sagui-7b-instruct-v0.1. Hugging Face model card.

Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R. (2022). Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems, pages 27730–27744.

Philippy, F., Guo, S., and Haddadan, S. (2023). Identifying the correlation between language distance and cross-lingual transfer in a multilingual representation space. In Proceedings of the 5th Workshop on Research in Computational Linguistic Typology and Multilingual NLP, pages 22–29.

Polignano, M., De Gemmis, M., and Basile, P. (2026). Advanced natural-based interaction for the ITAlian language: LLaMAntino-3-ANITA. Scientific Reports, 16:4764.

Shen, L., Tan, W., Chen, S., Chen, Y., Zhang, J., Xu, H., Zheng, B., Koehn, P., and Khashabi, D. (2024). The language barrier: Dissecting safety challenges of llms in multilingual contexts. In Findings of the Association for Computational Linguistics: ACL 2024, pages 2668–2680.

Singhania, A., Dupuy, C., Mangale, S., and Namboori, A. (2025). Multi-lingual multi-turn automated red teaming for llms.

Souly, A., Lu, Q., Bowen, D., Trinh, T., Hsieh, E., Pandey, S., Abbeel, P., Svegliato, J., Emmons, S., Watkins, O., and Toyer, S. (2024). A StrongREJECT for empty jailbreaks. In Advances in Neural Information Processing Systems, volume 37.

Yong, Z.-X., Menghini, C., and Bach, S. H. (2023). Low-resource languages jailbreak GPT-4. arXiv preprint arXiv:2310.02446.
Publicado
01/09/2026
SILVA, Gabriel de Jesus Coelho da; WESTPHALL, Carlos Becker. Além da Distância Linguística: Jailbreaks Multilíngues em LLMs Especializados por Idioma. In: SIMPÓSIO BRASILEIRO DE CIBERSEGURANÇA (SBSEG), 26. , 2026, Armação dos Búzios/RJ. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 48-62. DOI: https://doi.org/10.5753/sbseg.2026.27096.

Artigos mais lidos do(s) mesmo(s) autor(es)