Beyond Aggregate Scores: A Diagnostic Analysis of Portuguese LLM Evaluation Suites
Resumo
Benchmarks play a central role in evaluating large language models (LLMs), providing standardized comparisons across models, tasks, and adaptation strategies. However, aggregate benchmark scores often compress performance into a single number, obscuring important variation across tasks. This issue is especially relevant for less-resourced languages like Portuguese, where benchmarks may combine native tasks with translated datasets spanning heterogeneous categories and subareas. In this work, we examine this issue using PoETa v2, a broad Portuguese evaluation benchmark, as a diagnostic setting to evaluate three Qwen3 1.7B variants under different Portuguese adaptation settings. Our results show that, although overall scores differ by less than one percentage point across models, disaggregated analyses reveal substantially different patterns across task origin, category, and subarea. In particular, gains on native Portuguese tasks may coexist with losses on translated tasks, while different adaptation corpora redistribute performance in distinct ways. Taken together, our findings highlight the importance of complementing aggregate scores with multi-level analyses when evaluating Portuguese LLMs, particularly in the context of language-specific adaptation.Referências
Almeida, T. S., Nogueira, R., e Pedrini, H. (2025a). Curió-edu 7b: Examining data selection impacts in llm continued pretraining. arXiv preprint arXiv:2512.12770.
Almeida, T. S., Pires, R., Abonizio, H., Nogueira, R., e Pedrini, H. (2025b). PoETa v2: Toward more robust evaluation of large language models in Portuguese. arXiv preprint arXiv:2511.17808.
Assis, G., Freitas, C., e Paes, A. (2025). Exploring brazil’s llm fauna: Investigating the generative performance of large language models in portuguese. Journal of the Brazilian Computer Society, 31(1):939–971.
Chaves Rodrigues, R. (2023). Lessons learned from the evaluation of Portuguese language models. Master’s thesis, University of Malta.
Corrêa, N. K., Sen, A., Falk, S., e Fatimah, S. (2025). Tucano: Advancing neural text generation for Portuguese. Patterns, 6(1):101325.
Crespo, M. C. R. M., Rocha, M. L. d. S. J., Sturzeneker, M. L., Serras, F. R., Mello, G. L. d., Costa, A. S., Palma, M. F., Mesquita, R. M., Finger, M., Sousa, M. C. P. d., Namiuti, C., e Monte, V. M. d. (2023). Carolina: A general corpus of contemporary Brazilian Portuguese with provenance, typology and versioning information.
Cruz-Castañeda, W. A. e Amadeus, M. (2025). Large languages models in brazilian portuguese: A chronological survey. Journal of the Brazilian Computer Society, 31(1):1167–1186.
Garcia, G. L., de Medeiros, I. S., e de Oliveira, V. G. (2024). Introducing bode: A fine-tuned large language model for portuguese prompt-based tasks. arXiv preprint arXiv:2401.02909.
Gururangan, S., Marasović, A., Swayamdipta, S., Lo, K., Beltagy, I., Downey, D., e Smith, N. A. (2020). Don’t stop pretraining: Adapt language models to domains and tasks. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8342–8360.
Huang, H., Tang, T., Zhang, D., Zhao, X., Song, T., Xia, Y., e Wei, F. (2023). Not all languages are created equal in LLMs: Improving multilingual capability by cross-lingual-thought prompting. In Bouamor, H., Pino, J., e Bali, K., editors, Findings of the Association for Computational Linguistics: EMNLP 2023, pages 12365–12394, Singapore. Association for Computational Linguistics.
Larcher, C., Piau, M., Finardi, P., Gengo, P., Esposito, P., e Caridá, V. (2023). Cabrita: Closing the gap for foreign languages. arXiv preprint arXiv:2308.11878.
Paes, A., Vianna, D., e Rodrigues, J. (2026). Modelos de linguagem. In Caseli, H. M. e Nunes, M. G. V., editors, Processamento de Linguagem Natural: Conceitos, Técnicas e Aplicações em Português, volume 2, book chapter 1. BPLN, 4 edition.
Pires, R., Abonizio, H., Almeida, T. S., e Nogueira, R. (2023). Sabiá: Portuguese large language models. arXiv preprint arXiv:2304.07880.
Real, L., Carvalho, A., e Silva, A. S. d. (2026). Avaliação de grandes modelos de linguagem: Fundamentos, métodos tradicionais e desafios atuais. In Caseli, H. M. e Nunes, M. G. V., editors, Processamento de Linguagem Natural: Conceitos, Técnicas e Aplicações em Português, volume 2, book chapter 4. BPLN, 4 edition.
Rodriguez, P., Barrow, J., Hoyle, A., Lalor, J. P., Jia, R., e Boyd-Graber, J. (2021). Evaluation examples are not equally informative: How should that change NLP leaderboards? In Zong, C., Xia, F., Li, W., e Navigli, R., editors, Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 4486–4503, Online. Association for Computational Linguistics.
Severino, J. V. B. et al. (2025). Benchmarking open-source large language models on Portuguese revalida multiple-choice questions. BMJ Health Care Informatics, 32(1):e101195.
Yang, A., Li, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Gao, C., Huang, C., Lv, C., Zheng, C., Liu, D., Zhou, F., Huang, F., Hu, F., Ge, H., Wei, H., Lin, H., Tang, J., Yang, J., Tu, J., Zhang, J., Yang, J., Yang, J., Zhou, J., Zhou, J., Lin, J., Dang, K., Bao, K., Yang, K., Yu, L., Deng, L., Li, M., Xue, M., Li, M., Zhang, P., Wang, P., Zhu, Q., Men, R., Gao, R., Liu, S., Luo, S., Li, T., Tang, T., Yin, W., Ren, X., Wang, X., Zhang, X., Ren, X., Fan, Y., Su, Y., Zhang, Y., Zhang, Y., Wan, Y., Liu, Y., Wang, Z., Cui, Z., Zhang, Z., Zhou, Z., e Qiu, Z. (2025). Qwen3 technical report.
Almeida, T. S., Pires, R., Abonizio, H., Nogueira, R., e Pedrini, H. (2025b). PoETa v2: Toward more robust evaluation of large language models in Portuguese. arXiv preprint arXiv:2511.17808.
Assis, G., Freitas, C., e Paes, A. (2025). Exploring brazil’s llm fauna: Investigating the generative performance of large language models in portuguese. Journal of the Brazilian Computer Society, 31(1):939–971.
Chaves Rodrigues, R. (2023). Lessons learned from the evaluation of Portuguese language models. Master’s thesis, University of Malta.
Corrêa, N. K., Sen, A., Falk, S., e Fatimah, S. (2025). Tucano: Advancing neural text generation for Portuguese. Patterns, 6(1):101325.
Crespo, M. C. R. M., Rocha, M. L. d. S. J., Sturzeneker, M. L., Serras, F. R., Mello, G. L. d., Costa, A. S., Palma, M. F., Mesquita, R. M., Finger, M., Sousa, M. C. P. d., Namiuti, C., e Monte, V. M. d. (2023). Carolina: A general corpus of contemporary Brazilian Portuguese with provenance, typology and versioning information.
Cruz-Castañeda, W. A. e Amadeus, M. (2025). Large languages models in brazilian portuguese: A chronological survey. Journal of the Brazilian Computer Society, 31(1):1167–1186.
Garcia, G. L., de Medeiros, I. S., e de Oliveira, V. G. (2024). Introducing bode: A fine-tuned large language model for portuguese prompt-based tasks. arXiv preprint arXiv:2401.02909.
Gururangan, S., Marasović, A., Swayamdipta, S., Lo, K., Beltagy, I., Downey, D., e Smith, N. A. (2020). Don’t stop pretraining: Adapt language models to domains and tasks. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 8342–8360.
Huang, H., Tang, T., Zhang, D., Zhao, X., Song, T., Xia, Y., e Wei, F. (2023). Not all languages are created equal in LLMs: Improving multilingual capability by cross-lingual-thought prompting. In Bouamor, H., Pino, J., e Bali, K., editors, Findings of the Association for Computational Linguistics: EMNLP 2023, pages 12365–12394, Singapore. Association for Computational Linguistics.
Larcher, C., Piau, M., Finardi, P., Gengo, P., Esposito, P., e Caridá, V. (2023). Cabrita: Closing the gap for foreign languages. arXiv preprint arXiv:2308.11878.
Paes, A., Vianna, D., e Rodrigues, J. (2026). Modelos de linguagem. In Caseli, H. M. e Nunes, M. G. V., editors, Processamento de Linguagem Natural: Conceitos, Técnicas e Aplicações em Português, volume 2, book chapter 1. BPLN, 4 edition.
Pires, R., Abonizio, H., Almeida, T. S., e Nogueira, R. (2023). Sabiá: Portuguese large language models. arXiv preprint arXiv:2304.07880.
Real, L., Carvalho, A., e Silva, A. S. d. (2026). Avaliação de grandes modelos de linguagem: Fundamentos, métodos tradicionais e desafios atuais. In Caseli, H. M. e Nunes, M. G. V., editors, Processamento de Linguagem Natural: Conceitos, Técnicas e Aplicações em Português, volume 2, book chapter 4. BPLN, 4 edition.
Rodriguez, P., Barrow, J., Hoyle, A., Lalor, J. P., Jia, R., e Boyd-Graber, J. (2021). Evaluation examples are not equally informative: How should that change NLP leaderboards? In Zong, C., Xia, F., Li, W., e Navigli, R., editors, Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 4486–4503, Online. Association for Computational Linguistics.
Severino, J. V. B. et al. (2025). Benchmarking open-source large language models on Portuguese revalida multiple-choice questions. BMJ Health Care Informatics, 32(1):e101195.
Yang, A., Li, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Gao, C., Huang, C., Lv, C., Zheng, C., Liu, D., Zhou, F., Huang, F., Hu, F., Ge, H., Wei, H., Lin, H., Tang, J., Yang, J., Tu, J., Zhang, J., Yang, J., Yang, J., Zhou, J., Zhou, J., Lin, J., Dang, K., Bao, K., Yang, K., Yu, L., Deng, L., Li, M., Xue, M., Li, M., Zhang, P., Wang, P., Zhu, Q., Men, R., Gao, R., Liu, S., Luo, S., Li, T., Tang, T., Yin, W., Ren, X., Wang, X., Zhang, X., Ren, X., Fan, Y., Su, Y., Zhang, Y., Zhang, Y., Wan, Y., Liu, Y., Wang, Z., Cui, Z., Zhang, Z., Zhou, Z., e Qiu, Z. (2025). Qwen3 technical report.
Publicado
19/10/2026
Como Citar
HOCHGRAF, Gustavo; ASSIS, Gabriel; PAES, Aline; COZMAN, Fabio G..
Beyond Aggregate Scores: A Diagnostic Analysis of Portuguese LLM Evaluation Suites. In: SIMPÓSIO BRASILEIRO DE TECNOLOGIA DA INFORMAÇÃO E DA LINGUAGEM HUMANA (STIL), 17. , 2026, Cuiabá/MT.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2026
.
p. 209-217.
DOI: https://doi.org/10.5753/stil.2026.26539.
