A Comparison of Commonly Used Automatic Evaluation Metrics for Open-Ended Tasks in Portuguese
Resumo
Open-ended tasks in Natural Language processing are characterized by the absence of a fixed set of outputs or short, predetermined responses. Typical examples include open-domain question answering and summarization. The automatic evaluation of these tasks is challenging, and existing metrics suffer from important limitations. Thus, many works have already been proposed to evaluate the performance of such metrics. However, the great majority of these works were conducted in English, leaving a gap for analyses in other languages, such as Portuguese. This work investigates the performance of some of the most popular automated intrinsic evaluation metrics for open-ended tasks by analyzing their correlation with human judgment. Using both the Summarization and Question Answering tasks, this study compares traditional n-gram-based metrics with metrics based on contextual embeddings generated by Pre-trained Language Models (PLMs), performing all analyses using models and datasets available for Portuguese. Results indicate that while PLM-based metrics generally show slightly higher correlation with human perception, their performance is highly reliant on the quality of the models used. Still, none of the evaluated metrics displayed a strong correlation with human judgments, suggesting the need for more robust metrics.Referências
Banerjee, S. and Lavie, A. (2005). METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures, pages 65–72.
Deutsch, D., Dror, R., and Roth, D. (2021). A statistical analysis of summarization evaluation metrics using resampling methods. Transactions of the Association for Computational Linguistics, 9:1132–1146.
Deutsch, D., Dror, R., and Roth, D. (2022). Re-examining system-level correlations of automatic summarization evaluation metrics. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 6038–6052.
Fabbri, A. R., Kryściński, W., McCann, B., et al. (2021). SummEval: Re-evaluating summarization evaluation. Transactions of the Association for Computational Linguistics, 9:391–409.
Farea, A., Yang, Z., Duong, K., et al. (2025). Evaluation of question answering systems: Complexity of judging a natural language. ACM Comput. Surv.
Hasan, T., Bhattacharjee, A., Islam, M. S., et al. (2021). Xl-sum: Large-scale multilingual abstractive summarization for 44 languages. arXiv preprint arXiv:2106.13822.
Junqueira, J. d. R. and Moreira, V. P. (2026). The inadequacy of automatic evaluation metrics in question answering: A case-study in Portuguese. In Souza, M., de Dios-Flores, I., Santos, D., Freitas, L., Souza, J. W. d. C., and Ribeiro, E., editors, Proceedings of the 17th International Conference on Computational Processing of Portuguese (PROPOR 2026) - Vol. 1, pages 551–561, Salvador, Brazil. Association for Computational Linguistics.
Koehn, P. and Monz, C. (2006). Manual and automatic evaluation of machine translation between European languages. In Proceedings on the Workshop on Statistical Machine Translation, pages 102–121.
Krippendorff, K. (2011). Computing krippendorff’s alpha-reliability.
Krishna, K., Roy, A., and Iyyer, M. (2021). Hurdles to progress in long-form question answering. In Proceedings of the 2021 Conference of the North American Chapter of the ACL, pages 4940–4957.
Leite, B., Osório, T. F., and Cardoso, H. L. (2024). Fairytaleqa translated: Enabling educational question and answer generation in less-resourced languages. In Technology Enhanced Learning for Inclusive and Equitable Quality Education, pages 222–236.
Lewis, P., Perez, E., Piktus, A., et al. (2020). Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in neural information processing systems, 33:9459–9474.
Lin, C.-Y. (2004). ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pages 74–81.
Liu, C.-W., Lowe, R., Serban, I., et al. (2016). How NOT to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation. In Su, J., Duh, K., and Carreras, X., editors, Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2122–2132, Austin, Texas. Association for Computational Linguistics.
Liu, Y., Iter, D., Xu, Y., et al. (2023). G-eval: NLG evaluation using gpt-4 with better human alignment. In Bouamor, H., Pino, J., and Bali, K., editors, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 2511–2522, Singapore. Association for Computational Linguistics.
Nallapati, R., Zhou, B., dos Santos, C., et al. (2016). Abstractive text summarization using sequence-to-sequence RNNs and beyond. In Proceedings of the 20th SIGNLL Conference on Computational Natural Language Learning, pages 280–290.
NIST, N. (2006). Nist 2006 machine translation evaluation official results.
Paiola, P. H., de Rosa, G. H., and Papa, J. P. (2022). Deep learning-based abstractive summarization for brazilian portuguese texts. In BRACIS 2022: Intelligent Systems, pages 479–493.
Papineni, K., Roukos, S., Ward, T., et al. (2002). Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311–318.
Rajpurkar, P., Zhang, J., Lopyrev, K., et al. (2016). Squad: 100,000+ questions for machine comprehension of text. arXiv preprint arXiv:1606.05250.
Schober, P., Boer, C., and Schwarte, L. A. (2018). Correlation coefficients: appropriate use and interpretation. Anesthesia & analgesia, 126(5):1763–1768.
Shannon, C. E. (1948). A mathematical theory of communication. The Bell system technical journal, 27(3):379–423.
Souza, F., Nogueira, R., and Lotufo, R. (2020). Bertimbau: Pretrained bert models for brazilian portuguese. In Intelligent Systems: 9th Brazilian Conference, BRACIS 2020, page 403–417.
Team, G., Kamath, A., Ferret, J., et al. (2025). Gemma 3 technical report. arXiv preprint arXiv:2503.19786.
Xu, Y., Wang, D., Yu, M., et al. (2022). Fantastic questions and where to find them: Fairytaleqa – an authentic dataset for narrative comprehension.
Yuan, W., Neubig, G., and Liu, P. (2021). Bartscore: Evaluating generated text as text generation. Advances in neural information processing systems, 34:27263–27277.
Zhang, T., Kishore, V., Wu, F., et al. (2019). Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675.
Zhao, W., Peyrard, M., Liu, F., et al. (2019). Moverscore: Text generation evaluating with contextualized embeddings and earth mover distance. arXiv preprint arXiv:1902.02622.
Deutsch, D., Dror, R., and Roth, D. (2021). A statistical analysis of summarization evaluation metrics using resampling methods. Transactions of the Association for Computational Linguistics, 9:1132–1146.
Deutsch, D., Dror, R., and Roth, D. (2022). Re-examining system-level correlations of automatic summarization evaluation metrics. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 6038–6052.
Fabbri, A. R., Kryściński, W., McCann, B., et al. (2021). SummEval: Re-evaluating summarization evaluation. Transactions of the Association for Computational Linguistics, 9:391–409.
Farea, A., Yang, Z., Duong, K., et al. (2025). Evaluation of question answering systems: Complexity of judging a natural language. ACM Comput. Surv.
Hasan, T., Bhattacharjee, A., Islam, M. S., et al. (2021). Xl-sum: Large-scale multilingual abstractive summarization for 44 languages. arXiv preprint arXiv:2106.13822.
Junqueira, J. d. R. and Moreira, V. P. (2026). The inadequacy of automatic evaluation metrics in question answering: A case-study in Portuguese. In Souza, M., de Dios-Flores, I., Santos, D., Freitas, L., Souza, J. W. d. C., and Ribeiro, E., editors, Proceedings of the 17th International Conference on Computational Processing of Portuguese (PROPOR 2026) - Vol. 1, pages 551–561, Salvador, Brazil. Association for Computational Linguistics.
Koehn, P. and Monz, C. (2006). Manual and automatic evaluation of machine translation between European languages. In Proceedings on the Workshop on Statistical Machine Translation, pages 102–121.
Krippendorff, K. (2011). Computing krippendorff’s alpha-reliability.
Krishna, K., Roy, A., and Iyyer, M. (2021). Hurdles to progress in long-form question answering. In Proceedings of the 2021 Conference of the North American Chapter of the ACL, pages 4940–4957.
Leite, B., Osório, T. F., and Cardoso, H. L. (2024). Fairytaleqa translated: Enabling educational question and answer generation in less-resourced languages. In Technology Enhanced Learning for Inclusive and Equitable Quality Education, pages 222–236.
Lewis, P., Perez, E., Piktus, A., et al. (2020). Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in neural information processing systems, 33:9459–9474.
Lin, C.-Y. (2004). ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pages 74–81.
Liu, C.-W., Lowe, R., Serban, I., et al. (2016). How NOT to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation. In Su, J., Duh, K., and Carreras, X., editors, Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2122–2132, Austin, Texas. Association for Computational Linguistics.
Liu, Y., Iter, D., Xu, Y., et al. (2023). G-eval: NLG evaluation using gpt-4 with better human alignment. In Bouamor, H., Pino, J., and Bali, K., editors, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 2511–2522, Singapore. Association for Computational Linguistics.
Nallapati, R., Zhou, B., dos Santos, C., et al. (2016). Abstractive text summarization using sequence-to-sequence RNNs and beyond. In Proceedings of the 20th SIGNLL Conference on Computational Natural Language Learning, pages 280–290.
NIST, N. (2006). Nist 2006 machine translation evaluation official results.
Paiola, P. H., de Rosa, G. H., and Papa, J. P. (2022). Deep learning-based abstractive summarization for brazilian portuguese texts. In BRACIS 2022: Intelligent Systems, pages 479–493.
Papineni, K., Roukos, S., Ward, T., et al. (2002). Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pages 311–318.
Rajpurkar, P., Zhang, J., Lopyrev, K., et al. (2016). Squad: 100,000+ questions for machine comprehension of text. arXiv preprint arXiv:1606.05250.
Schober, P., Boer, C., and Schwarte, L. A. (2018). Correlation coefficients: appropriate use and interpretation. Anesthesia & analgesia, 126(5):1763–1768.
Shannon, C. E. (1948). A mathematical theory of communication. The Bell system technical journal, 27(3):379–423.
Souza, F., Nogueira, R., and Lotufo, R. (2020). Bertimbau: Pretrained bert models for brazilian portuguese. In Intelligent Systems: 9th Brazilian Conference, BRACIS 2020, page 403–417.
Team, G., Kamath, A., Ferret, J., et al. (2025). Gemma 3 technical report. arXiv preprint arXiv:2503.19786.
Xu, Y., Wang, D., Yu, M., et al. (2022). Fantastic questions and where to find them: Fairytaleqa – an authentic dataset for narrative comprehension.
Yuan, W., Neubig, G., and Liu, P. (2021). Bartscore: Evaluating generated text as text generation. Advances in neural information processing systems, 34:27263–27277.
Zhang, T., Kishore, V., Wu, F., et al. (2019). Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675.
Zhao, W., Peyrard, M., Liu, F., et al. (2019). Moverscore: Text generation evaluating with contextualized embeddings and earth mover distance. arXiv preprint arXiv:1902.02622.
Publicado
19/10/2026
Como Citar
FAÉ, Eduardo D.; BALREIRA, Dennis Giovani; MOREIRA, Viviane Pereira.
A Comparison of Commonly Used Automatic Evaluation Metrics for Open-Ended Tasks in Portuguese. In: SIMPÓSIO BRASILEIRO DE TECNOLOGIA DA INFORMAÇÃO E DA LINGUAGEM HUMANA (STIL), 17. , 2026, Cuiabá/MT.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2026
.
p. 126-138.
DOI: https://doi.org/10.5753/stil.2026.26578.
