Evaluating Portuguese Tokenizers as Morpheme Sequence in Relation to LLM Downstream Performance
Resumo
Sub-word tokenizers allow us to handle open vocabulary problems using a relatively small set of tokens. Despite its ease of use, it usually relies on a data-driven approach that does not directly employ linguistic or morphologic features for text tokenization. [Bostrom and Durrett 2020] and [Hofmann et al. 2021] demonstrate that morphemes improve LLM performance on English texts. In this work, we explore the hypothesis that tokenizers that are capable of producing a token sequence better aligned with a morpheme sequence can improve LLMs performance on Brazilian Portuguese. To evaluate how the presence of morphemes can impact LLMs for Brazilian Portuguese, we propose MorphEval-PT, a new evaluation procedure based on the psycholinguistic concept of morphological models of word processing. We build new BPE and Unigram vocabularies that are evaluated on MorphEval-PT and, in order to validate the impact of morphemes on LLMs performance, train new LLMs from scratch and evaluate its performance on downstream tasks. Consistently, BPE demonstrates a higher precision score than Unigram in its ability to represent morphemes, as well as better performance on every downstream task. These promising results indicate that accessing the ability of tokenizers to represent morphemes is an important feature in the development of LLMs for Brazilian Portuguese and that MorphEval-PT is a good and lightweight method to improve LLM performance before any pre-training.Referências
Batsuren, K., Bella, G., and Giunchiglia, F. (2021). Morphynet: a large multilingual database of derivational and inflectional morphology. In Proceedings of the 18th SIGMORPHON Workshop on Computational Research in Phonetics, Phonology, and Morphology, pages 39–48. Association for Computational Linguistics.
Beesley, K. R. and Karttunen, L. (2003). Finite-state morphology. CSLI Publications.
Bostrom, K. and Durrett, G. (2020). Byte pair encoding is suboptimal for language model pretraining. In Findings of the Association for Computational Linguistics: EMNLP 2020. Association for Computational Linguistics.
Crespo, M. C. R. M., de Souza Jeannine Rocha, M. L., Sturzeneker, M. L., Serras, F. R., de Mello, G. L., Costa, A. S., Palma, M. F., Mesquita, R. M., de Paula Guets, R., da Silva, M. M., Finger, M., de Sousa, M. C. P., Namiuti, C., and do Monte, V. M. (2023). Carolina: a general corpus of contemporary brazilian portuguese with provenance, typology and versioning information.
Creutz, M. and Lagus, K. (2002). Unsupervised discovery of morphemes. In Proceedings of the ACL-02 Workshop on Morphological and Phonological Learning, pages 21–30. Association for Computational Linguistics.
Creutz, M. and Linden, B. (2004). Morpheme segmentation gold standards for finnish and english.
de Alencar, L. F., Cuconato, B., and Rademaker, A. (2018). Morphobr: An open source large-coverage full-form lexicon for morphological analysis of portuguese. Texto Livre: Linguagem e Tecnologia, 11(3):1–25.
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2018). Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805.
Fonseca, E. R., dos Santos, L. B., Criscuolo, M., and Aluísio, S. M. (2016). Visao geral da avaliação de similaridade semântica e inferência textual. Linguamática, 8(2):3–13.
Gonçalves, C. A. (2019). Morfologia. Parábola, 1st edition.
He, P., Liu, X., Gao, J., and Chen, W. (2021). Deberta: Decoding-enhanced bert with disentangled attention. In International Conference on Learning Representations.
Hofmann, V., Pierrehumbert, J., and Schütze, H. (2021). Superbizarre is not superb: Derivational morphology improves BERT‘s interpretation of complex words. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3594–3608. Association for Computational Linguistics.
Hou, J., Katinskaia, A., Vu, A.-D., and Yangarber, R. (2023). Effects of sub-word segmentation on performance of transformer language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 7413–7425. Association for Computational Linguistics.
Kudo, T. (2018). Subword regularization: Improving neural network translation models with multiple subword candidates. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 66–75. Association for Computational Linguistics.
Leminen, A., Smolka, E., Duñabeitia, J. A., and Pliatsikas, C. (2019). Morphological processing in the brain: The good (inflection), the bad (derivation) and the ugly (compounding). Cortex, 116:4–44.
Li, X., Meng, Y., Sun, X., Han, Q., Yuan, A., and Li, J. (2019). Is word segmentation necessary for deep learning of Chinese representations? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics.
Mota, C., Carvalho, P., and Barreiro, A. (2016). Port4NooJ v3.0: Integrated linguistic resources for Portuguese NLP. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC‘16), pages 1264–1269. European Language Resources Association (ELRA).
Nouri, J. and Yangarber, R. (2016). A novel evaluation method for morphological segmentation. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC’16), pages 3102–3109. European Language Resources Association (ELRA).
Nouri, J. and Yangarber, R. (2017). Learning morphology of natural language as a finite-state grammar. In Statistical Language and Speech Processing, pages 44–57, Cham. Springer International Publishing.
Park, H. H., Zhang, K. J., Haley, C., Steimel, K., Liu, H., and Schwartz, L. (2021). Morphology matters: A multilingual language modeling analysis. Transactions of the Association for Computational Linguistics, 9:261–276.
Plag, I. (2018). Word-formation in English. Cambridge university press.
Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I. (2018). Improving language understanding by generative pre-training.
Real, L., Fonseca, E., and Gonçalo Oliveira, H. (2020). The assin 2 shared task: a quick overview. In International Conference on Computational Processing of the Portuguese Language, pages 406–412. Springer.
Santos, R., Rodrigues, J., Gomes, L., Silva, J., Branco, A., Cardoso, H. L., Osório, T. F., and Leite, B. (2024). Fostering the ecosystem of open neural encoders for portuguese with albertina pt-* family.
Schuster, M. and Nakajima, K. (2012). Japanese and korean voice search. In 2012 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 5149–5152.
Sennrich, R., Haddow, B., and Birch, A. (2016). Neural machine translation of rare words with subword units. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1715–1725. Association for Computational Linguistics.
Vargas, F., Carvalho, I., Rodrigues de Góes, F., Pardo, T., and Benevenuto, F. (2022). HateBR: A large expert annotated corpus of Brazilian Instagram comments for offensive language and hate speech detection. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 7174–7183, Marseille, France. European Language Resources Association.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. (2017). Attention is all you need. In Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc.
Yang, A., Li, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Gao, C., Huang, C., Lv, C., Zheng, C., Liu, D., Zhou, F., Huang, F., Hu, F., Ge, H., Wei, H., Lin, H., Tang, J., Yang, J., Tu, J., Zhang, J., Yang, J., Yang, J., Zhou, J., Zhou, J., Lin, J., Dang, K., Bao, K., Yang, K., Yu, L., Deng, L., Li, M., Xue, M., Li, M., Zhang, P., Wang, P., Zhu, Q., Men, R., Gao, R., Liu, S., Luo, S., Li, T., Tang, T., Yin, W., Ren, X., Wang, X., Zhang, X., Ren, X., Fan, Y., Su, Y., Zhang, Y., Zhang, Y., Wan, Y., Liu, Y., Wang, Z., Cui, Z., Zhang, Z., Zhou, Z., and Qiu, Z. (2025). Qwen3 technical report.
Beesley, K. R. and Karttunen, L. (2003). Finite-state morphology. CSLI Publications.
Bostrom, K. and Durrett, G. (2020). Byte pair encoding is suboptimal for language model pretraining. In Findings of the Association for Computational Linguistics: EMNLP 2020. Association for Computational Linguistics.
Crespo, M. C. R. M., de Souza Jeannine Rocha, M. L., Sturzeneker, M. L., Serras, F. R., de Mello, G. L., Costa, A. S., Palma, M. F., Mesquita, R. M., de Paula Guets, R., da Silva, M. M., Finger, M., de Sousa, M. C. P., Namiuti, C., and do Monte, V. M. (2023). Carolina: a general corpus of contemporary brazilian portuguese with provenance, typology and versioning information.
Creutz, M. and Lagus, K. (2002). Unsupervised discovery of morphemes. In Proceedings of the ACL-02 Workshop on Morphological and Phonological Learning, pages 21–30. Association for Computational Linguistics.
Creutz, M. and Linden, B. (2004). Morpheme segmentation gold standards for finnish and english.
de Alencar, L. F., Cuconato, B., and Rademaker, A. (2018). Morphobr: An open source large-coverage full-form lexicon for morphological analysis of portuguese. Texto Livre: Linguagem e Tecnologia, 11(3):1–25.
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2018). Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805.
Fonseca, E. R., dos Santos, L. B., Criscuolo, M., and Aluísio, S. M. (2016). Visao geral da avaliação de similaridade semântica e inferência textual. Linguamática, 8(2):3–13.
Gonçalves, C. A. (2019). Morfologia. Parábola, 1st edition.
He, P., Liu, X., Gao, J., and Chen, W. (2021). Deberta: Decoding-enhanced bert with disentangled attention. In International Conference on Learning Representations.
Hofmann, V., Pierrehumbert, J., and Schütze, H. (2021). Superbizarre is not superb: Derivational morphology improves BERT‘s interpretation of complex words. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3594–3608. Association for Computational Linguistics.
Hou, J., Katinskaia, A., Vu, A.-D., and Yangarber, R. (2023). Effects of sub-word segmentation on performance of transformer language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 7413–7425. Association for Computational Linguistics.
Kudo, T. (2018). Subword regularization: Improving neural network translation models with multiple subword candidates. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 66–75. Association for Computational Linguistics.
Leminen, A., Smolka, E., Duñabeitia, J. A., and Pliatsikas, C. (2019). Morphological processing in the brain: The good (inflection), the bad (derivation) and the ugly (compounding). Cortex, 116:4–44.
Li, X., Meng, Y., Sun, X., Han, Q., Yuan, A., and Li, J. (2019). Is word segmentation necessary for deep learning of Chinese representations? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics.
Mota, C., Carvalho, P., and Barreiro, A. (2016). Port4NooJ v3.0: Integrated linguistic resources for Portuguese NLP. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC‘16), pages 1264–1269. European Language Resources Association (ELRA).
Nouri, J. and Yangarber, R. (2016). A novel evaluation method for morphological segmentation. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC’16), pages 3102–3109. European Language Resources Association (ELRA).
Nouri, J. and Yangarber, R. (2017). Learning morphology of natural language as a finite-state grammar. In Statistical Language and Speech Processing, pages 44–57, Cham. Springer International Publishing.
Park, H. H., Zhang, K. J., Haley, C., Steimel, K., Liu, H., and Schwartz, L. (2021). Morphology matters: A multilingual language modeling analysis. Transactions of the Association for Computational Linguistics, 9:261–276.
Plag, I. (2018). Word-formation in English. Cambridge university press.
Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I. (2018). Improving language understanding by generative pre-training.
Real, L., Fonseca, E., and Gonçalo Oliveira, H. (2020). The assin 2 shared task: a quick overview. In International Conference on Computational Processing of the Portuguese Language, pages 406–412. Springer.
Santos, R., Rodrigues, J., Gomes, L., Silva, J., Branco, A., Cardoso, H. L., Osório, T. F., and Leite, B. (2024). Fostering the ecosystem of open neural encoders for portuguese with albertina pt-* family.
Schuster, M. and Nakajima, K. (2012). Japanese and korean voice search. In 2012 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 5149–5152.
Sennrich, R., Haddow, B., and Birch, A. (2016). Neural machine translation of rare words with subword units. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1715–1725. Association for Computational Linguistics.
Vargas, F., Carvalho, I., Rodrigues de Góes, F., Pardo, T., and Benevenuto, F. (2022). HateBR: A large expert annotated corpus of Brazilian Instagram comments for offensive language and hate speech detection. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 7174–7183, Marseille, France. European Language Resources Association.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. (2017). Attention is all you need. In Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc.
Yang, A., Li, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Gao, C., Huang, C., Lv, C., Zheng, C., Liu, D., Zhou, F., Huang, F., Hu, F., Ge, H., Wei, H., Lin, H., Tang, J., Yang, J., Tu, J., Zhang, J., Yang, J., Yang, J., Zhou, J., Zhou, J., Lin, J., Dang, K., Bao, K., Yang, K., Yu, L., Deng, L., Li, M., Xue, M., Li, M., Zhang, P., Wang, P., Zhu, Q., Men, R., Gao, R., Liu, S., Luo, S., Li, T., Tang, T., Yin, W., Ren, X., Wang, X., Zhang, X., Ren, X., Fan, Y., Su, Y., Zhang, Y., Zhang, Y., Wan, Y., Liu, Y., Wang, Z., Cui, Z., Zhang, Z., Zhou, Z., and Qiu, Z. (2025). Qwen3 technical report.
Publicado
19/10/2026
Como Citar
MELLO, Guilherme L.; FINGER, Marcelo.
Evaluating Portuguese Tokenizers as Morpheme Sequence in Relation to LLM Downstream Performance. In: SIMPÓSIO BRASILEIRO DE TECNOLOGIA DA INFORMAÇÃO E DA LINGUAGEM HUMANA (STIL), 17. , 2026, Cuiabá/MT.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2026
.
p. 245-257.
DOI: https://doi.org/10.5753/stil.2026.26562.
