Unsupervised Subword Segmentation for POS Tagging in Low-Resource Agglutinative Languages
Resumo
While Part-of-speech tagging is considered a well-understood task in the literature, most work has focused on indo-european languages with fusional morphology. For low-resource agglutinative languages, such as brazillian indigenous languages, the task is commonly constrained by small annotated corpora, high lexical sparsity, and morphological patterns that are poorly represented by word-level models. This paper investigates the impact of unsupervised subword segmentation techniques on POS tagging for brazillian indigenous languages. We compare a word-level baseline with Byte Pair Encoding, Morfessor, and FlatCat on Bororo, Nheengatu, and Tupinamba corpora, including a controlled experiment on the size of the training corpus. Our findings suggest that unsupervised segmentation can reduce sparsity in low-resource POS tagging, although its benefit depends on the language, corpus size, and segmentation method.
Palavras-chave:
POS tagging, low-resource, agglutinative, tokenization, indigenous
Referências
de Alencar, L. F. (2024). A Universal Dependencies treebank for Nheengatu. In Proceedings of the 16th International Conference on Computational Processing of Portuguese, volume 2, pages 37–54.
Eskander, R., Lowry, C., Khandagale, S., Klavans, J., Polinsky, M., and Muresan, S. (2022). Unsupervised stem-based cross-lingual part-of-speech tagging for morphologically rich low-resource languages. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4061–4072, Seattle, United States. Association for Computational Linguistics.
Gerardi, F. F. (2023). UD Tupinamba-TuDeT. Universal Dependencies treebank documentation.
Gerardi, F. F., Toribio, L., and Sollberger, D. (2023). UD Bororo-BDT. Universal Dependencies treebank documentation.
Grönroos, S.-A., Virpioja, S., Smit, P., and Kurimo, M. (2014). Morfessor FlatCat: An HMM-based method for unsupervised and semi-supervised learning of morphology. In Proceedings of COLING 2014, pages 1177–1185.
Huang, Z., Xu, W., and Yu, K. (2015). Bidirectional LSTM-CRF models for sequence tagging. arXiv preprint arXiv:1508.01991.
Keren, O., Avinari, T., Tsarfaty, R., and Levy, O. (2022). Breaking character: Are subwords good enough for MRLs after all? CoRR, abs/2204.04748.
Libovický, J. and Helcl, J. (2024). Lexically grounded subword segmentation. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 7403–7420, Miami, Florida, USA. Association for Computational Linguistics.
Mager, M., Gutierrez-Vasques, X., Sierra, G., and Meza-Ruiz, I. (2018). Challenges of language technologies for the indigenous languages of the Americas. In Proceedings of the 27th International Conference on Computational Linguistics, pages 55–69, Santa Fe, New Mexico, USA. Association for Computational Linguistics.
Mager, M., Oncevay, A., Mager, E., Kann, K., and Vu, T. (2022). BPE vs. morphological segmentation: A case study on machine translation of four polysynthetic languages. In Findings of the Association for Computational Linguistics: ACL 2022, pages 961–971, Dublin, Ireland. Association for Computational Linguistics.
Özçift, A., Akarsu, K., Yumuk, F., and Söylemez, C. (2021). Advancing natural language processing (nlp) applications of morphologically rich languages with bidirectional encoder representations from transformers (bert): an empirical case study for turkish. Automatika: časopis za automatiku, mjerenje, elektroniku, računarstvo i komunikacije, 62(2):226–238.
Saleva, J. and Lignos, C. (2021). The effectiveness of morphology-aware segmentation in low-resource neural machine translation. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Student Research Workshop, pages 164–174, Online. Association for Computational Linguistics.
Sennrich, R., Haddow, B., and Birch, A. (2016). Neural machine translation of rare words with subword units. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, pages 1715–1725.
Souza, F., Nogueira, R., and Lotufo, R. (2019). Portuguese named entity recognition using bert-crf. arXiv preprint arXiv:1909.10649.
Tonja, A., Balouchzahi, F., Butt, S., Kolesnikova, O., Ceballos, H., Gelbukh, A., and Solorio, T. (2024). Nlp progress in indigenous latin american languages. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 6972–6987.
Tsarfaty, R., Seddah, D., Kübler, S., and Nivre, J. (2013). Parsing morphologically rich languages: Introduction to the special issue. Computational linguistics, 39(1):15–22.
Virpioja, S., Smit, P., Grönroos, S.-A., and Kurimo, M. (2013). Morfessor 2.0: Python implementation and extensions for Morfessor Baseline. Technical report, Aalto University.
Eskander, R., Lowry, C., Khandagale, S., Klavans, J., Polinsky, M., and Muresan, S. (2022). Unsupervised stem-based cross-lingual part-of-speech tagging for morphologically rich low-resource languages. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4061–4072, Seattle, United States. Association for Computational Linguistics.
Gerardi, F. F. (2023). UD Tupinamba-TuDeT. Universal Dependencies treebank documentation.
Gerardi, F. F., Toribio, L., and Sollberger, D. (2023). UD Bororo-BDT. Universal Dependencies treebank documentation.
Grönroos, S.-A., Virpioja, S., Smit, P., and Kurimo, M. (2014). Morfessor FlatCat: An HMM-based method for unsupervised and semi-supervised learning of morphology. In Proceedings of COLING 2014, pages 1177–1185.
Huang, Z., Xu, W., and Yu, K. (2015). Bidirectional LSTM-CRF models for sequence tagging. arXiv preprint arXiv:1508.01991.
Keren, O., Avinari, T., Tsarfaty, R., and Levy, O. (2022). Breaking character: Are subwords good enough for MRLs after all? CoRR, abs/2204.04748.
Libovický, J. and Helcl, J. (2024). Lexically grounded subword segmentation. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 7403–7420, Miami, Florida, USA. Association for Computational Linguistics.
Mager, M., Gutierrez-Vasques, X., Sierra, G., and Meza-Ruiz, I. (2018). Challenges of language technologies for the indigenous languages of the Americas. In Proceedings of the 27th International Conference on Computational Linguistics, pages 55–69, Santa Fe, New Mexico, USA. Association for Computational Linguistics.
Mager, M., Oncevay, A., Mager, E., Kann, K., and Vu, T. (2022). BPE vs. morphological segmentation: A case study on machine translation of four polysynthetic languages. In Findings of the Association for Computational Linguistics: ACL 2022, pages 961–971, Dublin, Ireland. Association for Computational Linguistics.
Özçift, A., Akarsu, K., Yumuk, F., and Söylemez, C. (2021). Advancing natural language processing (nlp) applications of morphologically rich languages with bidirectional encoder representations from transformers (bert): an empirical case study for turkish. Automatika: časopis za automatiku, mjerenje, elektroniku, računarstvo i komunikacije, 62(2):226–238.
Saleva, J. and Lignos, C. (2021). The effectiveness of morphology-aware segmentation in low-resource neural machine translation. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Student Research Workshop, pages 164–174, Online. Association for Computational Linguistics.
Sennrich, R., Haddow, B., and Birch, A. (2016). Neural machine translation of rare words with subword units. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, pages 1715–1725.
Souza, F., Nogueira, R., and Lotufo, R. (2019). Portuguese named entity recognition using bert-crf. arXiv preprint arXiv:1909.10649.
Tonja, A., Balouchzahi, F., Butt, S., Kolesnikova, O., Ceballos, H., Gelbukh, A., and Solorio, T. (2024). Nlp progress in indigenous latin american languages. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 6972–6987.
Tsarfaty, R., Seddah, D., Kübler, S., and Nivre, J. (2013). Parsing morphologically rich languages: Introduction to the special issue. Computational linguistics, 39(1):15–22.
Virpioja, S., Smit, P., Grönroos, S.-A., and Kurimo, M. (2013). Morfessor 2.0: Python implementation and extensions for Morfessor Baseline. Technical report, Aalto University.
Publicado
19/10/2026
Como Citar
PEREIRA, Jonas Oliveira; SOUSA, Lilian Teixeira de; SOUZA, Marlo.
Unsupervised Subword Segmentation for POS Tagging in Low-Resource Agglutinative Languages. In: SIMPÓSIO BRASILEIRO DE TECNOLOGIA DA INFORMAÇÃO E DA LINGUAGEM HUMANA (STIL), 17. , 2026, Cuiabá/MT.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2026
.
p. 588-593.
DOI: https://doi.org/10.5753/stil.2026.29489.
