Extending Enhanced Universal Dependencies to Nheengatu: Semi-Automated Control and Raising Annotation
Resumo
This paper presents the first Enhanced Universal Dependencies annotation effort for a Brazilian Indigenous language, focusing on XCOMP constructions in Nheengatu, an endangered Tupian language. We describe a semiautomatic pipeline for the UD Nheengatu-CompLin treebank to annotate control and raising predicates and propagated subjects. The method combines theoretical discussions on XCOMP constructions with corpus-based analysis of Nheengatu syntax to create rule-based enhancement heuristics. The work also provides annotation guidelines and expands EUD resources for Indigenous and low-resource languages within the Universal Dependencies framework.Referências
Abeillé, A. (2024). Control and raising. In Müller, S., Abeillé, A., Borsley, R. D., and Koenig, J.-P., editors, Head-Driven Phrase Structure Grammar: The Handbook, number 9 in Empirically Oriented Theoretical Morphology and Syntax, pages 519–570. Language Science Press, Berlin, 2 edition.
Alencar, L. F. d., Silva, H. L. B., Gurgel, J. L., and Alexandre, D. M. (2025). Per aspera ad astra: Improving the automatic evaluation of a Universal Dependencies treebank for a low-resourced language. Revista de Estudos da Linguagem. In press.
Avila, M. T. (2021). Proposta de dicionário nheengatu-português. PhD thesis, Faculdade de Filosofia, Letras e Ciências Humanas da Universidade de São Paulo.
Bresnan, J., Asudeh, A., Toivonen, I., and Wechsler, S. (2015). Lexical-functional syntax, volume 16. John Wiley & Sons.
Casasnovas, A. (2006). Noções de língua geral ou nheengatú: gramática, lendas e vocabulário. Editora da Universidade Federal do Amazonas; Faculdade Salesiana Dom Bosco, Manaus, 2 edition.
Church, K. and Liberman, M. (2021). The future of computational linguistics: On beyond alchemy. Frontiers in Artificial Intelligence, 4:1–18.
de Alencar, L. F. (2023). Yauti: A tool for morphosyntactic analysis of Nheengatu within the Universal Dependencies framework. In Anais do XIV Simpósio Brasileiro de Tecnologia da Informação e da Linguagem Humana, pages 135–145, Porto Alegre, RS, Brasil. SBC.
de Alencar, L. F. (2024). A Universal Dependencies treebank for Nheengatu. In Gamallo, P., Claro, D., Teixeira, A. J. S., Real, L., García, M., Oliveira, H. G., and Amaro, R., editors, Proceedings of the 16th International Conference on Computational Processing of Portuguese, volume 2, pages 37–54, Santiago de Compostela, Galicia, Spain. Association for Computational Linguistics.
de Alencar, L. F. (2025). Enhancing a nheengatu morphosyntactic analyzer for word formation and non-standard language. In Anais do XVI Simpósio Brasileiro de Tecnologia da Informação e da Linguagem Humana, pages 13–28, Porto Alegre, RS, Brasil. SBC.
de Alencar, L. F. and Rademaker, A. (2022). Modelação da valência verbal numa gramática computacional do português no formalismo hpsg. Domínios de Lingu@gem, 16(4):1339–1400.
de Amorim, A. B. (1928). Lendas em Nheêngatu e em Portuguez. Revista do Instituto Historico e Geographico Brasileiro, 154(100):9–475. Tomo 100, vol. 154 (2º de 1926).
de Marneffe, M.-C., Ginter, F., Goldberg, Y., Hajič, J., Manning, C., McDonald, R., Nivre, J., Petrov, S., Pyysalo, S., Schuster, S., Silveira, N., Tsarfaty, R., Tyers, F., and Zeman, D. (2025). xcomp: open clausal complement. [link].
Di Felippo, A., Postali, C., Ceregatto, G., Gazana, L., Silva, E., Roman, N., and Pardo, T. (2021). Descrião preliminar do corpus DANTEStocks: Diretrizes de segmentação para anotação segundo Universal Dependencies. In Ruiz, E. E. S. and Torrent, T. T., editors, Proceedings of the 13th Brazilian Symposium in Information and Human Language Technology, pages 335–343, Porto Alegre, Brazil. Association for Computational Linguistics.
Duran, M., Lopes, L., das Graças Nunes, M., and Pardo, T. (2023). The dawn of the porttinari multigenre treebank: Introducing its journalistic portion. In Anais do XIV Simpósio Brasileiro de Tecnologia da Informação e da Linguagem Humana, pages 115–124, Porto Alegre, RS, Brasil. SBC.
Guillaume, B. and Perrier, G. (2021). Graph rewriting for enhanced Universal Dependencies. In Oepen, S., Sagae, K., Tsarfaty, R., Bouma, G., Seddah, D., and Zeman, D., editors, Proceedings of the 17th International Conference on Parsing Technologies and the IWPT 2021 Shared Task on Parsing into Enhanced Universal Dependencies (IWPT 2021), pages 175–183, Online. Association for Computational Linguistics.
Hartt, C. F. (1938). Notas sobre a língua geral, ou tupímoderno do Amazonas. Anais da Biblioteca Nacional do Rio de Janeiro, LI:305–390. [1929].
Kroeger, P. (2004). Analyzing Syntax: A Lexical-Functional Approach. Cambridge University Press, Cambridge.
Navarro, E. d. A. (2012). O último refúgio da língua geral no Brasil. Estudos Avançados, 26(76):245–254.
Nivre, J. et al. (2016). Universal Dependencies v1: A multilingual treebank collection. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC’16), pages 1659–1666, Portorož, Slovenia. European Language Resources Association (ELRA).
Pagano, A. S., Duran, M. S., and Pardo, T. A. S. (2023). Enhanced dependencies para o português brasileiro. In Pardo, T. A. S., Duran, M. S., and Lopes, L., editors, Proceedings of the 2nd Edition of the Universal Dependencies Brazilian Festival, pages 461–470, Belo Horizonte, Brazil. Association for Computational Linguistics.
Pugh, R., Huerta Mendez, M., Sasaki, M., and Tyers, F. (2022). Universal Dependencies for western sierra Puebla Nahuatl. In Calzolari, N., Béchet, F., Blache, P., Choukri, K., Cieri, C., Declerck, T., Goggi, S., Isahara, H., Maegaard, B., Mariani, J., Mazo, H., Odijk, J., and Piperidis, S., editors, Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 5011–5020, Marseille, France. European Language Resources Association.
Pugh, R. and Tyers, F. (2024). A Universal Dependencies treebank for Highland Puebla Nahuatl. In Duh, K., Gomez, H., and Bethard, S., editors, Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 1393–1403, Mexico City, Mexico. Association for Computational Linguistics.
Rodrigues, J. B. (1890). Poranduba amazonense ou kochiyma-uara porandub, 1872-1887. Typ. de G. Leuzinger & Filhos, Rio de Janeiro.
Sanches Duran, M., de Souza, E. A., das Graças Volpe Nunes, M., Pagano, A. S., and Pardo, T. A. S. (2025). Extending the enhanced Universal Dependencies – addressing subjects in pro-drop languages. In Bouma, G. and Çöltekin, Ç., editors, Proceedings of the Eighth Workshop on Universal Dependencies (UDW, SyntaxFest 2025), pages 143–152, Ljubljana, Slovenia. Association for Computational Linguistics.
Schuster, S. and Manning, C. D. (2016). Enhanced English Universal Dependencies: An improved representation for natural language understanding tasks. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC’16), pages 2371–2378, Portorož, Slovenia. European Language Resources Association (ELRA).
Souza, E., Duran, M., Nunes, M. d. G. V., Sampaio, G., Belasco, G., and Pardo, T. (2024). Automatic annotation of enhanced Universal Dependencies for Brazilian Portuguese. In Claro, D. B. and Pagano, A., editors, Proceedings of the 15th Brazilian Symposium in Information and Human Language Technology, pages 22–31, Belém do Pará, Brazil. Association for Computational Linguistics.
Thomas, G. (2019). Universal Dependencies for Mbyá Guaraní. In Proceedings of the Third Workshop on Universal Dependencies (UDW, SyntaxFest 2019), pages 70–77, Paris, France. Association for Computational Linguistics.
Universal Dependencies (2026). Enhanced Universal Dependencies syntax overview. [link]. Accessed: 2026-05-01.
Zeman, D. et al. (2025). Universal dependencies 2.17. Accessed: 2026-05-14.
Alencar, L. F. d., Silva, H. L. B., Gurgel, J. L., and Alexandre, D. M. (2025). Per aspera ad astra: Improving the automatic evaluation of a Universal Dependencies treebank for a low-resourced language. Revista de Estudos da Linguagem. In press.
Avila, M. T. (2021). Proposta de dicionário nheengatu-português. PhD thesis, Faculdade de Filosofia, Letras e Ciências Humanas da Universidade de São Paulo.
Bresnan, J., Asudeh, A., Toivonen, I., and Wechsler, S. (2015). Lexical-functional syntax, volume 16. John Wiley & Sons.
Casasnovas, A. (2006). Noções de língua geral ou nheengatú: gramática, lendas e vocabulário. Editora da Universidade Federal do Amazonas; Faculdade Salesiana Dom Bosco, Manaus, 2 edition.
Church, K. and Liberman, M. (2021). The future of computational linguistics: On beyond alchemy. Frontiers in Artificial Intelligence, 4:1–18.
de Alencar, L. F. (2023). Yauti: A tool for morphosyntactic analysis of Nheengatu within the Universal Dependencies framework. In Anais do XIV Simpósio Brasileiro de Tecnologia da Informação e da Linguagem Humana, pages 135–145, Porto Alegre, RS, Brasil. SBC.
de Alencar, L. F. (2024). A Universal Dependencies treebank for Nheengatu. In Gamallo, P., Claro, D., Teixeira, A. J. S., Real, L., García, M., Oliveira, H. G., and Amaro, R., editors, Proceedings of the 16th International Conference on Computational Processing of Portuguese, volume 2, pages 37–54, Santiago de Compostela, Galicia, Spain. Association for Computational Linguistics.
de Alencar, L. F. (2025). Enhancing a nheengatu morphosyntactic analyzer for word formation and non-standard language. In Anais do XVI Simpósio Brasileiro de Tecnologia da Informação e da Linguagem Humana, pages 13–28, Porto Alegre, RS, Brasil. SBC.
de Alencar, L. F. and Rademaker, A. (2022). Modelação da valência verbal numa gramática computacional do português no formalismo hpsg. Domínios de Lingu@gem, 16(4):1339–1400.
de Amorim, A. B. (1928). Lendas em Nheêngatu e em Portuguez. Revista do Instituto Historico e Geographico Brasileiro, 154(100):9–475. Tomo 100, vol. 154 (2º de 1926).
de Marneffe, M.-C., Ginter, F., Goldberg, Y., Hajič, J., Manning, C., McDonald, R., Nivre, J., Petrov, S., Pyysalo, S., Schuster, S., Silveira, N., Tsarfaty, R., Tyers, F., and Zeman, D. (2025). xcomp: open clausal complement. [link].
Di Felippo, A., Postali, C., Ceregatto, G., Gazana, L., Silva, E., Roman, N., and Pardo, T. (2021). Descrião preliminar do corpus DANTEStocks: Diretrizes de segmentação para anotação segundo Universal Dependencies. In Ruiz, E. E. S. and Torrent, T. T., editors, Proceedings of the 13th Brazilian Symposium in Information and Human Language Technology, pages 335–343, Porto Alegre, Brazil. Association for Computational Linguistics.
Duran, M., Lopes, L., das Graças Nunes, M., and Pardo, T. (2023). The dawn of the porttinari multigenre treebank: Introducing its journalistic portion. In Anais do XIV Simpósio Brasileiro de Tecnologia da Informação e da Linguagem Humana, pages 115–124, Porto Alegre, RS, Brasil. SBC.
Guillaume, B. and Perrier, G. (2021). Graph rewriting for enhanced Universal Dependencies. In Oepen, S., Sagae, K., Tsarfaty, R., Bouma, G., Seddah, D., and Zeman, D., editors, Proceedings of the 17th International Conference on Parsing Technologies and the IWPT 2021 Shared Task on Parsing into Enhanced Universal Dependencies (IWPT 2021), pages 175–183, Online. Association for Computational Linguistics.
Hartt, C. F. (1938). Notas sobre a língua geral, ou tupímoderno do Amazonas. Anais da Biblioteca Nacional do Rio de Janeiro, LI:305–390. [1929].
Kroeger, P. (2004). Analyzing Syntax: A Lexical-Functional Approach. Cambridge University Press, Cambridge.
Navarro, E. d. A. (2012). O último refúgio da língua geral no Brasil. Estudos Avançados, 26(76):245–254.
Nivre, J. et al. (2016). Universal Dependencies v1: A multilingual treebank collection. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC’16), pages 1659–1666, Portorož, Slovenia. European Language Resources Association (ELRA).
Pagano, A. S., Duran, M. S., and Pardo, T. A. S. (2023). Enhanced dependencies para o português brasileiro. In Pardo, T. A. S., Duran, M. S., and Lopes, L., editors, Proceedings of the 2nd Edition of the Universal Dependencies Brazilian Festival, pages 461–470, Belo Horizonte, Brazil. Association for Computational Linguistics.
Pugh, R., Huerta Mendez, M., Sasaki, M., and Tyers, F. (2022). Universal Dependencies for western sierra Puebla Nahuatl. In Calzolari, N., Béchet, F., Blache, P., Choukri, K., Cieri, C., Declerck, T., Goggi, S., Isahara, H., Maegaard, B., Mariani, J., Mazo, H., Odijk, J., and Piperidis, S., editors, Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 5011–5020, Marseille, France. European Language Resources Association.
Pugh, R. and Tyers, F. (2024). A Universal Dependencies treebank for Highland Puebla Nahuatl. In Duh, K., Gomez, H., and Bethard, S., editors, Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 1393–1403, Mexico City, Mexico. Association for Computational Linguistics.
Rodrigues, J. B. (1890). Poranduba amazonense ou kochiyma-uara porandub, 1872-1887. Typ. de G. Leuzinger & Filhos, Rio de Janeiro.
Sanches Duran, M., de Souza, E. A., das Graças Volpe Nunes, M., Pagano, A. S., and Pardo, T. A. S. (2025). Extending the enhanced Universal Dependencies – addressing subjects in pro-drop languages. In Bouma, G. and Çöltekin, Ç., editors, Proceedings of the Eighth Workshop on Universal Dependencies (UDW, SyntaxFest 2025), pages 143–152, Ljubljana, Slovenia. Association for Computational Linguistics.
Schuster, S. and Manning, C. D. (2016). Enhanced English Universal Dependencies: An improved representation for natural language understanding tasks. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC’16), pages 2371–2378, Portorož, Slovenia. European Language Resources Association (ELRA).
Souza, E., Duran, M., Nunes, M. d. G. V., Sampaio, G., Belasco, G., and Pardo, T. (2024). Automatic annotation of enhanced Universal Dependencies for Brazilian Portuguese. In Claro, D. B. and Pagano, A., editors, Proceedings of the 15th Brazilian Symposium in Information and Human Language Technology, pages 22–31, Belém do Pará, Brazil. Association for Computational Linguistics.
Thomas, G. (2019). Universal Dependencies for Mbyá Guaraní. In Proceedings of the Third Workshop on Universal Dependencies (UDW, SyntaxFest 2019), pages 70–77, Paris, France. Association for Computational Linguistics.
Universal Dependencies (2026). Enhanced Universal Dependencies syntax overview. [link]. Accessed: 2026-05-01.
Zeman, D. et al. (2025). Universal dependencies 2.17. Accessed: 2026-05-14.
Publicado
19/10/2026
Como Citar
ALEXANDRE, Dominick Maia; ALENCAR, Leonel Figueiredo de.
Extending Enhanced Universal Dependencies to Nheengatu: Semi-Automated Control and Raising Annotation. In: SIMPÓSIO BRASILEIRO DE TECNOLOGIA DA INFORMAÇÃO E DA LINGUAGEM HUMANA (STIL), 17. , 2026, Cuiabá/MT.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2026
.
p. 1-13.
DOI: https://doi.org/10.5753/stil.2026.26653.
