Do pequeno Asteróide B-612 ao universo: a anotação de O Pequeno Príncipe segundo o modelo Universal Dependencies
Resumo
Relatamos, neste artigo, o processo de anotação do icônico livro O Pequeno Príncipe segundo o modelo internacional Universal Dependencies, como parte de um esforço de expansão dos dados anotados atualmente disponíveis para o português do Brasil. Em especial, apresentamos o protocolo de anotação e as decisões linguísticas envolvidas, além de uma análise quantitativa do córpus.
Referências
Anchiêta, R. and Pardo, T. (2018). Towards AMR-BR: A SemBank for Brazilian Portuguese language. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018) (pp. 449–460). European Language Resources Association (ELRA).
Andrade, P. L. C. de, Silva, R. M. and Pardo, T. A. S. (2026). Caracterização lexical e sintática de notícias falsas em português produzidas por humanos e por máquinas. In Proceedings of the 17th International Conference on Computational Processing of Portuguese (PROPOR 2026) - Vol. 2 (pp. 148–158). Salvador, Brazil.
Banarescu, L., Bonial, C., Cai, S., Georgescu, M., Griffitt, K., Hermjakob, U., Knight, K., Koehn, P., Palmer, M. and Schneider, N. (2013). Abstract meaning representation for sembanking. In Proceedings of the 7th Linguistic Annotation Workshop and Interoperability with Discourse (pp. 178–186). Sofia, Bulgaria.
Di Felippo, A. and Roman, N. T. (2025). DANTEStocks: A multi-layered annotated corpus of stock market tweets for Brazilian Portuguese. Revista Brasileira de Linguística Aplicada, 25(1), 1–32.
Duran, M. S., Nunes, M. G. V., Lopes, L. and Pardo, T. A. S. (2022). Manual de anotação como recurso de processamento de linguagem natural: O modelo Universal Dependencies em língua portuguesa. Domínios de Lingu@gem, 16(4).
Duran, M., Lopes, L., Nunes, M. G. and Pardo, T. (2023). The dawn of the Porttinari multigenre treebank: Introducing its journalistic portion. In Proceedings of the 14th Brazilian Symposium in Information and Human Language Technology (pp. 124–133). Belo Horizonte, Brazil.
Duran, M. S., de Souza, E. A., Nunes, M. G. V., Pagano, A. S. and Pardo, T. A. S. (2025). Extending the enhanced Universal Dependencies: Addressing subjects in pro-drop languages. In Proceedings of the Eighth Workshop on Universal Dependencies (UDW, SyntaxFest 2025) (pp. 143–152). Ljubljana, Slovenia.
Fontes, M. G. (2012). A clivagem do constituinte interrogativo em sentenças interrogativas do português brasileiro: uma abordagem diacrônica. Signum: Estudos da Linguagem, 15(3), 149–170.
Freitas, C. and Pardo, T. A. S. (2025). PropBanks e representações semânticas: O que temos, o que queremos e o que podemos. Linguamática, 17(2), 3–31.
Guibon, G., Courtin, M., Gerdes, K. and Guillaume, B. (2020). When collaborative treebank curation meets graph grammars. In Proceedings of the Twelfth Language Resources and Evaluation Conference (pp. 5291–5300). Marseille, France.
Leal, S. E.; Duran, M. S.; Scarton, C. E.; Hartmann, N.; Aluísio, S. M. (2023) NILC-Metrix: assessing the complexity of written and spoken language in Brazilian Portuguese. Lang Resources & Evaluation
Lopes, L., Duran, M. S. and Pardo, T. A. S. (2023). Verifica-UD: A verifier for Universal Dependencies annotation for Portuguese. In Proceedings of the 2nd Edition of the Universal Dependencies Brazilian Festival (pp. 451–460). Belo Horizonte, Brazil.
Lopes, L. and Pardo, T. (2024). Towards Portparser: A highly accurate parsing system for Brazilian Portuguese following the Universal Dependencies framework. In Proceedings of the 16th International Conference on Computational Processing of Portuguese - Vol. 1 (pp. 401–410). Santiago de Compostela, Galicia/Spain.
Lopes, L., Nunes, M. G. V., Duran, M. S. and Pardo, T. A. S. (2025). A sintaxe no tribunal: Apresentando e explorando um corpus jurídico em português anotado sintaticamente segundo o modelo Universal Dependencies. In Proceedings of the XVI Symposium in Information and Human Language Technology (STIL) (pp. 220–232). Fortaleza/CE, Brazil.
de Marneffe, M.-C., Manning, C. D., Nivre, J. and Zeman, D. (2021). Universal Dependencies. Computational Linguistics, 47(2), 255–308.
Morais, M. A. T.; Salles, H. M. M. (2024). Resistência do dativo de primeira pessoa na batalha (quase) perdida dos clíticos pronominais do português brasileiro. Revista de Estudos da Linguagem, [S. l.], v. 30, n. 4, p. 1621–1657, 2024.
Oliveira, L., Claro, D. B. and Souza, M. (2023). DptOIE: A Portuguese open information extraction based on dependency analysis. Artificial Intelligence Review, 56, 7015–7046.
Othero, G. A.; Lazzari, M. (2024) Sujeitos e objetos nulos em português brasileiro: correlações e mudança. Revista de Estudos da Linguagem, [S. l.], v. 30, n. 4, p. 1831–1854, 2024.
Pagano, A. S., Rassi, A. and Pagano, A. C. S. (2026). A ordem e a função das palavras em uma sentença: Sintaxe. In H. M. Caseli and M. G. V. Nunes (Eds.), Processamento de linguagem natural: Conceitos, técnicas e aplicações em português (Vol. 1, 4th ed., chap. 6). BPLN.
Pardo, T., Duran, M., Lopes, L., Di Felippo, A., Roman, N. and Nunes, M. G. (2021). Porttinari: A large multi-genre treebank for Brazilian Portuguese. In Proceedings of the 13th Brazilian Symposium in Information and Human Language Technology (pp. 1–10). Porto Alegre, Brazil.
Qi, P.; Zhang, Y.; Zhang, Y.; Bolton, J.; Manning, C. D. (2020). Stanza: A Python Natural Language Processing Toolkit for Many Human Languages. In Association for Computational Linguistics (ACL) System Demonstrations. 2020. [link]
Rademaker, A., Chalub, F., Real, L., Freitas, C., Bick, E. and de Paiva, V. (2017). Universal Dependencies for Portuguese. In Proceedings of the Fourth International Conference on Dependency Linguistics (Depling 2017) (pp. 197–206). Pisa, Italy: Linköping University Electronic Press.
Souza, E., Silveira, A., Cavalcanti, T., Castro, M. and Freitas, C. (2021). PetroGold: Corpus padrão ouro para o domínio do petróleo. In Proceedings of the 13th Brazilian Symposium in Information and Human Language Technology (pp. 29–38). Porto Alegre, Brazil.
Straka, M. (2018). UDPipe 2.0 Prototype at CoNLL 2018 UD Shared Task. In: Proceedings of CoNLL 2018: The SIGNLL Conference on Computational Natural Language Learning, pp. 197-207, Association for Computational Linguistics, Stroudsburg, PA, USA.
Andrade, P. L. C. de, Silva, R. M. and Pardo, T. A. S. (2026). Caracterização lexical e sintática de notícias falsas em português produzidas por humanos e por máquinas. In Proceedings of the 17th International Conference on Computational Processing of Portuguese (PROPOR 2026) - Vol. 2 (pp. 148–158). Salvador, Brazil.
Banarescu, L., Bonial, C., Cai, S., Georgescu, M., Griffitt, K., Hermjakob, U., Knight, K., Koehn, P., Palmer, M. and Schneider, N. (2013). Abstract meaning representation for sembanking. In Proceedings of the 7th Linguistic Annotation Workshop and Interoperability with Discourse (pp. 178–186). Sofia, Bulgaria.
Di Felippo, A. and Roman, N. T. (2025). DANTEStocks: A multi-layered annotated corpus of stock market tweets for Brazilian Portuguese. Revista Brasileira de Linguística Aplicada, 25(1), 1–32.
Duran, M. S., Nunes, M. G. V., Lopes, L. and Pardo, T. A. S. (2022). Manual de anotação como recurso de processamento de linguagem natural: O modelo Universal Dependencies em língua portuguesa. Domínios de Lingu@gem, 16(4).
Duran, M., Lopes, L., Nunes, M. G. and Pardo, T. (2023). The dawn of the Porttinari multigenre treebank: Introducing its journalistic portion. In Proceedings of the 14th Brazilian Symposium in Information and Human Language Technology (pp. 124–133). Belo Horizonte, Brazil.
Duran, M. S., de Souza, E. A., Nunes, M. G. V., Pagano, A. S. and Pardo, T. A. S. (2025). Extending the enhanced Universal Dependencies: Addressing subjects in pro-drop languages. In Proceedings of the Eighth Workshop on Universal Dependencies (UDW, SyntaxFest 2025) (pp. 143–152). Ljubljana, Slovenia.
Fontes, M. G. (2012). A clivagem do constituinte interrogativo em sentenças interrogativas do português brasileiro: uma abordagem diacrônica. Signum: Estudos da Linguagem, 15(3), 149–170.
Freitas, C. and Pardo, T. A. S. (2025). PropBanks e representações semânticas: O que temos, o que queremos e o que podemos. Linguamática, 17(2), 3–31.
Guibon, G., Courtin, M., Gerdes, K. and Guillaume, B. (2020). When collaborative treebank curation meets graph grammars. In Proceedings of the Twelfth Language Resources and Evaluation Conference (pp. 5291–5300). Marseille, France.
Leal, S. E.; Duran, M. S.; Scarton, C. E.; Hartmann, N.; Aluísio, S. M. (2023) NILC-Metrix: assessing the complexity of written and spoken language in Brazilian Portuguese. Lang Resources & Evaluation
Lopes, L., Duran, M. S. and Pardo, T. A. S. (2023). Verifica-UD: A verifier for Universal Dependencies annotation for Portuguese. In Proceedings of the 2nd Edition of the Universal Dependencies Brazilian Festival (pp. 451–460). Belo Horizonte, Brazil.
Lopes, L. and Pardo, T. (2024). Towards Portparser: A highly accurate parsing system for Brazilian Portuguese following the Universal Dependencies framework. In Proceedings of the 16th International Conference on Computational Processing of Portuguese - Vol. 1 (pp. 401–410). Santiago de Compostela, Galicia/Spain.
Lopes, L., Nunes, M. G. V., Duran, M. S. and Pardo, T. A. S. (2025). A sintaxe no tribunal: Apresentando e explorando um corpus jurídico em português anotado sintaticamente segundo o modelo Universal Dependencies. In Proceedings of the XVI Symposium in Information and Human Language Technology (STIL) (pp. 220–232). Fortaleza/CE, Brazil.
de Marneffe, M.-C., Manning, C. D., Nivre, J. and Zeman, D. (2021). Universal Dependencies. Computational Linguistics, 47(2), 255–308.
Morais, M. A. T.; Salles, H. M. M. (2024). Resistência do dativo de primeira pessoa na batalha (quase) perdida dos clíticos pronominais do português brasileiro. Revista de Estudos da Linguagem, [S. l.], v. 30, n. 4, p. 1621–1657, 2024.
Oliveira, L., Claro, D. B. and Souza, M. (2023). DptOIE: A Portuguese open information extraction based on dependency analysis. Artificial Intelligence Review, 56, 7015–7046.
Othero, G. A.; Lazzari, M. (2024) Sujeitos e objetos nulos em português brasileiro: correlações e mudança. Revista de Estudos da Linguagem, [S. l.], v. 30, n. 4, p. 1831–1854, 2024.
Pagano, A. S., Rassi, A. and Pagano, A. C. S. (2026). A ordem e a função das palavras em uma sentença: Sintaxe. In H. M. Caseli and M. G. V. Nunes (Eds.), Processamento de linguagem natural: Conceitos, técnicas e aplicações em português (Vol. 1, 4th ed., chap. 6). BPLN.
Pardo, T., Duran, M., Lopes, L., Di Felippo, A., Roman, N. and Nunes, M. G. (2021). Porttinari: A large multi-genre treebank for Brazilian Portuguese. In Proceedings of the 13th Brazilian Symposium in Information and Human Language Technology (pp. 1–10). Porto Alegre, Brazil.
Qi, P.; Zhang, Y.; Zhang, Y.; Bolton, J.; Manning, C. D. (2020). Stanza: A Python Natural Language Processing Toolkit for Many Human Languages. In Association for Computational Linguistics (ACL) System Demonstrations. 2020. [link]
Rademaker, A., Chalub, F., Real, L., Freitas, C., Bick, E. and de Paiva, V. (2017). Universal Dependencies for Portuguese. In Proceedings of the Fourth International Conference on Dependency Linguistics (Depling 2017) (pp. 197–206). Pisa, Italy: Linköping University Electronic Press.
Souza, E., Silveira, A., Cavalcanti, T., Castro, M. and Freitas, C. (2021). PetroGold: Corpus padrão ouro para o domínio do petróleo. In Proceedings of the 13th Brazilian Symposium in Information and Human Language Technology (pp. 29–38). Porto Alegre, Brazil.
Straka, M. (2018). UDPipe 2.0 Prototype at CoNLL 2018 UD Shared Task. In: Proceedings of CoNLL 2018: The SIGNLL Conference on Computational Natural Language Learning, pp. 197-207, Association for Computational Linguistics, Stroudsburg, PA, USA.
Publicado
19/10/2026
Como Citar
DURAN, Magali Sanches; LOPES, Lucelene; NUNES, Maria das Graças Volpe; PARDO, Thiago Alexandre Salgueiro.
Do pequeno Asteróide B-612 ao universo: a anotação de O Pequeno Príncipe segundo o modelo Universal Dependencies. In: SIMPÓSIO BRASILEIRO DE TECNOLOGIA DA INFORMAÇÃO E DA LINGUAGEM HUMANA (STIL), 17. , 2026, Cuiabá/MT.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2026
.
p. 464-475.
DOI: https://doi.org/10.5753/stil.2026.33340.
