DANTE – A Treebank of Tweets in Brazilian Portuguese
Resumo
In this article, we present the current state of DANTE (Dependency-ANalysed corpora of TwEets), a treebank focusing on tweets (currently X posts) written in Brazilian Portuguese, featuring four subcorpora varying in size, domain, and level of annotation under the Universal Dependencies framework. This article aims to inform the Natural Language Processing and Linguistic communities about currently available corpora, as well as what to expect from the treebank in the near future, while also bringing together the tools and language resources developed within the DANTE project.
Referências
Banarescu, L., Bonial, C., Cai, S., Georgescu, M., Griffitt, K., Hermjakob, U., Knight, K., Koehn, P., Palmer, M., and Schneider, N. (2019). Abstract meaning representation (amr) 1.2.6 specification. Technical report, AMR Project. Version 1.2.6.
Barberia, L., Schmalz, P., Roman, N. T., Lombard, B., and de Sousa, T. M. (2025). It’s about what and how you say it: A corpus with stance and sentiment annotation for covid-19 vaccines posts on x/twitter by brazilian political elites. In Hämäläinen, M., Öhman, E., Bizzoni, Y., Miyagawa, S., and Alnajjar, K., editors, Proceedings of the 5th International Conference on Natural Language Processing for Digital Humanities (NLP4DH 2025), pages 365–376, Albuquerque, USA. Association for Computational Linguistics.
Barberia, L. G., de Santana Schmalz, P. H., and Roman, N. T. (2023). When tweets get viral - a deep learning approach for stance analysis of covid-19 vaccines tweets by brazilian political elites. In Proceedings of the 14th Symposium in Information and Human Language Technology (STIL 2023), Belo Horizonte, MG - Brazil.
Barbosa, B. and Felippo, A. D. (2025). Nounbank.ds: a lexical repository of nominal frames from stock market tweets in brazilian portuguese. In Anais do XVI Simpósio Brasileiro de Tecnologia da Informação e da Linguagem Humana, pages 29–41, Porto Alegre, RS, Brasil. SBC.
Barbosa, B. K. S. (2024). Descrição sintático-semântica de nomes predicadores em tweets do mercado financeiro em português. Dissertação (mestrado em linguística), Universidade Federal de São Carlos, São Carlos, SP.
Branco, A., Silva, J. R., Gomes, L., and António Rodrigues, J. (2022). Universal grammatical dependencies for Portuguese with CINTIL data, LX processing and CLARIN support. In Calzolari, N., Béchet, F., Blache, P., Choukri, K., Cieri, C., Declerck, T., Goggi, S., Isahara, H., Maegaard, B., Mariani, J., Mazo, H., Odijk, J., and Piperidis, S., editors, Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 5617–5626, Marseille, France. European Language Resources Association.
Ceregatto, G. and Felippo, A. D. (2025). Dantestocks-amr em construção: Avanços e desafios na anotação semântica de tweets financeiros. In Anais do XVI Simpósio Brasileiro de Tecnologia da Informação e da Linguagem Humana, pages 608–617, Porto Alegre, RS, Brasil. SBC.
da Silva, F. J. V., Roman, N. T., and Carvalho, A. M. B. R. (2020). Stock market tweets annotated with emotions. Corpora, 15(3):343–354. Online ISSN: 1755-1676.
de Carvalho, W. P. and Roman, N. T. (2025). Classifying emotions in tweets from the financial market: A bert-based approach. In Proceedings of the 15th International Conference on Recent Advances in Natural Language Processing (RANLP 2025), pages 218–226, Varna, Bulgaria. INCOMA Ltd., Shoumen, Bulgaria.
de Marneffe, M.-C., Manning, C. D., Nivre, J., and Zeman, D. (2021). Universal Dependencies. Computational Linguistics, 47(2):255–308.
Di-Felippo, A., das Graças Nunes, M., and Barbosa, B. (2024a). A dependency treebank of tweets in brazilian portuguese: Syntactic annotation issues and approach. In Anais do XV Simpósio Brasileiro de Tecnologia da Informação e da Linguagem Humana, pages 192–201, Porto Alegre, RS, Brasil. SBC.
Di-Felippo, A., Nunes, M. d. G. V., and Barbosa, B. K. d. S. (2024b). Diretrizes de anotação de relações de dependência em tweets do mercado financeiro. Technical Report n. 446, Relatório Técnico – ICMC, USP, São Carlos.
Di-Felippo, A., Postali, C., Ceregatto, G., Gazana, L., Silva, E., Roman, N., and Pardo, T. (2021). Descrição preliminar do corpus DANTEStocks: diretrizes de segmentação para anotação segundo Universal Dependencies. In Anais do XIII Simpósio Brasileiro de Tecnologia da Informação e da Linguagem Humana, pages 335–343, Porto Alegre, RS, Brasil. SBC.
Di-Felippo, A., Postali, C., Ceregatto, G., Gazana, L. S., and Roman, N. T. (2022). Diretrizes de anotação de pos tags em tweets do mercado financeiro: Orientações para anotação em língua portuguesa segundo a abordagem universal dependencies. Technical Report n. 438, Relatório Técnico – ICMC, USP, São Carlos-SP.
Di-Felippo, A. and Roman, N. T. (2025). Dantestocks: A multi-layered annotated corpus of stock market tweets for brazilian portuguese. Revista Brasileira de Linguística Aplicada, 25(1).
Di-Felippo, A., Roman, N. T., Barbosa, B. K. S., Oliveira, G. P. d., and Scandarolli, C. L. (2026). Lexical and orthographic variation in Portuguese financial tweets: Annotation, analysis, and implications for embedding models. In Souza, M., de Dios-Flores, I., Santos, D., Freitas, L., Souza, J. W. d. C., and Ribeiro, E., editors, Proceedings of the 17th International Conference on Computational Processing of Portuguese (PROPOR 2026) - Vol. 1, pages 628–637, Salvador, Brazil. Association for Computational Linguistics.
Di-Felippo, A., Roman, N. T., da Silva Barbosa, B. K., and Pardo, T. A. S. (2024c). Genipapo – a multigenre dependency parser for brazilian portuguese. In Proceedings of the 15th Brazilian Symposium in Information and Human Language Technology (STIL 2024), pages 257–266, Belém PA, Brazil.
Duran, M., Lopes, L., Nunes, M. d. G. V., and Pardo, T. A. S. (2023). The dawn of the Porttinari multigenre treebank: introducing its journalistic portion. In Proceedings of the XIV Brazilian Symposium in Information and Human Language Technology (STIL), pages 115–124, Porto Alegre, RS, Brasil. SBC.
Duran, M. S. (2021). Manual de anotação de relações de dependência: orientações para anotação de relações de dependência sintática em língua portuguesa, seguindo as diretrizes da abordagem universal dependencies (ud). Technical Report n. 435, Relatório Técnico – ICMC, USP, São Carlos.
Felippo, A. D., Roman, N. T., Pardo, T. A. S., and de Moura, L. P. (2023). The dantestocks corpus: an analysis of the distribution of universal dependencies-based part-of-speech tags. Revista da Abralin, 22(2):249–271.
Jurafsky, D. and Martin, J. H. (2026). Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition with Language Models. 3rd edition. Online manuscript released January 6, 2026.
Krumm, J., Davies, N., and Narayanaswami, C. (2008). User-generated content. IEEE Pervasive Computing, 7(4):10–11.
Lopes, L., Duran, M., Fernandes, P., and Pardo, T. (2022). PortiLexicon-UD: a Portuguese lexical resource according to Universal Dependencies model. In Calzolari, N., Béchet, F., Blache, P., Choukri, K., Cieri, C., Declerck, T., Goggi, S., Isahara, H., Maegaard, B., Mariani, J., Mazo, H., Odijk, J., and Piperidis, S., editors, Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 6635–6643, Marseille, France. European Language Resources Association.
Meyers, A. (2007). Annotation guidelines for nombank: Noun argument structure for propbank. Technical report, NomBank Project, S.l.
Mota, C. and Santos, D. (2008). Desafios na avaliação conjunta do reconhecimento de entidades mencionadas: O Segundo HAREM. Linguateca.
Palmer, M., Gildea, D., and Kingsbury, P. (2005). The proposition bank: An annotated corpus of semantic roles. Computational Linguistics, 31(1):71–106.
Piai, L., Di-Felippo, A., and Roman, N. T. (2025). Named entities in stock market tweets: A fine-grained and linguistically-motivated annotation. In Proceedings of the 16th Brazilian Symposium in Information and Human Language Technology: 10th Portuguese Description Journey (JDP 2025), pages 654–663, Fortaleza, CE, Brazil.
Plutchik, R. (2001). The nature of emotions: Human emotions have deep evolutionary roots, a fact that may explain their complexity and provide tools for clinical practice. American Scientist, 89(4):344–350.
Pradhan, S., Bonn, J., Myers, S., Conger, K., O’Gorman, T., Gung, J., Wright-Bettner, K., and Palmer, M. (2022). Propbank comes of age—larger, smarter, and more diverse. In Nastase, V., Pavlick, E., Pilehvar, M. T., Camacho-Collados, J., and Raganato, A., editors, Proceedings of the 11th Joint Conference on Lexical and Computational Semantics, pages 278–288, Seattle, Washington. Association for Computational Linguistics.
Qi, P., Zhang, Y., Zhang, Y., Bolton, J., and Manning, C. D. (2020). Stanza: A python natural language processing toolkit for many human languages. In Celikyilmaz, A. and Wen, T.-H., editors, Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, pages 101–108, Online. Association for Computational Linguistics.
Rademaker, A., Chalub, F., Real, L., Freitas, C., Bick, E., and de Paiva, V. (2017). Universal Dependencies for Portuguese. In Proceedings of the Fourth International Conference on Dependency Linguistics (Depling 2017), pages 197–206, Pisa, Italy. Linköping University Electronic Press.
Sanguinetti, M., Bosco, C., Cassidy, L., and et al. (2023). Treebanking user-generated content: a ud based overview of guidelines, corpora and unified recommendations. Language Resources Evaluation, 57:493–544.
Silva, E., Pardo, T., Roman, N., and Fellipo, A. (2021). Universal dependencies for tweets in brazilian portuguese: Tokenization and part of speech tagging. In Anais do XVIII Encontro Nacional de Inteligência Artificial e Computacional, pages 434–445, Porto Alegre, RS, Brasil. SBC.
Silva, E. H., Pardo, T. A. S., and Roman, N. T. (2023). Etiquetagem morfossintática multigênero para o português do brasil segundo o modelo ”universal dependencies”. In Proceedings of the 14th Symposium in Information and Human Language Technology (STIL 2023), Belo Horizonte, MG - Brazil.
Souza, E., Silveira, A., Cavalcanti, T., Castro, M., and Freitas, C. (2021). Petrogold – corpus padrão ouro para o domínio do petróleo. In Anais do XIII Simpósio Brasileiro de Tecnologia da Informação e da Linguagem Humana, pages 29–38, Porto Alegre, RS, Brasil. SBC.
Straka, M., Hajič, J., and Straková, J. (2016). UDPipe: Trainable pipeline for processing CoNLL-U files performing tokenization, morphological analysis, POS tagging and parsing. In Calzolari, N., Choukri, K., Declerck, T., Goggi, S., Grobelnik, M., Maegaard, B., Mariani, J., Mazo, H., Moreno, A., Odijk, J., and Piperidis, S., editors, Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC’16), pages 4290–4297, Portorož, Slovenia. European Language Resources Association (ELRA).
Zeman, D. e. a. (2017). CoNLL 2017 shared task: Multilingual parsing from raw text to Universal Dependencies. In Proceedings of the CoNLL 2017 Shared Task: Multilingual Parsing from Raw Text to Universal Dependencies, pages 1–19, Vancouver, Canada. ACL.
Zerbinati, M. M., Roman, N. T., and Felippo, A. D. (2024). A corpus of stock market tweets annotated with named entities. In Proceedings of the 16th International Conference on Computational Processing of Portuguese (PROPOR 2024), pages 276–284, Santiago de Compostela, Galicia/Spain.
