Developing Multilingual Multiword Expressions (MWEs) Resources: An Annotated Corpus of Fables following the PARSEME guidelines
Resumo
This paper introduces a Brazilian Portuguese corpus of fables developed as part of a broader multilingual initiative. It comprises the narratives and their concluding moral statements (epimythia), annotated for morphosyntax and multiword expressions (MWEs) according to the Universal Dependencies (UD) and PARSEME guidelines, respectively. The analysis reveals that 53.59% of the sentences contain at least one MWE, with inherently reflexive verbs (IRVs) constituting the most frequent category. MWEs are also distributed differently: narratives contain a higher proportion of adverbial MWEs, whereas adpositional MWEs are frequent in epimythia. Future work will enrich the corpus through semantic annotation based on Uniform Meaning Representation (UMR).Referências
Chun, J. and Xue, N. (2024). Uniform meaning representation parsing as a pipelined approach. In Ustalov, D., Gao, Y., Panchenko, A., Tutubalina, E., Nikishina, I., Ramesh, A., Sakhovskiy, A., Usbeck, R., Penn, G., and Valentino, M., editors, Proceedings of TextGraphs-17: Graph-based Methods for Natural Language Processing, pages 40–52, Bangkok, Thailand. Association for Computational Linguistics.
Dezotti, M. C. C., Berliner, E., and Duarte, A. (2013). Esopo: fábulas completas. Cosac Naify, São Paulo, Brazil.
Duran, M. (2021). Manual de anotação de pos tags: orientações para anotação de etiquetas morfossintáticas em língua portuguesa, segundo as diretrizes da abordagem universal dependencies. Technical Report Relatório Técnico No. 434, Universidade de São Paulo, ICMC. Accessed: 2026-04-10.
Duran, M. (2022). Manual de anotação de relações de dependência: versão revisada e estendida. Technical Report Relatório Técnico No. 440, Universidade de São Paulo, ICMC. Accessed: 2026-04-30.
Guibon, G. et al. (2020). When collaborative treebank curation meets graph grammars. In Proceedings of the Twelfth Language Resources and Evaluation Conference (LREC), pages 5291–5300.
Neves, M. H. M. (2014). A gramática pela fábula. ou: A fábula pela gramática. Linguística, 30(1):165–196.
Nivre, J. et al. (2020). Universal dependencies v2: An evergrowing multilingual treebank collection. arXiv preprint arXiv:2004.10643.
PARSEME Initiative (2026). Parseme corpora wiki: Citing parseme. [link]. Accessed: 2026-05-19.
Pasquer, C., Savary, A., Ramisch, C., and Antoine, J.-Y. (2020). Verbal multiword expression identification: Do we need a sledgehammer to crack a nut? In Proceedings of the 28th International Conference on Computational Linguistics (COLING 2020), Barcelona, Spain. HAL Id: hal-03013636.
Ramisch, C. et al. (2018). A corpus study of verbal multiword expressions in brazilian portuguese. In Computational Processing of the Portuguese Language: 13th International Conference, PROPOR 2018, Lecture Notes in Artificial Intelligence, Cham, Switzerland. Springer International Publishing.
Ramisch, R., Ramisch, C., and Villavicencio, A. (2026). Expressões multipalavras: fogo de palha ou osso duro de roer? In Caseli, H. M. and Nunes, M. G. V., editors, Processamento de Linguagem Natural: Conceitos, Técnicas e Aplicações em Português, volume 1, book chapter 5. BPLN, 4 edition.
Savary, A. et al. (2017). The parseme shared task on automatic identification of verbal multiword expressions. In Proceedings of the 13th Workshop on Multiword Expressions, pages 31–47. Association for Computational Linguistics.
Savary, A. et al. (2018). Parseme multilingual corpus of verbal multiword expressions. In Markantonatou, S., Ramisch, C., Savary, A., and Vincze, V., editors, Multiword Expressions at Length and in Depth: Extended Papers from the MWE 2017 Workshop, pages 87–147. Language Science Press, Berlin.
Savary, A., Scholivet, M., Ramisch, C., Nakamura, T., Bilinski, E., Stymne, S., Giouli, V., Markantonatou, S., Pais, V., Mitrofan, M., Estève, L., Guillaume, B., Mititelu, V. B., Čibej, J., Hernández, R. D., Fendel, V., Gantar, P., Kanishcheva, O., Krstev, C., Liebeskind, C., Lobzhanidze, I., Marković, A. M., Nešpore-Bērzkalne, G., Pagano, A. S., Shamsfard, M., Stankovic, R., Tajalli, V., Tiberius, C., and Padhye, A. (2026). Parseme 2.0 multilingual corpus of multiword expressions. In Piperidis, S., Bel, N., van den Heuvel, H., Ide, N., Krek, S., and Toral, A., editors, Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026), pages 4819–4834, Palma, Mallorca, Spain. European Language Resources Association (ELRA).
Scholivet, M., Savary, A., Ramisch, C., Bilinski, E., Nakamura, T., Mitrofan, M., and Pais, V. (2026). Edition 2.0 of the PARSEME shared task on multilingual identification and paraphrasing of multiword expressions. In Ojha, A. K., Mititelu, V. B., Constant, M., Stoyanova, I., Doğruöz, A. S., and Rademaker, A., editors, Proceedings of the 22nd Workshop on Multiword Expressions (MWE 2026), pages 254–275, Rabat, Marocco. Association for Computational Linguistics.
Souza, E., Silveira, A., Cavalcanti, T., Castro, M., and Freitas, C. (2021). PetroGold – corpus padrão ouro para o dominio do petroleo. In Ruiz, E. E. S. and Torrent, T. T., editors, Proceedings of the 13th Brazilian Symposium in Information and Human Language Technology, pages 29–38, Porto Alegre, Brazil. Association for Computational Linguistics.
Straka, M., Hajič, J., and Straková, J. (2016). Udpipe: Trainable pipeline for processing conll-u files performing tokenization, morphological analysis, pos tagging and parsing. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC), pages 4290–4297.
van Gompel, M. (2024). Folia linguistic annotation tool (version 0.11.5). KNAW Humanities Cluster CLST, Radboud University.
Dezotti, M. C. C., Berliner, E., and Duarte, A. (2013). Esopo: fábulas completas. Cosac Naify, São Paulo, Brazil.
Duran, M. (2021). Manual de anotação de pos tags: orientações para anotação de etiquetas morfossintáticas em língua portuguesa, segundo as diretrizes da abordagem universal dependencies. Technical Report Relatório Técnico No. 434, Universidade de São Paulo, ICMC. Accessed: 2026-04-10.
Duran, M. (2022). Manual de anotação de relações de dependência: versão revisada e estendida. Technical Report Relatório Técnico No. 440, Universidade de São Paulo, ICMC. Accessed: 2026-04-30.
Guibon, G. et al. (2020). When collaborative treebank curation meets graph grammars. In Proceedings of the Twelfth Language Resources and Evaluation Conference (LREC), pages 5291–5300.
Neves, M. H. M. (2014). A gramática pela fábula. ou: A fábula pela gramática. Linguística, 30(1):165–196.
Nivre, J. et al. (2020). Universal dependencies v2: An evergrowing multilingual treebank collection. arXiv preprint arXiv:2004.10643.
PARSEME Initiative (2026). Parseme corpora wiki: Citing parseme. [link]. Accessed: 2026-05-19.
Pasquer, C., Savary, A., Ramisch, C., and Antoine, J.-Y. (2020). Verbal multiword expression identification: Do we need a sledgehammer to crack a nut? In Proceedings of the 28th International Conference on Computational Linguistics (COLING 2020), Barcelona, Spain. HAL Id: hal-03013636.
Ramisch, C. et al. (2018). A corpus study of verbal multiword expressions in brazilian portuguese. In Computational Processing of the Portuguese Language: 13th International Conference, PROPOR 2018, Lecture Notes in Artificial Intelligence, Cham, Switzerland. Springer International Publishing.
Ramisch, R., Ramisch, C., and Villavicencio, A. (2026). Expressões multipalavras: fogo de palha ou osso duro de roer? In Caseli, H. M. and Nunes, M. G. V., editors, Processamento de Linguagem Natural: Conceitos, Técnicas e Aplicações em Português, volume 1, book chapter 5. BPLN, 4 edition.
Savary, A. et al. (2017). The parseme shared task on automatic identification of verbal multiword expressions. In Proceedings of the 13th Workshop on Multiword Expressions, pages 31–47. Association for Computational Linguistics.
Savary, A. et al. (2018). Parseme multilingual corpus of verbal multiword expressions. In Markantonatou, S., Ramisch, C., Savary, A., and Vincze, V., editors, Multiword Expressions at Length and in Depth: Extended Papers from the MWE 2017 Workshop, pages 87–147. Language Science Press, Berlin.
Savary, A., Scholivet, M., Ramisch, C., Nakamura, T., Bilinski, E., Stymne, S., Giouli, V., Markantonatou, S., Pais, V., Mitrofan, M., Estève, L., Guillaume, B., Mititelu, V. B., Čibej, J., Hernández, R. D., Fendel, V., Gantar, P., Kanishcheva, O., Krstev, C., Liebeskind, C., Lobzhanidze, I., Marković, A. M., Nešpore-Bērzkalne, G., Pagano, A. S., Shamsfard, M., Stankovic, R., Tajalli, V., Tiberius, C., and Padhye, A. (2026). Parseme 2.0 multilingual corpus of multiword expressions. In Piperidis, S., Bel, N., van den Heuvel, H., Ide, N., Krek, S., and Toral, A., editors, Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026), pages 4819–4834, Palma, Mallorca, Spain. European Language Resources Association (ELRA).
Scholivet, M., Savary, A., Ramisch, C., Bilinski, E., Nakamura, T., Mitrofan, M., and Pais, V. (2026). Edition 2.0 of the PARSEME shared task on multilingual identification and paraphrasing of multiword expressions. In Ojha, A. K., Mititelu, V. B., Constant, M., Stoyanova, I., Doğruöz, A. S., and Rademaker, A., editors, Proceedings of the 22nd Workshop on Multiword Expressions (MWE 2026), pages 254–275, Rabat, Marocco. Association for Computational Linguistics.
Souza, E., Silveira, A., Cavalcanti, T., Castro, M., and Freitas, C. (2021). PetroGold – corpus padrão ouro para o dominio do petroleo. In Ruiz, E. E. S. and Torrent, T. T., editors, Proceedings of the 13th Brazilian Symposium in Information and Human Language Technology, pages 29–38, Porto Alegre, Brazil. Association for Computational Linguistics.
Straka, M., Hajič, J., and Straková, J. (2016). Udpipe: Trainable pipeline for processing conll-u files performing tokenization, morphological analysis, pos tagging and parsing. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC), pages 4290–4297.
van Gompel, M. (2024). Folia linguistic annotation tool (version 0.11.5). KNAW Humanities Cluster CLST, Radboud University.
Publicado
19/10/2026
Como Citar
GUIMARÃES, Letícia Guedes; PAGANO, Adriana; CONEGLIAN, André V. Lopes; VILLAVICENCIO, Aline.
Developing Multilingual Multiword Expressions (MWEs) Resources: An Annotated Corpus of Fables following the PARSEME guidelines. In: SIMPÓSIO BRASILEIRO DE TECNOLOGIA DA INFORMAÇÃO E DA LINGUAGEM HUMANA (STIL), 17. , 2026, Cuiabá/MT.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2026
.
p. 200-208.
DOI: https://doi.org/10.5753/stil.2026.26591.
