Expressões multipalavras no discurso de fóruns sobre saúde mental: desenvolvimento de um corpus em inglês e português brasileiro

  • Gabriela Berndt UFMG
  • Sofia Serrano UFMG
  • Mariana Jin UFMG
  • Adriana Pagano UFMG

Resumo


Expressões multipalavras (multiword expressions, MWEs), isto é, sequências de palavras que apresentam algum grau de idiossincrasia, são centrais tanto para os estudos linguísticos quanto para o Processamento de Linguagem Natural (PLN). Este artigo apresenta a criação de um novo corpus de MWEs em postagens em inglês e português brasileiro de comunidades do Reddit dedicadas à saúde mental — um domínio ainda sub-representado nos recursos existentes de MWEs, predominantemente baseados em textos jornalísticos. O corpus foi compilado manualmente, anotado automaticamente com informações morfossintáticas e sintáticas por meio do parser UDPipe, revisado no Arborator Grew e anotado manualmente na plataforma FLAT, seguindo as diretrizes do PARSEME. A análise combina a distribuição de frequência por categoria de MWE e uma investigação contrastiva interlinguística. Os resultados evidenciam especificidades das MWEs nesse domínio e construções que funcionam como recursos convencionalizados para expressar sentimentos, avaliações e estados subjetivos no discurso sobre saúde mental nas redes sociais.

Referências

Baldwin, T. and Kim, S. N. (2010). Multiword expressions. In Indurkha, N.andDamerau, F. J., editor, Handbook of Natural Language Processing. CRC Press.

Calzolari, N., Fillmore, C. J., Grishman, R., Ide, N., Lenci, A., MacLeod, C., and Zampolli, A. (2002). Towards best practice for multiword expressions in computational lexicons. In González Rodríguez, M. and Suarez Araujo, C. P., editors, Proceedings of the Third International Conference on Language Resources and Evaluation (LREC’02), Las Palmas, Canary Islands - Spain. uropean Language Resources Association (ELRA).

Coll-Florit, M.; Oliver, A. R. S. (2021). Metaphors of mental illness: a corpus-based approach analysing first-person accounts of patients and mental health professionals. Cultura, Lenguaje y Representación, 25:85–104.

Constant, M., Eryiǧit, G., Monti, J., van der Plas, L., Ramisch, C., Rosner, M., and Todirascu, A. (2021). Multiword expression processing: A survey. Computational Linguistics, 43(4):837–892.

Cordeiro, S., Villavicencio, A., Idiart, M., and Ramisch, C. (2019). Unsupervised compositionality prediction of nominal compounds. Computational Linguistics, 45(1):1–57.

Duran, M., Lopes, L., das Graças Nunes, M., and Pardo, T. (2023). The dawn of the porttinari multigenre treebank: Introducing its journalistic portion. In Anais do XIV Simpósio Brasileiro de Tecnologia da Informação e da Linguagem Humana, pages 115–124, Porto Alegre, RS, Brasil. SBC.

Firth, J. R. (1957). A Synopsis of Linguistic Theory, pages 1–32. Blackwell, Oxford.

Matthiessen, C. M. I. M.; Teruya, K. L. M. (2010). Key Terms in Systemic Functional Linguistics. Continuum, London.

Ramisch, C. (2015). Multiword Expressions Acquisition: A Generic and Open Framework. springer.

Ramisch, C., Savary, A., Guillaume, B., Waszczuk, J., Candito, M., Vaidya, A., Barbu Mititelu, V., Bhatia, A., Iñurrieta, U., Giouli, V., Güngör, T., Jiang, M., Lichte, T., Liebeskind, C., Monti, J., Ramisch, R., Stymne, S., Walsh, A., and Xu, H. (2020). Edition 1.2 of the PARSEME shared task on semi-supervised identification of verbal multiword expressions. In Proceedings of the Joint Workshop on Multiword Expressions and Electronic Lexicons.

Ramisch, C. and Villavicencio, A. (2018). Computational treatment of multiword expressions. In Mitkov, R., editor, The Oxford Handbook of Computational Linguistics. Oxford University Press, 2nd edition. DOI: 10.1093/oxfordhb/9780199573691.013.56.

Ramisch, R., Ramisch, C., and Villavicencio, A. (2024). Expressões multipalavras: fogo de palha ou osso duro de roer? In Caseli, H. and das Graças Volpe Nunes, M., editors, Processamento de Linguagem Natural: Conceitos, Técnicas e Aplicações em Português, chapter 5. BPLN, 3 edition. [link].

Savary, A., Ben Khelil, C., Ramisch, C., Giouli, V., Barbu Mititelu, V., Hadj Mohamed, N., Krstev, C., Liebeskind, C., Xu, H., Stymne, S., Güngör, T., Pickard, T., Guillaume, B., Bejček, E., Bhatia, A., Candito, M., Gantar, P., Iñurrieta, U., Gatt, A., Kovalevskaite, J., Lichte, T., Ljubešić, N., Monti, J., Parra Escartín, C., Shamsfard, M., Stoyanova, I., Vincze, V., and Walsh, A. (2023). PARSEME corpus release 1.3. In Bhatia, A., Evang, K., Garcia, M., Giouli, V., Han, L., and Taslimipoor, S., editors, Proceedings of the 19th Workshop on Multiword Expressions (MWE 2023), pages 24–35, Dubrovnik, Croatia. Association for Computational Linguistics.

Savary, A., Ramisch, C., Cordeiro, S., Sangati, F., Vincze, V., QasemiZadeh, B., Candito, M., Cap, F., Giouli, V., Stoyanova, I., and Doucet, A. (2017). The PARSEME shared task on automatic identification of verbal multiword expressions. In Markantonatou, S., Ramisch, C., Savary, A., and Vincze, V., editors, Proceedings of the 13th Workshop on Multiword Expressions (MWE 2017), pages 31–47, Valencia, Spain. Association for Computational Linguistics.

Savary, A., Scholivet, M., Ramisch, C., Nakamura, T., Bilinski, E., Stymne, S., Giouli, V., Markantonatou, S., Pais, V., Mitrofan, M., Estève, L., Guillaume, B., Mititelu, V. B., Čibej, J., Hernández, R. D., Fendel, V., Gantar, P., Kanishcheva, O., Krstev, C., Liebeskind, C., Lobzhanidze, I., Marković, A. M., Nešpore-Bērzkalne, G., Pagano, A. S., Shamsfard, M., Stankovic, R., Tajalli, V., Tiberius, C., and Padhye, A. (2026). Parseme 2.0 multilingual corpus of multiword expressions. In Piperidis, S., Bel, N., van den Heuvel, H., Ide, N., Krek, S., and Toral, A., editors, Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026), pages 4819–4834, Palma, Mallorca, Spain. European Language Resources Association (ELRA).

Scholivet, M., Savary, A., Ramisch, C., Bilinski, E., Nakamura, T., Mitrofan, M., and Pais, V. (2026). Edition 2.0 of the PARSEME shared task on multilingual identification and paraphrasing of multiword expressions. In Ojha, A. K., Mititelu, V. B., Constant, M., Stoyanova, I., Doğruöz, A. S., and Rademaker, A., editors, Proceedings of the 22nd Workshop on Multiword Expressions (MWE 2026), pages 254–275, Rabat, Marocco. Association for Computational Linguistics.

Straka, M., Hajič, J., and Straková, J. (2016). UDPipe: Trainable pipeline for processing CoNLL-U files performing tokenization, morphological analysis, POS tagging and parsing. In Calzolari, N., Choukri, K., Declerck, T., Goggi, S., Grobelnik, M., Maegaard, B., Mariani, J., Mazo, H., Moreno, A., Odijk, J., and Piperidis, S., editors, Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC’16), pages 4290–4297, Portorož, Slovenia. European Language Resources Association (ELRA).

Talmy, L. (2000). Toward a Cognitive Semantics. MIT Press,.

Van Gompel, M., Başar, E., Neumann, A., and van der Klis, M. (2024). Proycon/flat: V0.11.5.

Wilkens, R., Caseli, H., Neris, V., and Villavicencio, A. (2026). The visible and the latent linguistic clues of mental health in Brazilian Portuguese textual posts. In Souza, M., de Dios-Flores, I., Santos, D., Freitas, L., Souza, J. W. d. C., and Ribeiro, E., editors, Proceedings of the 17th International Conference on Computational Processing of Portuguese (PROPOR 2026) - Vol. 2, pages 135–147, Salvador, Brazil. Association for Computational Linguistics.

Zeldes, A. (2017). The GUM corpus: Creating multilayer resources in the classroom. Language Resources and Evaluation, 51(3):581–612.
Publicado
19/10/2026
BERNDT, Gabriela; SERRANO, Sofia; JIN, Mariana; PAGANO, Adriana. Expressões multipalavras no discurso de fóruns sobre saúde mental: desenvolvimento de um corpus em inglês e português brasileiro. In: SIMPÓSIO BRASILEIRO DE TECNOLOGIA DA INFORMAÇÃO E DA LINGUAGEM HUMANA (STIL), 17. , 2026, Cuiabá/MT. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 41-49. DOI: https://doi.org/10.5753/stil.2026.26603.