Empirical Study Towards RBAMR-News: a Rule-Based AMR Parser for News in Portuguese

  • Maria Julia Bernardo Comarim USP / UFSCar
  • Ariani Di-Felippo USP / UFSCar

Resumo


This work evaluates AMR graph construction rules, originally developed for literary texts, in Portuguese news data to support building a symbolic AMR parser for this genre. Using the AMRNews-PT gold-standard corpus, we conduct an ablation analysis to assess rule contributions to AMR representation quality. The results reveal a hierarchical organization in which a small set of core structural rules drives most parsing performance, while additional rules provide finer semantic refinements. This rule hierarchy serves as the basis for building the parser and enabling the annotation of a larger news corpus.

Referências

Afonso, S., Bick, E., Haber, R., and Santos, D. (2002). Floresta sintá(c)tica: A treebank for portuguese. In Proceedings of the Third International Conference on Language Resources and Evaluation (LREC’02), Las Palmas, Canary Islands, Spain. European Language Resources Association (ELRA).

Agrawal, R. and Srikant, R. (1994). Fast algorithms for mining association rules in large databases. In Proceedings of the 20th International Conference on Very Large Data Bases, page 487–499, San Francisco, CA, USA. Morgan Kaufmann Publishers Inc.

Alva-Manchego, F. E. and Rosa, J. L. G. (2012). Semantic role labeling for brazilian portuguese: A benchmark. In Advances in Artificial Intelligence – IBERAMIA 2012, volume 7637 of LNCS, pages 481–490, Berlin, Heidelberg. Springer.

Anchiêta, R. T. and Pardo, T. A. S. (2018). A rule-based amr parser for portuguese. In Simari, G. R., Fermé, E., Gutiérrez Segura, F., and Rodríguez Melquiades, J. A., editors, Advances in Artificial Intelligence - IBERAMIA 2018, pages 341–353.

Anchiêta, R. T. and Pardo, T. A. S. (2022). Abstract meaning representation parsing for the brazilian portuguese language. In Pinheiro, V., Gamallo, P., Amaro, R., Scarton, C., Batista, F., Silva, D., Magro, C., and Pinto, H., editors, Computational Processing of the Portuguese Language, pages 429–434, Cham. Springer International Publishing.

Bai, X., Chen, Y., and Zhang, Y. (2022). Graph pre-training for AMR parsing and generation. In Muresan, S., Nakov, P., and Villavicencio, A., editors, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6001–6015, Dublin, Ireland. Association for Computational Linguistics.

Banarescu, L., Bonial, C., Cai, S., Georgescu, M., Griffitt, K., Hermjakob, U., Knight, K., Koehn, P., Palmer, M., and Schneider, N. (2013). Abstract Meaning Representation for sembanking. In Pareja-Lora, A., Liakata, M., and Dipper, S., editors, Proceedings of the 7th Linguistic Annotation Workshop and Interoperability with Discourse, pages 178–186, Sofia, Bulgaria. Association for Computational Linguistics.

Bevilacqua, M., Blloshmi, R., and Navigli, R. (2021). One spring to rule them both: Symmetric amr semantic parsing and generation without a complex pipeline. Proceedings of the AAAI Conference on Artificial Intelligence, 35(14):12564–12573.

Bick, E. (2000). The Parsing System Palavras: Automatic Grammatical Analysis of Portuguese in a Constraint Grammar Framework. Aarhus University Press, Aarhus.

Cai, S. and Knight, K. (2013). Smatch: an evaluation metric for semantic feature structures. In Schuetze, H., Fung, P., and Poesio, M., editors, Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 748–752, Sofia, Bulgaria. Association for Computational Linguistics.

Ceregatto, G. (2025). Representação formal de significado: o caso dos tweets do mercado financeiro. Dissertação (mestrado em linguística), Universidade Federal de São Carlos, São Carlos.

Damonte, M., Cohen, S. B., and Satta, G. (2017). An incremental parser for abstract meaning representation. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics (EACL), pages 536–546.

de Marneffe, M.-C., Manning, C. D., Nivre, J., and Zeman, D. (2021). Universal Dependencies. Computational Linguistics, 47(2):255–308.

Duran, M. S. and Aluísio, S. M. (2012). Propbank-br: a Brazilian treebank annotated with semantic role labels. In Calzolari, N., Choukri, K., Declerck, T., Doğan, M. U., Maegaard, B., Mariani, J., Moreno, A., Odijk, J., and Piperidis, S., editors, Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC’12), pages 1862–1867, Istanbul, Turkey. ELRA.

Ettinger, A., Hwang, J., Pyatkin, V., Bhagavatula, C., and Choi, Y. (2023). “you are an expert linguistic annotator”: Limits of LLMs as analyzers of Abstract Meaning Representation. In Bouamor, H., Pino, J., and Bali, K., editors, Findings of the Association for Computational Linguistics: EMNLP 2023, pages 8250–8263, Singapore. Association for Computational Linguistics.

Freitas, C. (2024). Anotação de papéis semânticos no corpus porttinari-base: procedimentos, resultados e análises. Technical Report 450, Instituto de Ciências Matemáticas e de Computação, Universidade de São Paulo, São Carlos-SP.

Freitas, C. and Pardo, T. A. S. (2024). Propbank e anotação de papéis semânticos para a língua portuguesa: O que há de novo? In Proceedings of the 15th Symposium in Information and Human Language Technology (STIL), pages 118–128, Belém-PA, Brazil.

Freitas, C. and Pardo, T. A. S. (2025). Propbanks e representações semânticas: o que temos, o que queremos e o que podemos. LinguaMÁTICA, 17(2):1–29.

Hartmann, N., Fonseca, E., Shulby, C., Treviso, M., Silva, J., and Aluísio, S. (2017). Portuguese word embeddings: Evaluating on word analogies and natural language tasks. In Proceedings of the 11th Brazilian Symposium in Information and Human Language Technology, pages 122–131.

Hartmann, N. S., Duran, M. S., and Aluísio, S. M. (2016). Automatic semantic role labeling on non-revised syntactic trees of journalistic texts. In International Conference on Computational Processing of the Portuguese Language, pages 202–212.

Heinecke, J. (2023). metamorphosed, a graphical editor for abstract meaning representation. In Proceedings of the 19th Joint ACL-ISO Workshop on Interoperable Semantics (ISA-19), pages 27–32, Nancy, France. Association for Computational Linguistics.

Hermjakob, U. (2013). Amr editor: A tool to build abstract meaning representations. [link]. Accessed: 25 Aug 2024.

Ho, S. H. (2025). Evaluation of finetuned llms in amr parsing. arXiv preprint arXiv:2508.05028. 27 pages, 32 figures.

Inácio, M. L., Cabezudo, M. A. S., Ramisch, R., Di-Felippo, A., and Pardo, T. A. S. (2023). The amr-pt corpus and the semantic annotation of challenging sentences from journalistic and opinion texts. DELTA: Documentação de Estudos em Linguística Teórica e Aplicada, 39(3):1–31.

Jurafsky, D. and Martin, J. H. (2026). Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition, with Language Models. 3rd edition. Online manuscript released January 6, 2026.

Li, Y. and Fowlie, M. (2025). Gpt makes a poor amr parser. Journal for Language Technology and Computational Linguistics, 38(2):43–76.

Lopes, L. and Pardo, T. (2024). Towards portparser - a highly accurate parsing system for Brazilian Portuguese following the Universal Dependencies framework. In Gamallo, P., Claro, D., Teixeira, A., Real, L., Garcia, M., Oliveira, H. G., and Amaro, R., editors, Proceedings of the 16th International Conference on Computational Processing of Portuguese - Vol. 1, pages 401–410, Santiago de Compostela, Galicia/Spain. Association for Computational Lingustics.

Mansouri, B. (2025). Survey of abstract meaning representation: Then, now, future. Meyes, R., Lu, M., de Puiseau, C. W., and Meisen, T. (2019). Ablation studies in artificial neural networks. CoRR, abs/1901.08644. Oliveira, S., Loureiro, D., and Jorge, A. (2021). Improving portuguese semantic role labeling with transformers and transfer learning. In 2021 IEEE 8th International Conference on Data Science and Advanced Analytics (DSAA), pages 1–9.

Palmer, M., Gildea, D., and Kingsbury, P. (2005). The Proposition Bank: An annotated corpus of semantic roles. Computational Linguistics, 31(1):71–106.

Sadeddine, Z., Opitz, J., and Suchanek, F. (2024). A survey of meaning representations – from theory to practical utility. In Duh, K., Gomez, H., and Bethard, S., editors, Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 2877–2892, Mexico City, Mexico. Association for Computational Linguistics.

Sanches Duran, M. and Aluísio, S. (2015). Automatic generation of a lexical resource to support semantic role labeling in Portuguese. In Palmer, M., Boleda, G., and Rosso, P., editors, Proceedings of the Fourth Joint Conference on Lexical and Computational Semantics, pages 216–221, Denver, Colorado. Association for Computational Linguistics.

Sobrevilla Cabezudo, M. A. and Pardo, T. (2019). Towards a general Abstract Meaning Representation corpus for Brazilian Portuguese. In Friedrich, A., Zeyrek, D., and Hoek, J., editors, Proceedings of the 13th Linguistic Annotation Workshop, pages 236–244, Florence, Italy. Association for Computational Linguistics.

Souza, F., Nogueira, R., and Lotufo, R. (2020). Bertimbau: Pretrained bert models for brazilian portuguese. In Proceedings of the 9th Brazilian Conference on Intelligent Systems (BRACIS), pages 403–417, Cham. Springer.

Straka, M., Straková, J., and Gamba, F. (2024). ÚFAL LatinPipe at EvaLatin 2024: Morphosyntactic analysis of Latin. In Sprugnoli, R. and Passarotti, M., editors, Proceedings of the Third Workshop on Language Technologies for Historical and Ancient Languages (LT4HALA), pages 207–214, Torino, Italia. ELRA and ICCL.

Vasylenko, P., Huguet Cabot, P. L., Martínez Lorenzo, A. C., and Navigli, R. (2023). Incorporating graph information in transformer-based AMR parsing. In Rogers, A., Boyd-Graber, J., and Okazaki, N., editors, Findings of the Association for Computational Linguistics: ACL 2023, pages 1995–2011, Toronto, Canada. Association for Computational Linguistics.

Wang, C., Xue, N., and Pradhan, S. (2015). Boosting transition-based AMR parsing with refined actions and auxiliary analyzers. In Zong, C. and Strube, M., editors, Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing, pages 857–862, Beijing, China. ACL.

Wein, S. and Opitz, J. (2024). A survey of AMR applications. In Al-Onaizan, Y., Bansal, M., and Chen, Y.-N., editors, Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 6856–6875, Miami, Florida, USA. Association for Computational Linguistics.

Zhou, J., Naseem, T., Fernandez Astudillo, R., Lee, Y.-S., Florian, R., and Roukos, S. (2021). Structure-aware fine-tuning of sequence-to-sequence transformers for transition-based AMR parsing. In Moens, M.-F., Huang, X., Specia, L., and Yih, S. W.-t., editors, Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6279–6290, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
Publicado
19/10/2026
COMARIM, Maria Julia Bernardo; DI-FELIPPO, Ariani. Empirical Study Towards RBAMR-News: a Rule-Based AMR Parser for News in Portuguese. In: SIMPÓSIO BRASILEIRO DE TECNOLOGIA DA INFORMAÇÃO E DA LINGUAGEM HUMANA (STIL), 17. , 2026, Cuiabá/MT. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 86-99. DOI: https://doi.org/10.5753/stil.2026.26504.