Evaluating LLM-based Triple Extraction for Knowledge Graph Fact-Checking in Portuguese

  • Roney Lira de Sales Santos UFBA
  • Lucas dos Santos UFBA
  • João Pedro Holanda Souza UNOPAR

Resumo


Automatic fake news detection in Portuguese is still often treated as a text classification task, without explicitly representing the factual veracity of the claims. In this work, we evaluate LLM-based triple extraction in a knowledge graph (KG)-based fact-checking system. We replace the Open Information Extraction component of a previous approach with triples generated by Sabiá 4, while keeping the remaining pipeline unchanged. The graph is built only from triples extracted from true news articles and is used as factual support to verify new instances. The experiments use true and fake news articles across four evaluation settings. In the closed setting with the complete KG, LLM-based extraction achieves an F1 score of 0.9992, outperforming the OIE-based configuration. However, in more restrictive evaluation settings, expressive triples require entity normalization, relation standardization, and semantic consolidation to provide factual support beyond direct evidence.

Referências

Auer, S., Bizer, C., Kobilarov, G., Lehmann, J., Cyganiak, R., and Ives, Z. (2007). DBpedia: A nucleus for a web of open data. In The Semantic Web, volume 4825 of Lecture Notes in Computer Science, pages 722–735, Berlin, Heidelberg. Springer.

Banko, M., Cafarella, M. J., Soderland, S., Broadhead, M., and Etzioni, O. (2007). Open information extraction from the web. In Proceedings of the 20th International Joint Conference on Artificial Intelligence, pages 2670–2676. Morgan Kaufmann Publishers Inc.

Chavarro, J. P., Carvalho, D. A., and Rezende, S. O. (2023). FakeTrueBR: Um corpus brasileiro de notícias falsas. In Anais da XVIII Escola Regional de Banco de Dados, pages 53–62, Porto Alegre, RS, Brasil. SBC.

Ciampaglia, G. L., Shiralkar, P., Rocha, L. M., Bollen, J., Menczer, F., and Flammini, A. (2015). Computational fact checking from knowledge networks. PLOS ONE, 10(6):e0128193.

Dijkstra, E. W. (1959). A note on two problems in connexion with graphs. Numerische Mathematik, 1:269–271.

Garcia, G. L., Afonso, L. C. S., and Papa, J. P. (2022). FakeRecogna: A new brazilian corpus for fake news detection. In Pinheiro, V., Gamallo, P., Amaro, R., Scarton, C., Batista, F., Silva, D., Magro, C., and Pinto, H., editors, Computational Processing of the Portuguese Language, volume 13208 of Lecture Notes in Computer Science, pages 57–67, Cham. Springer.

Geifman, Y. and El-Yaniv, R. (2017). Selective classification for deep neural networks. In Advances in Neural Information Processing Systems, volume 30.

Guo, Z., Schlichtkrull, M., and Vlachos, A. (2022). A survey on automated fact-checking. Transactions of the Association for Computational Linguistics, 10:178–206.

Hagberg, A. A., Schult, D. A., and Swart, P. J. (2008). Exploring network structure, dynamics, and function using NetworkX. In Proceedings of the 7th Python in Science Conference, pages 11–15.

Laitz, T., Almeida, T. S., Abonizio, H., Malaquias Junior, R., Bonás, G. K., Piau, M., Larcher, C., Pires, R., and Nogueira, R. (2026). Sabiá-4 technical report.

Mihindukulasooriya, N., Tiwari, A., Galkin, M., Costabello, L., and Simperl, E. (2025). Automatic prompt optimization for knowledge graph construction with large language models. In Proceedings of the Workshops of the VLDB 2025 Conference.

Monteiro, R. A., Santos, R. L. S., Pardo, T. A. S., Almeida, T. A., Ruiz, E. E. S., and Vale, O. A. (2018). Contributions to the study of fake news in portuguese: New corpus and automatic detection results. In Computational Processing of the Portuguese Language, volume 11122 of Lecture Notes in Computer Science, pages 324–334, Cham. Springer.

OpenAI, Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., Avila, R., Babuschkin, I., Balaji, S., Balcom, V., Baltescu, P., Bao, H., Bavarian, M., Belgum, J., Bello, I., Berdine, J., Bernadett-Shapiro, G., Berner, C., Bogdonoff, L., Boiko, O., Boyd, M., Brakman, A.-L., Brockman, G., Brooks, T., Brundage, M., Button, K., Cai, T., Campbell, R., Cann, A., Carey, B., Carlson, C., Carmichael, R., Chan, B., Chang, C., Chantzis, F., Chen, D., Chen, S., Chen, R., Chen, J., Chen, M., Chess, B., Cho, C., Chu, C., Chung, H. W., Cummings, D., Currier, J., Dai, Y., Decareaux, C., Degry, T., Deutsch, N., Deville, D., Dhar, A., Dohan, D., Dowling, S., Dunning, S., Ecoffet, A., Eleti, A., Eloundou, T., Farhi, D., Fedus, L., Felix, N., Fishman, S. P., Forte, J., Fulford, I., Gao, L., Georges, E., Gibson, C., Goel, V., Gogineni, T., Goh, G., Gontijo-Lopes, R., Gordon, J., Grafstein, M., Gray, S., Greene, R., Gross, J., Gu, S. S., Guo, Y., Hallacy, C., Han, J., Harris, J., He, Y., Heaton, M., Heidecke, J., Hesse, C., Hickey, A., Hickey, W., Hoeschele, P., Houghton, B., Hsu, K., Hu, S., Hu, X., Huizinga, J., Jain, S., Jain, S., Jang, J., Jiang, A., Jiang, R., Jin, H., Jin, D., Jomoto, S., Jonn, B., Jun, H., Kaftan, T., Łukasz Kaiser, Kamali, A., Kanitscheider, I., Keskar, N. S., Khan, T., Kilpatrick, L., Kim, J. W., Kim, C., Kim, Y., Kirchner, J. H., Kiros, J., Knight, M., Kokotajlo, D., Łukasz Kondraciuk, Kondrich, A., Konstantinidis, A., Kosic, K., Krueger, G., Kuo, V., Lampe, M., Lan, I., Lee, T., Leike, J., Leung, J., Levy, D., Li, C. M., Lim, R., Lin, M., Lin, S., Litwin, M., Lopez, T., Lowe, R., Lue, P., Makanju, A., Malfacini, K., Manning, S., Markov, T., Markovski, Y., Martin, B., Mayer, K., Mayne, A., McGrew, B., McKinney, S. M., McLeavey, C., McMillan, P., McNeil, J., Medina, D., Mehta, A., Menick, J., Metz, L., Mishchenko, A., Mishkin, P., Monaco, V., Morikawa, E., Mossing, D., Mu, T., Murati, M., Murk, O., Mély, D., Nair, A., Nakano, R., Nayak, R., Neelakantan, A., Ngo, R., Noh, H., Ouyang, L., O’Keefe, C., Pachocki, J., Paino, A., Palermo, J., Pantuliano, A., Parascandolo, G., Parish, J., Parparita, E., Passos, A., Pavlov, M., Peng, A., Perelman, A., de Avila Belbute Peres, F., Petrov, M., de Oliveira Pinto, H. P., Michael, Pokorny, Pokrass, M., Pong, V. H., Powell, T., Power, A., Power, B., Proehl, E., Puri, R., Radford, A., Rae, J., Ramesh, A., Raymond, C., Real, F., Rimbach, K., Ross, C., Rotsted, B., Roussez, H., Ryder, N., Saltarelli, M., Sanders, T., Santurkar, S., Sastry, G., Schmidt, H., Schnurr, D., Schulman, J., Selsam, D., Sheppard, K., Sherbakov, T., Shieh, J., Shoker, S., Shyam, P., Sidor, S., Sigler, E., Simens, M., Sitkin, J., Slama, K., Sohl, I., Sokolowsky, B., Song, Y., Staudacher, N., Such, F. P., Summers, N., Sutskever, I., Tang, J., Tezak, N., Thompson, M. B., Tillet, P., Tootoonchian, A., Tseng, E., Tuggle, P., Turley, N., Tworek, J., Uribe, J. F. C., Vallone, A., Vijayvergiya, A., Voss, C., Wainwright, C., Wang, J. J., Wang, A., Wang, B., Ward, J., Wei, J., Weinmann, C., Welihinda, A., Welinder, P., Weng, J., Weng, L., Wiethoff, M., Willner, D., Winter, C., Wolrich, S., Wong, H., Workman, L., Wu, S., Wu, J., Wu, M., Xiao, K., Xu, T., Yoo, S., Yu, K., Yuan, Q., Zaremba, W., Zellers, R., Zhang, C., Zhang, M., Zhao, S., Zheng, T., Zhuang, J., Zhuk, W., and Zoph, B. (2024). Gpt-4 technical report.

Santos, L. d., Santos, M. R. E., Souza, Y. S., Sousa, J. P. H., and Santos, R. L. d. S. (2026). Exploring knowledge graphs for automatic fake news detection in Portuguese. In Proceedings of the 17th International Conference on Computational Processing of Portuguese (PROPOR 2026) - Vol. 1, pages 446–455, Salvador, Brazil. Association for Computational Linguistics.

Santos, R. L. d. S. (2022). Detecção automática de notícias falsas em português. Tese (doutorado em ciências – ciências de computação e matemática computacional), Universidade de São Paulo, São Carlos.

Santos, R. L. d. S. and Pardo, T. A. S. (2020). Fact-checking for portuguese: Knowledge graph and google search-based methods. In Quaresma, P., Vieira, R., Aluísio, S., Moniz, H., Batista, F., and Gonçalves, T., editors, Computational Processing of the Portuguese Language, volume 12037 of Lecture Notes in Computer Science, pages 195–205, Cham. Springer.

Silva, R. M., Santos, R. L., Almeida, T. A., and Pardo, T. A. (2020). Towards automatically filtering fake news in portuguese. Expert Systems with Applications, 146:113199.

Stanovsky, G. and Dagan, I. (2018). Supervised open information extraction. In Proceedings of NAACL-HLT, pages 885–895.

Thorne, J., Vlachos, A., Christodoulopoulos, C., and Mittal, A. (2018). FEVER: A large-scale dataset for fact extraction and VERification. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 809–819, New Orleans, Louisiana. Association for Computational Linguistics.

Van Cauter, Z., Demeester, T., Vercruyssen, V., and Ongenae, F. (2024). Ontology-guided knowledge graph construction from maintenance short texts. In Proceedings of the 1st Workshop on Knowledge Augmented Large Language Models, pages 79–88. Association for Computational Linguistics.

Vosoughi, S., Roy, D., and Aral, S. (2018). The spread of true and false news online. Science, 359(6380):1146–1151.

Xu, D., Chen, W., Peng, W., Zhang, C., Xu, T., Zhao, X., Wu, X., Zheng, Y., and Wang, E. (2024). Large language models for generative information extraction: A survey. Frontiers of Computer Science, 18(6):186357.
Publicado
19/10/2026
SANTOS, Roney Lira de Sales; SANTOS, Lucas dos; SOUZA, João Pedro Holanda. Evaluating LLM-based Triple Extraction for Knowledge Graph Fact-Checking in Portuguese. In: SIMPÓSIO BRASILEIRO DE TECNOLOGIA DA INFORMAÇÃO E DA LINGUAGEM HUMANA (STIL), 17. , 2026, Cuiabá/MT. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 361-374. DOI: https://doi.org/10.5753/stil.2026.26618.