IDEA-C2: uma abordagem híbrida de modelagem conceitual apoiada por um modelo de linguagem e um metamodelo de dados no contexto de Comando e Controle
Resumo
A tese propõe o IDEA-C2, uma abordagem supervisionada híbrida que combina textos doutrinários, Recursos Semânticos e um metamodelo de alto nível para anotar corpora e ajustar Modelos de Linguagem (ML). Diferente das abordagens orientadas por dados (data-driven), que constroem Modelos de domínio (DM) subsimbólicos, e das orientadas por teoria (theory-driven), que enfrentam desafios na extração de classes e relações de textos, o IDEA-C2 emprega pré-anotação heurística e gera Knowledge Graphs (KG) flexíveis para consultas exploratórias e realização de inferências, apoiando a construção do DM. Avaliada em seis experimentos, a abordagem IDEA-C2 alcanc¸ou 95% de precisão em classes e 76% em relações na pré-anotação, além de uma precisão e cobertura acima de 85% no ML ajustado ao contexto. Em experimento com 28 participantes, 40% das classes e relações do KG se mostraram similares aos DM construídos de modo tradicional. Os resultados demonstram a utilidade e viabilidade do IDEA-C2 na geração de artefatos e na aplicação prática.
Palavras-chave:
Modelagem conceitual, híbrida, modelo de domínio, modelo de linguagem
Referências
Fries, J. A., Steinberg, E., Khattar, S., Fleming, S. L., Posada, J., Callahan, A., and Shah, N. H. (2021). Ontology-driven weak supervision for clinical entity classification in electronic health records. Nature communications, 12(1):2017.
Guizzardi, G., Pastor, O., and Storey, V. C. (2023). Thinking Fast and Slow in Software Engineering . IEEE Software, 40(06):139–142.
Hogan, A., Blomqvist, E., Cochez, M., et al. (2021). Knowledge graphs. ACM Computing Surveys, 54(4).
Lee, J., Yoon, W., Kim, S., Kim, D., Kim, S., So, C. H., and Kang, J. (2019). BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, 36(4):1234–1240.
Liu, P., Qian, L., Zhao, X., and Tao, B. (2023). The construction of knowledge graphs in the aviation assembly domain based on a joint knowledge extraction model. IEEE Access, 11:26483–26495.
Luan, Y., He, L., Ostendorf, M., and Hajishirzi, H. (2018). Multi-task identification of entities, relations, and coreference for scientific knowledge graph construction. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 3219–3232, Brussels, Belgium. Association for Computational Linguistics.
Mintz, M., Bills, S., Snow, R., and Jurafsky, D. (2009). Distant supervision for relation extraction without labeled data. In Proceedings of the Joint Conference of the 47th Annual Meeting of the ACL and the 4th International Joint Conference on Natural Language Processing of the AFNLP, pages 1003–1011, Suntec, Singapore. Association for Computational Linguistics.
Rosa, G. F., Avelino, J. O., Cavalcanti, M. C., and Duarte, J. C. (2025). TForMIX: A method that combines LLM and Multidimensional Modeling for Technological Foresight. IEEE Access, 13:153320–153339.
Saba, W. S. (2023). Stochastic llms do not understand language: Towards symbolic, explainable and ontologically based llms. In Almeida, J. P. A., Borbinha, J., Guizzardi, G., Link, S., and Zdravkovic, J., editors, Conceptual Modeling, pages 3–19, Cham. Springer Nature Switzerland.
Souza, F., Nogueira, R., and Lotufo, R. (2020). Bertimbau: Pretrained bert models for brazilian portuguese. In Cerri, R. and Prati, R. C., editors, Intelligent Systems, pages 403–417, Cham. Springer International Publishing.
Yang, J., Han, S. C., and Poon, J. (2022). A survey on extraction of causal relations from natural language text. Knowledge and Information Systems, 64(5):1161–1186.
Zhao, Q., Huang, H., and Ding, H. (2021). Study on military regulations knowledge construction based on knowledge graph. In 2021 7th International Conference on Big Data and Information Analytics (BigDIA), pages 180–184.
Zhou, J., Li, X., Wang, S., and Song, X. (2022). Ner-based military simulation scenario development process. The Journal of Defense Modeling and Simulation, 20(4):563–575.
Guizzardi, G., Pastor, O., and Storey, V. C. (2023). Thinking Fast and Slow in Software Engineering . IEEE Software, 40(06):139–142.
Hogan, A., Blomqvist, E., Cochez, M., et al. (2021). Knowledge graphs. ACM Computing Surveys, 54(4).
Lee, J., Yoon, W., Kim, S., Kim, D., Kim, S., So, C. H., and Kang, J. (2019). BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, 36(4):1234–1240.
Liu, P., Qian, L., Zhao, X., and Tao, B. (2023). The construction of knowledge graphs in the aviation assembly domain based on a joint knowledge extraction model. IEEE Access, 11:26483–26495.
Luan, Y., He, L., Ostendorf, M., and Hajishirzi, H. (2018). Multi-task identification of entities, relations, and coreference for scientific knowledge graph construction. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 3219–3232, Brussels, Belgium. Association for Computational Linguistics.
Mintz, M., Bills, S., Snow, R., and Jurafsky, D. (2009). Distant supervision for relation extraction without labeled data. In Proceedings of the Joint Conference of the 47th Annual Meeting of the ACL and the 4th International Joint Conference on Natural Language Processing of the AFNLP, pages 1003–1011, Suntec, Singapore. Association for Computational Linguistics.
Rosa, G. F., Avelino, J. O., Cavalcanti, M. C., and Duarte, J. C. (2025). TForMIX: A method that combines LLM and Multidimensional Modeling for Technological Foresight. IEEE Access, 13:153320–153339.
Saba, W. S. (2023). Stochastic llms do not understand language: Towards symbolic, explainable and ontologically based llms. In Almeida, J. P. A., Borbinha, J., Guizzardi, G., Link, S., and Zdravkovic, J., editors, Conceptual Modeling, pages 3–19, Cham. Springer Nature Switzerland.
Souza, F., Nogueira, R., and Lotufo, R. (2020). Bertimbau: Pretrained bert models for brazilian portuguese. In Cerri, R. and Prati, R. C., editors, Intelligent Systems, pages 403–417, Cham. Springer International Publishing.
Yang, J., Han, S. C., and Poon, J. (2022). A survey on extraction of causal relations from natural language text. Knowledge and Information Systems, 64(5):1161–1186.
Zhao, Q., Huang, H., and Ding, H. (2021). Study on military regulations knowledge construction based on knowledge graph. In 2021 7th International Conference on Big Data and Information Analytics (BigDIA), pages 180–184.
Zhou, J., Li, X., Wang, S., and Song, X. (2022). Ner-based military simulation scenario development process. The Journal of Defense Modeling and Simulation, 20(4):563–575.
Publicado
08/09/2026
Como Citar
AVELINO, Jones O.; CORDEIRO, Kelli F.; CAVALCANTI, Maria Cláudia.
IDEA-C2: uma abordagem híbrida de modelagem conceitual apoiada por um modelo de linguagem e um metamodelo de dados no contexto de Comando e Controle. In: CONCURSO DE TESES E DISSERTAÇÕES (CTDBD) - SIMPÓSIO BRASILEIRO DE BANCO DE DADOS (SBBD), 41. , 2026, São Carlos/SP.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2026
.
p. 384-388.
DOI: https://doi.org/10.5753/sbbd_estendido.2026.249503.
