Query-CNAT: Uma Ferramenta de Busca Semântica Híbrida para Bancos de Dados Relacionais Médicos
Resumo
A recuperação de informações em bancos de dados médicos é desafiadora devido à especificidade clínica. Embora LLMs mitiguem a rigidez sintática, sua natureza probabilística apresenta riscos de alucinações. Este artigo demonstra o Query-CNAT, uma ferramenta extrativa de busca semântica híbrida que traduz intenções médicas em subconjuntos tabulares. Combinando embeddings biomédicos especializados (BioBERT-PT e BioWordVec) e otimização evolutiva, a aplicação recupera dados precisos sem depender de APIs generativas externas. Operando localmente, garante conformidade com a LGPD. A demonstração1 ilustra sua eficácia em um banco de dados epidemiológico real.
Palavras-chave:
Busca Semântica Híbrida, Bancos de Dados Relacionais Médicos, Embeddings Biomédicos, Algoritmos Genéticos, Recuperação de Informação
Referências
Zhao, W.X., et al.: A survey of large language models. Frontiers of Computer Science 20(12), 2012627 (2026). DOI: 10.1007/s11704-026-60308-3
Kalyan, K., Rajasekharan, A., Sangeetha, S.: AMMUS: a survey of transformer-based pretrained models in natural language processing. arXiv preprint arXiv:2108.05542 (2021).
Gad, A.F.: PyGAD: an intuitive genetic algorithm Python library. Multimedia Tools Appl. (2023). DOI: 10.1007/s11042-023-17167-y
Grundy, S.M., et al.: 2018 AHA/ACC guideline on the management of blood cholesterol. Journal of the American College of Cardiology 73(24), e285–e350 (2019). DOI: 10.1016/j.jacc.2018.11.003
Mirjalili, S., Dong, J., Lewis, A.: Nature-Inspired Optimizers: Theories, Literature Reviews and Applications. Springer (2020). DOI: 10.1007/978-3-030-12127-3
Gao, Y., et al.: Retrieval-augmented generation for large language models: a survey. arXiv preprint arXiv:2312.10997 (2023).
Johnson, A.E.W., et al.: MIMIC-III, a freely accessible critical care database. Scientific Data 3, 160035 (2016). DOI: 10.1038/sdata.2016.35
Ji, Z., et al.: Survey of hallucination in natural language generation. ACM Comput. Surv. 55(12), 1–38 (2023). DOI: 10.1145/3571730
Zhang, Y., et al.: BioWordVec, improving biomedical word embeddings with subword information and MeSH. Scientific Data 6(1), 52 (2019). DOI: 10.1038/s41597-019-0055-0
Lee, J., et al.: BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics 36(4), 1234–1240 (2020). DOI: 10.1093/bioinformatics/btz682
Schneider, E., et al.: BioBERTpt - A Portuguese Neural Language Model for Clinical Named Entity Recognition. In: Proc. Clinical Natural Language Processing Workshop. pp. 65–72 (2020). DOI: 10.18653/v1/2020.clinicalnlp-1.7
Katoch, S., Chauhan, S.S., Kumar, V.: A review on genetic algorithm: past, present, and future. Multimedia Tools Appl. 80, 8091–8126 (2021). DOI: 10.1007/s11042-020-10139-6
Pourreza, M., Rafiei, D.: DIN-SQL: decomposed in-context learning of text-to-SQL with self-correction. In: Adv. Neural Inf. Process. Syst. vol. 36, pp. 36339–36348 (2023).
U.S. Dept. Health Human Serv.: HIPAA of 1996. Pub. L. 104-191 (1996).
Presidência da República do Brasil: LGPD, Lei nº 13.709. Diário Oficial da União (2018).
Kalyan, K., Rajasekharan, A., Sangeetha, S.: AMMUS: a survey of transformer-based pretrained models in natural language processing. arXiv preprint arXiv:2108.05542 (2021).
Gad, A.F.: PyGAD: an intuitive genetic algorithm Python library. Multimedia Tools Appl. (2023). DOI: 10.1007/s11042-023-17167-y
Grundy, S.M., et al.: 2018 AHA/ACC guideline on the management of blood cholesterol. Journal of the American College of Cardiology 73(24), e285–e350 (2019). DOI: 10.1016/j.jacc.2018.11.003
Mirjalili, S., Dong, J., Lewis, A.: Nature-Inspired Optimizers: Theories, Literature Reviews and Applications. Springer (2020). DOI: 10.1007/978-3-030-12127-3
Gao, Y., et al.: Retrieval-augmented generation for large language models: a survey. arXiv preprint arXiv:2312.10997 (2023).
Johnson, A.E.W., et al.: MIMIC-III, a freely accessible critical care database. Scientific Data 3, 160035 (2016). DOI: 10.1038/sdata.2016.35
Ji, Z., et al.: Survey of hallucination in natural language generation. ACM Comput. Surv. 55(12), 1–38 (2023). DOI: 10.1145/3571730
Zhang, Y., et al.: BioWordVec, improving biomedical word embeddings with subword information and MeSH. Scientific Data 6(1), 52 (2019). DOI: 10.1038/s41597-019-0055-0
Lee, J., et al.: BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics 36(4), 1234–1240 (2020). DOI: 10.1093/bioinformatics/btz682
Schneider, E., et al.: BioBERTpt - A Portuguese Neural Language Model for Clinical Named Entity Recognition. In: Proc. Clinical Natural Language Processing Workshop. pp. 65–72 (2020). DOI: 10.18653/v1/2020.clinicalnlp-1.7
Katoch, S., Chauhan, S.S., Kumar, V.: A review on genetic algorithm: past, present, and future. Multimedia Tools Appl. 80, 8091–8126 (2021). DOI: 10.1007/s11042-020-10139-6
Pourreza, M., Rafiei, D.: DIN-SQL: decomposed in-context learning of text-to-SQL with self-correction. In: Adv. Neural Inf. Process. Syst. vol. 36, pp. 36339–36348 (2023).
U.S. Dept. Health Human Serv.: HIPAA of 1996. Pub. L. 104-191 (1996).
Presidência da República do Brasil: LGPD, Lei nº 13.709. Diário Oficial da União (2018).
Publicado
08/09/2026
Como Citar
SILVA DE LIMA, Denilson José do Bom Jesus; DE PONTES FILHO, Rui Nóbrega; SILVA LIMA, Allan Miller; LIMA MARTINS, Denis Mayr; DE LIMA NETO, Fernando Buarque.
Query-CNAT: Uma Ferramenta de Busca Semântica Híbrida para Bancos de Dados Relacionais Médicos. In: DEMONSTRAÇÕES E APLICAÇÕES - SIMPÓSIO BRASILEIRO DE BANCO DE DADOS (SBBD), 41. , 2026, São Carlos/SP.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2026
.
p. 227-232.
DOI: https://doi.org/10.5753/sbbd_estendido.2026.249483.
