Query-CNAT: Uma Ferramenta de Busca Semântica Híbrida para Bancos de Dados Relacionais Médicos

Resumo


A recuperação de informações em bancos de dados médicos é desafiadora devido à especificidade clínica. Embora LLMs mitiguem a rigidez sintática, sua natureza probabilística apresenta riscos de alucinações. Este artigo demonstra o Query-CNAT, uma ferramenta extrativa de busca semântica híbrida que traduz intenções médicas em subconjuntos tabulares. Combinando embeddings biomédicos especializados (BioBERT-PT e BioWordVec) e otimização evolutiva, a aplicação recupera dados precisos sem depender de APIs generativas externas. Operando localmente, garante conformidade com a LGPD. A demonstração1 ilustra sua eficácia em um banco de dados epidemiológico real.
Palavras-chave: Busca Semântica Híbrida, Bancos de Dados Relacionais Médicos, Embeddings Biomédicos, Algoritmos Genéticos, Recuperação de Informação

Referências

Zhao, W.X., et al.: A survey of large language models. Frontiers of Computer Science 20(12), 2012627 (2026). DOI: 10.1007/s11704-026-60308-3

Kalyan, K., Rajasekharan, A., Sangeetha, S.: AMMUS: a survey of transformer-based pretrained models in natural language processing. arXiv preprint arXiv:2108.05542 (2021).

Gad, A.F.: PyGAD: an intuitive genetic algorithm Python library. Multimedia Tools Appl. (2023). DOI: 10.1007/s11042-023-17167-y

Grundy, S.M., et al.: 2018 AHA/ACC guideline on the management of blood cholesterol. Journal of the American College of Cardiology 73(24), e285–e350 (2019). DOI: 10.1016/j.jacc.2018.11.003

Mirjalili, S., Dong, J., Lewis, A.: Nature-Inspired Optimizers: Theories, Literature Reviews and Applications. Springer (2020). DOI: 10.1007/978-3-030-12127-3

Gao, Y., et al.: Retrieval-augmented generation for large language models: a survey. arXiv preprint arXiv:2312.10997 (2023).

Johnson, A.E.W., et al.: MIMIC-III, a freely accessible critical care database. Scientific Data 3, 160035 (2016). DOI: 10.1038/sdata.2016.35

Ji, Z., et al.: Survey of hallucination in natural language generation. ACM Comput. Surv. 55(12), 1–38 (2023). DOI: 10.1145/3571730

Zhang, Y., et al.: BioWordVec, improving biomedical word embeddings with subword information and MeSH. Scientific Data 6(1), 52 (2019). DOI: 10.1038/s41597-019-0055-0

Lee, J., et al.: BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics 36(4), 1234–1240 (2020). DOI: 10.1093/bioinformatics/btz682

Schneider, E., et al.: BioBERTpt - A Portuguese Neural Language Model for Clinical Named Entity Recognition. In: Proc. Clinical Natural Language Processing Workshop. pp. 65–72 (2020). DOI: 10.18653/v1/2020.clinicalnlp-1.7

Katoch, S., Chauhan, S.S., Kumar, V.: A review on genetic algorithm: past, present, and future. Multimedia Tools Appl. 80, 8091–8126 (2021). DOI: 10.1007/s11042-020-10139-6

Pourreza, M., Rafiei, D.: DIN-SQL: decomposed in-context learning of text-to-SQL with self-correction. In: Adv. Neural Inf. Process. Syst. vol. 36, pp. 36339–36348 (2023).

U.S. Dept. Health Human Serv.: HIPAA of 1996. Pub. L. 104-191 (1996).

Presidência da República do Brasil: LGPD, Lei nº 13.709. Diário Oficial da União (2018).
Publicado
08/09/2026
SILVA DE LIMA, Denilson José do Bom Jesus; DE PONTES FILHO, Rui Nóbrega; SILVA LIMA, Allan Miller; LIMA MARTINS, Denis Mayr; DE LIMA NETO, Fernando Buarque. Query-CNAT: Uma Ferramenta de Busca Semântica Híbrida para Bancos de Dados Relacionais Médicos. In: DEMONSTRAÇÕES E APLICAÇÕES - SIMPÓSIO BRASILEIRO DE BANCO DE DADOS (SBBD), 41. , 2026, São Carlos/SP. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 227-232. DOI: https://doi.org/10.5753/sbbd_estendido.2026.249483.