SABRE: A Text-to-Database Search Architecture for Relational Databases
Resumo
Relational database queries depend on SQL and prior knowledge of the schema. This work proposes SABRE, a plain-text search architecture for relational databases based on the vectorization of tuples and queries and on retrieval by semantic similarity. The evaluation used a pilot study for selecting the embedding model and experiments with the TPC-H benchmark, comparing the relational and vector modes in terms of retrieval quality, latency, and end-to-end cost. The vector approach achieved a Recall@K of 0.8459, a Precision@K of 0.8039, and an NDCG@K of 0.8886, indicating good coverage of relevant tuples and adequate ranking quality. However, SQL execution showed lower latency in 21 of the 22 queries, and the vector cost was dominated by the database search stage. The results indicate that SABRE is suitable as a complementary mechanism for semantic retrieval and candidate generation, but not as a direct substitute for exact SQL execution in analytical workloads such as TPC-H.
Referências
Bordawekar, R. and Shmueli, O. (2017). Using word embedding to enable semantic queries in relational databases. In Proceedings of the 1st Workshop on Data Management for End-to-End Machine Learning, DEEM’17, New York, NY, USA. Association for Computing Machinery.
Chen, C., Jin, C., Zhang, Y., Podolsky, S., Wu, C., Wang, S.-P., Hanson, E., Sun, Z., Walzer, R., and Wang, J. (2024a). Singlestore-v: An integrated vector database system in singlestore. Proceedings of the VLDB Endowment, 17(12).
Chen, J., Xiao, S., Zhang, P., Luo, K., Lian, D., and Liu, Z. (2024b). M3-embedding: Multi-linguality, multi-functionality, multi-granularity text embeddings through self-knowledge distillation. In Findings of the association for computational linguistics: ACL 2024, pages 2318–2335.
Church, K. W. (2017). Word2vec. Natural Language Engineering, 23(1):155–162.
Corsi, A. L. M. (2025). Estudo de caso de um sistema de gerenciamento de banco de dados vetorial para manipulação de dados de redes sociais online.
de Mendonça, A. L. C., Barioni, M. C. N., and Razente, H. (2024). Consultas analíticas por similaridade em sgbd relacionais. In Brazilian e-Science Workshop (BreSci), pages 48–55. SBC.
Douze, M., Guzhva, A., Deng, C., Johnson, J., Szilvasy, G., Mazaré, P.-E., Lomeli, M., Hosseini, L., and Jégou, H. (2025). The faiss library. IEEE Transactions on Big Data.
Freitag, M., Bandle, M., Schmidt, T., Kemper, A., and Neumann, T. (2020). Adopting worst-case optimal joins in relational database systems. Proceedings of the VLDB Endowment, 13(12):1891–1904.
Garg, N., Schiebinger, L., Jurafsky, D., and Zou, J. (2018). Word embeddings quantify 100 years of gender and ethnic stereotypes. Proceedings of the National Academy of Sciences, 115(16):E3635–E3644.
Gollapudi, S., Karia, N., Sivashankar, V., Krishnaswamy, R., Begwani, N., Raz, S., Lin, Y., Zhang, Y., Mahapatro, N., Srinivasan, P., et al. (2023). Filtered-diskann: Graph algorithms for approximate nearest neighbor search with filters. In Proceedings of the ACM Web Conference 2023, pages 3406–3416.
Greiner-Petter, A., Youssef, A., Ruas, T., Miller, B. R., Schubotz, M., Aizawa, A., and Gipp, B. (2020). Math-word embedding in math search and semantic extraction. Scientometrics, 125(3):3017–3046.
He, D., Nakandala, S., Banda, D., Sen, R., Saur, K., Park, K., Curino, C., Camacho-Rodríguez, J., Karanasos, K., and Interlandi, M. Query processing on tensor computation runtimes.
Khalifeh, F., Taheri, M., Mansoori, E., and Fakhrahmad, M. (2025). Enhancing keyword search in relational databases with word embeddings. IEEE Access.
Kumar, P. (2024). Large language models (llms): survey, technical frameworks, and future challenges. Artificial Intelligence Review, 57(10):260.
Liu, J., Zhang, Y., Jin, C., Gupta, A., Liu, S., and Wang, J. (2026). Fast vector search in postgresql: A decoupled approach. In Conference on Innovative Data Systems Research (CIDR).
Mama, R. and Machkour, M. (2025). Semantic and fuzzy integration: A new approach to efficient and flexible querying of relational databases. International Journal of Advanced Computer Science & Applications, 16(5).
Mathur, S. and Chhabra, A. (2024). Vector search algorithms: A brief survey. In 2024 4th International Conference on Ubiquitous Computing and Intelligent Information Systems (ICUIS), pages 365–371. IEEE.
Pan, J. J., Wang, J., and Li, G. (2024a). Survey of vector database management systems. The VLDB Journal, 33(5):1591–1615.
Pan, J. J., Wang, J., and Li, G. (2024b). Vector database management techniques and systems. In Companion of the 2024 International Conference on Management of Data, pages 597–604.
Petrola, L., Brayner, A., and Franco, W. (2025). Heuristic-guided text-to-sql translation with llms: Optimizing natural language interfaces for relational databases. In Simpósio Brasileiro de Banco de Dados (SBBD), pages 126–139. SBC.
Riccardo Cappuzzo, P. P. e. S. T. (2020). Criação de incorporações de conjuntos de dados relacionais heterogêneos para tarefas de integração de dados. Anais da Conferência Internacional ACM SIGMOD de 2020 sobre Gerenciamento de Dados.
Wei, C., Wu, B., Wang, S., Lou, R., Zhan, C., Li, F., and Cai, Y. (2020). Analyticdb-v: A hybrid analytical engine towards query fusion for structured and unstructured data. Proc. VLDB Endow., 13(12):3152–3165.
Yong, K. K. and Teh, P. C. (2025). Evaluating gpu-accelerated structured query engines using tpc-h: A comparative study with duckdb and rapids. In 2025 IEEE 11th International Conference on Smart Instrumentation, Measurement and Applications (ICSIMA), pages 144–149. IEEE.
