A Hybrid RAG–Text2SQL Architecture for Auditable Question Answering in Brazilian Portuguese Legal-Administrative Collections
Resumo
Question answering over Brazilian Portuguese legal-administrative collections must support semantic information needs and structured intents such as counting, filtering, aggregation, and ranking. Retrieval-augmented generation (RAG) grounds answers in text, but it does not reliably execute relational operations. Standalone Text2SQL can execute auditable queries, but lacks textual grounding and explanatory coverage. This paper presents a hybrid architecture that runs semantic retrieval and a constrained Text2SQL pipeline in parallel and synchronizes both outputs into a single answer. The structured branch includes column routing, schema restriction, normalized derived columns, WHERE-hints, read-only execution policies, and decomposition for multi-intent questions. We formalize the evaluation protocol for structured and semantic queries, define the proposed metrics, and clarify the role of the LLM-based judge used for semantic assessment. On a curated TCE-GO corpus, the hybrid architecture outperforms a RAG-only baseline on structured queries while remaining competitive on semantic ones. The contribution lies in showing how semantic retrieval and constrained SQL execution can be combined for Portuguese legal-administrative QA with explicit safety, auditability, transparency, and evaluation policies.
Palavras-chave:
Retrieval-Augmented Generation, Text-to-SQL, Hybrid Question Answering, Legal Natural Language Processing, Auditable Artificial Intelligence, Brazilian Portuguese, Legal-Administrative Documents, Semantic Retrieval
Referências
Chalkidis, I., Fergadiotis, M., Androutsopoulos, I., and Aletras, N. (2020). LEGAL-BERT: The muppets straight out of law school. arXiv preprint arXiv:2010.02559.
Chalkidis, I., Jana, A., Hartung, D., Bommarito, M., Androutsopoulos, I., Katz, D. M., and Aletras, N. (2022). Lexglue: A benchmark dataset for legal language understanding in english. In Proceedings of ACL.
Gao, Y., Xiong, Y., Gao, X., Jia, K., Pan, J., Bi, Y., Dai, Y., Sun, J., Wang, M., and Wang, H. (2023). Retrieval-augmented generation for large language models: A survey. arXiv preprint arXiv:2312.10997.
Gururangan, S., Marasović, A., Swayamdipta, S., Lo, K., Beltagy, I., Downey, D., and Smith, N. A. (2020). Don’t stop pretraining: Adapt language models to domains and tasks. In Proceedings of ACL.
Izacard, G. and Grave, E. (2021). Leveraging passage retrieval with generative models for open domain question answering. In Proceedings of EACL.
Katz, D. M., Hartung, D., Gerlach, L., Jana, A., and II, M. J. B. (2023). Natural language processing in the legal domain. arXiv preprint arXiv:2302.12039.
Khattab, O. and Zaharia, M. (2020). ColBERT: Efficient and effective passage search via contextualized late interaction over BERT. In Proceedings of SIGIR.
Khot, T., Trivedi, H., Finlayson, M., Fu, Y., Richardson, K., Clark, P., and Sabharwal, A. (2023). Decomposed prompting: A modular approach for solving complex tasks. In International Conference on Learning Representations (ICLR).
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W., Rocktäschel, T., Riedel, S., and Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. In Advances in Neural Information Processing Systems.
Li, J., Hui, B., Qu, G., Yang, J., Li, B., Li, B., Wang, B., Qin, B., Geng, R., Huo, N., et al. (2023). Can llm already serve as a database interface? a big bench for large-scale database grounded text-to-sqls. In Advances in Neural Information Processing Systems (NeurIPS), volume 36.
Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., and Liang, P. (2024). Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics.
Pourreza, M. and Rafiei, D. (2023). Din-sql: Decomposed in-context learning of text-to-sql with self-correction. In Advances in Neural Information Processing Systems (NeurIPS), volume 36.
Reimers, N. and Gurevych, I. (2019). Sentence-bert: Sentence embeddings using siamese bert-networks. In Proceedings of EMNLP-IJCNLP.
Scholak, T., Schucher, N., and Bahdanau, D. (2021). PICARD: Parsing incrementally for constrained auto-regressive decoding from language models. In Proceedings of EMNLP.
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y. (2023). ReAct: Synergizing reasoning and acting in language models. In International Conference on Learning Representations (ICLR).
Zhang, K., Lin, X., Wang, Y., Zhang, X., Sun, F., Cen, J., Tan, H., Jiang, X., and Shen, H. (2023). Redsql: A retrieval-augmented framework for text-to-sql generation. In Findings of the Association for Computational Linguistics: EMNLP 2023.
Chalkidis, I., Jana, A., Hartung, D., Bommarito, M., Androutsopoulos, I., Katz, D. M., and Aletras, N. (2022). Lexglue: A benchmark dataset for legal language understanding in english. In Proceedings of ACL.
Gao, Y., Xiong, Y., Gao, X., Jia, K., Pan, J., Bi, Y., Dai, Y., Sun, J., Wang, M., and Wang, H. (2023). Retrieval-augmented generation for large language models: A survey. arXiv preprint arXiv:2312.10997.
Gururangan, S., Marasović, A., Swayamdipta, S., Lo, K., Beltagy, I., Downey, D., and Smith, N. A. (2020). Don’t stop pretraining: Adapt language models to domains and tasks. In Proceedings of ACL.
Izacard, G. and Grave, E. (2021). Leveraging passage retrieval with generative models for open domain question answering. In Proceedings of EACL.
Katz, D. M., Hartung, D., Gerlach, L., Jana, A., and II, M. J. B. (2023). Natural language processing in the legal domain. arXiv preprint arXiv:2302.12039.
Khattab, O. and Zaharia, M. (2020). ColBERT: Efficient and effective passage search via contextualized late interaction over BERT. In Proceedings of SIGIR.
Khot, T., Trivedi, H., Finlayson, M., Fu, Y., Richardson, K., Clark, P., and Sabharwal, A. (2023). Decomposed prompting: A modular approach for solving complex tasks. In International Conference on Learning Representations (ICLR).
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W., Rocktäschel, T., Riedel, S., and Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. In Advances in Neural Information Processing Systems.
Li, J., Hui, B., Qu, G., Yang, J., Li, B., Li, B., Wang, B., Qin, B., Geng, R., Huo, N., et al. (2023). Can llm already serve as a database interface? a big bench for large-scale database grounded text-to-sqls. In Advances in Neural Information Processing Systems (NeurIPS), volume 36.
Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., and Liang, P. (2024). Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics.
Pourreza, M. and Rafiei, D. (2023). Din-sql: Decomposed in-context learning of text-to-sql with self-correction. In Advances in Neural Information Processing Systems (NeurIPS), volume 36.
Reimers, N. and Gurevych, I. (2019). Sentence-bert: Sentence embeddings using siamese bert-networks. In Proceedings of EMNLP-IJCNLP.
Scholak, T., Schucher, N., and Bahdanau, D. (2021). PICARD: Parsing incrementally for constrained auto-regressive decoding from language models. In Proceedings of EMNLP.
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y. (2023). ReAct: Synergizing reasoning and acting in language models. In International Conference on Learning Representations (ICLR).
Zhang, K., Lin, X., Wang, Y., Zhang, X., Sun, F., Cen, J., Tan, H., Jiang, X., and Shen, H. (2023). Redsql: A retrieval-augmented framework for text-to-sql generation. In Findings of the Association for Computational Linguistics: EMNLP 2023.
Publicado
08/09/2026
Como Citar
SILVA, Josiel Pantaleão Cardoso; DE OLIVEIRA, Sávio Salvarino Teles; BRAKES, Matheus Fares Costa; PRESA, João Paulo Cavalcante.
A Hybrid RAG–Text2SQL Architecture for Auditable Question Answering in Brazilian Portuguese Legal-Administrative Collections. In: SIMPÓSIO BRASILEIRO DE BANCO DE DADOS (SBBD), 41. , 2026, São Carlos/SP.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2026
.
p. 224-237.
ISSN 2763-8979.
DOI: https://doi.org/10.5753/sbbd.2026.249192.
