Avaliação Sistemática de Text-to-SQL em Português: Um Benchmark Unificado Multi-Domínio
Resumo
A tarefa Text-to-SQL tem avançado significativamente com o uso de Grandes Modelos de Linguagem (LLMs), mas sua avaliação em português ainda é limitada. Este trabalho propõe um benchmark unificado que consolida quatro bases públicas de diferentes domínios e avalia seis LLMs distintos a partir das métricas EX e CM sob um protocolo experimental padronizado. Os resultados mostram que, embora os modelos superem 70% de EX no Spider, o melhor desempenho no DATASUS alcança apenas 27,5%. A discrepância entre CM e EX também evidencia limitações de generalização composicional. Esses resultados sugerem que avaliações isoladas podem superestimar a capacidade dos modelos e destacam a necessidade de benchmarks integrados e multi-domínio.
Referências
de Carvalho, L. F. C., Júnior, P. S. d. S., and de Oliveira, H. T. A. (2025). Benchmarking large language models for text-to-sql in brazilian portuguese and english. In Simpósio Brasileiro de Tecnologia da Informação e da Linguagem Humana (STIL), pages 101–112.
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L. (2023). Qlora: Efficient finetuning of quantized llms. Advances in neural information processing systems, 36:10088–10115.
Dou, L., Gao, Y., Pan, M., Wang, D., Che, W., Zhan, D., and Lou, J.-G. (2023). Multispider: towards benchmarking multilingual text-to-sql semantic parsing. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 37, pages 12745–12753.
Fróes, K. and Braghetto, K. (2025). Exploring temporal text-to-sql challenges in brazilian portuguese: Lessons from educational data. In Proceedings of the Brazilian Symposium on Data Bases (SBBD), pages 963–969.
Gan, Y., Chen, X., Huang, Q., and Purver, M. (2022). Measuring and improving compositional generalization in text-to-SQL via component alignment. In Findings of the Association for Computational Linguistics: NAACL 2022, pages 831–843. Association for Computational Linguistics (ACL).
Gao, D., Wang, H., Li, Y., Sun, X., Qian, Y., Ding, B., and Zhou, J. (2024). Text-to-sql empowered by large language models: A benchmark evaluation. Proc. VLDB Endow., 17(5):1132–1145.
Grattafiori, A., Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Vaughan, A., et al. (2024). The llama 3 herd of models. arXiv preprint arXiv:2407.21783.
Hui, B., Yang, J., Cui, Z., Yang, J., Liu, D., Zhang, L., Liu, T., Zhang, J., Yu, B., Lu, K., et al. (2024). Qwen2. 5-coder technical report. arXiv preprint arXiv:2409.12186.
José, M. A. and Cozman, F. G. (2021). mrat-sql+ gap: a portuguese text-to-sql transformer. In Brazilian Conference on Intelligent Systems, pages 511–525. Springer.
Jose, M. A. and Cozman, F. G. (2023). A multilingual translator to sql with database schema pruning to improve self-attention. International Journal of Information Technology, 15(6):3015–3023.
Lei, F., Chen, J., Ye, Y., Cao, R., Shin, D., Su, H., Suo, Z., Gao, H., Hu, W., Yin, P., et al. (2025). Spider 2.0: Evaluating language models on real-world enterprise text-to-sql workflows. In International Conference on Learning Representations (ICLR), pages 28691–28735.
Li, J., Hui, B., Qu, G., Yang, J., Li, B., Li, B., Wang, B., Qin, B., Geng, R., Huo, N., et al. (2023). Can llm already serve as a database interface? a big bench for large-scale database grounded text-to-sqls. Advances in Neural Information Processing Systems, 36:42330–42357.
Moraes, M., Figueiredo, I., Marques, V., Santos, J., and Manssour, I. H. (2025). From questions to answers: A natural language interface for datasus hospitalization data. In Escola Regional de Aprendizado de Máquina e Inteligência Artificial da Região Sul (ERAMIA-RS), pages 296–299. SBC.
Pedroso, B. C., Pereira, M. R., and Pereira, D. A. (2025). Performance evaluation of llms in the text-to-sql task in portuguese. In Simpósio Brasileiro de Sistemas de Informação (SBSI), pages 260–269.
Petrola, L. and Franco, W. (2025). Heuristic-guided text-to-sql translation with llms: Optimizing natural language interfaces for relational databases. In Proceedings of the Brazilian Symposium on Data Bases (SBBD), pages 128–141.
Pham, K. T., Nguyen, T. H., Jo, J., Nguyen, Q. V. H., and Nguyen, T. T. (2025). Multilingual text-to-sql: Benchmarking the limits of language models with collaborative language agents. In Australasian Database Conference, pages 108–123. Springer.
Qwen Team (2025). Qwen2.5 technical report. arXiv preprint arXiv:2412.15115.
Song, Y., Wang, G., Li, S., and Lin, B. Y. (2025). The good, the bad, and the greedy: Evaluation of llms should not ignore non-determinism. In Proceedings of the Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 4195–4206.
Yamate, B. Y., Neubauer, T. R., Fantinato, M., and Peres, S. M. (2025). Text-to-sql oriented to the process mining domain: A pt-en dataset for query translation. arXiv preprint arXiv:2509.09684.
Yu, T., Zhang, R., Yang, K., Yasunaga, M., Wang, D., Li, Z., Ma, J., Li, I., Yao, Q., Roman, S., et al. (2018). Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-sql task. In Proceedings of the Conference on Empirical Methods in Natural Language Processing, pages 3911–3921.
