AtlasSQL-BR: A Brazilian Portuguese Geospatial Text-to-SQL Dataset with Spatial Hierarchies
Resumo
The Text-to-SQL task, which translates natural language into SQL queries, democratizes database access by enabling non-experts to query and analyze data effectively. While the field has seen significant progress, state-of-the-art models still fall short in processing geospatial queries due to a lack of diverse spatial data in existing training sets. This issue is compounded by the overwhelming dominance of English-language resources, leaving languages like Brazilian Portuguese largely underexplored. To bridge these gaps, we introduce AtlasSQL-BR, a novel, real-world geospatial Text-to-SQL dataset in Brazilian Portuguese. Built upon school census data and geographic boundary hierarchies, this dataset provides a robust benchmark for geospatial queries.
Referências
Chen, T., Xu, B., Zhang, C., and Guestrin, C. (2016). Training deep nets with sublinear memory cost. [link].
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., et al. (2024). Scaling instruction-finetuned language models. Journal of Machine Learning Research, 25(70):1–53.
Coelho, G. M. C., Nascimento, E. R. S., Izquierdo, Y. T., García, G. M., Feijó, L., Lemos, M., Garcia, R. L. S., de Oliveira, A. R., Pinheiro, J. P., and Casanova, M. A. (2024). Improving the accuracy of Text-to-SQL tools based on large language models for real-world relational databases. In Database and Expert Systems Applications, pages 93–107. Springer Nature Switzerland.
Deng, N., Chen, Y., and Zhang, Y. (2022). Recent advances in text-to-SQL: A survey of what we have and what we expect. In Proceedings of the 29th International Conference on Computational Linguistics (COLING), Gyeongju, Republic of Korea. International Committee on Computational Linguistics.
Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L. (2022). GPT3.int8(): 8-bit matrix multiplication for transformers at scale. [link].
Gao, D., Wang, H., Li, Y., Sun, X., Qian, Y., Ding, B., and Zhou, J. (2024). Text-to-SQL empowered by large language models: A benchmark evaluation. Proceedings of the VLDB Endowment, 17(5):1132–1145.
Grattafiori, A., Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Fan, A., et al. (2024). The Llama 3 herd of models. [link].
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W. (2022). LoRA: Low-rank adaptation of large language models. [link].
Jiang, Y. and Yang, C. (2024). Is ChatGPT a good geospatial data analyst? Exploring the integration of natural language into structured query language within a spatial database. ISPRS International Journal of Geo-Information, 13(1):26.
Kalamkar, D., Mudigere, D., Mellempudi, N., Das, D., Banerjee, K., Wang, S., Gu, B., Wang, Y., Chen, H., Kaul, B., and Fang, A. (2019). A study of BFLOAT16 for deep learning training.
Katsogiannis-Meimarakis, G. and Koutrika, G. (2023). A survey on deep learning approaches for text-to-SQL. The VLDB Journal, 32:905–936.
Kirkpatrick, J., Pascanu, R., Rabinowitz, N., Veness, J., Desjardins, G., Rusu, A. A., Milan, K., Quan, J., Ramalho, T., Grabska-Barwinska, A., Hassabis, D., Clopath, C., Kumaran, D., and Hadsell, R. (2017). Overcoming catastrophic forgetting in neural networks. Proceedings of the National Academy of Sciences, 114(13):3521–3526.
Li, J., Hui, B., Qu, G., Yang, J., Li, B., Li, B., Wang, B., Qin, B., Geng, R., Huo, N., Zhou, X., Ma, C., Li, G., Chang, K. C., Huang, F., Cheng, R., and Li, Y. (2023). Can LLM already serve as a database interface? a big bench for large-scale database grounded text-to-SQLs. In Proceedings of the 37th International Conference on Neural Information Processing Systems, NeurIPS ’23. Curran Associates Inc.
Liu, M., Wang, X., Xu, J., Lu, H., and Tong, Y. (2025). NALSpatial: A natural language interface for spatial databases. IEEE Transactions on Knowledge and Data Engineering, 37(4):2056–2070.
Loshchilov, I. and Hutter, F. (2019). Decoupled weight decay regularization. In International Conference on Learning Representations (ICLR).
Nan, L., Zhao, Y., Zou, W., Ri, N., Tae, J., Zhang, E., Cohan, A., and Radev, D. (2023). Enhancing text-to-SQL capabilities of large language models: A study on prompt design strategies. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 14935–14956, Singapore. ACL.
Nascimento, E. R., García, G., Izquierdo, Y. T., Feijó, L., Coelho, G. M. C., de Oliveira, A. R., Lemos, M., Garcia, R. L. S., Leme, L. A. P. P., and Casanova, M. A. (2025). LLM-Based Text-to-SQL for Real-World databases. SN Computer Science, 6(2):130.
Radford, A. and Narasimhan, K. (2018). Improving language understanding by generative pre-training. [link].
Raiaan, M. A. K., Mukta, M. S. H., Fatema, K., Fahad, N. M., Sakib, S., Mim, M. M. J., Ahmad, J., Ali, M. E., and Azam, S. (2024). A review on large language models: Architectures, applications, taxonomies, open issues and challenges. IEEE Access, 12:26839–26874.
Stolze, K. (2003). SQL/MM spatial: The standard to manage spatial data in a relational database system. In BTW 2003–Datenbanksysteme für Business, Technologie und Web, Tagungsband der 10. BTW Konferenz, pages 247–264. Gesellschaft für Informatik eV.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017). Attention is all you need. Advances in neural information processing systems, 30.
Wu, J., Gan, W., Chao, H.-C., and Yu, P. S. (2024). Geospatial big data: Survey and challenges. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 17:17007–17020.
Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D., Huang, F., Dong, G., et al. (2024). Qwen2.5 technical report. [link].
Yu, T., Zhang, R., Yang, K., Yasunaga, M., Wang, D., Li, Z., Ma, J., Li, I., Yao, Q., Roman, S., Zhang, Z., and Radev, D. (2018). Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-SQL task. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 3911–3921. ACL.
Zelle, J. M. and Mooney, R. J. (1996). Learning to parse database queries using inductive logic programming. In Proceedings of the 13th National Conference on Artificial Intelligence (AAAI).
Zhang, B., Ye, Y., Du, G., Hu, X., Li, Z., Yang, S., Liu, C. H., Zhao, R., Li, Z., and Mao, H. (2024a). Benchmarking the Text-to-SQL capability of large language models: A comprehensive evaluation. [link].
Zhang, X., Xiao, F., Yan, L., and Zhang, Z. (2024b). GS-SQL: Modeling spatial semantics in spatial text-to-sql. In 2024 International Joint Conference on Neural Networks (IJCNN), pages 1–7.
Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., Min, Y., Zhang, B., Zhang, J., Dong, Z., Du, Y., Yang, C., Chen, Y., Chen, Z., Jiang, J., Ren, R., Li, Y., Tang, X., Liu, Z., Liu, P., Nie, J.-Y., and Wen, J.-R. (2023). A survey of large language models. [link].
Zhong, R., Yu, T., and Klein, D. (2020). Semantic evaluation for text-to-SQL with distilled test suites. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4883–4893. ACL.
Zhong, V., Xiong, C., and Socher, R. (2017). Seq2SQL: Generating structured queries from natural language using reinforcement learning. [link].
