The Impact of Language and Portuguese Intralinguistic Variation on NL2SQL

  • Ricardo Trainotti Rabonato UNIFESP
  • Lilian Berton UNIFESP

Resumo


Natural Language to Structured Query Language (NL2SQL) systems have advanced with large language models (LLMs), yet evaluations remain centered on English, leaving open how these systems behave across languages and language varieties. We aim to contribute to this by analyzing two state-of-the-art LLMs (LLaMA 3.3 70B and Qwen 3.6 35B, both zero-shot) on the WikiSQL benchmark in English, Brazilian Portuguese (PT-BR), and European Portuguese (PT-PT), keeping database schemas and inference configuration fixed. Translations were validated through distance metrics (cosine similarity 0.966), contrastive linguistic analysis confirming varietal features (e.g., cleft constructions 36.8× more frequent in PT-PT), and multi-model judgment (85% / 81% semantic equivalence, 95.5% PT-BR and 96.0% PT-PT inter-judge agreement). Results show a marked drop from English (47.25% EX) to Portuguese (PT-BR: 40.03%; PT-PT: 39.34%), with the PT-BR vs. PT-PT difference reaching statistical significance (p = 0.019). English-only successes outnumber Portuguese-only successes nearly 10:1. A replication on a second model (Qwen 3.6 35B) confirms the English-Portuguese gap but suggests that intralinguistic effects may be model-dependent. Stratified analyses show semantic complexity and quantification amplify degradation, particularly for COUNT queries. These findings point to persistent English bias in NL2SQL and measurable effects of intralinguistic variation, even between closely related varieties.

Referências

Azevedo, M. M. (2005). Portuguese: A linguistic introduction. Cambridge University Press.

Chen, M. et al. (2021). Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374.

Fan, J., Gu, Z., Zhang, S., Zhang, Y., Chen, Z., Cao, L., Li, G., Madden, S., Du, X., and Tang, N. (2024). Combining small language models and large language models for zero-shot nl2sql. Proceedings of the VLDB Endowment, 17(11):2750–2763.

Katsogiannis-Meimarakis, G. and Koutrika, G. (2023). A survey on deep learning approaches for text-to-sql. The VLDB Journal, 32(4):905–936.

Liu, X., Shen, S., Li, B., Ma, P., Jiang, R., Zhang, Y., Fan, J., Li, G., Tang, N., and Luo, Y. (2024). A survey of nl2sql with large language models: Where are we, and where are we going? arXiv preprint arXiv:2408.05109.

Meta AI (2024). Llama 3.3 70b model card. [link]. Accessed: 2025-12-10.

Pires, T., Schlinger, E., and Garrette, D. (2019). How multilingual is multilingual bert? In Proceedings of ACL.

Preda, D., Osório, T., and Cardoso, H. L. (2024). Across the atlantic: Distinguishing between european and brazilian portuguese dialects. In Proceedings of the 16th International Conference on Computational Processing of Portuguese-Vol. 1, pages 353–363.

Qwen Team (2026). Qwen3.6-35B-A3B: Agentic coding power, now open to all.

Rademaker, A., Chalub, F., Real, L., Freitas, C., Bick, E., and de Paiva, V. (2017). Universal Dependencies for Portuguese. In Montemagni, S. and Nivre, J., editors, Proceedings of the Fourth International Conference on Dependency Linguistics (Depling 2017), pages 197–206, Pisa, Italy. Linköping University Electronic Press.

Raffel, C. et al. (2020). Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21:1–67.

Scholak, T., Schucher, N., and Bahdanau, D. (2021). Picard: Parsing incrementally for constrained auto-regressive decoding from language models.

Shi, P., Zhang, R., Bai, H., and Lin, J. (2022). Xricl: Cross-lingual retrieval-augmented in-context learning for cross-lingual text-to-sql semantic parsing. arXiv preprint arXiv:2210.13693.

Souza, F., Nogueira, R., and Lotufo, R. (2020). BERTimbau: pretrained BERT models for Brazilian Portuguese. In 9th Brazilian Conference on Intelligent Systems, BRACIS, Rio Grande do Sul, Brazil, October 20-23 (to appear).

Touvron, H. et al. (2023). Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.

Wagner Filho, J. A., Wilkens, R., Idiart, M., and Villavicencio, A. (2018). The brWaC corpus: A new open resource for Brazilian Portuguese. In Calzolari, N., Choukri, K., Cieri, C., Declerck, T., Goggi, S., Hasida, K., Isahara, H., Maegaard, B., Mariani, J., Mazo, H., Moreno, A., Odijk, J., Piperidis, S., and Tokunaga, T., editors, Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan. European Language Resources Association (ELRA).

Wang, B. et al. (2020). Rat-sql: Relation-aware schema encoding and linking for text-to-sql parsers. In Proceedings of ACL, pages 7567–7578.

Xu, X., Liu, C., and Song, D. (2017). Sqlnet: Generating structured queries from natural language without reinforcement learning.

Yu, T., Zhang, R., Yang, K., et al. (2018). Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-sql task. In Proceedings of EMNLP, pages 3911–3921.

Zhao, P., Shao, M., Cui, J., and Zhou, H. (2025). A survey of nl2sql research: From foundations to frontiers. In 2025 2nd International Conference on Intelligent Perception and Pattern Recognition (IPPR), pages 396–402. IEEE.

Zhong, V., Xiong, C., and Socher, R. (2017). Seq2sql: Generating structured queries from natural language using reinforcement learning.
Publicado
19/10/2026
RABONATO, Ricardo Trainotti; BERTON, Lilian. The Impact of Language and Portuguese Intralinguistic Variation on NL2SQL. In: SIMPÓSIO BRASILEIRO DE TECNOLOGIA DA INFORMAÇÃO E DA LINGUAGEM HUMANA (STIL), 17. , 2026, Cuiabá/MT. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 317-330. DOI: https://doi.org/10.5753/stil.2026.26611.