Text to SPARQL via LLMs fine-tuned with rewards and reasoning

  • Marcos Gôlo USP
  • Paulo Viviurka do Carmo V., Leipzig
  • Edgard Marx V., Leipzig
  • Ricardo Marcacini USP

Resumo


Large language models have made notable progress in translating natural-language questions into SPARQL queries for knowledge graphs. Recent studies from the first Text2SPARQL challenge have explored zero-shot and few-shot scenarios, as well as in-context learning techniques. With state-of-the-art (SoTA) results, the challenge exposes significant gaps, including limited use of open-source models and a scarcity of open-source fine-tuning, especially for reasoning. In this sense, we propose a new method for the Text2SPARQL task that combines pre-fine-tuning with explicit reasoning and a subsequent reward-based fine-tuning stage. First, we perform a fine-tuning to generate both a reasoning chain and the corresponding SPARQL query. Next, we apply reinforcement learning guided by reward functions. Our experiments show that the proposed model outperforms all baselines and six SoTA methods.

Palavras-chave: Text2SPARQL, Knowledge Graph Question Answering, Qwen3-Reasoning, Rewards Finetuning

Referências

Ali, W., Saleem, M., Yao, B., Hogan, A., and Ngomo, A.-C. N. A survey of rdf stores & sparql engines for querying knowledge graphs. The VLDB Journal, 2022.

Berezin, D., Avdeev, R., and Somov, O. Airi team in text2sparql challenge: Text-to-sparql executor for question-answering over knowledge graphs. In First International TEXT2SPARQL Challenge. CEUR, 2025.

Brei, F., B́’uhmann, L., Frey, J., Gerber, D., Meyer, L.-P., Stadler, C., and Bulert, K. Aruqula - an llm based text2sparql approach using react and knowledge graph exploration utilities. In First International TEXT2SPARQL Challenge. CEUR, 2025.

Chang, Y., Wang, X., Wang, J., Wu, Y., Yang, L., Zhu, K., Chen, H., Yi, X., Wang, C., Wang, Y., et al. A survey on evaluation of large language models. ACM Transactions on Intelligent Systems and Technology, 2023.

Dorsch, R., Henselmann, D., and Harth, A. Graf von data: A knowledge graph question answering agent for organisational usage. In First International TEXT2SPARQL Challenge. CEUR, 2025.

Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. arXiv preprint arXiv:2501.12948, 2025.

Ji, S., Pan, S., Cambria, E., Marttinen, P., and Yu, P. S. A survey on knowledge graphs: Representation, acquisition, and applications. IEEE transactions on neural networks and learning systems 33 (2): 494–514, 2021.

Lehmann, J., Isele, R., Jakob, M., Jentzsch, A., Kontokostas, D., Mendes, P. N., Hellmann, S., Morsey, M., Van Kleef, P., Auer, S., et al. Dbpedia–a large-scale, multilingual knowledge base extracted from wikipedia. Semantic web 6 (2): 167–195, 2015.

Liu, X., Zheng, Y., Du, Z., Ding, M., Qian, Y., Yang, Z., and Tang, J. Gpt understands, too. AI Open, 2023.

Marx, E., do Carmo, P., Gôlo, M., and Trump, S., editors. Preface of the First International TEXT2SPARQL Challenge (TEXT2SPARQL’25). CEUR, 2025.

Perevalov, A. and Andreas, B. Text-to-sparql goes beyond english: Multilingual question answering over knowledge graphs through human-inspired reasoning. In First International TEXT2SPARQL Challenge. CEUR, 2025.

Perevalov, A., Both, A., and Ngonga Ngomo, A.-C. Multilingual question answering systems for knowledge graphs– a survey. Semantic Web 15 (5): 2089–2124, 2024.

Shao, Z., Wang, P., Zhu, Q., Xu, R., Song, J., Bi, X., Zhang, H., Zhang, M., Li, Y., et al. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300, 2024.

Soru, T., Joshi, S., Tiwari, S., Shahinmoghadam, M., and Panchbhai, A. Question answering over dbpedia with fine-tuned autoregressive models. In First International TEXT2SPARQL Challenge. CEUR, 2025.

Wardenga, J. G. and Kafer, T. Leveraging data shapes in large language model contexts for question answering on public and private knowledge graphs. In First International TEXT2SPARQL Challenge. CEUR, 2025.

Yang, A., Li, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Gao, C., Huang, C., Lv, C., et al. Qwen3 technical report. arXiv preprint arXiv:2505.09388, 2025.
Publicado
19/10/2026
GÔLO, Marcos; CARMO, Paulo Viviurka do; MARX, Edgard; MARCACINI, Ricardo. Text to SPARQL via LLMs fine-tuned with rewards and reasoning. In: SYMPOSIUM ON KNOWLEDGE DISCOVERY, MINING AND LEARNING (KDMILE), 14. , 2026, Cuiabá/MT. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 233-240. ISSN 2763-8944. DOI: https://doi.org/10.5753/kdmile.2026.30058.