Benchmarking Agentic Capabilities of LLMs in Data Science: Prompt Analysis and Performance of Open-Weight Models
Resumo
This work analyzes three LLM benchmarks for data science: DS-1000, InfiAgent-DABench, and DataSciBench. The focus is on the evolution of evaluated tasks, from code generation in constrained settings to agent-oriented scenarios involving planning, tool use, code execution, and error correction. The analysis covers prompt complexity, the data science areas represented in each benchmark, and the performance of open-weight models. The methodology combines prompt difficulty analysis, topic modeling with BERTopic, and local evaluation of models in the 24B to 30B parameter range. The results show that newer benchmarks include more complex prompts, broader coverage of the data science workflow, and stronger dependence on capabilities associated with LLM-based agents. In the performance evaluation, Qwen3.6-27B and GLM-4.7-30B achieved competitive results in tasks that require tool use, code execution, and iterative correction.
Referências
Campello, R. J. G. B., Moulavi, D., and Sander, J. (2013). Density-based clustering based on hierarchical density estimates. In Advances in Knowledge Discovery and Data Mining, pages 160–172.
Chang, Y., Wang, X., Wang, J., et al. (2024). A survey on evaluation of large language models. ACM Transactions on Intelligent Systems and Technology, 15(3):1–45.
Chen, M., Tworek, J., Jun, H., et al. (2021). Evaluating large language models trained on code. arXiv preprint arXiv:2107.03374.
Chiang, W.-L., Zheng, L., Sheng, Y., et al. (2024). Chatbot arena: An open platform for evaluating llms by human preference. In Proceedings of the 41st International Conference on Machine Learning, ICML’24.
Grootendorst, M. (2022). Bertopic: Neural topic modeling with a class-based TF-IDF procedure. arXiv preprint arXiv:2203.05794.
Hong, S., Lin, Y., Liu, B., et al. (2025). Data interpreter: An LLM agent for data science. In Findings of the Association for Computational Linguistics: ACL 2025, pages 19796–19821.
Hu, X., Zhao, Z., Wei, S., et al. (2024). Infiagent-dabench: Evaluating agents on data analysis tasks. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of ICML’24, pages 19544–19572.
Jolliffe, I. T. (2002). Principal Component Analysis. Springer, 2 edition.
Lai, Y., Li, C., Wang, Y., et al. (2023). Ds-1000: A natural and reliable benchmark for data science code generation. In Proceedings of the 40th International Conference on Machine Learning, ICML’23.
Nikolakopoulos, A., Litke, A., Psychas, A., et al. (2025). Exploring the potential of offline LLMs in data science: A study on code generation for data analysis. IEEE Access, 13:64087–64114.
Raji, I. D., Denton, E., Bender, E. M., Hanna, A., and Paullada, A. (2021). AI and the everything in the whole wide world benchmark. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2), NIPS ’21.
Salton, G., Wong, A., and Yang, C.-S. (1975). A vector space model for automatic indexing. Communications of the ACM, 18(11):613–620.
Schick, T., Dwivedi-Yu, J., Dessı̀, R., et al. (2023). Toolformer: Language models can teach themselves to use tools. In Proceedings of the 37th International Conference on Neural Information Processing Systems, NIPS ’23.
Spärck Jones, K. (1972). A statistical interpretation of term specificity and its application in retrieval. Journal of Documentation, 28(1):11–21.
Vaswani, A., Shazeer, N., Parmar, N., et al. (2017). Attention is all you need. In Advances in Neural Information Processing Systems, volume 30.
Wang, L., Ma, C., Feng, X., et al. (2024). A survey on large language model based autonomous agents. Frontiers of Computer Science, 18(6):186345.
Wooldridge, M. (2009). An Introduction to MultiAgent Systems. John Wiley & Sons, 2 edition.
Yao, S., Zhao, J., Yu, D., et al. (2023). React: Synergizing reasoning and acting in language models. In The Eleventh International Conference on Learning Representations.
Zhang, D., Zhoubian, S., Cai, M., et al. (2026). DataSciBench: An LLM agent benchmark for data science. In Findings of the Association for Computational Linguistics: ACL 2026, pages 3685–3728.
Zhang, Y., Lin, Y., Khan, A., and Wan, H. (2025). Large language model prompt datasets: An in-depth analysis and insights. arXiv preprint arXiv:2510.09316.
