DATAEXPLAINER: Comprehensible Agentic Data Science with Consumer-Grade Hardware
Resumo
Data Science (DS), a complex field often involving advanced knowledge of statistics, mathematics, and programming, has historically been restricted to deeply knowledgeable human professionals. Recent advancements in artificial reasoning and the expanding toolset of large language models (LLMs) have challenged this paradigm. However, most LLM-based approaches for automated DS are still computationally demanding, time-consuming, and not especially easy to follow. In this work, we expand on previous advancements to present DATAEXPLAINER, a framework for automated DS that provides step-by-step solutions and surpasses the average Kaggle user while using a single consumer-grade GPU and less than 2 hours of processing.
Referências
Chandel, S., Clement, C. B., Serrato, G., and Sundaresan, N. (2022). Training and evaluating a jupyter notebook data science assistant. CoRR, abs/2201.12901.
Chi, Y., Lin, Y., Hong, S., Pan, D., Fei, Y., Mei, G., Liu, B., Pang, T., Kwok, J., Zhang, C., Liu, B., and Wu, C. (2024). Sela: Tree-search enhanced llm agents for automated machine learning. arXiv preprint arXiv:2410.17238.
DeepMind (2025). Gemini — deepmind.google. [link]. [Accessed 19-02-2025].
Gemma Team (2025a). Gemma 3 technical report. arXiv preprint arXiv:2503.19786.
Gemma Team (2025b). Gemma formatting and system instructions | Google AI for Developers — ai.google.dev. [link]. [Accessed 13-02-2026].
Grosnit, A., Maraval, A., Doran, J., Paolo, G., Thomas, A., Beevi, R. S. H. N., Gonzalez, J., Khandelwal, K., Iacobacci, I., Benechehab, A., Cherkaoui, H., El-Hili, Y. A., Shao, K., Hao, J., Yao, J., Kegl, B., Bou-Ammar, H., and Wang, J. (2024). Large language models orchestrating structured reasoning achieve kaggle grandmaster level. arXiv preprint arXiv:2411.03562.
Guo, S., Deng, C., Wen, Y., Chen, H., Chang, Y., and Wang, J. (2024). DS-agent: Automated data science by empowering large language models with case-based reasoning. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, pages 16813–16848. PMLR.
Hong, S., Lin, Y., Liu, B., Liu, B., Wu, B., Zhang, C., Li, D., Chen, J., Zhang, J., Wang, J., Zhang, L., Zhang, L., Yang, M., Zhuge, M., Guo, T., Zhou, T., Tao, W., Tang, R., Lu, X., Zheng, X., Liang, X., Fei, Y., Cheng, Y., Ni, Y., Gou, Z., Xu, Z., Luo, Y., and Wu, C. (2025). Data interpreter: An LLM agent for data science. In Che, W., Nabende, J., Shutova, E., and Pilehvar, M. T., editors, Findings of the Association for Computational Linguistics: ACL 2025, pages 19796–19821, Vienna, Austria. Association for Computational Linguistics.
Hu, X., Zhao, Z., Wei, S., Chai, Z., Ma, Q., Wang, G., Wang, X., Su, J., Xu, J., Zhu, M., Cheng, Y., Yuan, J., Li, J., Kuang, K., Yang, Y., Yang, H., and Wu, F. (2024). Infiagentdabench: Evaluating agents on data analysis tasks. In International Conference on Machine Learning (ICML). JMLR.org.
Huang, Q., Vora, J., Liang, P., and Leskovec, J. (2024a). MLAgentBench: Evaluating language agents on machine learning experimentation. In Salakhutdinov, R., Kolter, Z., Heller, K., Weller, A., Oliver, N., Scarlett, J., and Berkenkamp, F., editors, Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, pages 20271–20309. PMLR.
Huang, X., Liu, W., Chen, X., Wang, X., Wang, H., Lian, D., Wang, Y., Tang, R., and Chen, E. (2024b). Understanding the planning of llm agents: A survey.
Hui, B., Yang, J., Cui, Z., Yang, J., Liu, D., Zhang, L., Liu, T., Zhang, J., Yu, B., Dang, K., et al. (2024). Qwen2. 5-coder technical report. arXiv preprint arXiv:2409.12186.
Hutter, F., Kotthoff, L., and Vanschoren, J. (2019). Automated Machine Learning: Methods, Systems, Challenges. Springer Cham, 1st edition.
Jacovi, A., Caciularu, A., Goldman, O., and Goldberg, Y. (2023). Stop uploading test data in plain text: Practical strategies for mitigating data contamination by evaluation benchmarks. In The 2023 Conference on Empirical Methods in Natural Language Processing.
Jiang, Z., Schmidt, D., Srikanth, D., Xu, D., Kaplan, I., Jacenko, D., and Wu, Y. (2025). Aide: Ai-driven exploration in the space of code. arXiv preprint arXiv:2502.13138.
Lai, Y., Li, C., Wang, Y., Zhang, T., Zhong, R., Zettlemoyer, L., Yih, W.-t., Fried, D., Wang, S., and Yu, T. (2023). Ds-1000: a natural and reliable benchmark for data science code generation. In Proceedings of the 40th International Conference on Machine Learning, ICML’23. JMLR.org.
Li, Z., Zang, Q., Ma, D., Guo, J., Zheng, T., minghao liu, Niu, X., Yue, X., Wang, Y., Yang, J., Liu, J., Zhong, W., Zhou, W., Huang, W., and Zhang, G. (2025). Autokaggle: A multi-agent framework for autonomous data science competitions.
OpenAI Team (2025). gpt-oss-120b & gpt-oss-20b model card. arXiv preprint arXiv:2508.10925.
Provost, F. J. and Fawcett, T. (2013). Data science and its relationship to big data and data-driven decision making. Big data, 1 1:51–9.
Qwen3 Team (2025). Qwen3 technical report. arXiv preprint arXiv:2505.09388.
Russell, S. and Norvig, P. (2022). Artificial Intelligence: A Modern Approach. Pearson, USA, 4rd edition. Global Edition.
Taylor, P. (2024). Data growth worldwide 2010-2028 | statista — statista.com. [link]. [Accessed 24-02-2025].
Wang, C., Lee, B., Drucker, S., Marshall, D., and Gao, J. (2025). Data formulator 2: Iterative creation of data visualizations, with ai transforming data along the way. arXiv preprint arXiv:2408.16119.
Zhang, L., Zhang, Y., Ren, K., Li, D., and Yang, Y. (2024). MLCopilot: Unleashing the power of large language models in solving machine learning tasks. In Graham, Y. and Purver, M., editors, Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2931–2959, St. Julian’s, Malta. Association for Computational Linguistics.
Zhu, Q., Guo, D., Shao, Z., Yang, D., Wang, P., Xu, R., Wu, Y., Li, Y., Gao, H., Ma, S., et al. (2024). Deepseek-coder-v2: Breaking the barrier of closed-source models in code intelligence. arXiv preprint arXiv:2406.11931.
