We Keep Evaluating LLMs, But Are We Learning Anything?
Resumo
Large language model (LLM) research has become increasingly dominated by benchmark-oriented evaluations that measure performance on predefined tasks and datasets. While such evaluations are useful for tracking progress, they often provide limited scientific insight into how these systems reason, generalize, fail, or support real-world applications. In this article, we argue that LLM research should move beyond incremental leaderboard improvements and adopt a broader scientific agenda inspired by the historical evolution of machine learning research. Drawing analogies with established research areas in machine learning, we propose a set of guidelines for conducting more rigorous and impactful LLM research. These guidelines include investigating internal reasoning mechanisms, developing calibration and uncertainty estimation methods, addressing bias and fairness systematically, studying generalization to novel domains, exploring benchmark-free evaluation paradigms, improving methodological transparency, and, more ambitiously, using LLMs as instruments for scientific discovery rather than merely as objects of evaluation.
Palavras-chave:
Large language models, Benchmark evaluation, Bias and fairness, Generalization, Scientific discovery
Referências
Barocas, S. and Selbst, A. D. (2016). Big data’s disparate impact. California Law Review, 104(3):671–732.
Bengio, Y., Goodfellow, I., Courville, A., et al. (2017). Deep learning, volume 1. MIT press Cambridge, MA, USA.
Chang, Y., Wang, X., Wang, J., Wu, Y., Yang, L., Zhu, K., Chen, H., Yi, X., Wang, C., Wang, Y., et al. (2024). A survey on evaluation of large language models. ACM transactions on intelligent systems and technology, 15(3):1–45.
Fehring, L., Frings, J., Rust, P., Kempny, C., Thürmann, P. A., and Meister, S. (2025). Extension of the consolidated criteria for reporting qualitative research guideline to large language models (coreq+ llm): Protocol for a multiphase study. JMIR Research Protocols, 14(1):e78682.
Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q. (2017). On calibration of modern neural networks. In International conference on machine learning, pages 1321–1330. PMLR.
Li, Y. et al. (2025). Multi agent large language models for biomedical hypothesis generation in drug combination discovery. iScience, 28(12):113984.
Lipton, Z. C. and Steinhardt, J. (2019). Troubling trends in machine learning scholarship: Some ml papers suffer from flaws that could mislead the public and stymie future research. Queue, 17(1):45–77.
Madry, A., Makelov, A., Schmidt, L., Tsipras, D., and Vladu, A. (2017). Towards deep learning models resistant to adversarial attacks. stat, 1050:9.
McIntosh, T. R., Susnjak, T., Arachchilage, N., Liu, T., Xu, D., Watters, P., and Halgamuge, M. N. (2025). Inadequacies of large language model benchmarks in the era of generative artificial intelligence. IEEE Transactions on Artificial Intelligence.
Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., and Galstyan, A. (2021). A survey on bias and fairness in machine learning. ACM Computing Surveys, 54(6):1–35.
Meng, K., Bau, D., Andonian, A., and Belinkov, Y. (2022). Locating and editing factual associations in gpt. Advances in neural information processing systems, 35:17359–17372.
Ovadia, Y., Fertig, E., Ren, J., et al. (2019). Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift. Advances in Neural Information Processing Systems (NeurIPS), 32.
Shi, W., Ajith, A., Xia, M., Huang, Y., Liu, D., Blevins, T., Chen, D., and Zettlemoyer, L. (2024). Detecting pretraining data from large language models. In International Conference on Learning Representations, volume 2024, pages 51826–51843.
Springer, J. M., Goyal, S., Wen, K., Kumar, T., Yue, X., Malladi, S., Neubig, G., and Raghunathan, A. (2025). Overtrained language models are harder to fine-tune. arXiv preprint arXiv:2503.19206.
Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., et al. (2023). Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. Transactions on machine learning research.
Tan, H., Zhan, S., Jia, F., Zheng, H.-T., and Chan, W. K. V. (2025). A hierarchical framework for measuring scientific paper innovation via large language models. Information Sciences, page 122787.
Wang, H., Fu, T., Du, Y., Gao, W., Huang, K., Liu, Z., Chandak, P., Liu, S., Van Katwyk, P., Deac, A., et al. (2023). Scientific discovery in the age of artificial intelligence. Nature, 620(7972):47–60.
Weiss, K., Khoshgoftaar, T. M., and Wang, D. (2016). A survey of transfer learning. Journal of Big data, 3(1):9.
Wiggins, W. F. and Tejani, A. S. (2022). On the opportunities and risks of foundation models for natural language processing in radiology. Radiology: Artificial Intelligence, 4(4):e220119.
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T., Cao, Y., and Narasimhan, K. (2023). Tree of thoughts: Deliberate problem solving with large language models. Advances in neural information processing systems, 36:11809–11822.
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y. (2022). React: Synergizing reasoning and acting in language models. arXiv preprint arXiv:2210.03629.
Bengio, Y., Goodfellow, I., Courville, A., et al. (2017). Deep learning, volume 1. MIT press Cambridge, MA, USA.
Chang, Y., Wang, X., Wang, J., Wu, Y., Yang, L., Zhu, K., Chen, H., Yi, X., Wang, C., Wang, Y., et al. (2024). A survey on evaluation of large language models. ACM transactions on intelligent systems and technology, 15(3):1–45.
Fehring, L., Frings, J., Rust, P., Kempny, C., Thürmann, P. A., and Meister, S. (2025). Extension of the consolidated criteria for reporting qualitative research guideline to large language models (coreq+ llm): Protocol for a multiphase study. JMIR Research Protocols, 14(1):e78682.
Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q. (2017). On calibration of modern neural networks. In International conference on machine learning, pages 1321–1330. PMLR.
Li, Y. et al. (2025). Multi agent large language models for biomedical hypothesis generation in drug combination discovery. iScience, 28(12):113984.
Lipton, Z. C. and Steinhardt, J. (2019). Troubling trends in machine learning scholarship: Some ml papers suffer from flaws that could mislead the public and stymie future research. Queue, 17(1):45–77.
Madry, A., Makelov, A., Schmidt, L., Tsipras, D., and Vladu, A. (2017). Towards deep learning models resistant to adversarial attacks. stat, 1050:9.
McIntosh, T. R., Susnjak, T., Arachchilage, N., Liu, T., Xu, D., Watters, P., and Halgamuge, M. N. (2025). Inadequacies of large language model benchmarks in the era of generative artificial intelligence. IEEE Transactions on Artificial Intelligence.
Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., and Galstyan, A. (2021). A survey on bias and fairness in machine learning. ACM Computing Surveys, 54(6):1–35.
Meng, K., Bau, D., Andonian, A., and Belinkov, Y. (2022). Locating and editing factual associations in gpt. Advances in neural information processing systems, 35:17359–17372.
Ovadia, Y., Fertig, E., Ren, J., et al. (2019). Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift. Advances in Neural Information Processing Systems (NeurIPS), 32.
Shi, W., Ajith, A., Xia, M., Huang, Y., Liu, D., Blevins, T., Chen, D., and Zettlemoyer, L. (2024). Detecting pretraining data from large language models. In International Conference on Learning Representations, volume 2024, pages 51826–51843.
Springer, J. M., Goyal, S., Wen, K., Kumar, T., Yue, X., Malladi, S., Neubig, G., and Raghunathan, A. (2025). Overtrained language models are harder to fine-tune. arXiv preprint arXiv:2503.19206.
Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., et al. (2023). Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. Transactions on machine learning research.
Tan, H., Zhan, S., Jia, F., Zheng, H.-T., and Chan, W. K. V. (2025). A hierarchical framework for measuring scientific paper innovation via large language models. Information Sciences, page 122787.
Wang, H., Fu, T., Du, Y., Gao, W., Huang, K., Liu, Z., Chandak, P., Liu, S., Van Katwyk, P., Deac, A., et al. (2023). Scientific discovery in the age of artificial intelligence. Nature, 620(7972):47–60.
Weiss, K., Khoshgoftaar, T. M., and Wang, D. (2016). A survey of transfer learning. Journal of Big data, 3(1):9.
Wiggins, W. F. and Tejani, A. S. (2022). On the opportunities and risks of foundation models for natural language processing in radiology. Radiology: Artificial Intelligence, 4(4):e220119.
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T., Cao, Y., and Narasimhan, K. (2023). Tree of thoughts: Deliberate problem solving with large language models. Advances in neural information processing systems, 36:11809–11822.
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., and Cao, Y. (2022). React: Synergizing reasoning and acting in language models. arXiv preprint arXiv:2210.03629.
Publicado
08/09/2026
Como Citar
NASCIMENTO, Dimas Cassimiro; SILVA, Daliton da.
We Keep Evaluating LLMs, But Are We Learning Anything?. In: SIMPÓSIO BRASILEIRO DE BANCO DE DADOS (SBBD), 41. , 2026, São Carlos/SP.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2026
.
p. 840-846.
ISSN 2763-8979.
DOI: https://doi.org/10.5753/sbbd.2026.249427.
