What Does It Take to Research with AI? A Rapid Review of Competencies to Train LLM-Literate Researchers

  • Danilo Monteiro Ribeiro CESAR School
  • Ronnie de Souza Santos CESAR School / University of Calgary
  • Rodrigo Siqueira CESAR School
  • Breno Andrade CESAR School
  • Rafael Batista Duarte CESAR School / Universidade de Pernambuco (UPE)
  • Gilberto Hida CESAR School
  • Julia Alencar CESAR School
  • Gustavo Pinto Universidade Federal do Pará (UFPA)

Resumo


The growing adoption of Large Language Models in scientific research has created a need to understand what competencies researchers and graduate students require to use these tools critically and responsibly. This rapid review analyzed 194 articles retrieved from Elicit and Google Scholar (2022–2025), from which 40 were selected for competency extraction and thematic analysis following independent dual screening (Gwet AC1: 0.76–0.83). Eight competencies were identified. The most prevalent was domain expertise and oversight of AI outputs (Σn = 123), encompassing subject-matter mastery, systematic skepticism, source verification, and researcher accountability. Other key competencies include metacognition and decision making about AI use (Σn = 55), ethics and academic integrity (Σn = 53), prompt engineering for research (Σn = 38), and reproducibility of AI use (Σn = 29). AI literacy and technical knowledge (Σn = 16) was explicitly identified as a risk factor when absent, with domain expertise treated as a prerequisite for meaningful critical evaluation. The findings suggest that preparing researchers to use LLMs goes beyond technical instruction, requiring an integrated set of epistemic, ethical, and methodological competencies centered on human accountability for the knowledge produced. These results have direct implications for the design of graduate programs and AI literacy initiatives.
Palavras-chave: Large Language Models, Research Competencies, Rapid Review

Referências

Baltes, S., Angermeir, F., Arora, C., Barón, M. M., Chen, C., Böhme, L., Calefato, F., Ernst, N., Falessi, D., Fitzgerald, B., et al. (2025). Guidelines for empirical studies in software engineering involving large language models. arXiv preprint arXiv:2508.15503.

Blau, W., Cerf, V. G., Enriquez, J., Francisco, J. S., Gasser, U., Gray, M. L., Greaves, M., Grosz, B. J., Jamieson, K. H., Haug, G. H., et al. (2024). Protecting scientific integrity in an age of generative ai.

Cartaxo, B., Pinto, G., and Soares, S. (2020). Rapid reviews in software engineering. In Contemporary empirical methods in software engineering, pages 357–384. Springer.

Chen, L., Chen, P., and Lin, Z. (2020). Artificial intelligence in education: A review. IEEE access, 8:75264–75278.

Cohen, J. (1968). Weighted kappa: Nominal scale agreement provision for scaled disagreement or partial credit. Psychological bulletin, 70(4):213.

Cruzes, D. S. and Dyba, T. (2011). Recommended steps for thematic synthesis in software engineering. In 2011 international symposium on empirical software engineering and measurement, pages 275–284. IEEE.

Gwet, K. L. (2008). Computing inter-rater reliability and its variance in the presence of high agreement. British Journal of Mathematical and Statistical Psychology, 61(1):29–48.

Kirova, V. D., Ku, C. S., Laracy, J. R., and Marlowe, T. J. (2024). Software engineering education must adapt and evolve for an llm environment. In Proceedings of the 55th ACM technical symposium on computer science education v. 1, pages 666–672.

Kitchenham, B., Brereton, O. P., Budgen, D., Turner, M., Bailey, J., and Linkman, S. (2009). Systematic literature reviews in software engineering–a systematic literature review. Information and software technology, 51(1):7–15.

Kitchenham, B. A., Dyba, T., and Jorgensen, M. (2004). Evidence-based software engineering. In Proceedings. 26th International Conference on Software Engineering, pages 273–281. IEEE.

Kozov, V., Ivanova, G., and Atanasova, D. (2024). Practical application of ai and large language models in software engineering education. International Journal of Advanced Computer Science & Applications, 15(1).

Krippendorff, K. (2018). Content analysis: An introduction to its methodology. Sage publications.

Landis JRKoch, G. (1977). The measurement of observer agreement for categorical data. Biometrics, 33(1):159174.

Lissack, M. and Meagher, B. (2024). Navigating the future of large language models in scientific research: Opportunities, challenges, and ethical considerations. Challenges, and Ethical Considerations (September 02, 2024).

Meerah, T. S. M., Osman, K., Zakaria, E., Ikhsan, Z. H., Krish, P., Lian, D. K. C., and Mahmod, D. (2012). Measuring graduate students research skills. Procedia-Social and Behavioral Sciences, 60:626–629.

Meyer, J. G., Urbanowicz, R. J., Martin, P. C., O'Connor, K., Li, R., Peng, P.-C., Bright, T. J., Tatonetti, N., Won, K. J., Gonzalez-Hernandez, G., et al. (2023). Chatgpt and large language models in academia: opportunities and challenges. BioData mining, 16(1):20.

Ng, D. T. K., Leung, J. K. L., Chu, S. K. W., and Qiao, M. S. (2021). Conceptualizing ai literacy: An exploratory review. Computers and Education: Artificial Intelligence, 2:100041.

Pizard, S., Acerenza, F., Otegui, X., Moreno, S., Vallespir, D., and Kitchenham, B. (2021). Training students in evidence-based software engineering and systematic reviews: a systematic review and empirical study. Empirical Software Engineering, 26(3):50.

Rahman, M., Terano, H. J. R., Rahman, N., Salamzadeh, A., and Rahaman, S. (2023). Chatgpt and academic research: A review and recommendations based on practical examples. Journal of Education, Management and Development Studies, 3(1):1–12.

Reddy, C. K. and Shojaee, P. (2025). Towards scientific discovery with generative ai: Progress, opportunities, and challenges. In Proceedings of the AAAI conference on artificial intelligence, volume 39, pages 28601–28609.

Reyes, C. E. G. and Morales, L. D. G. (2021). Research competencies mediated by technologies: A systematic mapping of the literature. Education in the Knowledge Society (EKS), 22:e23897–e23897.

Santos, R. d. S., Santos, I., Bento, M., Destefanis, G., Magalhães, C., and Wessel, M. (2026). Llm use, cheating, and academic integrity in software engineering education. International Conference on the Foundations of Software Engineering (FSE).

Trinkenreich, B., Calefato, F., Blincoe, K., Wivestad, V. T., Alves, A. P. S., Araújo, J. C., Araújo, M. C., Tell, P., Kalinowski, M., Zimmermann, T., et al. (2026). Taking a pulse on how generative ai is reshaping the software engineering research landscape. arXiv preprint arXiv:2604.11184.

Trinkenreich, B., Calefato, F., Hanssen, G., Blincoe, K., Kalinowski, M., Pezzè, M., Tell, P., and Storey, M.-A. (2025). Get on the train or be left on the station: Using llms for software engineering research. In Proceedings of the 33rd ACM International Conference on the Foundations of Software Engineering, pages 1503–1507.

Wohlin, C. (2007). Empirical software engineering: Teaching methods and conducting studies. In Empirical Software Engineering Issues. Critical Assessment and Future Directions: International Workshop, Dagstuhl Castle, Germany, June 26-30, 2006. Revised Papers, pages 135–142. Springer.
Publicado
05/10/2026
RIBEIRO, Danilo Monteiro; SANTOS, Ronnie de Souza; SIQUEIRA, Rodrigo; ANDRADE, Breno; DUARTE, Rafael Batista; HIDA, Gilberto; ALENCAR, Julia; PINTO, Gustavo. What Does It Take to Research with AI? A Rapid Review of Competencies to Train LLM-Literate Researchers. In: SIMPÓSIO BRASILEIRO DE INFORMÁTICA NA EDUCAÇÃO (SBIE), 37. , 2026, Goiânia/GO. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 2100-2114. DOI: https://doi.org/10.5753/sbie.2026.28206.