A Multi-Agent System for Detection and Prioritization of Research Gaps in Scientific Literature
Resumo
Identifying promising research gaps remains a significant challenge in scientific research. As the volume of published literature increases, determining which questions remain unresolved and which approaches have been previously explored becomes more complex. This study introduces a multi-agent system designed to detect and prioritize candidate research gaps from either a research theme or a corpus of scientific papers. The system systematically maps the literature, identifies evidence-based research opportunities, evaluates their feasibility, and generates a prioritized report to support researchers during the literature review process. Instead of relying on a single language model, the workflow is divided into specialized components with clearly defined responsibilities. Deterministic modules handle tasks such as resource retrieval and ranking, while large language model (LLM)-based agents handle tasks requiring semantic interpretation, including evidence synthesis, classification, and critique. Candidate gaps are identified from two independent sources of evidence: infrequent but meaningful associations between scientific concepts extracted from an entity cooccurrence graph, and explicit statements of limitations or future work reported by original authors. Each candidate is linked to the corresponding supporting passages, assigned a confidence level, and excluded if the available evidence is insufficient. Experiments conducted on three literature corpora demonstrate that 90% of the reported candidates are supported by verbatim evidence extracted from the source documents. In contrast, a single-prompt baseline produces no evidence-grounded suggestions and generates more generic or trivial recommendations. Although the proposed system identifies fewer candidates, the resulting reports offer higher precision and greater traceability.Referências
Abd-alrazaq, A., Nashwan, A. J., Shah, Z., et al. (2024). Machine learning–based approach for identifying research gaps: COVID-19 as a case study. JMIR Formative Research, 8:e49411.
Agarwal, S., Laradji, I. H., Charlin, L., et al. (2024). LitLLM: A toolkit for scientific literature review. arXiv:2402.01788.
Asai, A., He, J., Shao, R., et al. (2024). OpenScholar: Synthesizing scientific literature with retrieval-augmented LMs. arXiv:2411.14199.
Baek, J., Jauhar, S. K., Cucerzan, S., et al. (2025). ResearchAgent: Iterative research idea generation over scientific literature with large language models. In Proc. NAACL.
Beltagy, I., Lo, K., and Cohan, A. (2019). SciBERT: A pretrained language model for scientific text. In Proc. EMNLP.
Bouma, G. (2009). Normalized (pointwise) mutual information in collocation extraction. In Proc. GSCL, pages 31–40.
Grootendorst, M. (2022). BERTopic: Neural topic modeling with a class-based TF-IDF procedure. arXiv:2203.05794.
Hong, S., Zhuge, M., Chen, J., et al. (2024). MetaGPT: Meta programming for a multi-agent collaborative framework. In Proc. ICLR.
Lála, J., O’Donoghue, O., Shtedritski, A., et al. (2023). PaperQA: Retrieval-augmented generative agent for scientific research. arXiv:2312.07559.
LangChain, Inc. (2024). LangGraph: A library for building stateful, multi-actor applications with LLMs. [link]. Accessed: 2026-06-22.
Li, G., Hammoud, H. A. A. K., Itani, H., et al. (2023). CAMEL: Communicative agents for “mind” exploration of large language model society. In Proc. NeurIPS.
Liang, W., Zhang, Y., Cao, H., et al. (2024). Can large language models provide useful feedback on research papers? a large-scale empirical analysis. NEJM AI, 1(8).
Liu, R. and Shah, N. B. (2023). ReviewerGPT? an exploratory study on using large language models for paper reviewing. arXiv:2306.00622.
Lu, C., Lu, C., Lange, R. T., et al. (2024). The AI scientist: Towards fully automated open-ended scientific discovery. arXiv:2408.06292.
Madaan, A., Tandon, N., Gupta, P., et al. (2023). Self-refine: Iterative refinement with self-feedback. In Proc. NeurIPS.
Reimers, N. and Gurevych, I. (2019). Sentence-BERT: Sentence embeddings using siamese BERT-networks. In Proc. EMNLP.
Salem, N. M., White, E., Bada, M., et al. (2025). GAPMAP: Mapping scientific knowledge gaps in biomedical literature using large language models. arXiv:2510.25055.
Shinn, N., Cassano, F., Berman, E., et al. (2023). Reflexion: Language agents with verbal reinforcement learning. In Proc. NeurIPS.
Si, C., Yang, D., and Hashimoto, T. (2024). Can LLMs generate novel research ideas? a large-scale human study with 100+ NLP researchers. arXiv:2409.04109.
Swanson, D. R. (1986). Undiscovered public knowledge. The Library Quarterly, 56(2):103–118.
Taylor, R., Kardas, M., Cucurull, G., et al. (2022). Galactica: A large language model for science. arXiv:2211.09085.
van Eck, N. J. and Waltman, L. (2010). Software survey: VOSviewer, a computer program for bibliometric mapping. Scientometrics, 84(2):523–538.
Wadden, D., Lin, S., Lo, K., et al. (2020). Fact or fiction: Verifying scientific claims. In Proc. EMNLP.
Wang, Y., Guo, Q., Yao, W., et al. (2024). AutoSurvey: Large language models can automatically write surveys. In Proc. NeurIPS.
Wei, J., Wang, X., Schuurmans, D., et al. (2022). Chain-of-thought prompting elicits reasoning in large language models. In Proc. NeurIPS.
Wooldridge, M. (2009). An Introduction to MultiAgent Systems. John Wiley & Sons, 2nd edition.
Wu, Q., Bansal, G., Zhang, J., et al. (2023). AutoGen: Enabling next-gen LLM applications via multi-agent conversation. arXiv:2308.08155.
Yao, S., Zhao, J., Yu, D., et al. (2023). ReAct: Synergizing reasoning and acting in language models. In Proc. ICLR.
Agarwal, S., Laradji, I. H., Charlin, L., et al. (2024). LitLLM: A toolkit for scientific literature review. arXiv:2402.01788.
Asai, A., He, J., Shao, R., et al. (2024). OpenScholar: Synthesizing scientific literature with retrieval-augmented LMs. arXiv:2411.14199.
Baek, J., Jauhar, S. K., Cucerzan, S., et al. (2025). ResearchAgent: Iterative research idea generation over scientific literature with large language models. In Proc. NAACL.
Beltagy, I., Lo, K., and Cohan, A. (2019). SciBERT: A pretrained language model for scientific text. In Proc. EMNLP.
Bouma, G. (2009). Normalized (pointwise) mutual information in collocation extraction. In Proc. GSCL, pages 31–40.
Grootendorst, M. (2022). BERTopic: Neural topic modeling with a class-based TF-IDF procedure. arXiv:2203.05794.
Hong, S., Zhuge, M., Chen, J., et al. (2024). MetaGPT: Meta programming for a multi-agent collaborative framework. In Proc. ICLR.
Lála, J., O’Donoghue, O., Shtedritski, A., et al. (2023). PaperQA: Retrieval-augmented generative agent for scientific research. arXiv:2312.07559.
LangChain, Inc. (2024). LangGraph: A library for building stateful, multi-actor applications with LLMs. [link]. Accessed: 2026-06-22.
Li, G., Hammoud, H. A. A. K., Itani, H., et al. (2023). CAMEL: Communicative agents for “mind” exploration of large language model society. In Proc. NeurIPS.
Liang, W., Zhang, Y., Cao, H., et al. (2024). Can large language models provide useful feedback on research papers? a large-scale empirical analysis. NEJM AI, 1(8).
Liu, R. and Shah, N. B. (2023). ReviewerGPT? an exploratory study on using large language models for paper reviewing. arXiv:2306.00622.
Lu, C., Lu, C., Lange, R. T., et al. (2024). The AI scientist: Towards fully automated open-ended scientific discovery. arXiv:2408.06292.
Madaan, A., Tandon, N., Gupta, P., et al. (2023). Self-refine: Iterative refinement with self-feedback. In Proc. NeurIPS.
Reimers, N. and Gurevych, I. (2019). Sentence-BERT: Sentence embeddings using siamese BERT-networks. In Proc. EMNLP.
Salem, N. M., White, E., Bada, M., et al. (2025). GAPMAP: Mapping scientific knowledge gaps in biomedical literature using large language models. arXiv:2510.25055.
Shinn, N., Cassano, F., Berman, E., et al. (2023). Reflexion: Language agents with verbal reinforcement learning. In Proc. NeurIPS.
Si, C., Yang, D., and Hashimoto, T. (2024). Can LLMs generate novel research ideas? a large-scale human study with 100+ NLP researchers. arXiv:2409.04109.
Swanson, D. R. (1986). Undiscovered public knowledge. The Library Quarterly, 56(2):103–118.
Taylor, R., Kardas, M., Cucurull, G., et al. (2022). Galactica: A large language model for science. arXiv:2211.09085.
van Eck, N. J. and Waltman, L. (2010). Software survey: VOSviewer, a computer program for bibliometric mapping. Scientometrics, 84(2):523–538.
Wadden, D., Lin, S., Lo, K., et al. (2020). Fact or fiction: Verifying scientific claims. In Proc. EMNLP.
Wang, Y., Guo, Q., Yao, W., et al. (2024). AutoSurvey: Large language models can automatically write surveys. In Proc. NeurIPS.
Wei, J., Wang, X., Schuurmans, D., et al. (2022). Chain-of-thought prompting elicits reasoning in large language models. In Proc. NeurIPS.
Wooldridge, M. (2009). An Introduction to MultiAgent Systems. John Wiley & Sons, 2nd edition.
Wu, Q., Bansal, G., Zhang, J., et al. (2023). AutoGen: Enabling next-gen LLM applications via multi-agent conversation. arXiv:2308.08155.
Yao, S., Zhao, J., Yu, D., et al. (2023). ReAct: Synergizing reasoning and acting in language models. In Proc. ICLR.
Publicado
19/10/2026
Como Citar
SILVA, Rodrigo; LEMOS, Bruno; FALCUCCI, Bruna; MAIA, Guilherme.
A Multi-Agent System for Detection and Prioritization of Research Gaps in Scientific Literature. In: WORKSHOP-ESCOLA DE SISTEMAS DE AGENTES, SEUS AMBIENTES E APLICAÇÕES (WESAAC), 20. , 2026, Cuiabá/MT.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2026
.
p. 312-323.
ISSN 2326-5434.
DOI: https://doi.org/10.5753/wesaac.2026.32053.
