MixRAG: a Low-cost alternative to a Graph RAG for scientific papers

  • Luis Cesar de Azevedo Universidade Federal do ABC (UFABC) https://orcid.org/0000-0001-6783-8009
  • Arthur P. Machado Universidade Estadual de Campinas (UNICAMP)
  • Ronaldo C. Prati Universidade Federal do ABC (UFABC)

Resumo


Retrieval-Augmented Generation (RAG) combines Large Language Models (LLMs) with retrieval systems to enhance the relevance and accuracy of AI-generated responses. For complex texts, such as academic papers, GraphRAG, a variant of RAG, further refines context quality by integrating Knowledge Graphs (KGs) into the retrieval process. Unlike traditional RAG, GraphRAG uses KGs to model entities, relationships, and hierarchies explicitly, capturing richer semantic connections (e.g., causal links in research findings or experimental dependencies). This structured, graph-based approach enables LLMs to reason over interconnected concepts, improving both the accuracy and coherence of outputs in technical or nuanced domains. However, the use of LLMs for constructing KGs becomes costly for large text sets, as the process requires repeated LLM processing of text to extract and structure entities and relationships. In this paper, we introduce MixRAG, a cost-effective RAG/GraphRAG hybrid approach. MixRAG optimizes resource usage by strategically generating a KG only from document abstracts, while storing the remainder as raw text in the vector database. This selective structuring balances the efficiency of a simple RAG with the informational depth of GraphRAG. Our experiments show that MixRAG reduces KG construction costs by more than 95%, while improving the response quality.
Palavras-chave: Retrieval-Augmented Generation, GraphRAG, Knowledge Graphs, Large Language Models, Cost Optimization

Referências

Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al. (2023). Gpt-4 technical report. arXiv preprint arXiv:2303.08774.

Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020). Language models are few-shot learners. NIPS’2020, 33:1877–1901.

Dahl, M., Magesh, V., Suzgun, M., and Ho, D. E. (2024). Large legal fictions: Profiling legal hallucinations in large language models. J. Leg. Anal,, 16(1):64–93.

Grattafiori, A., Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Vaughan, A., et al. (2024). The llama 3 herd of models. arXiv preprint arXiv:2407.21783.

Gu, J., Jiang, X., Shi, Z., Tan, H., Zhai, X., Xu, C., Li, W., Shen, Y., Ma, S., Liu, H., et al. (2024). A survey on llm-as-a-judge. The Innovation.

Guo, Z., Xia, L., Yu, Y., Ao, T., and Huang, C. (2024). Lightrag: Simple and fast retrieval-augmented generation. ArXiv, abs/2410.05779.

Han, H., Shomer, H., Wang, Y., Lei, Y., Guo, K., Hua, Z., Long, B., Liu, H., and Tang, J. (2025). Rag vs. graphrag: A systematic evaluation and key insights. ArXiv, abs/2502.11371.

Han, H., Wang, Y., Shomer, H., Guo, K., Ding, J., Lei, Y., Halappanavar, M., Rossi, R. A., Mukherjee, S., Tang, X., He, Q., Hua, Z., Long, B., Zhao, T., Shah, N., Javari, A., Xia, Y., and Tang, J. (2024). Retrieval-augmented generation with graphs (graphrag). ArXiv, abs/2501.00309.

Hogan, A., Blomqvist, E., Cochez, M., d’Amato, C., Melo, G. D., Gutierrez, C., Kirrane, S., Gayo, J. E. L., Navigli, R., Neumaier, S., et al. (2021). Knowledge graphs. ACM Comput Surv., 54(4):1–37.

Huang, L., Yu, W., Ma, W., Zhong, W., Feng, Z., Wang, H., Chen, Q., Peng, W., Feng, X., Qin, B., and Liu, T. (2023). A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Trans. Inf. Syst., 43:1 – 55.

Lee, M.-C., Zhu, Q., Mavromatis, C., Han, Z., Adeshina, S., Ioannidis, V. N., Rangwala, H., and Faloutsos, C. (2025). Hybgrag: Hybrid retrieval-augmented generation on textual and relational knowledge bases. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), page 879–893. Association for Computational Linguistics.

Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., et al. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. In NIPS’2020, volume 33.

Matsumoto, N., Moran, J., Choi, H., Hernandez, M. E., Venkatesan, M., Wang, P., and Moore, J. H. (2024). Kragen: a knowledge graph-enhanced rag framework for biomedical problem solving using large language models. Bioinformatics, 40(6):btae353.

Min, C., Bansal, S., Pan, J., Keshavarzi, A., Mathew, R., and Kannan, A. V. (2025). Towards practical graphrag: Efficient knowledge graph construction and hybrid retrieval at scale. arXiv preprint arXiv:2507.03226.

Munikoti, S., Acharya, A., Wagle, S., and Horawalavithana, S. (2023). Atlantic: Structure-aware retrieval-augmented language model for interdisciplinary science. arXiv preprint arXiv:2311.12289.

Peng, B., Zhu, Y., Liu, Y., Bo, X., Shi, H., Hong, C., Zhang, Y., and Tang, S. (2024). Graph retrieval-augmented generation: A survey. ACM Trans. Inf. Syst.

Yin, S., Fu, C., Zhao, S., Li, K., Sun, X., Xu, T., and Chen, E. (2023). A survey on multimodal large language models. Natl. Sci. Rev., 11.

Yu, H. Q. and McQuade, F. (2025). RAG-KG-IL: A multi-agent hybrid framework for reducing hallucinations and enhancing llm reasoning through rag and incremental knowledge graph learning integration. arXiv preprint arXiv:2503.13514.

Zhang, B. and Soh, H. (2024). Extract, define, canonicalize: An llm-based framework for knowledge graph construction. In Proceedings of the 2024 conference on empirical methods in natural language processing, pages 9820–9836.

Zhang, Q., Chen, S., Bei, Y.-Q., Yuan, Z., Zhou, H., Hong, Z., Dong, J., Chen, H., Chang, Y., and Huang, X. (2025). A survey of graph retrieval-augmented generation for customized large language models. ArXiv, abs/2501.13958.

Zhuang, L., Chen, S., Xiao, Y., Zhou, H., Zhang, Y., Chen, H., Zhang, Q., and Huang, X. (2025). Linearrag: Linear graph retrieval augmented generation on large-scale corpora. ArXiv, abs/2510.10114.
Publicado
08/09/2026
DE AZEVEDO, Luis Cesar; MACHADO, Arthur P.; PRATI, Ronaldo C.. MixRAG: a Low-cost alternative to a Graph RAG for scientific papers. In: BRAZILIAN E-SCIENCE WORKSHOP (BRESCI), 20. , 2026, São Carlos/SP. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 81-88. ISSN 2763-8774. DOI: https://doi.org/10.5753/bresci.2026.249338.