Graph-Driven Exam Gen: Plataforma GraphRAG para Geração de Itens com Ancoragem na Matriz de Referência do ENEM

  • Jeová Leite Universidade de Pernambuco (UPE)
  • Arthur Lopes Universidade de Pernambuco (UPE)
  • Marcos Monteiro Universidade de Pernambuco (UPE)
  • Ig Ibert Bittencourt Universidade Federal de Alagoas (UFAL) https://orcid.org/0000-0001-5676-2280
  • Emanuel Marques Queiroga Universidade Federal de Santa Maria (UFSM)
  • Aêda Monalliza Cunha de Sousa Universidade de Pernambuco (UPE)

Resumo


Sistemas de geração automática de itens baseados em LLMs produzem questões linguisticamente corretas, mas sem garantia de alinhamento à Matriz de Referência do ENEM [Tan et al. 2025]. A proposta modela a Matriz como grafo de propriedades Neo4j injetado como contexto curricular estruturado (GraphRAG) e compara três condições de geração (contexto do grafo, few-shot sem grafo, sem contexto) com GLM-5.1 a temperatura 0. Os 150 itens (10 habilidades × 3 × 5) apresentaram conformidade estrutural, e o contexto do grafo produziu maior coerência e especificidade lexical que as demais condições. A qualidade pedagógica, porém, ainda depende da avaliação por especialistas proposta como próxima etapa.
Palavras-chave: Geração Automática de Itens, Grafos de Conhecimento, ENEM

Referências

Agrawal, G., Kumarage, T., Alghamdi, Z., and Liu, H. (2024). Can knowledge graphs reduce hallucinations in LLMs? A survey. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3947–3960. Association for Computational Linguistics.

Amanlou, M., Moghaddam, E. S., Jafari, Y. A., Noori, M., Farsi, F., and Bahrak, B. (2026). KNIGHT: Knowledge graph-driven multiple-choice question generation with adaptive hardness calibration. In Proceedings of the Third Conference on Parsimony and Learning (CPAL 2026). arXiv:2602.20135.

Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., et al. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 33:1877–1901.

Chen, C. H. and Shiu, M. F. (2025). KAQG: A knowledge-graph-enhanced RAG for difficulty-controlled question generation. IEEE Access. Early access.

Han, H., Ma, L., Wang, Y., Shomer, H., Lei, Y., Qi, Z., Guo, K., Hua, Z., Long, B., Liu, H., Aggarwal, C. C., and Tang, J. (2025). RAG vs. GraphRAG: A systematic evaluation and key insights. arXiv preprint, arXiv:2502.11371.

INEP (2009). Matriz de referência para o ENEM 2009. Technical report, Instituto Nacional de Estudos e Pesquisas Educacionais Anísio Teixeira (INEP/MEC), Brasília.

Jaloto, A. and Primi, R. (2024). Enem de próxima geração com menos itens e alta confiabilidade usando CAT. Estudos em Avaliação Educacional, 35:e10142.

Kim, E., Li, S., Khalil, S., and Shin, H. J. (2025). STAIR-AIG: Optimizing the automated item generation process through human-AI collaboration for critical thinking assessment. In Proceedings of the 20th Workshop on Innovative Use of NLP for Building Educational Applications (BEA 2025), pages 920–930. Association for Computational Linguistics.

Kurdi, G., Leo, J., Parsia, B., Sattler, U., and Al-Emari, S. (2020). A systematic review of automatic question generation for educational purposes. International Journal of Artificial Intelligence in Education, 30(1):121–204.

Lai, J. W., Ho, S. Y., and Lim, F. S. (2026). Learning analytics for assessment preparation: Constructing graphs to guide higher-order question generation. In Proceedings of the 16th International Learning Analytics and Knowledge Conference (LAK 2026).

Lavrinovics, E., Biswas, R., Bjerva, J., and Hose, K. (2025). Knowledge graphs, large language models, and hallucinations: An NLP perspective. Web Semantics: Science, Services and Agents on the World Wide Web, 85:100844.

Li, M., Miao, S., and Li, P. (2025). Simple is effective: The roles of graphs and large language models in knowledge-graph-based retrieval-augmented generation. In Proceedings of the 13th International Conference on Learning Representations (ICLR 2025). arXiv:2410.20724.

Liang, H., Lin, Q., Han, Z., Ma, X., Wong, Z. H., Qiang, M., Sun, L., and Zhang, W. (2026). K12-KGraph: A curriculum-aligned knowledge graph for benchmarking and training educational LLMs. arXiv preprint, arXiv:2605.09635.

Lu, X. and Wang, X. (2024). Generative students: Using LLM-simulated student profiles to support question item evaluation. In Proceedings of the Eleventh ACM Conference on Learning @ Scale (L@S 2024). ACM.

Mucciaccia, S. S., Paixão, T. M., Mutz, F., Badue, C. S., De Souza, A. F., and Oliveira-Santos, T. (2025). Automatic multiple-choice question generation and evaluation systems based on LLM: A study case with university resolutions. In Proceedings of the 31st International Conference on Computational Linguistics, pages 2246–2260.

Quincozes, C. B., Molinos, D., Araújo, R. D., Quincozes, S., and Guedes, G. T. A. (2025). Engenharia de prompt para a geração automatizada de questões assistida por LLMs: Uma análise comparativa. In Anais do XXXVI Simpósio Brasileiro de Informática na Educação (SBIE 2025). SBC.

Robinson, I., Webber, J., and Eifrem, E. (2015). Graph Databases: New Opportunities for Connected Data. O'Reilly Media, Sebastopol, CA, 2nd edition.

Runge, A., Attali, Y., LaFlair, G. T., Park, Y., and Church, J. (2024). A generative AI-driven interactive listening assessment task. Frontiers in Artificial Intelligence, 7.

Santos, M. M., Barros, A. P., Santos, E., Silva, J. G. d., Isotani, S., Bittencourt, I. I., Macario, V., Rodrigues, L., and Dermeval, D. (2025). Geração de questões com LLMs leves: Um estudo inicial sobre a percepção de educadores. In Anais do XXXVI Simpósio Brasileiro de Informática na Educação (SBIE 2025), pages 1635–1646. SBC.

Scaria, N., Chenna, S. D., and Subramani, D. N. (2024). Automated educational question generation at different bloom's skill levels using large language models: Strategies and evaluation. In Proceedings of the 25th International Conference on Artificial Intelligence in Education (AIED 2024), Lecture Notes in Computer Science, pages 165–179. Springer.

Seyler, D., Yahya, M., and Berberich, K. (2017). Knowledge questions from knowledge graphs. In Proceedings of the ACM SIGIR International Conference on the Theory of Information Retrieval (ICTIR 2017), pages 11–18, Amsterdam, The Netherlands.

Tan, B., Armoush, N., Mazzullo, E., Bulut, O., and Gierl, M. J. (2025). A review of automatic item generation techniques leveraging large language models. International Journal of Assessment Tools in Education, 12(2):317–340.

Wang, Y., Wei, T., Li, Q., and Zeng, L. (2026). Beyond static question banks: Dynamic knowledge expansion via LLM-automated graph construction and adaptive generation. Proceedings of the VLDB Endowment, 14(1). arXiv:2602.00020.

Wei, Y., Carvalho, P., and Stamper, J. (2025). KCluster: An LLM-based clustering approach to knowledge component discovery. In Proceedings of the 18th International Conference on Educational Data Mining (EDM 2025). International Educational Data Mining Society. arXiv:2505.06469.
Publicado
05/10/2026
LEITE, Jeová; LOPES, Arthur; MONTEIRO, Marcos; BITTENCOURT, Ig Ibert; QUEIROGA, Emanuel Marques; DE SOUSA, Aêda Monalliza Cunha. Graph-Driven Exam Gen: Plataforma GraphRAG para Geração de Itens com Ancoragem na Matriz de Referência do ENEM. In: SIMPÓSIO BRASILEIRO DE INFORMÁTICA NA EDUCAÇÃO (SBIE), 37. , 2026, Goiânia/GO. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 2687-2697. DOI: https://doi.org/10.5753/sbie.2026.27847.