Agentic Multi-Document Summarization with Anchoring and Consolidation of Parallel Events: An Experiment with the Four Gospels
Resumo
Esta pesquisa investiga a sumarização multidocumento (MDS) de narrativas longas, problema relevante do Processamento de Linguagem Natural (PLN) devido às dificuldades de integração semântica e preservação da coerência temporal entre fontes paralelas, com aplicações em narrativas históricas e documentais. Posicionamo-nos especificamente no sub-problema de Consolidação Narrativa (NC), que estende a MDS clássica ao priorizar integridade cronológica e completude informacional em vez de compressão. Desenvolvemos e avaliamos um protótipo agentizado de MDS com ancoragem temporal e consolidação abstrativa de eventos paralelos, aplicado aos quatro Evangelhos em formato XML. O sistema compreende uma pipeline modular com quatro agentes autônomos: ancoragem de eventos; consolidação abstrativa; organização de narrativa e fluência narrativa. Implementados em Python com LangChain e LangGraph, modelando eventos como grafos temporais. A validação por métricas automáticas (Kendall’s Tau, ROUGE-L e BERTScore) e análise qualitativa por especialista indica ganhos em fidelidade semântica e fluidez narrativa, superando LLMs de estado da arte em regime zero-shot (BERTScore: 0,9379 vs. 0,9013 do melhor baseline). Os resultados demonstram a viabilidade técnica da orquestração agentizada como abordagem confiável para consolidação multidocumento.
Palavras-chave:
Sumarização Multidocumento, Agentes Autônomos, Grandes Modelos de Linguagem, Ancoragem Temporal, Processamento de Linguagem Natural
Referências
Barzilay, R., McKeown, K. R., and Elhadad, M. (1999). Information fusion in the context of multidocument summarization. In Proceedings of the 37th Annual Meeting of the Association for Computational Linguistics, pages 550–557, College Park, Maryland, EUA.
Bedi, H., Patil, S., Hingmire, S., and Palshikar, G. (2017). Event timeline generation from history textbooks. In Proceedings of the 4th Workshop on Natural Language Processing Techniques for Educational Applications (NLPTEA 2017), pages 69–77, Taipei, Taiwan.
Biblica, Inc. (2011). The holy bible, new international version. Acesso em: 16 out. 2025.
Cunha, A. (2025). Semana da paixão unificada: PLN e ancoragem temporal na harmonização dos evangelhos. Anais do IV Congresso Brasileiro de Humanismo Solidário na Ciência.
Erkan, G. and Radev, D. R. (2004). LexRank: Graph-based lexical centrality as salience in text summarization. Journal of Artificial Intelligence Research, 22:457–479.
Finger, R. A., Cortes, E. G., Rigo, S. J., and de O. Ramos, G. (2025). Narrative consolidation: Formulating a new task for unifying multi-perspective accounts.
Kendall, M. G. (1938). A new measure of rank correlation. Biometrika, 30(1/2):81–93.
LangChain Team (2024). LangGraph documentation. Acesso em: 2025.
Lin, C.-Y. (2004). ROUGE: A package for automatic evaluation of summaries. In Proceedings of the Workshop on Text Summarization Branches Out, pages 74–81, Barcelona, Espanha.
Ma, C., Zhang, W. E., Guo, M., Wang, H., and Sheng, Q. Z. (2023). Multi-document summarization via deep learning techniques: A survey. ACM Computing Surveys, 55(5):1–37.
Mamidala, K. K. and Sanampudi, S. K. (2021). A novel framework for multi-document temporal summarization (MDTS). Emerging Science Journal, 5(2):184–190.
Petersen, W. L. (1994). Tatian’s Diatessaron: Its Creation, Dissemination, Significance, and History in Scholarship. Brill, Leiden.
Russell, S. J. and Norvig, P. (2021). Artificial Intelligence: A Modern Approach. Pearson, Hoboken, 4 edition.
Santana, B., Campos, R., Amorim, E., Jorge, A., Silvano, P., and Nunes, S. (2023). A survey on narrative extraction from textual data. Artificial Intelligence Review, 56(8):8393–8435.
Song, J., Akhter, M. E., Atzil-Slonim, D., and Liakata, M. (2025). Temporal reasoning for timeline summarisation in social media. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 28085–28101, Viena, Áustria.
Wooldridge, M. (2009). An Introduction to Multiagent Systems. Wiley, Chichester, 2 edition.
Xiao, W., Beltagy, I., Carenini, G., and Cohan, A. (2022). PRIMERA: Pyramid-based masked sentence pre-training for multi-document summarization. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 4016–4036, Dublin, Irlanda.
Yasunaga, M., Zhang, R., Meelu, K., Pareek, A., Srinivasan, K., and Radev, D. (2017). Graph-based neural multi-document summarization. In Proceedings of the 21st Conference on Computational Natural Language Learning (CoNLL 2017), pages 452–462, Vancouver, Canadá.
Yu, Y., Jatowt, A., Doucet, A., Sugiyama, K., and Yoshikawa, M. (2021). Multi-timeline summarization (MTLS): Improving timeline summarization by generating multiple summaries. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 377–387, Online.
Zhang, J., Zhao, Y., Saleh, M., and Liu, P. J. (2020a). PEGASUS: Pre-training with extracted gap-sentences for abstractive summarization. In Proceedings of the 37th International Conference on Machine Learning, pages 11328–11339, Viena, Áustria.
Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., and Artzi, Y. (2020b). BERTScore: Evaluating text generation with BERT. In International Conference on Learning Representations.
Bedi, H., Patil, S., Hingmire, S., and Palshikar, G. (2017). Event timeline generation from history textbooks. In Proceedings of the 4th Workshop on Natural Language Processing Techniques for Educational Applications (NLPTEA 2017), pages 69–77, Taipei, Taiwan.
Biblica, Inc. (2011). The holy bible, new international version. Acesso em: 16 out. 2025.
Cunha, A. (2025). Semana da paixão unificada: PLN e ancoragem temporal na harmonização dos evangelhos. Anais do IV Congresso Brasileiro de Humanismo Solidário na Ciência.
Erkan, G. and Radev, D. R. (2004). LexRank: Graph-based lexical centrality as salience in text summarization. Journal of Artificial Intelligence Research, 22:457–479.
Finger, R. A., Cortes, E. G., Rigo, S. J., and de O. Ramos, G. (2025). Narrative consolidation: Formulating a new task for unifying multi-perspective accounts.
Kendall, M. G. (1938). A new measure of rank correlation. Biometrika, 30(1/2):81–93.
LangChain Team (2024). LangGraph documentation. Acesso em: 2025.
Lin, C.-Y. (2004). ROUGE: A package for automatic evaluation of summaries. In Proceedings of the Workshop on Text Summarization Branches Out, pages 74–81, Barcelona, Espanha.
Ma, C., Zhang, W. E., Guo, M., Wang, H., and Sheng, Q. Z. (2023). Multi-document summarization via deep learning techniques: A survey. ACM Computing Surveys, 55(5):1–37.
Mamidala, K. K. and Sanampudi, S. K. (2021). A novel framework for multi-document temporal summarization (MDTS). Emerging Science Journal, 5(2):184–190.
Petersen, W. L. (1994). Tatian’s Diatessaron: Its Creation, Dissemination, Significance, and History in Scholarship. Brill, Leiden.
Russell, S. J. and Norvig, P. (2021). Artificial Intelligence: A Modern Approach. Pearson, Hoboken, 4 edition.
Santana, B., Campos, R., Amorim, E., Jorge, A., Silvano, P., and Nunes, S. (2023). A survey on narrative extraction from textual data. Artificial Intelligence Review, 56(8):8393–8435.
Song, J., Akhter, M. E., Atzil-Slonim, D., and Liakata, M. (2025). Temporal reasoning for timeline summarisation in social media. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 28085–28101, Viena, Áustria.
Wooldridge, M. (2009). An Introduction to Multiagent Systems. Wiley, Chichester, 2 edition.
Xiao, W., Beltagy, I., Carenini, G., and Cohan, A. (2022). PRIMERA: Pyramid-based masked sentence pre-training for multi-document summarization. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 4016–4036, Dublin, Irlanda.
Yasunaga, M., Zhang, R., Meelu, K., Pareek, A., Srinivasan, K., and Radev, D. (2017). Graph-based neural multi-document summarization. In Proceedings of the 21st Conference on Computational Natural Language Learning (CoNLL 2017), pages 452–462, Vancouver, Canadá.
Yu, Y., Jatowt, A., Doucet, A., Sugiyama, K., and Yoshikawa, M. (2021). Multi-timeline summarization (MTLS): Improving timeline summarization by generating multiple summaries. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 377–387, Online.
Zhang, J., Zhao, Y., Saleh, M., and Liu, P. J. (2020a). PEGASUS: Pre-training with extracted gap-sentences for abstractive summarization. In Proceedings of the 37th International Conference on Machine Learning, pages 11328–11339, Viena, Áustria.
Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., and Artzi, Y. (2020b). BERTScore: Evaluating text generation with BERT. In International Conference on Learning Representations.
Publicado
19/10/2026
Como Citar
CUNHA, André F. de C. C.; SENA, Carlos A.; BARBOSA, Jacson R..
Agentic Multi-Document Summarization with Anchoring and Consolidation of Parallel Events: An Experiment with the Four Gospels. In: WORKSHOP-ESCOLA DE SISTEMAS DE AGENTES, SEUS AMBIENTES E APLICAÇÕES (WESAAC), 20. , 2026, Cuiabá/MT.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2026
.
p. 37-48.
ISSN 2326-5434.
DOI: https://doi.org/10.5753/wesaac.2026.30734.
