A Composite Metric for Abstractive Summarization Evaluation: Integrating Lexical, Semantic, and Factual Dimensions
Resumo
Automatic evaluation of abstractive summarization remains challenging because traditional lexical-overlap metrics often fail to capture semantic adequacy and factual consistency. This work proposes a composite evaluation metric integrating lexical, semantic, and factual dimensions. Metrics from each family are combined using a hierarchical weighted formulation, with both intra-family and inter-family weights optimized via Bayesian Optimization using the SummEval benchmark [Fabbri et al., 2021]. The resulting metric assigns 19% weight to lexical overlap, 58% to semantic similarity, and 23% to factuality, and achieves a Spearman correlation of 0.488 on the test set, outperforming the best individual metric, BERTScore-Precision (0.467). These results indicate that combining complementary evaluation dimensions yields a more robust and human-aligned signal for abstractive summarization.
Palavras-chave:
Automatic Text Summarization, Bayesian Optimization, Composite Metric, Evaluation Metrics, Natural Language Processing, Text Mining
Referências
Akiba, T., Sano, S., Yanase, T., Ohta, T., and Koyama, M. Optuna: A next-generation hyperparameter optimization framework. In Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. Anchorage, AK, USA, pp. 2623–2631, 2019.
Banerjee, S. and Lavie, A. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and Summarization. Ann Arbor, MI, USA, pp. 65–72, 2005.
Fabbri, A. R., Kryściński, W., McCann, B., Xiong, C., Socher, R., and Radev, D. R. SummEval: Re-evaluating summarization evaluation. Transactions of the Association for Computational Linguistics vol. 9, pp. 391–416, 2021.
Kryściński, W., McCann, B., Xiong, C., and Socher, R. Evaluating the factual consistency of abstractive text summarization. In Proceedings of the Conference on Empirical Methods in Natural Language Processing. Online, pp. 9332–9346, 2020.
Lin, C.-Y. ROUGE: A package for automatic evaluation of summaries. In Proceedings of the ACL Workshop on Text Summarization Branches Out. Barcelona, Spain, pp. 74–81, 2004.
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J. BLEU: a method for automatic evaluation of machine translation. In Proceedings of the Annual Meeting of the Association for Computational Linguistics. Philadelphia, PA, USA, pp. 311–318, 2002.
Pereira, V. pyautosummarizer. [link], 2023.
Reimers, N. and Gurevych, I. Sentence-BERT: Sentence embeddings using siamese BERT-networks. In Proceedings of the Conference on Empirical Methods in Natural Language Processing and the International Joint Conference on Natural Language Processing. Hong Kong, China, pp. 3982–3992, 2019.
Scialom, T., Dray, R., Delbé, C., Galerne, A., Cohen, S., Dragoni, F., Boucheron, J., Déloison, L., Galley, M., Sène, K., Sekhari, A., and Teton, F. QuestEval: Summarization asks for fact-based evaluation. In Proceedings of the Conference on Empirical Methods in Natural Language Processing. Online and Punta Cana, Dominican Republic, pp. 659–670, 2021.
Spearman, C. The proof and measurement of association between two things. The American Journal of Psychology 15 (1): 72–101, 1904.
Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., and Artzi, Y. BERTScore: Evaluating text generation with BERT. In Proceedings of the International Conference on Learning Representations. Online, 2020.
Banerjee, S. and Lavie, A. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and Summarization. Ann Arbor, MI, USA, pp. 65–72, 2005.
Fabbri, A. R., Kryściński, W., McCann, B., Xiong, C., Socher, R., and Radev, D. R. SummEval: Re-evaluating summarization evaluation. Transactions of the Association for Computational Linguistics vol. 9, pp. 391–416, 2021.
Kryściński, W., McCann, B., Xiong, C., and Socher, R. Evaluating the factual consistency of abstractive text summarization. In Proceedings of the Conference on Empirical Methods in Natural Language Processing. Online, pp. 9332–9346, 2020.
Lin, C.-Y. ROUGE: A package for automatic evaluation of summaries. In Proceedings of the ACL Workshop on Text Summarization Branches Out. Barcelona, Spain, pp. 74–81, 2004.
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J. BLEU: a method for automatic evaluation of machine translation. In Proceedings of the Annual Meeting of the Association for Computational Linguistics. Philadelphia, PA, USA, pp. 311–318, 2002.
Pereira, V. pyautosummarizer. [link], 2023.
Reimers, N. and Gurevych, I. Sentence-BERT: Sentence embeddings using siamese BERT-networks. In Proceedings of the Conference on Empirical Methods in Natural Language Processing and the International Joint Conference on Natural Language Processing. Hong Kong, China, pp. 3982–3992, 2019.
Scialom, T., Dray, R., Delbé, C., Galerne, A., Cohen, S., Dragoni, F., Boucheron, J., Déloison, L., Galley, M., Sène, K., Sekhari, A., and Teton, F. QuestEval: Summarization asks for fact-based evaluation. In Proceedings of the Conference on Empirical Methods in Natural Language Processing. Online and Punta Cana, Dominican Republic, pp. 659–670, 2021.
Spearman, C. The proof and measurement of association between two things. The American Journal of Psychology 15 (1): 72–101, 1904.
Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., and Artzi, Y. BERTScore: Evaluating text generation with BERT. In Proceedings of the International Conference on Learning Representations. Online, 2020.
Publicado
19/10/2026
Como Citar
OUVERNEY, Thiago; PEREIRA, Valdecy.
A Composite Metric for Abstractive Summarization Evaluation: Integrating Lexical, Semantic, and Factual Dimensions. In: SYMPOSIUM ON KNOWLEDGE DISCOVERY, MINING AND LEARNING (KDMILE), 14. , 2026, Cuiabá/MT.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2026
.
p. 17-24.
ISSN 2763-8944.
DOI: https://doi.org/10.5753/kdmile.2026.27518.
