A Comparative Analysis of Generative Error Correction Strategies for Long-Form and Noisy Audio

  • Yanna Torres Gonçalves Universidade Federal do Ceará (UFC)
  • Vincenzo Silva Fadda Universidade Federal do Ceará (UFC)
  • Ticiana L. Coelho da Silva Universidade Federal do Ceará (UFC)
  • José A. Fernandes de Macêdo Universidade Federal do Ceará (UFC)

Resumo


Automatic Speech Recognition (ASR) degrades substantially under real-world acoustic conditions, motivating Generative Error Correction (GER) with Large Language Models (LLMs). This study evaluates Single-ASR and Multi-ASR GER, with and without retrieval-augmented generation (RAG) strategies, across 13,500 audio samples under controlled degradation. GER reduces overall Word Error Rate by up to 43.9%, reaching 45.6% under severe degradation, but increases error on clean audio due to over-correction. Maximal Marginal Relevance (MMR) retrieval outperforms random sampling under severe conditions, while Multi-ASR remains competitive but is more conditionsensitive. These results call for adaptive, input-aware correction strategies.
Palavras-chave: Generative Error Correction (GER), Automatic Speech Recognition (ASR), Acoustic Degradation, Noise robustness

Referências

Abouelenin, A., Ashfaq, A., Atkinson, A., Awadalla, H., Bach, N., Bao, J., Benhaim, A., Cai, M., Chaudhary, V., Chen, C., et al. (2025). Phi-4-mini technical report: Compact yet powerful multimodal language models via mixture-of-loras. arXiv preprint arXiv:2503.01743.

Goldstein, J. and Carbonell, J. G. (1998). Summarization:(1) using mmr for diversity-based reranking and (2) evaluating summaries. In TIPSTER TEXT PROGRAM PHASE III: Proceedings of a Workshop held at Baltimore, Maryland, October 13-15, 1998, pages 181–195.

Ma, R., Qian, M., Gales, M., and Knill, K. (2025). Asr error correction using large language models. IEEE Transactions on Audio, Speech and Language Processing, 33:1389–1401.

Polevoi, A., Kragin, A., and Loukachevitch, N. (2025). Ground truth-free wer prediction for asr via audio quality and model confidence features. In International Conference on Speech and Computer, pages 29–44. Springer.

Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I. (2022). Robust speech recognition via large-scale weak supervision.

Reddy, C. K., Beyrami, E., Pool, J., Cutler, R., Srinivasan, S., and Gehrke, J. (2019). A scalable noisy speech dataset and online subjective test framework. In Interspeech 2019, pages 1816–1820. ISCA.

Sekoyan, M., Koluguri, N. R., Tadevosyan, N., Zelasko, P., Bartley, T., Karpov, N., Balam, J., and Ginsburg, B. (2025). Canary-1b-v2 & parakeet-tdt-0.6 b-v3: Efficient and high-performance models for multilingual asr and ast. arXiv preprint arXiv:2509.14128.

Srivastav, V., Zheng, S., Bezzam, E., Bihan, E. L., Koluguri, N., Żelasko, P., Majumdar, S., Moumen, A., and Gandhi, S. (2025). Open asr leaderboard: Towards reproducible and transparent multilingual and long-form speech recognition evaluation.

Wei, V. J., Wang, W., Jiang, D., Song, Y., and Wang, L. (2025). Asr-ec benchmark: Evaluating large language models on chinese asr error correction. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track, pages 1567–1575.

Yang, A., Li, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Gao, C., Huang, C., Lv, C., Zheng, C., Liu, D., Zhou, F., Huang, F., Hu, F., Ge, H., Wei, H., Lin, H., Tang, J., Yang, J., Tu, J., Zhang, J., Yang, J., Yang, J., Zhou, J., Zhou, J., Lin, J., Dang, K., Bao, K., Yang, K., Yu, L., Deng, L., Li, M., Xue, M., Li, M., Zhang, P., Wang, P., Zhu, Q., Men, R., Gao, R., Liu, S., Luo, S., Li, T., Tang, T., Yin, W., Ren, X., Wang, X., Zhang, X., Ren, X., Fan, Y., Su, Y., Zhang, Y., Zhang, Y., Wan, Y., Liu, Y., Wang, Z., Cui, Z., Zhang, Z., Zhou, Z., and Qiu, Z. (2025). Qwen3 technical report.

Yang, C.-H. H., Gu, Y., Liu, Y.-C., Ghosh, S., Bulyko, I., and Stolcke, A. (2023). Generative speech recognition error correction with large language models and task-activating prompting. In 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), pages 1–8. IEEE.

Yang, C.-H. H., Park, T., Gong, Y., Li, Y., Chen, Z., Lin, Y.-T., Chen, C., Hu, Y., Dhawan, K., Żelasko, P., et al. (2024). Large language model based generative error correction: A challenge and baselines for speech recognition, speaker tagging, and emotion recognition. In 2024 IEEE Spoken Language Technology Workshop (SLT), pages 371–378. IEEE.

Zhang, Y., Li, M., Long, D., Zhang, X., Lin, H., Yang, B., Xie, P., Yang, A., Liu, D., Lin, J., Huang, F., and Zhou, J. (2025). Qwen3 embedding: Advancing text embedding and reranking through foundation models. arXiv preprint arXiv:2506.05176.
Publicado
08/09/2026
GONÇALVES, Yanna Torres; FADDA, Vincenzo Silva; COELHO DA SILVA, Ticiana L.; MACÊDO, José A. Fernandes de. A Comparative Analysis of Generative Error Correction Strategies for Long-Form and Noisy Audio. In: SIMPÓSIO BRASILEIRO DE BANCO DE DADOS (SBBD), 41. , 2026, São Carlos/SP. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 373-386. ISSN 2763-8979. DOI: https://doi.org/10.5753/sbbd.2026.249224.