Evaluation of Fréchet Audio Distance with Bass Sounds

  • Adelmo Assmar UFBA
  • Flávio Assis UFBA

Resumo


Objective metrics are important to determine the effectiveness of automatic audio processing, such as in denoising and music source separation. The Fréchet Audio Distance (FAD) is an objective metric that can be used to compare two sets of audios based on their statistical properties. An advantage of FAD when compared to traditional metrics is that FAD does not depend on the existence of the original clean version of processed or noisy audios to be evaluated. However, the effectiveness of the use of FAD has been divergent in the literature, requiring thus further investigation. In this paper we study the impact on FAD scores when different reference sets and models are used, for the restricted case of bass sounds. We studied the sensitivity of FAD to audio effects and noises and used it to evaluate denoising processes. We compared the results with a qualified subjective evaluation, carried out by people with experience with professional recording. We show that reference sets must be designed very specifically, as apparently equivalent sets might provide different results. We additionally highlight the usefulness of understanding FAD scores relatively, instead of absolutely.

Palavras-chave: Fréchet audio distance, automatic audio enhancement, audio evaluation metrics

Referências

Y. Mitsufuji, G. Fabbro, S. Uhlich, F.-R. Stöter, A. Défossez, M. Kim, W. Choi, C.-Y. Yu, and K.-W. Cheuk. Music demixing challenge 2021. Frontiers in Signal Processing, 1, Jan 2022.

R. Hennequin, A. Khlif, F. Voituret, and M. Moussallam. Spleeter: a fast and efficient music source separation tool with pre-trained models. Journal of Open Source Software, 5(50):2154, 2020. Deezer Research.

A. Défossez, N. Usunier, L. Bottou, and F. Bach. Music source separation in the waveform domain, 2021. arXiv:1911.13254.

F.-R. Stöter, S. Uhlich, A. Liutkus, and Y. Mitsufuji. Open-Unmix - a reference implementation for music source separation. Journal of Open Source Software, 2019.

A. Hines, J. Skoglund, A.C. Kokaram, and N. Harte. ViSQOL: an objective speech quality model. EURASIP Journal on Audio, Speech, and Music Processing, 2015(1):13, 2015.

A.W. Rix, J.G. Beerends, M.P. Hollier, and A.P. Hekstra. Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs. In 2001 IEEE International Conference on Acoustics, Speech, and Signal Processing. Proceedings (Cat. No.01CH37221), volume 2, pages 749–752 vol.2, 2001.

J.G. Beerends, C. Schmidmer, J.Berger, M.Obermann, R. Ullmann, J.Pomy, and M. Keyhl. Perceptual objective listening quality assessment (POLQA), the third generation ITU-T standard for end-to-end speech quality measurement part i—temporal alignment. Journal of the Audio Engineering Society, 61(6):366–384, june 2013.

E. Vincent, R. Gribonval, and C. Fevotte. Performance measurement in blind audio source separation. IEEE Transactions on Audio, Speech, and Language Processing, 14(4):1462–1469, 2006.

J.L. Roux, S. Wisdom, H. Erdogan, and J.R. Hershey. SDR - half-baked or well done?, 2018. arXiv:1811.02508.

N. Kandpal, O. Nieto, and Z. Jin. Music enhancement via image translation and vocoding. In ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 3124–3128, 2022.

E. Rumbold. A critical analysis of objective evaluation metrics for music source separation audio quality. Master’s thesis, Northwestern University - Computer Science Department, Aug. 2022. Master Thesis - Northwestern University - Computer Science Department - Technical Report NU-CS-2022-11.

K. Kilgour, M. Zuluaga, D. Roblek, and M. Sharifi. Fréchet audio distance: A metric for evaluating music enhancement algorithms. In Proceedings of INTERSPEECH, Graz, Austria, Sept. 2019.

F. Caspe, A. McPherson, and M. Sandler. DDX7: Differentiable FM synthesis of musical instrument sounds. In Proc. of the 23rd International Society for Music Information Retrieval Conference (ISMIR), 2022.

J. Imort, G. Fabbro, M.A. M. Ramírez, S. Uhlich, Y. Koyama, and Y. Mitsufuji. Distortion audio effects: Learning how to recover the clean signal, 2022. arXiv:2202.01664.

T.-M. Hung, B.-Y. Chen, Y.-T. Yeh, and Y.-H. Yang. A benchmarking initiative for audio-domain music generation using the freesound loop dataset, 2022. arXiv:2108.01576.

F. Kreuk, G. Synnaeve, A. Polyak, U. Singer, A. Défossez, J. Copet, D. Parikh, Y. Taigman, and Y. Adi. AudioGen: Textually guided audio generation. In The Eleventh International Conference on Learning Representations, 2023.

A. Gui, H. Gamper, S. Braun, and D. Emmanouilidou. Adapting frechet audio distance for generative music evaluation. In ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1331–1335, 2024.

M. Heusel, R. Hubert, U. Thomas, and N. Bernhard. GANs trained by a two time-scale update rule converge to a local nash equilibrium. Advances in Neural Information Processing Systems, pages 6626–6637, 2017.

S. Hershey, S. Chaudhuri, D.P.W. Ellis, J.F. Gemmeke, A. Jansen, R.C. Moore, M. Plakal, D. Platt, R.A. Saurous, B. Seybold, M. Slaney, R.J. Weiss, and K. Wilson. CNN architectures for large-scale audio classification. In 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 131–135, 2017.

D.C. Dowson and B.V Landau. The fréchet distance between multivariate normal distributions. Journal of Multivariate Analysis, 12(3):450–455, 1982.

A. Vinay and A. Lerch. Evaluating generative audio systems and their metrics. In Proc. of the 23rd Int. Society for Music Information Retrieval Conf. (ISMIR), Bengaluru, India, 2022.

A. Agostinelli, T.I. Denk, Z. Borsos, J. Engel, M. Verzetti, A. Caillon, Q. Huang, A. Jansen, A. Roberts, M. Tagliasacchi, M. Sharifi, N. Zeghidour, and C. Frank. MusicLM: Generating music from text, 2023. arXiv:2301.11325.

J. Copet, F. Kreuk, I. Gat, T. Remez, D. Kant, G. Synnaeve, Y. Adi, and A. Défossez. Simple and controllable music generation. In Proc. of the 37th International Conference on Neural Information Processing Systems, NIPS ’23, Red Hook, NY, USA, 2024. Curran Associates Inc.

B. Hayes, C. Saitis, and G. Fazekas. Neural waveshaping synthesis. In Proc. of the 22nd Int. Society for Music Information Retrieval Conf. (ISMIR), 2021.

J. Nistal, S. Lattner, and G. Richard. DarkGAN: Exploiting knowledge distillation for comprehensible audio synthesis with gans. In Proc. of the 22nd Int. Society for Music Information Retrieval Conf. (ISMIR), 2021.

S. Rouard and G. Hadjeres. CRASH: Raw audio score-based generative modeling for controllable high-resolution drum sound synthesis. In Proc. of the 22nd Int. Society for Music Information Retrieval Conf. (ISMIR), 2021.

E. Moliner, J. Lehtinen, and V. Välimäki. Solving audio inverse problems with a diffusion model. In ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1–5, 2023.

Z. Kong, W. Ping, J. Huang, K. Zhao, and B. Catanzaro. DiffWave: A versatile diffusion model for audio synthesis. In Proc. of the 9th International Conference on Learning Representations (ICLR), 2021.

J. Engel, L. Hantrakul, C. Gu, and A. Roberts. DDSP: Differentiable digital signal processing. In Proc. of the 8th International Conference on Learning Representations (ICLR), Addis Ababa, Ethiopia, 2020.

J. Engel, C. Resnick, A. Roberts, S. Dieleman, M. Norouzi, D. Eck, and K. Simonyan. Neural audio synthesis of musical notes with WaveNet autoencoders. In Proc. of the 34th International Conference on Machine Learning (PLMR), 2017.

E. Law, K. West, M. Mandel, M. Bay, and J.S. Downie. Evaluation of algorithms using games: the case of music annotation. In Proceedings of the 10th International Conference on Music Information Retrieval (ISMIR), 2009.

B. Elizalde, S. Deshmukh, M. Al Ismail, and H. Wang. Clap learning audio concepts from natural language supervision. In Proc. of the 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1–5. IEEE, 2023.

Y. Wu, K. Chen, T. Zhang, Y. Hui, T. Berg-Kirkpatrick, and S. Dubnov. Large-scale contrastive language-audio pre-training with feature fusion and keyword-to-caption augmentation. In IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP, 2023.

Y. Li, R. Yuan, G. Zhang, Y. Ma, X. Chen, H. Yin, C. Xiao, C. Lin, A. Ragni, E. Benetos, N. Gyenge, R. Dannenberg, R. Liu, W. Chen, G. Xia, Y. Shi, W. Huang, Z. Wang, Y. Guo, and J. Fu. MERT: acoustic music understanding model with large-scale self-supervised training, 2024. arXiv:2306.00107.

P. Manocha, Z. Jin, R. Zhang, and A. Finkelstein. CD-PAM: Contrastive learning for perceptual audio similarity. In ICASSP 2021, jun 2021.

A. Défossez, J. Copet, G. Synnaeve, and Y. Adi. High fidelity neural audio compression, 2022. arXiv:2210.13438.

R. Kumar, P. Seetharaman, A. Luebs, I. Kumar, and K. Kumar. High-fidelity audio compression with improved rvq-gan, 2023. arXiv:2306.06546.

S. Uhlich, M. Porcu, F. Giron, M. Enenkl, T. Kemp, N. Takahashi, and Y. Mitsufuji. Improving music source separation based on deep neural networks through data augmentation and network blending. In 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 261–265, 2017.

C. Bagwell et al. Sound eXchange (SoX), 2015. Available at: [link].

B. McFee, M. McVicar, D. Faronbi, I. Roman, et al. librosa/librosa: 0.10.2.post1, 2024. DOI: 10.5281/zenodo.11192913.

J. Abeßer, H. Lukashevich, and G. Schuller. Feature-based extraction of plucking and expression styles of the electric bass guitar. In 2010 IEEE International Conference on Acoustics, Speech and Signal Processing, pages 2290–2293, 2010.

Z. Rafii, A. Liutkus, F.-R. Stöter, S.I. Mimilakis, and R. Bittner. The MUSDB18 corpus for music separation, December 2017. DOI: 10.5281/zenodo.1117372.
Publicado
15/09/2025
ASSMAR, Adelmo; ASSIS, Flávio. Evaluation of Fréchet Audio Distance with Bass Sounds. In: SIMPÓSIO BRASILEIRO DE COMPUTAÇÃO MUSICAL (SBCM), 19. , 2025, Campinas/SP. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2025 . p. 74-81. DOI: https://doi.org/10.5753/sbcm.2025.12254.