Representation Matters: An Evaluation of MEL, CQT, VQT, and MCQT for Beat Tracking
Resumo
Neural beat-tracking models predominantly rely on Mel spectrograms. These representations often fail to capture harmonic nuances in complex genres (e.g., jazz, classical), leading to rhythmically ambiguous signal patterns and suboptimal performance. We challenge the architectural-centric paradigm in music analysis by (1) evaluating alternative time-frequency representations, and (2) proposing a dual-input architecture to synergize complementary spectrogram properties. Through controlled experiments, we compare Mel, Constant-Q Transform (CQT), and Variable-Q Transform (VQT) inputs. We then introduce MCQT: a dual-branch model that fuses Mel and CQT, validated using 8-fold cross-validation and ensemble inference. Replacing Mel with CQT reduced subharmonic errors by 22% in harmony-rich contexts. MCQT further outperformed single-input baselines, achieving state-of-the-art performance in annotation coverage (0.912 vs. Mel’s 0.910) and reducing subharmonic errors by 26% compared to Mel. It excelled in complex genres (e.g., 6.4% higher downbeat F1-score in classical). Strategic representation engineering, not just architectural complexity, is pivotal for robust beat tracking. MCQT’s fusion of complementary spectrograms establishes a new paradigm for MIR tasks, particularly in harmonically complex music.Referências
Anastasia Natsiou and Sean O’Leary. Audio representations for deep learning in sound synthesis: a review. arXiv preprint arXiv:2201.02490, 2022.
Friedrich Wolf-Monheim. Spectral and rhythm features for audio classification with deep convolutional neural networks, 2024.
Judith C. Brown. Calculation of a constant Q spectral transform. The Journal of the Acoustical Society of America, 89(1):425–434, 1991.
Christian Schörkhuber and Anssi Klapuri. Constant-q transform toolbox for music processing. In 7th Sound and Music Computing Conference, Barcelona, Spain, page 3, 2010.
Steven Davis and Paul Mermelstein. Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences. IEEE transactions on acoustics, speech, and signal processing, 28(4):357–366, 1980.
Sebastian Böck, Florian Krebs, and Gerhard Widmer. Joint beat and downbeat tracking with recurrent neural networks. In Proceedings of the 17th International Society for Music Information Retrieval Conference (ISMIR), pages 255–261. New York, NY, USA, 2016.
Rainer Kelz, Sebastian Böck, and Gerhard Widmer. An end-to-end neural network for polyphonic piano music transcription. In Proceedings of the 17th International Society for Music Information Retrieval Conference (ISMIR), pages 676–682. New York, NY, USA, 2016.
Yicheng Gu, Xueyao Zhang, Liumeng Xue, and Zhizheng Wu. Multi-scale sub-band constant-q transform discriminator for high-fidelity vocoder, 2023.
Hongming Guo, Ruibo Fu, Yizhong Geng, Shuai Liu, Shuchen Shi, Tao Wang, Chunyu Qiang, Chenxing Li, Ya Li, Zhengqi Wen, Yukun Liu, and Xuefei Liu. Mel-refine: A plug-and-play approach to refine melspectrogram in audio generation, 2024.
Micael Antunes da Silva. Explorando Texturas Sonoras: Análise e Composição com Suporte de Descritores de Áudio e Modelos Computacionais. Tese (doutorado), Universidade Estadual de Campinas (UNICAMP), Instituto de Artes, Campinas, SP, 2024. Disponível em: 20.500.12733/20336. Acesso em: 14 jun. 2025.
Francesco Foscarin, Jan Schlüter, and Gerhard Widmer. Beat this! accurate beat tracking without dbn postprocessing. [link], 2024.
Haytham Fayek. Speech processing for machine learning: Filter banks, mel-frequency cepstral coefficients (mfccs) and what’s in-between. [link], April 2016.
Simon Dixon. Automatic music transcription. In Music Data Mining, pages 267–294. CRC Press, 2007.
Sebastian Böck and Markus Schedl. Enhanced beat tracking with context-aware neural networks. In Proceedings of the 20th ACM international conference on Multimedia, pages 657–660. ACM, 2012.
Joaquim B. Cavalcante and Julio Hsu. Beat this: Multi-representation beat tracking implementation. [link], 2025.
Jon Gillick, Adam Roberts, Jesse Engel, Douglas Eck, and David Bamman. Learning to groove with inverse sequence transformations. In International Conference on Machine Learning (ICML), volume 97, pages 2269–2279, 2019.
Fabien Gouyon. A computational approach to rhythm description—Audio features for the computation of rhythm periodicity functions and their use in tempo induction and music content processing. PhD thesis, Universitat Pompeu Fabra, 2005.
Jason A Hockman, Matthew EP Davies, and Ichiro Fujinaga. One in the jungle: Downbeat detection in hardcore, jungle, and drum and bass. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR), 2012.
Stephen W Hainsworth and Malcolm D Macleod. Particle filtering applied to musical tempo tracking. In EURASIP Journal on Advances in Signal Processing, volume 2004, 2004.
Andre Holzapfel, Matthew EP Davies, José R Zapata, João Lobato Oliveira, and Fabien Gouyon. Selective sampling for beat tracking evaluation. IEEE Transactions on Audio, Speech, and Language Processing, 20(9), 2012.
Vasiliy Eremenko, Eren Demirel, Barış Bozkurt, and Xavier Serra. Audio-aligned jazz harmony dataset for automatic chord transcription and corpus-based research. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR), 2018.
Jonathan Driedger, Hendrik Schreiber, W Bas De Haas, and Meinard Müller. Towards automatically correcting tapped beat annotations for music recordings. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR), 2019.
Dave Foster and Simon Dixon. Filosax: A dataset of annotated jazz saxophone recordings. In International Society for Music Information Retrieval (ISMIR) Conference, 2021.
Qingyang Xi, Rachel M Bittner, Johan Pauwels, Xuzhou Ye, and Juan Pablo Bello. GuitarSet: A dataset for guitar transcription. PhD thesis, New York University, 2018.
Matthew EP Davies, Norberto Degara, and Mark D Plumbley. Evaluation methods for musical audio beat tracking algorithms. Technical report, Queen Mary University of London, 2009.
Masataka Goto, Hiroki Hashiguchi, Takuichi Nishimura, and Ryuichi Oka. Rwc music database: Popular, classical and jazz music databases. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR), 2002.
Florian Krebs, Sebastian Böck, and Gerhard Widmer. Rhythmic pattern modeling for beat and downbeat tracking in musical audio. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR), 2013.
Francesco Foscarin, Andrew Mcleod, Philippe Rigaux, Florent Jacquemard, and Masahiko Sakai. Asap: a dataset of aligned scores and performances for piano transcription. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR), 2020.
Oriol Nieto, Matthew McCallum, Matthew Davies, Andrew Robertson, Adam Stark, and Eran Egozy. The harmonix set: Beats, downbeats, and functional segment annotations of western popular music. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR), 2019.
Martín Rocamora, Luis Jure, et al. Beat and downbeat tracking based on rhythmic patterns applied to the uruguayan candombe drumming. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR), 2015.
Tadas Baltrušaitis, Chaitanya Ahuja, and Louis-Philippe Morency. Multimodal machine learning: A survey and taxonomy. IEEE transactions on pattern analysis and machine intelligence, 41(2):423–443, 2019.
Robert A Jacobs, Michael I Jordan, Steven J Nowlan, and Geoffrey E Hinton. Adaptive mixtures of local experts. Neural computation, 3(1):79–87, 1991.
George Tzanetakis, Georg Essl, and Perry Cook. Musical genre classification of audio signals, 2001.
Matthew EP Davies and Mark D Plumbley. Evaluation of the audio beat tracking system beatroot. In Journal of New Music Research, volume 38, pages 117–129, 2009.
Ching-Yu Chiu, Meinard Müller, Matthew E. P. Davies, Alvin Wen-Yu Su, and Yi-Hsuan Yang. An analysis method for metric-level switching in beat tracking. IEEE Signal Processing Letters, 29:2153–2157, 2022.
Colin Raffel, Brian McFee, Eric J Humphrey, Justin Salamon, Oriol Nieto, Dawen Liang, and Daniel PW Ellis. mir eval: A transparent implementation of common mir metrics. Proceedings of the 15th International Society for Music Information Retrieval Conference, pages 367–372, 2014.
Friedrich Wolf-Monheim. Spectral and rhythm features for audio classification with deep convolutional neural networks, 2024.
Judith C. Brown. Calculation of a constant Q spectral transform. The Journal of the Acoustical Society of America, 89(1):425–434, 1991.
Christian Schörkhuber and Anssi Klapuri. Constant-q transform toolbox for music processing. In 7th Sound and Music Computing Conference, Barcelona, Spain, page 3, 2010.
Steven Davis and Paul Mermelstein. Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences. IEEE transactions on acoustics, speech, and signal processing, 28(4):357–366, 1980.
Sebastian Böck, Florian Krebs, and Gerhard Widmer. Joint beat and downbeat tracking with recurrent neural networks. In Proceedings of the 17th International Society for Music Information Retrieval Conference (ISMIR), pages 255–261. New York, NY, USA, 2016.
Rainer Kelz, Sebastian Böck, and Gerhard Widmer. An end-to-end neural network for polyphonic piano music transcription. In Proceedings of the 17th International Society for Music Information Retrieval Conference (ISMIR), pages 676–682. New York, NY, USA, 2016.
Yicheng Gu, Xueyao Zhang, Liumeng Xue, and Zhizheng Wu. Multi-scale sub-band constant-q transform discriminator for high-fidelity vocoder, 2023.
Hongming Guo, Ruibo Fu, Yizhong Geng, Shuai Liu, Shuchen Shi, Tao Wang, Chunyu Qiang, Chenxing Li, Ya Li, Zhengqi Wen, Yukun Liu, and Xuefei Liu. Mel-refine: A plug-and-play approach to refine melspectrogram in audio generation, 2024.
Micael Antunes da Silva. Explorando Texturas Sonoras: Análise e Composição com Suporte de Descritores de Áudio e Modelos Computacionais. Tese (doutorado), Universidade Estadual de Campinas (UNICAMP), Instituto de Artes, Campinas, SP, 2024. Disponível em: 20.500.12733/20336. Acesso em: 14 jun. 2025.
Francesco Foscarin, Jan Schlüter, and Gerhard Widmer. Beat this! accurate beat tracking without dbn postprocessing. [link], 2024.
Haytham Fayek. Speech processing for machine learning: Filter banks, mel-frequency cepstral coefficients (mfccs) and what’s in-between. [link], April 2016.
Simon Dixon. Automatic music transcription. In Music Data Mining, pages 267–294. CRC Press, 2007.
Sebastian Böck and Markus Schedl. Enhanced beat tracking with context-aware neural networks. In Proceedings of the 20th ACM international conference on Multimedia, pages 657–660. ACM, 2012.
Joaquim B. Cavalcante and Julio Hsu. Beat this: Multi-representation beat tracking implementation. [link], 2025.
Jon Gillick, Adam Roberts, Jesse Engel, Douglas Eck, and David Bamman. Learning to groove with inverse sequence transformations. In International Conference on Machine Learning (ICML), volume 97, pages 2269–2279, 2019.
Fabien Gouyon. A computational approach to rhythm description—Audio features for the computation of rhythm periodicity functions and their use in tempo induction and music content processing. PhD thesis, Universitat Pompeu Fabra, 2005.
Jason A Hockman, Matthew EP Davies, and Ichiro Fujinaga. One in the jungle: Downbeat detection in hardcore, jungle, and drum and bass. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR), 2012.
Stephen W Hainsworth and Malcolm D Macleod. Particle filtering applied to musical tempo tracking. In EURASIP Journal on Advances in Signal Processing, volume 2004, 2004.
Andre Holzapfel, Matthew EP Davies, José R Zapata, João Lobato Oliveira, and Fabien Gouyon. Selective sampling for beat tracking evaluation. IEEE Transactions on Audio, Speech, and Language Processing, 20(9), 2012.
Vasiliy Eremenko, Eren Demirel, Barış Bozkurt, and Xavier Serra. Audio-aligned jazz harmony dataset for automatic chord transcription and corpus-based research. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR), 2018.
Jonathan Driedger, Hendrik Schreiber, W Bas De Haas, and Meinard Müller. Towards automatically correcting tapped beat annotations for music recordings. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR), 2019.
Dave Foster and Simon Dixon. Filosax: A dataset of annotated jazz saxophone recordings. In International Society for Music Information Retrieval (ISMIR) Conference, 2021.
Qingyang Xi, Rachel M Bittner, Johan Pauwels, Xuzhou Ye, and Juan Pablo Bello. GuitarSet: A dataset for guitar transcription. PhD thesis, New York University, 2018.
Matthew EP Davies, Norberto Degara, and Mark D Plumbley. Evaluation methods for musical audio beat tracking algorithms. Technical report, Queen Mary University of London, 2009.
Masataka Goto, Hiroki Hashiguchi, Takuichi Nishimura, and Ryuichi Oka. Rwc music database: Popular, classical and jazz music databases. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR), 2002.
Florian Krebs, Sebastian Böck, and Gerhard Widmer. Rhythmic pattern modeling for beat and downbeat tracking in musical audio. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR), 2013.
Francesco Foscarin, Andrew Mcleod, Philippe Rigaux, Florent Jacquemard, and Masahiko Sakai. Asap: a dataset of aligned scores and performances for piano transcription. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR), 2020.
Oriol Nieto, Matthew McCallum, Matthew Davies, Andrew Robertson, Adam Stark, and Eran Egozy. The harmonix set: Beats, downbeats, and functional segment annotations of western popular music. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR), 2019.
Martín Rocamora, Luis Jure, et al. Beat and downbeat tracking based on rhythmic patterns applied to the uruguayan candombe drumming. In Proceedings of the International Society for Music Information Retrieval Conference (ISMIR), 2015.
Tadas Baltrušaitis, Chaitanya Ahuja, and Louis-Philippe Morency. Multimodal machine learning: A survey and taxonomy. IEEE transactions on pattern analysis and machine intelligence, 41(2):423–443, 2019.
Robert A Jacobs, Michael I Jordan, Steven J Nowlan, and Geoffrey E Hinton. Adaptive mixtures of local experts. Neural computation, 3(1):79–87, 1991.
George Tzanetakis, Georg Essl, and Perry Cook. Musical genre classification of audio signals, 2001.
Matthew EP Davies and Mark D Plumbley. Evaluation of the audio beat tracking system beatroot. In Journal of New Music Research, volume 38, pages 117–129, 2009.
Ching-Yu Chiu, Meinard Müller, Matthew E. P. Davies, Alvin Wen-Yu Su, and Yi-Hsuan Yang. An analysis method for metric-level switching in beat tracking. IEEE Signal Processing Letters, 29:2153–2157, 2022.
Colin Raffel, Brian McFee, Eric J Humphrey, Justin Salamon, Oriol Nieto, Dawen Liang, and Daniel PW Ellis. mir eval: A transparent implementation of common mir metrics. Proceedings of the 15th International Society for Music Information Retrieval Conference, pages 367–372, 2014.
Publicado
15/09/2025
Como Citar
CAVALCANTE, Joaquim B.; HSU, Julio; BATISTA, Carlos; M. FILHO, Telmo; GAUDÊNCIO, Thaís; MALHEIROS, Yuri.
Representation Matters: An Evaluation of MEL, CQT, VQT, and MCQT for Beat Tracking. In: SIMPÓSIO BRASILEIRO DE COMPUTAÇÃO MUSICAL (SBCM), 19. , 2025, Campinas/SP.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2025
.
p. 165-172.
DOI: https://doi.org/10.5753/sbcm.2025.13275.
