From-Scratch Audio Classification with Multi-Branch Feature Fusion: Comparing CNN, Transformer, and Mamba
Resumo
From-scratch audio classification on weakly labeled data is essential for resource-constrained and domain-specific settings where large pretrained models are unavailable or prohibitively costly. We adopt multi-branch feature fusion as the core approach and compare CNN, Inception, Transformer, and Mamba architectures on the AudioSet balanced training split. Multi-branch fusion (per-feature encoders, merged before the head) improves all four architectures by 17–33% mAP. Five-fold CV shows Inception significantly outperforms the Transformer; Transformer and Mamba reach comparable multi-branch performance (p = 0.19), with preliminary results suggesting feature fusion may be at least as important as sequence model choice in this regime.
Palavras-chave:
Audio classification, Weakly labeled learning, Multi-branch feature fusion
Referências
Chen, S., Wu, Y., Wang, C., Liu, S., Tompkins, D., Chen, Z., and Wei, F. (2022). Beats: Audio pre-training with acoustic tokenizers. arXiv preprint arXiv:2212.09058.
Gemmeke, J. F., Ellis, D. P. W., Freedman, D., Jansen, A., Lawrence, W., Moore, R. C., Plakal, M., and Ritter, M. (2017). AudioSet: An ontology and human-labeled dataset for audio events. In Proc. IEEE ICASSP, pages 776–780.
Gong, Y., Chung, Y.-A., and Glass, J. (2021). AST: Audio spectrogram transformer. In Proc. Interspeech, pages 571–575.
Gu, A. and Dao, T. (2023). Mamba: Linear-time sequence modeling with selective state spaces. arXiv:2312.00752.
Kong, Q., Cao, Y., Iqbal, T., Wang, Y., Wang, W., and Plumbley, M. D. (2020). PANNs: Large-scale pretrained audio neural networks for audio pattern recognition. IEEE/ACM Trans. Audio, Speech, Lang. Process., 28:2880–2894.
Mohmmad, S. and Sanampudi, S. K. (2024). Exploring current research trends in sound event detection: a systematic literature review. Multimedia Tools and Applications, 83:84699–84741.
Moreira Souza, A., Moreira, G. A., and Pulcinelli, L. E. G. (2025). A Comparative Analysis of Denoising Methods for Deep Learning-Based Audio Event Detection in Noisy Agricultural Environments. In Anais do XL Simpósio Brasileiro de Banco de Dados (SBBD 2025), pages 942–948, Brasil. Sociedade Brasileira de Computação - SBC.
Mu, D., Zhang, Z., and Yue, H. (2024). MFF-EINV2: Multi-scale feature fusion across spectral-spatial-temporal domains for sound event localization and detection. In Proc. Interspeech, pages 92–96.
Schmid, F., Koutini, K., and Widmer, G. (2023). Efficient large-scale audio tagging via transformer-to-CNN knowledge distillation. In Proc. IEEE ICASSP, pages 1–5.
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A. (2015). Going deeper with convolutions. In Proc. IEEE CVPR, pages 1–9.
Turab, M., Kumar, T., Bendechache, M., and Saber, T. (2022). Investigating multi-feature selection and ensembling for audio classification. arXiv:2206.07511.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017). Attention is all you need. In Advances in Neural Information Processing Systems, volume 30, pages 5998–6008.
Yadav, S. and Tan, Z.-H. (2024). Audio mamba: Selective state spaces for self-supervised audio representations. In Proc. Interspeech, pages 552–556.
Zhang, Y., Huang, D., and Togneri, R. (2025). Pseudo strong labels from frame-level predictions for weakly supervised sound event detection. arXiv:2501.03740.
Gemmeke, J. F., Ellis, D. P. W., Freedman, D., Jansen, A., Lawrence, W., Moore, R. C., Plakal, M., and Ritter, M. (2017). AudioSet: An ontology and human-labeled dataset for audio events. In Proc. IEEE ICASSP, pages 776–780.
Gong, Y., Chung, Y.-A., and Glass, J. (2021). AST: Audio spectrogram transformer. In Proc. Interspeech, pages 571–575.
Gu, A. and Dao, T. (2023). Mamba: Linear-time sequence modeling with selective state spaces. arXiv:2312.00752.
Kong, Q., Cao, Y., Iqbal, T., Wang, Y., Wang, W., and Plumbley, M. D. (2020). PANNs: Large-scale pretrained audio neural networks for audio pattern recognition. IEEE/ACM Trans. Audio, Speech, Lang. Process., 28:2880–2894.
Mohmmad, S. and Sanampudi, S. K. (2024). Exploring current research trends in sound event detection: a systematic literature review. Multimedia Tools and Applications, 83:84699–84741.
Moreira Souza, A., Moreira, G. A., and Pulcinelli, L. E. G. (2025). A Comparative Analysis of Denoising Methods for Deep Learning-Based Audio Event Detection in Noisy Agricultural Environments. In Anais do XL Simpósio Brasileiro de Banco de Dados (SBBD 2025), pages 942–948, Brasil. Sociedade Brasileira de Computação - SBC.
Mu, D., Zhang, Z., and Yue, H. (2024). MFF-EINV2: Multi-scale feature fusion across spectral-spatial-temporal domains for sound event localization and detection. In Proc. Interspeech, pages 92–96.
Schmid, F., Koutini, K., and Widmer, G. (2023). Efficient large-scale audio tagging via transformer-to-CNN knowledge distillation. In Proc. IEEE ICASSP, pages 1–5.
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A. (2015). Going deeper with convolutions. In Proc. IEEE CVPR, pages 1–9.
Turab, M., Kumar, T., Bendechache, M., and Saber, T. (2022). Investigating multi-feature selection and ensembling for audio classification. arXiv:2206.07511.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017). Attention is all you need. In Advances in Neural Information Processing Systems, volume 30, pages 5998–6008.
Yadav, S. and Tan, Z.-H. (2024). Audio mamba: Selective state spaces for self-supervised audio representations. In Proc. Interspeech, pages 552–556.
Zhang, Y., Huang, D., and Togneri, R. (2025). Pseudo strong labels from frame-level predictions for weakly supervised sound event detection. arXiv:2501.03740.
Publicado
08/09/2026
Como Citar
GIACOMINI, Anderson H.; MAGALHÃES, André M. S.; SOUSA, Elaine P. M..
From-Scratch Audio Classification with Multi-Branch Feature Fusion: Comparing CNN, Transformer, and Mamba. In: SIMPÓSIO BRASILEIRO DE BANCO DE DADOS (SBBD), 41. , 2026, São Carlos/SP.
Anais [...].
Porto Alegre: Sociedade Brasileira de Computação,
2026
.
p. 994-1000.
ISSN 2763-8979.
DOI: https://doi.org/10.5753/sbbd.2026.249650.
