Efficiency-First Bioacoustic Audio Event Detection Under Resource-Constrained Systems
Resumo
Passive acoustic monitoring enables wildlife biodiversity assessment and precision livestock management, yet accurate detectors remain difficult to deploy on resource-limited systems. We present a controlled evaluation of AST, SSAMBA, and MAMBASPEC on two IoT-relevant bioacoustic datasets (aSwine, AnuraSet). MAMBASPEC matches AST within 0.7 pp ROC-AUCw on AnuraSet with 80× fewer parameters and 1,617× fewer FLOPs, while SSAMBA achieves statistically indistinguishable ROC-AUCw at a 12.8× smaller footprint. Friedman tests (p<10−3) and Pareto analysis indicate that MAMBASPEC is the efficiency-first choice under resource constraints, whereas SSAMBA is the accuracy-oriented choice when modest extra latency is acceptable.
Referências
Demšar, J. (2006). Statistical comparisons of classifiers over multiple data sets. Journal of Machine Learning Research, 7:1–30.
Erol, M. H., Senocak, A., Feng, J., and Chung, J. S. (2024). Audio Mamba: Bidirectional State Space Model for Audio Representation Learning. IEEE Signal Processing Letters, 31:2975–2979.
Gong, Y., Chung, Y.-A., and Glass, J. (2021). AST: Audio Spectrogram Transformer. In Proc. Interspeech, pages 571–575.
Gu, A. and Dao, T. (2023). Mamba: Linear-Time Sequence Modeling with Selective State Spaces. arXiv preprint arXiv:2312.00752.
Hagiwara, M., Hoffman, B., Liu, J.-Y., Cusimano, M., Effenberger, F., and Zacarian, K. (2023). BEANS: The Benchmark of Animal Sounds. In ICASSP 2023 – 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1–5.
Liang, J., Nolasco, I., Ghani, B., Phan, H., Benetos, E., and Stowell, D. (2024). Mind the domain gap: A systematic analysis on bioacoustic sound event detection. In 2024 32nd European Signal Processing Conference (EUSIPCO), pages 1257–1261.
Lima, J., Salles, R., Escobar, L., Géa, C., Fernandes, P. A., Pacitti, E., Porto, F., Coutinho, R., and Ogasawara, E. (2022). Towards a cloud-based framework for online and integrated event detection. In Simpósio Brasileiro de Banco de Dados (SBBD), pages 199–202. SBC.
Lin, J., Chen, W.-M., Lin, Y., cohn, j., Gan, C., and Han, S. (2020). MCUNet: Tiny deep learning on IoT devices. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H., editors, Advances in Neural Information Processing Systems, volume 33, pages 11711–11722. Curran Associates, Inc.
Rauch, L., Schwinger, R., Wirth, M., Heinrich, R., Huseljic, D., Herde, M., Lange, J., Kahl, S., Sick, B., Tomforde, S., and Scholz, C. (2025). BirdSet: A Large-Scale Dataset for Audio Classification in Avian Bioacoustics. In Yue, Y., Garg, A., Peng, N., Sha, F., and Yu, R., editors, International Conference on Learning Representations, pages 29482–29520.
Shams, S., Dindar, S. S., Jiang, X., and Mesgarani, N. (2024). SSAMBA: Self-Supervised Audio Representation Learning with Mamba State Space Model. In Proc. IEEE Spoken Language Technology Workshop (SLT), pages 1053–1059.
Souza, A. M., Kobayashi, L. L., Tassoni, L. A., Garbossa, C. A. P., Ventura, R. V., and Machado de Sousa, E. P. (2025). Deep learning solutions for audio event detection in a swine barn using environmental audio and weak labels. Applied Intelligence, 55(7).
Stowell, D. (2022). Computational bioacoustics with deep learning: a review and roadmap. PeerJ, 10:e13152.
Tang, C. and Baskiyar, S. (2025). State space models for bioacoustics: A comparative evaluation with Transformers. arXiv:2512.03563.
Vuilliomenet, A., Martínez Balvanera, S., Mac Aodha, O., Jones, K. E., and Wilson, D. (2026). acoupi: An open-source python framework for deploying bioacoustic AI models on edge devices. Methods in Ecology and Evolution, 17(1):67–76.
