BOSSA: Um Dataset Multimodal de Popularidade Musical Regional no Spotify Brasileiro

  • Gabriel H. Silva Universidade Federal de Ouro Preto (UFOP)
  • Leonam V. O. L. Almeida Universidade Federal de Ouro Preto (UFOP)
  • Helen C. S. C. Lima Universidade Federal de Ouro Preto (UFOP)
  • Filipe A. S. Moura Universidade Federal de Minas Gerais (UFMG)

Resumo


Este artigo apresenta o BOSSA (Brazilian Observatory of Streaming and Song Analytics), um conjunto de dados multimodal que cobre as faixas mais populares no Spotify Charts em 17 cidades brasileiras de janeiro de 2022 a dezembro de 2024. O BOSSA integra dados de ranking semanal por cidade para 16.148 faixas, shallow audio features (API ReccoBeats), representações profundas de áudio (MERT-v1-330M) e letras transcritas automaticamente (Whisper v3). Três análises exploratórias validam o repositório de dadost: uma projeção UMAP revela agrupamentos geográficos não supervisionados, a similaridade de cosseno entre cidades reflete proximidade cultural (e.g., Campinas-São Paulo, Belém-São Paulo) e uma correlação de Spearman confirma a complementaridade entre as duas modalidades de áudio. O BOSSA é o primeiro conjunto de dados brasileiro a oferecer essa combinação de modalidades com granularidade por cidade.

Palavras-chave: Recuperação de Informação Musical, Spotify, Processamento de Linguagem Natural

Referências

Bello, P. and Garcia, D. (2021). Cultural divergence in popular music: The increasing diversity of music consumption on Spotify across countries. Humanities and Social Sciences Communications, 8(1):182.

Berendes, H.-U., Schwär, S., and Müller, M. (2024). Lyrics transcription in western classical music with Whisper: A case study on Schubert’s winterreise. In Proceedings of the 3rd Workshop on NLP for Music and Audio (NLP4MusA), pages 11–16.

Bertoni, A. and Lemos, R. (2021). Três datasets criados a partir de um banco de canções populares brasileiras de sucesso e não-sucesso de 2014 a 2019. In Anais do III Dataset Showcase Workshop, pages 11–20, Porto Alegre, RS, Brasil. SBC.

Li, Y., Yuan, R., Zhang, G., Ma, Y., Chen, X., Yin, H., Xiao, C., Lin, C., Ragni, A., Benetos, E., Gyenge, N., Dannenberg, R., Liu, R., Chen, W., Xia, G., Shi, Y., Huang, W., Wang, Z., Guo, Y., and Fu, J. (2024). {MERT}: Acoustic music understanding model with large-scale self-supervised training. In The Twelfth International Conference on Learning Representations.

McInnes, L., Healy, J., and Melville, J. (2020). Umap: Uniform manifold approximation and projection for dimension reduction.

Moura, F. A. S. et al. (2024). Characterization of the Brazilian musical landscape: A study of regional preferences based on the Spotify charts. In Proceedings of the 30th Brazilian Symposium on Multimedia and the Web (WebMedia), pages 80–88. SBC.

Moura, F. A. S., Ferreira, C. H. G., and Lima, H. C. S. C. (2025). Beyond acoustic features: A data-driven analysis of music consumption in Brazil. Journal on Interactive Systems, 16(1).

Oliveira, G., Barbosa, G., Melo, B., Silva, M., Seufitelli, D., and Moro, M. (2021). Muhsic: An open dataset with temporal musical success information. In Anais do III Dataset Showcase Workshop, pages 65–76, Porto Alegre, RS, Brasil. SBC.

Oliveira, G., Vassio, L., Silva, A., and Moro, M. (2025). Modeling music popularity as an epidemic: insights from the brazilian market. In Anais do XIV Brazilian Workshop on Social Network Analysis and Mining, pages 79–92, Porto Alegre, RS, Brasil. SBC.

Policy, M. T. (2025). AI implications of spotify’s updated terms of use: Your data is their new oil.

Radford, A., Kim, J. W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I. (2023). Robust speech recognition via large-scale weak supervision. In Proceedings of the 40th International Conference on Machine Learning (ICML), pages 28492–28518.

Rompolas, G., Smpoukis, A., Kafeza, E., and Makris, C. (2024). Predicting song popularity through machine learning and sentiment analysis on social networks. In Artificial Intelligence Applications and Innovations (AIAI 2024), volume 715 of IFIP Advances in Information and Communication Technology. Springer.

Sebastian, N., Mayer, F., et al. (2024). Beyond beats: A recipe to song popularity? a machine learning approach.

Seufitelli, D., Oliveira, G., Silva, M., and Moro, M. (2023a). Mgd+: An enhanced music genre dataset with success-based networks. In Anais do V Dataset Showcase Workshop, pages 36–47, Porto Alegre, RS, Brasil. SBC.

Seufitelli, D. B., Oliveira, G. P., Silva, M. O., Scofield, C., and Moro, M. M. (2023b). Hit song science: A comprehensive survey and research directions. Journal of New Music Research, 52(1):41–72.

Venna, J. and Kaski, S. (2006). Local multidimensional scaling. Neural Networks, 19(6):889–899. Advances in Self Organising Maps - WSOM’05.

Way, L. D., Garcia, D., Gathright, J., et al. (2020). Local trends in global music streaming. In Proceedings of the International AAAI Conference on Web and Social Media.

Zhuo, L., Yuan, R., Pan, J., Ma, Y., Li, Y., Zhang, G., et al. (2023). LyricWhiz: Robust multilingual zero-shot lyrics transcription by whispering to ChatGPT. In Proceedings of the 24th International Society for Music Information Retrieval Conference (ISMIR), pages 343–351.
Publicado
08/09/2026
SILVA, Gabriel H.; ALMEIDA, Leonam V. O. L.; LIMA, Helen C. S. C.; MOURA, Filipe A. S.. BOSSA: Um Dataset Multimodal de Popularidade Musical Regional no Spotify Brasileiro. In: DATASET SHOWCASE WORKSHOP (DSW), 8. , 2026, São Carlos/SP. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 36-47. DOI: https://doi.org/10.5753/dsw.2026.249463.