ARTEUS: an Algorithmic Rating Tool for Educating Untrained Singers

  • Arthur Nicholas dos Santos Unicamp
  • Bruno Sanches Masiero Unicamp

Resumo


This study presents the development of a prototype software system for evaluating karaoke performances using a hybrid scoring methodology that integrates objective and non-intrusive subjective metrics. The proposed framework employs scientifically grounded techniques to assess vocal dynamics, timing, pitch, and timbral similarity, complemented by data-driven emotion recognition to capture the perceptual aspects of the performance. A key feature of the system is a composite scoring mechanism with customizable weightings, which enables flexible adaptation to user preferences in pedagogical contexts or entertainment scenarios. By leveraging artificial intelligence models, the system introduces a novel paradigm in performance evaluation, offering a structured yet extensible approach to quantifying aspects of artistic expression that are traditionally considered subjective. Designed with modularity and future scalability in mind, the framework can be expanded to include user-specific evaluation profiles with potential applications in both professional and recreational settings.

Referências

Laurence Green. The Serious Business of Song: Karaoke as Discipline and Industry in Japan, page 231–243. Amsterdam University Press, 2022.

X. Zhou and F. Tarocco. Karaoke: The Global Phenomenon. Reaktion Books, 2007.

Supreme Court of the Philippines. Roberto l. del rosario vs. court of appeals and janito corporation. [link], March 1996. G.R. No. 115106, March 15, 1996.

Heikki Ruismäki, Antti Juvonen, and Kimmo Lehtonen. Karaoke–the chance to be a star. The European Journal of Social & Behavioural Sciences, 7:1222–1233, 2013.

Karaoke World Championships. Kwc – world’s biggest global music contest. [link], 2025. Accessed: 2025-05-30.

Karaoke World Championships USA. Kwcusa official judging criteria – solo division. [link], 2025. Accessed: 2025-05-30.

United States Karaoke Association. Uska scoring system. [link], 2025. Accessed: 2025-05-30.

UltraStar Deluxe Team. Ultrastar deluxe – free open source karaoke game. [link], 2025. Accessed: 2025-05-30.

Smule Inc. Smule – social singing app. [link], 2025. Accessed: 2025-05-30.

Huan Zhang, Yi Jiang, Tao Jiang, and Peng Hu. Learn by referencing: Towards deep metric learning for singing assessment. In International Society for Music Information Retrieval Conference, 2021.

Wei-Ho Tsai and Hsin-Chieh Lee. Automatic evaluation of karaoke singing based on pitch, volume, and rhythm features. IEEE transactions on audio, speech, and language processing, 20(4):1233–1243, 2011.

Oscar Mayor, Jordi Bonada, and Alex Loscos. Performance analysis and scoring of the singing voice. In Proc. 35th AES Intl. Conf., London, UK, pages 1–7, 2009.

Frank Filipanits. Design and implementation of an auralization system with a spectrum-based temporal processing optimization. Master’s Research Project, University of Miami May, 1994.

Ennio Cruz da Costa. Acústica técnica. Bluscher, 2003.

Jont B. Allen and David A. Berkley. Image method for efficiently simulating small-room acoustics. The Journal of the Acoustical Society of America, 65:943–950, 4 1979.

Alan V. Oppenheim, Alan S. Willsky, and S. Hamid Nawab. Signals & Systems (2nd Ed.). Prentice-Hall, Inc., USA, 1996.

Franz Zotter and Matthias Frank. Ambisonics, volume 19. Springer International Publishing, 2019.

Richard F. Lyon. Human and Machine Hearing. Cambridge University Press, 5 2017.

A. Jansson, E. Humphrey, N. Montecchio, R. Bittner, A. Kumar, and T. Weyde. Singing voice separation with deep u-net convolutional networks. Unpublished, October 2017.

Romain Hennequin, Anis Khlif, Felix Voituret, and Manuel Moussallam. Spleeter: A fast and stateof-the art music source separation tool with pre-trained models. Late-Breaking/Demo ISMIR, 2019.

Renato Panda, Ricardo Malheiro, and Rui Pedro Paiva. Audio features for music emotion recognition: A survey. IEEE Transactions on Affective Computing, 14(1):68–88, 2023.

Pedro Benevenuto Valadares, Karen Gissell Rosero Jácome, Arthur Nicholas dos Santos, and Bruno Sanches Masiero. Song emotion recognition: a performance comparison between audio features and artificial neural networks. Revista Eletrônica de Iniciação Científica em Computação, 20(4), dez. 2022.

Steven R. Livingstone and Frank A. Russo. The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english. PLoS ONE, 13(5), 2018.

Andrew J. R. Simpson. Probabilistic binary-mask cocktail-party source separation in a convolutional deep neural network, 2015.
Publicado
15/09/2025
SANTOS, Arthur Nicholas dos; MASIERO, Bruno Sanches. ARTEUS: an Algorithmic Rating Tool for Educating Untrained Singers. In: SIMPÓSIO BRASILEIRO DE COMPUTAÇÃO MUSICAL (SBCM), 19. , 2025, Campinas/SP. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2025 . p. 6-13. DOI: https://doi.org/10.5753/sbcm.2025.13097.