PATRICIA: proof-of-concept implementation and validation of a real-time singing synthesizer

  • Leonardo A. Z. Brum UFBA
  • Eduardo A. L. Meneses Metalab – Société des arts technologiques
  • Edward D. Moreno UFS

Resumo


This paper describes the implementation and validation process of PATRICIA, a proof-of-concept prototype of a system that performs real-time singing voice synthesis (SVS) for the Brazilian Portuguese language. A technological mapping was conducted to study the latest industrial developments in real-time SVS and give directions for PATRICIA design and implementation. The implemented architecture is described, and for the validation process, a series of video recordings of the system working was made as a musical demonstration. This demonstration was subjected to an evaluation by musical educators. A performance analysis with two CPU usage indicators was also performed on two different devices. The results of the validation are discussed and future enhancements are pointed out to overcome the current limitations of the system.

Referências

Anastasia Georgaki. Virtual voices on hands: Prominent applications on the synthesis and control of the singing voice. In Journées d’informatique musicale, Paris, France, 2004.

Rong Li. Intelligent analysis of music education singing skills based on music waveform feature extraction. Mobile Information Systems, 2022(1):9747342, 2022.

Kimi Kärki. Vocaloid liveness? hatsune miku and the live production of japanese virtual idol concerts. In Researching Live Music, pages 127–140. Focal Press, 2021.

Jessica Tsun Lem Hui. Reconfiguring voice in the end: Virtuosity, technological affordance and the reversibility of hatsune miku in the intermundane. Cambridge Opera Journal, 34(3):364–379, 2022.

Hideki Kenmochi. Singing synthesis as a new musical instrument. In 2012 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 5385–5388, 2012.

Ji-Sang Hwang, Sang-Hoon Lee, and Seong-Whan Lee. Hiddensinger: High-quality singing voice synthesis via neural audio codec and latent diffusion models. Neural Networks, 181:106762, 2025.

Leonardo A. Z. Brum and Edward D. Moreno. State of art of real-time singing voice synthesis. In Anais do XVII Simpósio Brasileiro de Computação Musical, pages 50–57, Porto Alegre, RS, Brasil, 2019. SBC.

Shota Kagami, Keizo Hamano, Kazuki Kashiwaze, and Kazuhiko Yamamoto. Development of realtime japanese vocal keyboard. Information Processing Society of Japan INTERACTION, pages 837–842, 2012.

Hideki Kenmochi and Hayato Ohshita. VOCALOID - commercial singing synthesizer based on sample concatenation. In Proc. Interspeech 2007, pages 4009–4010, 2007.

Leonardo A. Z. Brum and Edward D. Moreno. Challenges and perspectives on real-time singing voice synthesis. Revista de Informática Teórica e Aplicada, 27(4).

Leonardo A. Z. Brum, Eduardo A. L. Meneses, and Edward D. Moreno. Patricia: a real-time singing synthesizer prototype for the brazilian portuguese language. In Proceedings of the International Computer Music Conference 2023, page 176 – 180, 2023.

James McCartney. Supercollider: a new real time synthesis language. In Proc. International Computer Music Conference (ICMC’96), pages 257–258, 1996.

Linda Sherrell. Evolutionary prototyping. Encyclopedia of Sciences and Religions, pages 803–803, 2013.

T. Dutoit, V. Pagel, N. Pierret, F. Bataille, and O. van der Vrecken. The mbrola project: towards a set of high quality speech synthesizers free of use for non commercial purposes. In Proceeding of Fourth International Conference on Spoken Language Processing. ICSLP ’96, volume 3, pages 1393–1396 vol.3, 1996.

Haruo Kubozono. The mora and syllable structure in japanese: Evidence from speech errors. Language and Speech, 32(3):249–278, 1989.

Kazuki Kashiwase. An over-the-shoulder keyboard that extends the potential for vocaloid performance. Yamaha Corporation, 2017. Accessed: 2023-01-29.

Manuel Bandeira. A versificação em língua portuguêsa. Enciclopédia Delta Larousse, pages 3239–3249, 1960.

Thaı̈s Cristófaro Silva and Camila Leite. Padrões sonoros emergentes:(oclusiva alveolar+ sibilante) no português brasileiro. Caderno de Letras, (24):15–36, 2015.

John C Wells et al. Sampa computer readable phonetic alphabet. Handbook of standards and resources for spoken language systems, 4:684–732, 1997.

Christian Benoı̂t, Martine Grice, and Valérie Hazan. The sus test: A method for the assessment of text-to-speech synthesis intelligibility using semantically unpredictable sentences. Speech Communication, 18(4):381–392, 1996.

Katherine Morton. Naturalness in synthetic speech. Proceedings of the Institute of Acoustics, 12(Part 10):125–132, 1990.

Ankur Joshi, Saket Kale, Satish Chandel, and D Kumar Pal. Likert scale: Explored and explained. British journal of applied science & technology, 7(4):396, 2015.

Robert C Streijl, Stefan Winkler, and David S Hands. Mean opinion score (mos) revisited: methods and applications, limitations and alternatives. Multimedia Systems, 22(2):213–227, 2016.

Peng Bai, Meizhen Zheng, and Xiaodong Shi. A survey of singing voice synthesis. In Kun Qian, Xin Wang, Qinglin Meng, and Mingzhi Chen, editors, Proceedings of the 10th Conference on Sound and Music Technology, pages 19–30, Singapore, 2025. Springer Nature Singapore.
Publicado
15/09/2025
BRUM, Leonardo A. Z.; MENESES, Eduardo A. L.; MORENO, Edward D.. PATRICIA: proof-of-concept implementation and validation of a real-time singing synthesizer. In: SIMPÓSIO BRASILEIRO DE COMPUTAÇÃO MUSICAL (SBCM), 19. , 2025, Campinas/SP. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2025 . p. 144-151. DOI: https://doi.org/10.5753/sbcm.2025.13208.