Toward Expressive Timbre Modeling: Convergence of DDSP and Neural Voice Synthesis
Resumo
This article investigates the convergence between Differentiable Digital Signal Processing (DDSP) and neural singing voice synthesis approaches, focusing on the modeling and control of vocal timbre. By integrating explainable acoustic structures with deep learning, DDSP-based models offer interpretable, efficient, and expressive synthesis, overcoming limitations of purely neural architectures such as WaveNet or Tacotron. Analyzing models like the Neural Parametric Singing Synthesizer (NPSS), the article proposes a conceptual and technical framework that highlights recent advances in the explicit separation of pitch, timbre, and prosody. The study discusses implementations using differentiable LPC and FIR filters, hybrid vocoders, harmonic + noise synthesis, and strategies for timbre transfer and singing voice conversion (SVC). It argues that hybridization between DSP and neural networks marks a new methodological frontier for vocal synthesis — one that is more transparent, controllable, and suitable for expressive, pedagogical, and interactive applications.
Referências
Ben Hayes, Jordie Shier, György Fazekas, Andrew McPherson, and Charalampos Saitis. A review of differentiable digital signal processing for music speech synthesis, 2023. Available at: [link]. Accessed on: 27 May 2025.
Aäron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu. WaveNet: A generative model for raw audio. In 9th ISCA Workshop on Speech Synthesis Workshop (SSW 9), page 125, 2016. Available at: [link]. Accessed on: 27 May 2025.
Yuxuan Wang, R.J. Skerry-Ryan, Daisy Stanton, Yonghui Wu, Ron J. Weiss, Navdeep Jaitly, Zongheng Yang, Ying Xiao, Zhifeng Chen, Samy Bengio, Quoc Le, Yannis Agiomyrgiannakis, Rob Clark, and Rif A. Saurous. Tacotron: Towards end-to-end speech synthesis. In Interspeech 2017, pages 4006–4010, 2017. Available at: [link] 2017/wang17n_interspeech.html. Accessed on: 27 May 2025.
Jesse Engel, Kumar Krishna Agrawal, Shuo Chen, Ishaan Gulrajani, Chris Donahue, and Adam Roberts. GAN-Synth: Adversarial neural audio synthesis. In International Conference on Learning Representations, 2019. Available at: [link]. Accessed on: 27 May 2025.
Soroush Mehri, Kundan Kumar, Ishaan Gulrajani, Rithesh Kumar, Shubham Jain, Jose Sotelo, Aaron Courville, and Yoshua Bengio. SampleRNN: An unconditional end-to-end neural audio generation model. In International Conference on Learning Representations, 2017. Available at: [link]. Accessed on: 27 May 2025.
Merlijn Blaauw and Jordi Bonada. A neural parametric singing synthesizer modeling timbre and expression from natural songs. Applied Sciences, 7(12):1313, 2017. Available at: [link]. Accessed on: 27 May 2025.
Ye-Xin Lu, Yang Ai, and Zhen-Hua Ling. Source-filter-based generative adversarial neural vocoder for high fidelity speech synthesis. In Ling Zhenhua, Gao Jianqing, Yu Kai, and Jia Jia, editors, Man-Machine Speech Communication, pages 68–80, Singapore, 2023. Springer Nature Singapore. Available at: [link]. Accessed on: 27 May 2025.
Hyeong-Seok Choi, Jinhyeok Yang, Juheon Lee, and Hyeongju Kim. NANSY++: Unified voice synthesis with neural analysis and synthesis. In International Conference on Learning Representations, 2023. Available at: [link]. Accessed on: 27 May 2025.
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae. HiFi-GAN: generative adversarial networks for efficient and high fidelity speech synthesis. In Proceedings of the 34th International Conference on Neural Information Processing Systems, number 1428 in NIPS’20, pages 17022 – 17033, Red Hook, NY, USA, 2020. Curran Associates Inc. Available at: [link]. Accessed on: 27 May 2025.
Eliya Nachmani and Lior Wolf. Unsupervised singing voice conversion. In Interspeech 2019, pages 2583–2587, 2019. Available at: [link]. Accessed on: 27 May 2025.
Shahan Nercessian. End-to-end zero-shot voice conversion using a DDSP vocoder. In 2021 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), pages 1–5, 2021. Available at: [link]. Accessed on: 27 May 2025.
Jinglin Liu, Chengxi Li, Yi Ren, Feiyang Chen, and Zhou Zhao. DiffSinger: Singing voice synthesis via shallow diffusion mechanism. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 36, pages 11020–11028, 2022. Available at: [link]. Accessed on: 27 May 2025.
Xin Wang, Shinji Takaki, and Junichi Yamagishi. Neural source-filter waveform models for statistical parametric speech synthesis. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 28:402–415, 2020. Available at: [link]. Accessed on: 27 May 2025.
Da-Yi Wu, Wen-Yi Hsiao, Fu-Rong Yang, Oscar D Friedman, Warren Jackson, Scott Bruzenak, Yi-Wen Liu, and Yi-Hsuan Yang. DDSP-based singing vocoders: A new subtractive-based synthesizer and a comprehensive evaluation. In Proceedings of the 23rd International Society for Music Information Retrieval Conference, pages 76–83. ISMIR, 2022. Available at: [link]. Accessed on: 27 May 2025.
Nercessian Shahan. Differentiable world synthesizer-based neural vocoder with application to end-to-end audio style transfer. Journal of the Audio Engineering Society, (10661), May 2023. Available at: [link]. Accessed on: 27 May 2025.
Ben Hayes, Charalampos Saitis, and Gyorgy Fazekas. Neural waveshaping synthesis. In Proceedings of the 22nd International Society for Music Information Retrieval Conference, pages 254–261. ISMIR, 2021. Available at: [link]. Accessed on: 27 May 2025.
Joseph Turian and Max Henry. I’m sorry for your loss: Spectrally-based audio distances are bad at pitch. In NeurIPS 2020 Workshop ICBINB, 2020. Available at: [link]. Accessed on: 27 May 2025.
Henrique Vaz. Estudo de modelagem: Naqshe khani. IV Encontro Internacional de Música Eletroacústica da SBME–UNESPAR, 2024. Available at: [link]. Accessed on: 31 Aug 2025.
Henrique Vaz. Mimosa pudica. Festival Escuta Aqui!, SBME, 2024. Available at: [link]. Accessed on: 31 Aug 2025.
Henrique Vaz. De silenti natura. Seleção de Músicas Experimentais, Académie Charles Cros, 2025. Available at: [link]. Accessed on: 31 Aug 2025.
Henrique Vaz. De silenti natura. Editado por Mappa Editions, 2024. Available at: [link]. Accessed on: 31 Aug 2025.
Henrique Vaz. 8 frames. Festival ARTELLIGENT, Caixa Cultural Salvador, 2024. Available at: [link]. Accessed on: 31 Aug 2025.
