conTorchionist: A flexible nomadic library for exploring machine listening/learning in multiple platforms, languages, and time-contexts

  • José Henrique Padovani UFMG
  • Vinícius Cesar de Oliveira UNICAMP

Resumo


The concept of ‘machine listening’ has profoundly shaped the development of interactive music systems, evolving from symbolic MIDI processing to a wide range of audio analysis techniques. While the recent proliferation of neural networks has greatly expanded these capabilities, it has also introduced complex processes that often function as opaque ‘black boxes’, limiting creative exploration. To address this, we present ‘conTorchionist’: a flexible, nomadic library designed to foster a more transparent approach to machine listening and learning. Built upon libtorch/PyTorch, it features a single shared core with dedicated interfaces for diverse environments, including Pure Data, Max, SuperCollider, and Python. This architecture unifies real-time and non-real-time workflows, leveraging libtorch’s strengths in GPU-accelerated audio analysis, signal processing, and neural network inference. Ultimately, conTorchionist provides an adaptable toolset that empowers researchers and artists to seamlessly bridge different platforms, languages, and time-contexts in their creative and investigative work.

Referências

Robert Rowe. Machine listening and composing with cypher. Computer Music Journal, pages 43–63, 1992.

Robert Rowe, Brad Garton, Peter Desain, Henkjan Honing, Roger Dannenberg, Dean Jacobs, Stephen Travis Pope, Miller Puckette, Cort Lippe, Zack Settel, and George Lewis. Editor’s Notes: Putting Max in Perspective. Computer Music Journal, 17(2):3–11, 1993.

George E. Lewis. Interacting with latter-day musical automata. Contemporary Music Review, 18(3):99–112, January 1999.

George E. Lewis. Too Many Notes: Computers, Complexity and Culture in ”Voyager”. Leonardo Music Journal, 10:33–39, 2000.

Nicholas M. Collins. Towards Autonomous Agents for Live Computer Music: Realtime Machine Listening and Interactive Music Systems. PhD thesis, Faculty of Music, University of Cambridge, Cambridge, 2006.

Nick Collins. Musical robots and listening machines. In Nick Collins and Julio Escrivan, editors, The Cambridge Companion to Electronic Music, pages 171–184. Cambridge University Press, 1ª edição edition, December 2007.

Nick Collins. Machine Listening in SuperCollider, pages 439–461. MIT Press, Cambridge, Mass, 2011.

Tristan Jehan. Creating Music by Listening. PhD thesis, Massachusetts Institute of Technology, School of Architecture and Planning . . . , 2005.

J. Stephen Downie. Music information retrieval. Annual review of information science and technology, 37(1):295–340, 2003.

Geoffroy Peeters. A large set of audio features for sound description (similarity and classification) in the CUIDADO project. pages 1–25, 2004.

Bozena Kostek. Perception-Based Data Processing in Acoustics: Applications to Music Information Retrieval and Psychophysiology of Hearing. Number v. 3 in Studies in Computational Intelligence. Springer, Berlin ; New York, 2005.

Anssi Klapuri and Manuel Davy. Signal Processing Methods for Music Transcription. Springer, New York, 2006.

Shipra J. Arora and Rishi Pal Singh. Automatic speech recognition: A review. International Journal of Computer Applications, 60(9), 2012.

Guy J Brown and DeLiang Wang. Computational Auditory Scene Analysis: Principles, Algorithms, and Applications. 2015.

Dong Yu and Li Deng. Automatic Speech Recognition: A Deep Learning Approach. Signals and Communication Technology. Springer London, London, 2015.

Miller Puckette. Pure data: Recent progress. In Proceedings of the Third Intercollege Computer Music Festival, pages 1–4. Citeseer, 1997.

Supercollider. SuperCollider Documentation. [link], 2025.

William Brent. Cepstral analysis tools for percussive timbre identification. In Proceedings of the 3rd International Pure Data Convention, pages 1–7, São Paulo, 2009.

Nick Collins. SCMIR: A SuperCollider Music Information Retrieval Library. International Computer Music Conference Proceedings, 2011, 2011.

Pierre Alexandre Tremblay, Owen Green, Gerard Roma, and Alexander Harker. From collections to corpora: Exploring sounds through fluid decomposition. In International Computer Music Conference and New York City Electroacoustic Music Festival, pages 223–228. International Computer Music Association, 2019.

Pierre Alexandre Tremblay, Gerard Roma, and Owen Green. Enabling programmatic data mining as musicking: The fluid corpus manipulation toolkit. Computer Music Journal, 45(2):9–23, 2021.

Ted Moore, James Bradbury, Pierre Alexandre Tremblay, and Owen Green. Making Machine Learning Musical: Reflections on a Year of Teaching FluCoMa. Journal SEAMUS, 32, 2021.

Pierre Alexandre Tremblay, Owen Green, Gerard Roma, James Bradbury, Ted Moore, Jacob Hart, and Alex Harker. Fluid corpus manipulation toolbox. 2022.

James Bradbury. Harnessing Content-Aware Programs for Computer-Aided Composition in a Studio-Based Workflow. PhD thesis, University of Huddersfield, 2021.

Brian McFee, Colin Raffel, Dawen Liang, Daniel PW Ellis, Matt McVicar, Eric Battenberg, and Oriol Nieto. Librosa: Audio and music signal analysis in python. SciPy, 2015:18–24, 2015.

Dmitry Bogdanov, N Wack, Emilia Gómez, Sankalp Gulati, Perfecto Herrera, Oscar Mayor, Gerard Roma, Justin Salamon, Jose Zapata, and Xavier Serra. ESSENTIA: An Audio Analysis Library for Music Information Retrieval. November 2013.

Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. PyTorch: An imperative style, high-performance deep learning library. In Proceedings of the 33rd International Conference on Neural Information Processing Systems, number 721, pages 8026–8037. Curran Associates Inc., Red Hook, NY, USA, December 2019.

Jeff Hwang, Moto Hira, Caroline Chen, Xiaohui Zhang, Zhaoheng Ni, Guangzhi Sun, Pingchuan Ma, Ruizhe Huang, Vineel Pratap, and Yuekai Zhang. TorchAudio 2.1: Advancing speech recognition, self-supervised learning, and audio processing components for PyTorch. In 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), pages 1–9. IEEE, 2023.

Yao-Yuan Yang, Moto Hira, Zhaoheng Ni, Artyom Astafurov, Caroline Chen, Christian Puhrsch, David Pollack, Dmitriy Genzel, Donny Greenberg, and Edward Z. Yang. Torchaudio: Building blocks for audio and speech processing. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 6982–6986. IEEE, 2022.

Antoine Caillon and Philippe Esling. RAVE: A variational autoencoder for fast and high-quality neural audio synthesis, December 2021.

Antoine Caillon and Philippe Esling. Streamable Neural Audio Synthesis With Non-Causal Convolutions, April 2022.
Publicado
15/09/2025
PADOVANI, José Henrique; OLIVEIRA, Vinícius Cesar de. conTorchionist: A flexible nomadic library for exploring machine listening/learning in multiple platforms, languages, and time-contexts. In: SIMPÓSIO BRASILEIRO DE COMPUTAÇÃO MUSICAL (SBCM), 19. , 2025, Campinas/SP. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2025 . p. 36-43. DOI: https://doi.org/10.5753/sbcm.2025.14270.