Overview of Neural Audio Synthesis Algorithms
Neural audio synthesis has seen great progress during the last decade, and currently underpins commercial applications such as text-to-speech synthesis, music generation or neural timbre transfer.
This talk will review some of the most common algorithms for neural audio synthesis, including autoregressive models, generative adversarial networks, variational autoencoders, differential DSP and diffusion models, with an emphasis on practical issues such as audio representation, capacity, computational cost, causality and real-time inference.
Several implementations will be demonstrated.
Gerard Roma
Gerard Roma is a researcher, musician, educator and software developer with an interest in creative applications of digital signal processing and machine learning. He is currently a lecturer in audio engineering at University of West London, where he leads the MSc courses on digital audio engineering and artificial intelligence for sound and music.