Cloning a musical instrument from 16 seconds of audio
Cloning a musical instrument from 16 seconds of audio
In 2020, Magenta released DDSP [1], a machine learning algorithm / python library which made it possible to generate good sounding instrument synthesizers from about 6-10 minutes of data. While working with DDSP for a project, we realised how it was actually quite hard to find 6-10 minute of clean recordings of monophonic instruments. In this project, we have combined the DDSP architecture with a domain adaptation technique from speech synthesis [2]. This domain adaptation technique works by pre-training our model on many different recordings from the Solos dataset [3] first and then fine-tuning parts of the model to the new recording. This allows us to produce decent sounding instrument synthesisers from as little as 16 seconds of target audio instead of 6-10 minutes. [1] https://arxiv.org/abs/2001.04643 [2] https://arxiv.org/abs/1802.06006 [3] https://arxiv.org/abs/2006.07931 We hope to publish a paper on the topic soon.
Share cardActual performance
Launch Intel predictions
Analyze your own launch →Correct prediction on native model
Similar products
Musical instrument practice with web audio
Orbital - New WebAudio Musical Instrument (Chrome for now)
Practice all the songs you know on your musical instrument
Musical madness
Mantel-top computerized musical chimes with MicroPython on an ESP-32
How tuned is your musical ear?
Tools for practicing a musical instrument as single HTML files
Musical Instrument + Realtime Pitch Detection for Songwriters
YouTube Musical Spectrum – audio visualizer with musical notes
Record in both 16:9 and 9:16 at the same time