We're heading to two conferences in November, and it's our first time contributing to either community. Felix submitted papers to both: one on generating drum kits at ISMIR, and one on our engine at the Audio Developer Conference. We're really looking forward to it.
ISMIR
ISMIR is the yearly conference of the International Society for Music Information Retrieval, and the main place where researchers who study music with computers meet. That covers a lot: analyzing recordings and scores, recommending music, transcribing it, and more and more, generating it. People come from universities, labs and companies like streaming services and instrument makers.
Our Text-to-Kit paper was accepted for the Late-Breaking/Demo session at ISMIR 2026 (Abu Dhabi and online, 8 to 12 November).
The idea: you describe a style, say "spacecraft foley, steel bulkheads and pressure valves", and get back sixteen drum pads you can play right away. Kicks, snares, hats, toms, percussion. They should sound like one kit, but not like the same sound sixteen times, and every hit needs a clean start and a tail that isn't cut off.
Most evaluation looks at generated sounds one at a time, which misses all of that. So we propose a benchmark that scores the kit as a whole, and we compared two ways of making one with Stable Audio 3 and a fine-tuned adapter.
We expected that rendering one long take and chopping it into hits would give the most consistent kit. It didn't. Rendering each pad separately, with a shared style prompt plus a short description of its role, gave roughly three times as many distinct sounds, almost no duplicates, and kits just as consistent as real ones. Chopping also tends to cut off the tails.
It's early work and we're upfront about that in the paper. The proper controlled study and a listening test come next. In the meantime you can try the process in the Text-to-Drumkit workbench.
ADC
The Audio Developer Conference is where the people who build audio software get together: plugin and synth developers, DAW teams, DSP engineers, people making music apps and hardware. The talks are very hands-on, about code, performance and the tools everyone uses.
We'll be there in Bristol (9 to 11 November) to talk about the Engine.
If you've built audio software, you know the problem: the web demo, the plugin and the hardware version slowly turn into separate codebases that don't sound quite the same. We write a module once, in TypeScript, and build the browser, plugin and hardware versions from it. Then we check that they actually produce the same audio.
We'd love to hear how other people handle realtime safety, latency and keeping targets in sync. We have some measurements and plenty of open questions.
Let's meet
If you work on generative audio, evaluation, plugins or DSP, we'd love to chat. We can do an online meeting during either week, or grab a coffee if you're there in person. Get in touch here and mention ISMIR or ADC.
See you in November.
