Music generators, voice synthesis, transcription engines and repair tools for bad recordings. Grouped by what you are trying to do rather than by which lab built the model.
Full songs with vocals from a text prompt, and the vocals are the part it does unusually well. Fine for demos, background beds and jokes; mixing quality still lags a real production.
The closer rival to Suno, with more control over sections and extensions. Musicians tend to prefer it because you can steer an arrangement instead of rerolling the whole track.
The best text to speech available, with voice cloning and dubbing that hold up in production. Watch the character limits, because long-form narration eats a plan quickly.
Transcription API with speaker labels, summaries and topic detection in one call. Pick it over rolling your own Whisper stack when you want the extras without building them.
Built for speed and streaming, which makes it the usual choice for live captions and voice agents. Accuracy on messy phone audio is a notch above most competitors.
Very strong on accents and non-English languages, and it can run on your own infrastructure. That last point is why regulated industries end up here.
Upload a rough voice recording and it comes back sounding like a studio mic. It is free, it takes ten seconds, and it can over-process music or room tone, so keep the original.
Splits a finished track into vocals, drums, bass and other stems with very few artefacts. Useful for remixing, karaoke and pulling a clean voice out of a mixed recording.
Stem separation aimed at musicians, with pitch shift, tempo change and chord detection alongside it. The mobile app makes it the practical choice for practising along to a song.
Run transcription on your own machine with no per-minute bill and no audio leaving your laptop. Slower than the hosted APIs and it needs a decent GPU for the larger models.
The Deezer stem splitter that started this whole category. Quality is behind the paid services now, but it is free, scriptable and fine for batch work.
Text to speech built around a studio timeline, so you can sync narration to slides and video without a separate editor. Voices are less lifelike than ElevenLabs but easier to direct.
Low-latency voice API aimed at conversational agents rather than narration. Worth comparing against ElevenLabs on price if you are generating speech at volume.
Strips background noise and echo from calls at the driver level, so it works in any meeting app. The free tier covers most people who just want the dog to stop being audible.
Automated levelling, loudness normalisation and noise reduction for podcasts. Less flashy than the newer AI cleanup tools and considerably more predictable.
Removes filler words, stutters and mouth noises from spoken recordings. Narrow by design, and a real time saver on long unscripted interviews.
Records each guest locally at full quality, so a bad connection does not ruin the take. That single feature is why remote podcasts and interviews end up here.