· By The Vocal Market
The localization industry has always had a blind spot: songs. Dialogue can be dubbed relatively cheaply by hiring voice actors in the target language. Songs require singers, studios, producers, and often rewrites of the lyrics to preserve meter and rhyme. The cost per minute of localized song content has historically been 5 to 20 times the cost per minute of localized dialogue. AI singing voice synthesis is changing that math. Not by replacing singers entirely (the quality ceiling still favors humans for hero content) but by dramatically reducing the cost of secondary and tertiary song localization: background tracks in games,...
· By The Vocal Market
Games used to solve NPC dialogue with pre-recorded voice actors and solve in-game music with licensed soundtracks. Both approaches worked, both were expensive, and both were rigid. If a side character needed a new line six months after launch, you hired the voice actor back. If a tavern scene needed a different song, you paid for a new license. Procedural content was beautiful in theory and unsustainable in practice. AI singing voice models have started to change that calculus. A game studio that can generate singing on demand, in specific voices, in specific styles, for specific in-world contexts, has new...
· By The Vocal Market
Karaoke and lyric-sync apps are a specific corner of the voice AI market with their own technical and legal requirements. The products that dominate the category (Smule, StarMaker, Yokee, Sing Karaoke by Smule, and a growing list of ChatGPT-era newcomers) all depend on training data that most developers cannot easily source: clean vocal performances with accurate pitch annotations, time-aligned lyrics, and known musical keys. This post is a guide for developers building in this space. It covers the specific ML tasks a karaoke or lyric-sync app needs to solve, the training data each task requires, and how to assemble a...
· By The Vocal Market
Voice cloning started as a speech problem. The early production systems (XTTS, ElevenLabs, PlayHT, Resemble) were built on speech datasets and optimized for TTS-style output. Then the market asked for singing. Users wanted to clone their own voice to sing over instrumentals. Artists wanted to generate harmonies in their own style. Developers wanted to build karaoke apps that could produce any song in any voice. The singing use case turned out to be harder and more legally fraught than the speech use case, and the gap in licensed training data became visible fast. This post is for product and ML...
· By The Vocal Market
If your ML background is in computer vision, NLP, or tabular data, the audio world comes with its own vocabulary and a few genuinely confusing distinctions. Most of the terms are borrowed from music production, where they have been stable for decades. A few of them have been reused by the ML community in ways that do not perfectly match the original meaning. This post is a short glossary to help you navigate. The glossary is organized roughly by topic: sources and types, processing states, container formats, and audio signal properties. For each term we give the definition, how ML...
· By The Vocal Market
"How much data do I need?" is the first question every ML team asks when starting a voice model project, and it is the question with the least satisfying answer. The real answer is "it depends," and the dependencies matter more than any single number. This post breaks down the actual training data requirements across the current landscape of singing voice models, from two-minute fine-tunes on top of pretrained bases up to multi-hundred-thousand-hour frontier systems. All numbers below come from published papers, GitHub documentation, or reported commercial specifications. Where a source is uncertain or contested, it is flagged. The two...