Glossary
What Is a Voice Model?
A voice model is a small AI trained on recordings of one voice so new audio can sound like that person. Here is the term, and how VibeSing trains and uses yours.
The short version
A voice model is a file (or a set of weights) that represents how one person sounds — pitch habits, vowel color, the grain on consonants — well enough that a generator or voice conversion system can produce new audio in that identity. It is not an mp3 of you saying a sentence. It is not a playlist. It is a compressed recipe for your timbre.
Once it exists, you can apply it to speech or, in cover apps, to a sung performance you never recorded. The words and melody can come from somewhere else. The "who" comes from the model.
Model vs. sample vs. clone
Samples are the recordings you captured. Three short English lines in a quiet room, if you are on VibeSing.
Voice model training is the compute step that turns those samples into the recipe.
Voice cloning is the product English for the same pipeline: get a model that sounds like someone. On this site, that someone is supposed to be you.
A generic TTS voice is also a voice model. It was trained on many speakers or on a studio actor, not on your bathroom takes. Convenient, anonymous, useless for "that's you on drivers license."
Community "artist models" you see on cover forums are voice models too — usually trained on leaked or scraped singing. VibeSing does not ship those. No cloning other people.
What a singing cover actually asks of the model
Speech samples do not contain belted high notes. The conversion system will still try to follow Beautiful Things. Sometimes it sounds like you. Sometimes it sounds like you stretched. That is a coverage problem in the model, not a moral failing. Cleaner, slightly more varied samples help. Performing the whole song does not. See recording clean samples.
The model does not "know" Olivia Rodrigo. It knows you. The song's identity stays in the source vocal and the instrumental.
How VibeSing uses yours
- Open Studio. Record the three prompts. There is no Voices locker, no "upload dataset zip," no picker of starter celebrities.
- Training runs in the background (10 credits). On the free plan that training is once, lifetime. Credits reset; the cap does not.
- Generate a clip (10 credits) against a demo or an upload you have rights to. Same model, many songs, until you are allowed to train again on a paid plan.
- Share page. Other people do not download your model when they tap Make your version — they train theirs.
Free: 100 credits a month, about ten songs. Is it free? Pricing for more.
Keep the model boringly personal. A voice model is an identity object. Treat it like a voiceprint, not like a meme pack. If Band Mode ever lets friends combine models into one render, it will still be each person recording their own — roadmap, not a reason to share weights in Discord.
Note: Band Mode (group covers) is on the VibeSing roadmap and isn't live yet — any Band Mode sections below describe the planned experience. Solo covers in your own voice, share pages, and vertical video export are live today.
Open Studio and make the only model you have the right to: yours.