Glossary
What Is Voice Conversion?
Voice conversion takes existing audio and makes it sound like a target voice while keeping the words and timing. Here is the term, and how VibeSing uses it for covers.
The short version
Voice conversion (VC) is an AI technique that takes audio in and produces audio out that preserves what was said or sung — the phonemes, the timing, often the melody — while changing who it sounds like. You start with a recording. You end with a recording that has a different timbre.
That is the opposite of text-to-speech, which starts with a script. It is also different from a cheap "voice changer" (pitch up, robot, chipmunk). Those are effects. Conversion is a trained mapping from one voice space to another.
If you have heard an AI cover where the original singer's performance is still there — breaths, runs, the same hook — but the tone is someone else, you heard voice conversion. RVC is one popular family of methods. Singing voice conversion is the same idea optimized for melody instead of speech.
What the model is actually swapping
Roughly, conversion tries to split content from identity:
- Content: which syllable, which pitch, how long the note is.
- Identity: timbre, the grain of the voice, the stuff that makes you sound like you on a voicemail.
A good conversion keeps content and replaces identity. A bad one smears both: words turn to mush, or the output still sounds like the original singer with a filter. Vocal isolation and stem separation matter because converting a full mix (drums and all) gives the model garbage to treat as "voice."
Voice conversion vs. voice cloning vs. TTS
Voice cloning is how you get the target identity — train a voice model on samples. Conversion is how you use it on a new performance.
TTS / vocal synthesis invents the performance from text (and, for singing systems, from a score). Fine for audiobooks. The wrong tool if you wanted APT. with its original arrangement intact.
Pitch shift is not conversion. Sliding a clip up two semitones does not give you a second person.
How VibeSing uses it
VibeSing is a conversion app wrapped in a short studio, not a research demo and not a text-to-song box.
- You record three short English samples in Studio. That trains your model. There is no Voices tab, no celebrity library, no "upload dataset" screen.
- You pick a demo or upload a song you have rights to.
- Generate runs isolation, conversion onto your model, and a remix you can share.
You never type the lyrics. The source vocal already contains them. That is why a speaking-trained clone can still "sing" Anti-Hero — the melody is coming from the track, not from your karaoke skills. Making a cover without singing is the consumer version of this paragraph.
Consent is structural: conversion is only as ethical as the target model. VibeSing trains on the person at the mic. Ethics is the longer argument.
Free: 100 credits a month (~10 songs), one lifetime training. Train 10, clip 10.
Open Studio if you wanted to hear conversion on a real song instead of a spectrogram.