← All terms

Glossary

What Is Singing Voice Conversion?

Singing voice conversion (SVC) maps a sung performance onto a target voice model while keeping melody and timing. Here is the term, and how cover apps like VibeSing use it.

The short version

Singing voice conversion (SVC) is voice conversion aimed at sung audio. The input is a vocal performance — pitch, vibrato, phrasing, lyrics as they were sung. The output is that same performance with a different identity: a trained voice model instead of the original singer.

Speech conversion can ignore melody. SVC cannot. Notes have to land. Breaths and legato have to look like singing, not like a voicemail on top of a beat. That is why cover apps care about this term even if they never show it on a button.

RVC is a well-known open-source line of work in this neighborhood. SVC is the problem. RVC is one approach. You do not need to pick one in VibeSing.

What has to be true for SVC to work

A vocal, not a mix. Conversion on a full track hears kick drums as consonants. Vocal isolation / stem separation / Demucs-class models sit upstream. None of that is a Studio control here; it is the plumbing.

A target identity with enough coverage. A model trained on three spoken English sentences will still be asked to follow vampire. Sometimes the hook is close enough. Sometimes high belts smear. SVC copies the source's pitch contour; it cannot invent a chest voice you never showed it.

Consent and rights. SVC is how you get "me on their song," not "me as them." Running SVC toward a scraped celebrity model is the lawsuit flavor of this technology. VibeSing trains the target on the person recording the prompts.

SVC vs. the other "AI singing" buttons

Text-to-song synthesizes a new vocal from lyrics (and maybe a melody prompt). No original performance required. Also no that production of Seven.

Auto-Tune / pitch correction is you singing, then snapped. SVC is not you singing.

Pitch shift changes F0 globally. It does not swap identity.

AI karaoke in ads often means SVC plus an instrumental, sold as karaoke. Scoring apps are still a different product.

How VibeSing uses it

The consumer flow hides the acronym:

  1. Three English samples in Studio. No Voices tab, no dataset upload.
  2. Demo track or an upload you have rights to.
  3. Generate — isolation, SVC-style conversion onto your model, mix, share page.

You do not paste lyrics. You do not pick "SVC vs. TTS." You do not get a celebrity. Friends remake from the page with their models. How to make an AI cover is the longer walkthrough; without singing is the honest one.

Free: 100 credits a month (~10 songs), one lifetime training. Train 10, clip 10.

Open Studio. SVC is why a spoken take can come back as a chorus. It is not a setting you toggle.

See it in action — try VibeSing free.

Clone your voice in 30 seconds and make your first AI cover song.

Open Studio