LyriSynclyrisync
How it worksAfter EffectsGuidesPricingContactSign in
01How it works02After Effects03Guides04Pricing05Contact06Sign in
All guides

How to sync lyrics to a song without an acapella

You can automatically sync lyrics against a finished song without stems or an acapella. The old advice said you needed an isolated vocal first. That has not been true for a while, and the reason also explains which tracks can still give an aligner trouble.

Why an acapella used to be the requirement

Automatic timing works by matching the words you supply against what it hears. On a finished mix the vocal is sitting underneath drums, bass, guitars and reverb, all of which occupy the same frequencies as a voice. A system listening to the full mix is trying to find speech in something that is mostly not speech, and it does badly.

So the workflow used to be: find or make an acapella first, then time against that. For most people making a lyric video that was a dead end, because they had the master and nothing else.

What changed

Source separation got good. Models trained to pull a vocal out of a finished mix are now reliable enough to be a normal preprocessing step rather than an experiment. The practical consequence is that the acapella requirement moved: you still need one, but the software makes it for you from the master.

It is worth being precise about what that does and does not fix. It removes the instrumental problem. It does not make an unclear vocal clear, and it does not tell the system anything about words it was never given.

The workflow

  • Start from the best master you have. A lossless file or a high bitrate MP3 is fine. There is no benefit to converting formats first.
  • Get your lyrics exactly right before you start. This matters more than anything else on this page. See below.
  • Let the tool separate the vocal and align. The separation happens automatically as part of timing.
  • Listen to the isolated vocal. If a section sounds mangled after separation, that is exactly where the timing will be weakest, and you now know where to look.
  • Play it through and correct. Expect to move some words. Treat the output as a strong start, not a finished file.

Your lyrics have to match the recording

This is the single most common cause of a bad result, and it has nothing to do with the audio. Automatic timing places the words you give it. It does not transcribe, and it will not invent a word you did not supply or skip one you did.

So a lyric sheet from the internet will often be wrong in ways that matter:

  • Ad libs and background vocals that are sung but not written down
  • A repeated chorus written once with a repeat instruction
  • Lyrics for the album version when you are working on the radio edit
  • Section labels like Verse or Chorus left in as if they were sung
  • A producer tag or spoken intro that is not in the text at all

Every one of these pushes the surrounding words out of place, because the system is trying to fit your text onto audio that contains something different. Ten minutes checking the lyrics against the actual recording is the highest value thing you can do.

What still makes a track hard

Being honest about this is more useful than pretending it is solved:

  • Dense or heavily processed mixes. Heavy compression, saturation and thick reverb all make the vocal harder to separate cleanly.
  • Several voices at once. Stacked harmonies and call and response sections give the system more than one thing to follow.
  • Whispered, shouted or growled delivery. Anything far from ordinary singing is harder.
  • Long instrumental breaks. Gaps in the vocal are where timing tends to drift, since there is nothing to anchor to.
  • The wrong language, or an unsupported one. English, French, German, Italian and Spanish are supported, and you pick the language when you upload. It decides how the words are pronounced for alignment, so a track aligned as the wrong language will come out wrong. Italian has had less testing than the other four. Anything else, including any language written in a script other than the Latin alphabet, is not supported.

What you get at the end

A start and end time for every word and every line, which you can export as captions for an editor, as data for your own animation, or as a lyrics file for a player. Which format depends on where it is going.

Next: word level vs line level timing and SRT vs LRC.

LyriSync, LyrisyncAfter EffectsLRC generatorSRT generatorKaraoke lyricsGuidesTermsPrivacyCreditsBy TJCreateContact