Word-level vs line-level lyric timing, explained
The distinction sounds like a detail and it decides your whole workflow. It sets which file format is useful to you, how much checking you have to do, and whether a particular animation is possible at all.
Line level
Line level timing gives you one start and one end per written line. It is what a subtitle is. The whole line appears, sits there, and goes away.
This is all you need for a large share of lyric videos. If the text cuts in as a block, or fades in as a block, nothing in your animation cares where the third word falls. Line level is also much faster to check by eye, because there are perhaps forty things to verify in a song rather than four hundred.
Word level
Word level gives you a start and end for every individual word. You need it the moment anything in your design responds to a specific word rather than the line as a whole:
- A karaoke highlight travelling along a line as it is sung
- Words appearing one at a time
- Kinetic typography where a word scales, cuts or moves on its own beat
- Anything cut to the vocal rhythm rather than to the bar
It is strictly more information, so you can always collapse word level down to line level later. The reverse is not true, which is a good argument for getting word level even if your first edit only uses lines.
The part people underestimate
Word level timing is more to check. A line that reads correctly can still contain a word sitting in the wrong place, and you will not see it until the highlight jumps. Budget real time for playing the track through and watching the words rather than reading them.
This is why any word level workflow needs an editor where you can drag a word and hear the result immediately. An aligner you cannot correct is not much use on real work, because no automatic system is right on every word of every track, and the ones it gets wrong tend to be the ones a viewer notices.
Where the two disagree most
Long held notes and melisma, where a single syllable is sung across several notes, are the classic case. Line level timing is untroubled: the line starts when it starts. Word level has to decide where one stretched word ends and the next begins, and that boundary is genuinely ambiguous even to a human listener.
Fast delivery is the other one. In rapid vocals the gaps between words shrink towards nothing, so a small error in one word boundary is visible as a highlight running ahead of or behind the voice, even though the line as a whole is perfectly placed.
A practical middle ground
Phrase level sits between the two. Instead of one cue per line or one per word, it breaks lines where the singer actually breathes, which often matches how you would cut the text on screen anyway. It is usually the most readable option for straight captions on a busy track, because a long line does not sit on screen for eight seconds and a short one does not flash past.
The sensible default for most work: get word level data, export line or phrase level cues for anything that is just text on screen, and use the word data for the parts that animate.
Build a word-following export with the karaoke lyrics generator. Related: SRT vs LRC, and which formats can carry word timings.