· By The Vocal Market
How Does Autotune Work? Pitch Correction Explained for Producers
Short version
A pitch corrector does three jobs, over and over, while the vocal plays. It measures the pitch of the voice, it decides which note in your key the singer meant, and it shifts the audio onto that note. Retune speed controls how fast that last step happens: 0 ms is the hard, stepped sound, and slower settings sound like singing. The difficult part is vibrato, because a good tuner has to correct the note without correcting away the movement around it. The rest of this article goes through each step with real pitch traces from a catalog vocal.
Most producers know what autotune sounds like long before they know what it's doing. You drop it on a vocal, pick a key, turn a knob, and either it sounds like a better take or it sounds like a robot. When it goes the wrong way, it's hard to fix a process you can't picture.
So this is the picture. Everything below uses our own plugin, TUNE, because we can show you exactly what it does inside, but the three steps are the same in every pitch corrector you'll come across. Once you know them, the knobs on all of them start to make sense.
The examples are three seconds of White Flag by CODA, a real vocal from the catalog, in F# major. The traces aren't drawings. They were measured from the audio that came out of the plugin, 100 points a second.
The three jobs
Strip away the interface and a pitch corrector runs a loop:
- Measure. What pitch is the voice on right now?
- Decide. Which note in the key was the singer going for?
- Move. Shift the audio by the difference, at the speed you've set.
TUNE runs that loop roughly every 6 milliseconds, about 170 times a second. Each step has its own problems, and nearly every control on a tuner exists to deal with one of them.
Look at how little of that line sits exactly on a note. That's normal. Good singers scoop into notes, drift a few cents, and hold long notes with vibrato. The question a tuner has to answer is which of those movements are mistakes and which are the performance.
Step 1: measuring the pitch
A sung note is a repeating waveform. The pitch you hear is how many times per second that shape repeats: A above middle C repeats 440 times a second, so it's 440 Hz. Each semitone up is about 6% faster, and each semitone is divided into 100 cents, which is the unit you'll see in every tuner.
To find the pitch, TUNE compares the incoming audio with a delayed copy of itself. When the delay is exactly one cycle long, the two copies line up almost perfectly, and that delay is the period of the note. The method is called YIN. It's been around for over twenty years and is still one of the most common ways to track a single voice, because it's accurate and doesn't need much audio to work with.
Octave errors
There's a catch. A waveform that lines up with itself after one cycle also lines up after two. Occasionally the best match is the double-length one, and the tracker reports a note an octave too low. You've heard this if a tuner has ever briefly thrown a note down an octave on a breathy passage.
TUNE handles it by scoring several candidate periods rather than trusting the first match. Anything that still jumps an octave away from its neighbours is pulled back toward the rest of the note instead of being passed straight to the shifter.
Why low voices need more time
To recognise a cycle, the tracker has to see at least a couple of them. A low male voice around 80 Hz has cycles 12.5 ms long, a soprano's are a quarter of that. The lowest note the tracker has to be ready for sets how much audio it needs before it can answer. TUNE Pro has a voice type setting for this: telling it roughly where the singer sits narrows the range it has to search.
S's, T's and breaths
Consonants and breaths don't repeat. An S is noise. It has no pitch, so there's nothing to correct, and a tuner that tries anyway produces the lispy, zippery artefacts people blame on autotune.
TUNE decides, frame by frame, whether the voice has a pitch at all. Where it doesn't, the audio crossfades back to the untouched signal. S's, T's and breaths come through exactly as they were recorded. In TUNE Pro, Gate sets how strict that decision is: turn it up and less of the take gets pitched, more passes dry. Floor is a separate noise gate that fades the output to silence below a level you set, which helps when there's room noise or headphone bleed on the track.
Step 2: deciding the note
This is where tuners differ most, and it's the step that decides whether a vocal sounds tuned or sounds processed.
The key comes first
Setting the key and scale gives the tuner its list of allowed notes. In F# major that's F#, G#, A#, B, C#, D# and F (strictly E#, but every DAW calls it F). TUNE shows the list right under the key selector, so you can see what it will tune to.
TUNE doesn't detect the key for you, and that's deliberate. We built key detection, tested it on vocals with a known key, and it got half of them wrong. It isn't fixable from a vocal alone: a melody that never sings a particular scale degree fits two keys equally well, and the key finders that work listen to the whole mix, which a plugin on the vocal track never hears. Setting it yourself takes five seconds, and guessing it wrong makes every note wrong.
Why snapping to the nearest note doesn't work
The simplest tuner takes each pitch measurement and snaps it to the closest allowed note. It works on a steady, well-sung note. It falls apart on vibrato.
Vibrato is the pitch swinging above and below a note five or six times a second, often by 40 or 50 cents either way. Snap every measurement to the nearest note and one of two things happens. Either the whole wobble gets flattened into a dead, straight line, or, if the wobble crosses the halfway point between two notes, the tuner flips back and forth between them and you get a warble that was never sung.
Note centre plus vibrato
So TUNE doesn't correct the raw pitch. It splits it into two parts: the centre of the note the singer is holding, and the vibrato moving around that centre. Only the centre is moved onto the scale. Then the vibrato is put back on top, by as much as the Vibrato setting says.
Finding the centre properly is the hard part. TUNE starts measuring at the moment a new note begins, ignores anything more than a semitone away from where the voice currently is, and when the pitch is clearly oscillating it measures how long one vibrato cycle takes and averages over whole cycles. That puts the centre in the middle of the wobble rather than wherever the wobble happened to be when the note started. In testing, Natural kept 98% of a vibrato 50 cents deep, while Hard Tune kept 14%.
This split is what the presets are really setting. Natural keeps all of the vibrato, Modern Pop keeps about half, Hard Tune keeps none. The correction of the note underneath is the same in all three.
Notes that sit right between two others
A voice hovering near the midpoint between two notes is the worst case for any tuner, because a few cents either way changes the decision. That's where you hear notes flip. Most of the time the note centre solves it, but SAUCE, which snaps every movement, needs something extra. It only switches note when the voice is clearly past the halfway point and the audio just ahead agrees, and a note held inside that grey zone is allowed to settle for about 120 ms before it moves. The result is the full hard sound without the random jumps.
What Pro adds to the decision
The Advanced controls in TUNE Pro all adjust this step:
- Humanize gives held notes their natural movement back while short notes still snap. On notes held longer than about 120 ms, the amount of vibrato kept rises until it's all back by around 400 ms. That's how a hard-tuned hook can still sound sung.
- Flex-Tune makes the correction weaker the further the voice is from a note. Slides and blue notes, sung well away from the scale, are left alone. Pitch sung close to a note is still corrected.
- Strength sets how far toward the target the correction goes.
- The note keyboard removes notes from the scale. If a slide between two notes keeps catching a third on the way, take that note out and TUNE goes straight past it.
- Wobble is the hard-tuned trill: the target alternates between the corrected note and the next note up in the scale, about 5.4 times a second.
Step 3: moving the voice
Retune speed
Once the tuner knows how far off the voice is, it has to decide how quickly to fix it. That's retune speed.
At 0 ms the correction arrives instantly. Every note snaps the moment it starts and every slide between notes becomes a jump, and that's the stepped sound people mean when they talk about hard autotune. At 42 ms, where the Natural preset sits, the correction eases in over the start of each note, the way a singer lands on a pitch. TUNE's range runs up to 200 ms, which only touches notes that stay off for a while.
One detail matters a lot here. In TUNE, retune only slows down the correction of the note. It never slows down the vibrato, which is put back separately. If both went through the same smoothing, a slow retune would blur the vibrato into a lazy wobble, and a fast one would chase it. Keeping them apart is why Natural can be quick on the note and still leave the vibrato alone.
Shifting without changing the voice
Now the audio has to actually move. The obvious way is to play it faster or slower, like changing the speed of a tape. That changes the pitch, but it also moves the formants: the resonances of the singer's throat and mouth that make a voice sound like that particular person. Push a note up that way and the singer sounds smaller and younger. Push it down and they sound bigger.
In its normal mode TUNE uses a method called PSOLA. It cuts the voice into small pieces, each one pitch cycle long, and places them closer together to raise the pitch or further apart to lower it. Each piece keeps its original shape, so the formants stay where they were sung and a corrected note sounds like the same person hit it. That's on in the free version.
If you want the voice to sound different, TUNE Pro has Character, which moves the formants on purpose by up to four semitones either way while leaving the pitch where it is.
Stereo vocals
Correcting the left and right channels separately sounds harmless, but the two sides can end up processed a fraction differently and partly cancel when the mix is summed to mono. TUNE corrects the middle of a stereo vocal and passes the width through untouched and time-aligned, so it stays mono-safe and keeps the stereo image.
Why autotune needs to look ahead
There's a timing problem hiding in all of this. The tracker needs a slice of audio to measure a pitch, so by the time it has an answer, that answer describes audio from a moment ago. If the correction were applied to the audio arriving now, it would always be late. On vibrato that's bad enough to push the pitch the wrong way.
The fix is to delay the audio. TUNE holds it back by about 78 ms, so every correction is applied to exactly the audio it was measured from, and the note-centre decision also gets to see a little of what comes next. A new note gets recognised cleanly at its start rather than a beat late.
You don't hear that delay while mixing. TUNE reports it to your DAW, and the DAW's delay compensation lines everything back up. What you bounce is identical to what you heard in real time, at any buffer size.
What LIVE mode trades
78 ms is fine in a mix and useless when someone is singing through the plugin, because they'd hear themselves late in their headphones. LIVE mode brings the delay down to about 3 ms, under the 5 ms Logic allows in Low Latency Mode.
Getting there means giving things up. There's no look-ahead in LIVE, so every decision is made from the newest audio only. The shifter changes too: LIVE uses a much faster one whose delay is about half a pitch cycle, and that one can't preserve formants. TUNE shows the Formant and Character controls as STUDIO only while LIVE is on.
So treat LIVE as a tracking tool. Record through it so the artist hears the effect, then switch it off for the mix and TUNE goes back to the full look-ahead and the formant-preserving shifter.
Reading TUNE's pitch display
The display in TUNE draws exactly what these figures show: grey for what was sung, colour for what comes out. It's worth watching while you set it up, because most tuning problems are visible before they're obvious by ear.
- The colour line lands on a note you didn't expect. Check the key first. If the key is right, the singer is closer to that note than you'd think, and taking it out on the note keyboard in Pro decides it.
- The colour line is flat where the singer had vibrato. You're on a preset that keeps little or none of it. Try Natural or Modern Pop.
- The line flips between two notes on a held note. The voice is sitting near the halfway point. A slower retune or a gentler preset usually settles it.
- Grey and colour are almost the same. That's a well-sung take on Natural. Nothing is wrong.
Which preset for which job
| Preset | Retune | Vibrato kept | Reach for it when |
|---|---|---|---|
| Natural | 42 ms | All | The tuning shouldn't be heard. Start here on a lead. |
| Gentle Polish | 70 ms, 70% strength | All | The take is nearly there already. |
| Modern Pop | 12 ms | About half | Tight and polished, or doubles that need to lock to the lead. |
| Hard Tune | 0 ms | None | You want the effect you know. |
| Expressive Hard | 0 ms, Humanize 60%, Flex-Tune 35% | On long notes | Hard snap on short notes, long notes still breathing. |
| Parallel Blend | 0 ms at 50% mix | In the dry half | The character of hard tune under the real take. |
| Live Tracking | 8 ms, LIVE on | Some | Someone is recording through it. |
| Wobble Trill (Pro) | Hard Tune, Wobble 70% | None | The hard-tuned trill between two notes. |
SAUCE sits on top of any of these. It overrides anything the preset would soften, so it's the same hard sound whichever preset you're on, and you can automate it on for the hook and off for the verses.
Frequently asked questions
Why does my autotune sound robotic?
Because retune is fast and the vibrato is being removed. That combination is the robotic sound, and it's exactly what Hard Tune is set up to do. If you want the vocal to sound sung, use a preset that keeps the vibrato, like Natural, and let the correction ease in.
Do I have to set the key?
Yes. A tuner can only move notes onto the scale you give it, and a vocal on its own doesn't reliably tell you the key. Every acapella on The Vocal Market lists its key on the page.
Can autotune fix a badly sung vocal?
It can put the notes in tune. It can't fix timing, tone or a performance that doesn't feel right, and the further off a note starts, the more obvious the correction becomes. A good take tuned lightly will always beat a weak take tuned hard.
Does autotune work on harmonies or a whole stack?
TUNE works on one voice at a time, because the tracker follows a single pitch. Put a separate instance on each vocal track. At about 1% of one CPU core per instance, that's not a problem.
Does autotune work on rap?
Yes. Rap has pitch in it, even when it isn't melodic, and the hard, stepped sound on rap vocals is one of the most common uses of autotune. Hard Tune or SAUCE with a fast retune is the usual starting point.
Why does autotune add latency?
The tracker needs audio to measure a pitch, and the audio is delayed so each correction lands on the audio it was measured from. Your DAW compensates for it during playback. For recording through the plugin, use a low-latency mode like TUNE's LIVE.
Is TUNE really free?
Yes. Free TUNE has the full correction engine, LIVE, SAUCE, formant preservation and seven of the eight presets, with no time limit and no audio dropouts. TUNE Pro, at €29, unlocks the Advanced controls so you can set them yourself instead of following the preset.
Try it on a real vocal
The quickest way to understand all of this is to watch it happen. Put TUNE on a dry vocal, set the key, and switch between Natural and Hard Tune while the pitch display runs. If you don't have a vocal to hand, every acapella in the catalog lists its key and BPM, so you can set TUNE before you've heard a note.
After TUNE, the rest of the vocal sound lives in the processing chain: compression, EQ and saturation. Our free vocal chain covers that, and our vocal mixing guide goes through what each stage does.
TUNE, the free autotune plugin
AU and VST3 on Mac, VST3 on Windows. Checkout is €0 and doesn't ask for a card.
Download TUNE free