Why we show a musical key on only 44% of tracks
Every track you split on StemX gets a tempo and a musical key, and so does any file you drop into our free key finder. For most tracks, we show no key — on purpose. Here is the measurement behind that choice, and why we think an honest blank beats a confident guess.
How good is automatic key detection?
We score key detection against GiantSteps, the standard public key benchmark: 604 tracks of electronic music with human key annotations, scored by the MIREX convention so the numbers compare with published work. The detector is S-KEY, Deezer's pre-trained key model.
Asked to name a key for every track, it is exactly right 64.4% of the time (95% CI 60.5–68.1%), with a MIREX weighted score of 0.699 — a long way up from the hand-rolled chroma-template detector it replaced, which managed 44.7%.
But 64% means one track in three gets the wrong key. For a field a DJ sorts or filters a library by, that is not good enough to show unconditionally.
A confidence you can actually gate on
The model scores all 24 keys, and the margin between its top answer and the runner-up is a real measure of how sure it is. We report that margin as key_confidence — which meant writing our own inference call, because the model's packaged entry point takes the top answer and throws the scores away.
With a confidence in hand, we swept the threshold and asked two questions at each step: on how many tracks would a key be shown, and how often would the shown key be right?
| Confidence floor | Key shown on | Right when shown |
|---|---|---|
| 0.00 | 100% | 64% |
| 0.25 | 79% | 72% |
| 0.50 | 65% | 77% |
| 0.65 | 52% | 78% |
| 0.70 | 44% | 80% |
| 0.75 | 34% | 80% |
| 0.80 | 25% | 80% |
Precision climbs to 80% and stops. Raising the floor past 0.70 throws away tracks the model would have gotten right and buys no further accuracy. So we ship at 0.70: a key is shown on 44% of tracks and is right about 80% of the time when it is.
Those two numbers only mean something together. "80% accurate" alone would be true and misleading; so would "64% accurate". They answer different questions.
We missed our own target
Our design goal was a key shown at around 85% precision. The sweep says this model does not get there at any threshold on this benchmark. We shipped anyway, at the best precision it does reach, and put the shortfall in our API documentation rather than leave it for someone to discover by counting wrong answers.
What the wrong answers look like
Not all misses are equal. Across all 604 tracks: 389 exact, 36 off by a fifth, 35 given the relative key (A minor for C major), 22 the parallel key (C minor for C major), and 122 simply wrong. A fifth or a relative key is still harmonically close — mixing on it usually works.
And mode is the reliable half: major versus minor is right 84.3% of the time, the root note only 68.0%. That surprised us. Our early tests on synthetic four-note passages said the opposite — root solid, mode shaky — and our docs briefly repeated it. Real music reversed the finding, which is the best argument we have against tuning audio analysis on synthetic input.
What a blank key means
When we show no key — "not confident enough to call this one" on the site, key: null from our API — either detection failed or it succeeded below the confidence floor and we withheld it. We do not tell you which: a withheld key sitting next to a low confidence score would only invite you to use the guess anyway.
Treat the key as a hint for sorting and pre-filling a field a human confirms — not as input to anything automated.
Limits
GiantSteps is electronic music only, in 2-minute previews. It suits the people most likely to want a key, and says nothing about orchestral or acoustic material.
Try the vocal remover
Upload a song and hear a free 60-second preview of all six stems. No account needed.
Remove vocals free