Skip to content

What the Machine Cannot Hear

The Pattern

Your pipeline measures the actual sound of your library with near-perfect coverage. Phase 4 pulls Plex's sonic analysis; loudness and loudness-range (LRA) are present on 99.98% of tracks — 7,409 of your rated ones carry both. It is the single best-covered variable in the entire model.

It predicts nothing.

loudness range (LRA) mean rating n
0–4 dB (compressed) 3.19 2,496
4–6 3.23 1,781
6–8 3.20 1,143
8–10 3.25 793
10+ dB (dynamic) 3.22 1,196

A spread of 0.06 across 7,409 tracks. Integrated loudness is barely better — 3.18, 3.25, 3.24, 3.20, and 3.08 for the most crushed band, which is the only hint of signal in the whole table and amounts to a quarter-star penalty on the loudest 366 tracks.

Now set that against the mood tags, which cover only 28% of the library and are not measured at all — they are inherited editorial judgements, words a critic once typed:

mean n
Brooding 3.47 355
Bittersweet 3.40 341
Energetic 2.99 224
Exuberant 2.90 99

A 0.57-point spread — ten times the loudness result, on a quarter of the data, from a variable nobody measured.

The best-measured thing in your model is useless. The worst-covered, least-rigorous, most subjective thing in it is the strongest predictor you have.

Historical Context

This is not a bug in Plex, and it is not a surprising result once you know where each number comes from.

Loudness and LRA are broadcast-engineering metrics. EBU R 128, adopted across European broadcasting from 2010, and ITU-R BS.1770 before it, exist to solve one specific industrial problem: television adverts were louder than programmes and viewers were reaching for the remote. The standard defines a gated integrated loudness measure so that a broadcaster can normalise everything to a common level. LRA was added as a companion statistic describing how much a programme's loudness varies.

These are the tools that ended the loudness war, and they did real good. But they were designed to answer "will this jolt the viewer?" — not "is this any good?" They describe the envelope of a signal, deliberately stripped of everything about what the signal contains. A crushed 128kbps dancehall single and a crushed post-metal album measure the same. Your Neurosis material and your Snow Patrol material sit in overlapping loudness bands and are separated by a full star and a half in your ratings.

The mood vocabulary has an almost opposite provenance. Brooding, Wistful, Freewheeling, Literate, Amiable/Good-Natured — the slashed compounds give it away as AllMusic's house taxonomy, built in the mid-1990s by paid staff critics for a print guide and a CD-ROM database, and passed down the metadata supply chain through Rovi and TiVo into your NAS. It has no acoustic basis whatsoever. It is a fossil of the last era in which someone could make a living writing a paragraph about every record released.

So the two variables represent two different theories of what music is. One treats a record as a signal with measurable properties. The other treats it as a thing a person had a reaction to and tried to name. Your ratings — which are also a person having a reaction and trying to name it — correlate with the second and are indifferent to the first. Stated that way it is almost tautological. It is still worth having measured, because the pipeline spends real time every week on the variable that does nothing.

The Recordings

The clearest demonstration sits in artists whose acoustic profiles are near-identical and whose ratings are not.

Portishead (3.83 across 24 electronic-tagged tracks) and Boards of Canada (3.00 across 10) occupy much the same territory by any engineering measure: mid-tempo, sample-based, heavily processed, moderate dynamic range, both nocturnal. No loudness statistic separates them. The mood tags do — Portishead's material carries Eerie and Enigmatic, both near the top of your scale; Boards of Canada's carries Nostalgic and Hypnotic with a pastoral warmth your ratings consistently decline.

Neurosis at 3.77 and Electric Wizard at 2.50 are the same test at the heavy end. Both are slow, loud, distorted, long-form and compressed. Every acoustic number the pipeline holds says they are the same music. Your ratings say one is the anchor of an entire lane and the other is closed. The mood columns say Brooding/Ominous versus Menacing/Druggy/Nihilistic — and that distinction, unmeasurable and entirely editorial, is the one that tracks what you actually did.

There is one genuine acoustic finding worth keeping: the 366 rated tracks in the most-crushed loudness band average 3.08 against a 3.22 baseline. Small, but real, and consistent with a collection weighted towards pre-2000 mastering. It is not enough to build anything on.

What's Missing

Nothing needs acquiring to act on this. Two operational points follow instead, and both are cheap.

Mood coverage at 28% is the binding constraint on the most useful column in the model — and the gap is not random. The reggae family has almost no mood coverage at all, which is precisely where the affect axis is most testable. Whatever the Phase 4 freshness gate is doing, widening mood coverage buys more than any further sonic work will.

And the sonic centroid ideas parked in the backlog — sonic radio, seeding discovery from acoustic similarity — should be understood as building on sand. Acoustic proximity does not predict your ratings. taste_score already knows this implicitly, which is why it works; the parked ideas would be reintroducing the useless variable as a primary key.

The acquisition targets below are the one place this essay can be tested rather than merely argued: four records that are acoustically alike and affectively opposite, and all four are verified absent.

Acquisition Targets

  • Bark Psychosis — Hex (1994) — the record that named post-rock; acoustically near-identical to material you rate 2.8, affectively identical to material you rate 3.8.
  • Labradford — Mi Media Naranja (1997) — minimal, low-dynamic-range drone that should be indistinguishable from Duster (2.95) by any acoustic measure and very distinguishable by mood.
  • Seefeel — Quique (1993) — shoegaze dissolving into ambient techno; a direct test of whether the Hypnotic/Eerie tags travel to material with no vocal at all.
  • Rachel's — Music for Egon Schiele (1996) — chamber post-rock, high dynamic range, deeply interior; the opposite acoustic profile to Neurosis and, the thesis predicts, a similar rating.