6 Commits
Author SHA1 Message Date
lyrathorpe acd042ad8d chore(release): v0.3.0 2026-08-24 16:28:00 +00:00
lyrathorpe 54e19992e4 Merge pull request 'feat: split unmatched listening by whether the library holds the title' (#6) from diag/split-misses-by-ownership into main
Build and publish container / build (push) Successful in 9m7s
Reviewed-on: #6
2026-08-24 17:19:09 +01:00
Emma Thorpe 2886c02a2e feat: split unmatched listening by whether the library holds the title
Build and publish container / build (pull_request) Successful in 8m41s
"Unmatched by an artist the library holds" was presented as the matcher's
misses. On real data it is not: owning one album by an artist says nothing about
owning a particular single of theirs, and most of that figure turned out to be
drum and bass tracks streamed but never bought.

Split it in two. A title the library holds under some other artist is an
attribution disagreement -- a remixer credited as the artist, a guest billed as
one -- and is a genuine miss worth fixing; the report now names the artist the
library files it under, which is the information needed to judge it. A title the
library does not hold at all under any artist was never bought, and no
improvement to matching will conjure it.

The distinction matters beyond presentation. That figure is the gate on the
cull, and a gate computed from a number that overstates the failure rate blocks
work that is actually safe to do.
2026-08-24 17:15:56 +01:00
lyrathorpe 56a06a6cd5 chore(release): v0.2.4 2026-08-24 16:09:38 +00:00
lyrathorpe a3d0689c0a Merge pull request 'fix: match bracketed guest credits and hyphenated version suffixes' (#5) from fix/normalise-bracketed-credits-and-dash-suffixes into main
Build and publish container / build (push) Successful in 11m42s
Reviewed-on: #5
2026-08-24 16:58:05 +01:00
Emma Thorpe 7588fee302 fix: match bracketed guest credits and hyphenated version suffixes
Build and publish container / build (pull_request) Successful in 10m45s
Two normalisation faults, both found by running a real coverage report's
unmatched list back through the normaliser. Between them they account for five
of the fifteen worst misses by play count.

The guest-credit pattern required whitespace immediately before the word, so it
caught "Yellowcard feat. Tay Jardine" but missed "Self vs Self (feat. In
Flames)" -- and the bracketed form is the more common of the two. An opening
bracket is now allowed in that position.

The trailing-version pattern matched the suffix as a run of non-hyphens, which
cannot cross a hyphen inside the suffix itself: "Gold Dust - Shy FX Re-Edit" and
"Back To Your Roots - Friction & K-Tee Remix" both survived untouched. Matched
lazily instead.

Four version words are added for how drum and bass marks its variants: vip,
bootleg, rework, extended. They only apply inside a bracket or after a trailing
dash, so the exposure is small, and ordinary titles carrying those words --
Editors, Mixed Emotions, Radio Ga Ga, Live and Let Die -- are pinned as tests
against exactly that.

The report also gains the figures that explain why the MBID tier contributes so
little. Two thirds of scrobbles carry a recording id, and only a twentieth of
them join on one: MusicBrainz holds a separate recording per release, and the
two sides rarely choose the same one. Counting the pairs that carried an id and
matched on name anyway measures that disagreement directly, and settles that the
weakness is not a bug in the join.
2026-08-24 16:56:13 +01:00
4 changed files with 183 additions and 14 deletions
+20 -5
View File
@@ -29,7 +29,14 @@ Two tiers, and no third.
| `name` | Normalised artist and title | Everything the first tier could not carry |
| `none` | — | Recorded as a miss, never guessed at |
The normalisation is the load-bearing part, because the two sides disagree in
The name tier does almost all of the work. A recording MBID is exact when it
lands, but MusicBrainz holds a separate recording per release, and Last.fm and
Lidarr rarely pick the same one: on a real library, two thirds of scrobbles
carry a recording id and barely a twentieth of them join on it. The report
counts how many carried an id and matched on name anyway, which is the measure
of that disagreement.
So the normalisation is the load-bearing part, because the two sides disagree in
predictable ways. It folds case and accents, drops guest credits (`Yellowcard
feat. Tay Jardine` against a tag of `Yellowcard`), strips a trailing
version suffix (`(Remastered 2011)`, `- Live`), expands `&`, and removes a
@@ -102,10 +109,18 @@ matcher. The line to watch is:
unmatched by an artist the library holds: N pairs, M plays
```
That is a track that was played, sitting beside a file it should have matched.
Those are the matcher's real misses, and every one is a candidate for being
wrongly called cold in stage four. The report lists the worst fifteen by play
count so they can be eyeballed.
That is a track that was played by an artist the library holds. It is then
split again, because owning an artist is a weak proxy for owning a track:
- **the title exists under another artist** — an attribution disagreement, a
remixer or a guest billed as the artist. These are the genuine misses, and
each is a candidate for being wrongly called cold in stage four. The report
names the artist the library files them under.
- **the title is nowhere in the library** — never bought. No amount of matching
conjures a file that does not exist.
Counting both as matcher failures overstates the problem and would over-block
the cull. The report lists the worst of each by play count.
## How the ingest works
+86 -8
View File
@@ -187,18 +187,28 @@ VERSION_WORDS = (
"anniversary",
"reissue",
"instrumental",
# Drum and bass and its neighbours mark versions their own way.
"vip",
"bootleg",
"rework",
"extended",
)
_VERSIONS = "|".join(VERSION_WORDS)
BRACKETED_VERSION = re.compile(
rf"\s*[\(\[][^\)\]]*\b(?:{_VERSIONS})\b[^\)\]]*[\)\]]\s*$", re.IGNORECASE
)
TRAILING_VERSION = re.compile(rf"\s+-\s+[^-]*\b(?:{_VERSIONS})\b.*$", re.IGNORECASE)
# The suffix is matched lazily rather than as a run of non-hyphens, because the
# thing being stripped frequently contains hyphens of its own -- "Gold Dust -
# Shy FX Re-Edit", "Back To Your Roots - Friction & K-Tee Remix".
TRAILING_VERSION = re.compile(rf"\s+-\s+.*?\b(?:{_VERSIONS})\b.*$", re.IGNORECASE)
# Last.fm routinely carries the guest credit in the artist field where the file
# tag holds only the primary artist -- "Yellowcard feat. Tay Jardine" against a
# tag of "Yellowcard". `with` is deliberately absent: it appears in far too many
# real titles to cut on sight.
GUEST_CREDIT = re.compile(r"\s+(?:feat|ft|featuring)\b.*$", re.IGNORECASE)
# Last.fm routinely carries the guest credit where the file tag holds only the
# primary artist -- "Yellowcard feat. Tay Jardine" against a tag of
# "Yellowcard", or "Self vs Self (feat. In Flames)" against "Self vs Self". The
# opening bracket has to be allowed for: requiring whitespace immediately before
# the word misses every bracketed credit, which is most of them. `with` is
# deliberately absent -- it appears in far too many real titles to cut on sight.
GUEST_CREDIT = re.compile(r"[\s(\[]+(?:feat|ft|featuring)\b.*$", re.IGNORECASE)
LEADING_ARTICLE = re.compile(r"^the\s+")
# Deleted rather than spaced, so "Don't" and "Dont" agree. Every other mark
@@ -1078,18 +1088,86 @@ def coverage_report(store):
row["plays"],
)
# Why the mbid tier performs the way it does. A recording id on both sides
# that still fails to join means the two disagree about which recording the
# song is -- MusicBrainz holds a separate recording per release, and Last.fm
# and Lidarr need not have picked the same one. That is not a fault to fix
# in the matcher; it is the reason the name tier has to carry the load.
library_with_mbid = store.scalar(
"SELECT COUNT(*) FROM lidarr_track WHERE recording_mbid IS NOT NULL"
)
logger.info(
"library tracks carrying a recording MBID: %d of %d (%.1f%%)",
library_with_mbid,
tracks,
100 * library_with_mbid / tracks,
)
disagreed = store.connection.execute(
"SELECT COUNT(*) AS pairs, COALESCE(SUM(plays), 0) AS plays FROM scrobble_key"
" WHERE track_mbid IS NOT NULL AND method = 'name'"
).fetchone()
logger.info(
"carried a recording MBID, joined on name instead: %d pairs, %d plays"
" -- both sides know the song, they disagree on which recording it is",
disagreed["pairs"],
disagreed["plays"],
)
suspect = store.connection.execute(
"SELECT COUNT(*) AS pairs, COALESCE(SUM(plays), 0) AS plays FROM scrobble_key k"
" WHERE k.track_id IS NULL"
" AND EXISTS (SELECT 1 FROM lidarr_artist a WHERE a.norm_name = k.norm_artist)"
).fetchone()
logger.info(
"unmatched by an artist the library holds: %d pairs, %d plays -- these are the"
" matcher's misses, not music you do not own",
"unmatched by an artist the library holds: %d pairs, %d plays",
suspect["pairs"],
suspect["plays"],
)
# Owning an artist is a weak proxy for owning a track, so that figure alone
# overstates the matcher's failings. Split it. A title the library holds
# under some other artist is an attribution disagreement -- a remixer
# credited as the artist, a guest billed as one -- and is a real miss. A
# title the library does not hold at all was simply never bought, and no
# amount of matching will conjure it.
attribution = store.connection.execute(
"SELECT COUNT(*) AS pairs, COALESCE(SUM(plays), 0) AS plays FROM scrobble_key k"
" WHERE k.track_id IS NULL"
" AND EXISTS (SELECT 1 FROM lidarr_artist a WHERE a.norm_name = k.norm_artist)"
" AND EXISTS (SELECT 1 FROM lidarr_track t WHERE t.norm_title = k.norm_track)"
).fetchone()
logger.info(
" of those, the title exists under another artist: %d pairs, %d plays"
" -- attribution disagreements, and the genuine misses",
attribution["pairs"],
attribution["plays"],
)
logger.info(
" the rest, %d pairs, %d plays: you own the artist but not the track",
suspect["pairs"] - attribution["pairs"],
suspect["plays"] - attribution["plays"],
)
mismatched = store.connection.execute(
"SELECT k.artist, k.track, k.plays,"
" (SELECT a.name FROM lidarr_track t"
" JOIN lidarr_artist a ON a.id = t.artist_id"
" WHERE t.norm_title = k.norm_track LIMIT 1) AS filed_under"
" FROM scrobble_key k"
" WHERE k.track_id IS NULL"
" AND EXISTS (SELECT 1 FROM lidarr_artist a WHERE a.norm_name = k.norm_artist)"
" AND EXISTS (SELECT 1 FROM lidarr_track t WHERE t.norm_title = k.norm_track)"
" ORDER BY k.plays DESC, k.artist LIMIT 10"
).fetchall()
for position, row in enumerate(mismatched, start=1):
logger.info(
" attribution %2d: %-45s %4d plays, filed under %s",
position,
f"{row['artist']} - {row['track']}"[:45],
row["plays"],
row["filed_under"],
)
with_files = store.scalar("SELECT COUNT(*) FROM lidarr_track WHERE has_file = 1")
played = store.scalar(
"SELECT COUNT(DISTINCT k.track_id) FROM scrobble_key k"
+1 -1
View File
@@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta"
[project]
name = "music-curator"
version = "0.2.3"
version = "0.3.0"
description = "Ingest a Last.fm listening history and curate a music library from it"
readme = "README.md"
requires-python = ">=3.11"
+76
View File
@@ -687,3 +687,79 @@ def test_lidarr_talks_to_a_real_server_through_keep_alive(http_server):
client.transport.close()
assert server.connections == 1
# Titles taken verbatim from a real coverage report's unmatched list. Each one
# was a genuine miss before the normaliser handled it.
@pytest.mark.parametrize(
("scrobbled", "tagged"),
[
("Self vs Self (feat. In Flames)", "Self vs Self"),
("Grime Battle of Hastings (feat. The Town Crier)", "Grime Battle of Hastings"),
("Gold Dust - Shy FX Re-Edit", "Gold Dust"),
("Back To Your Roots - Friction & K-Tee Remix", "Back To Your Roots"),
("Constellations - Forza Horizon 3 VIP", "Constellations"),
("Everyday (Netsky Remix)", "Everyday"),
("Voodoo People [Pendulum Remix] [Live At Brixton Academy]", "Voodoo People"),
],
)
def test_real_unmatched_titles_now_agree_with_their_tags(scrobbled, tagged):
assert music_curator.normalise(scrobbled) == music_curator.normalise(tagged)
@pytest.mark.parametrize(
"title",
[
"Dancing with Myself",
"(Don't Fear) The Reaper",
"Live and Let Die",
"Radio Ga Ga",
"Editors",
"Mixed Emotions",
"Vipassana",
],
)
def test_the_version_words_do_not_eat_ordinary_titles(title):
"""Every one of these contains a version word and must survive intact."""
assert music_curator.normalise(title) == music_curator.normalise(title.lower())
assert len(music_curator.normalise(title).split()) == len(title.split())
def test_a_bracketed_guest_credit_matches_the_bare_tag(tmp_path):
"""Whitespace-then-feat misses the bracketed form, which is most of them."""
store = indexed(tmp_path, [scrobble_of("Yellowcard", "Here I Am Alive (feat. Someone)")])
method, track_id = verdict(store, "Yellowcard", "Here I Am Alive (feat. Someone)")
assert method == "name"
assert track_id is not None
def test_misses_are_split_by_whether_the_library_holds_the_title(tmp_path):
"""Owning an artist is a weak proxy for owning a track. Counting both as
matcher failures overstates the problem and would over-block the cull."""
store = indexed(
tmp_path,
[
# The library holds "Hells Bells", but under AC/DC, not Yellowcard:
# an attribution disagreement, and a real miss.
scrobble_of("Yellowcard", "Hells Bells"),
# Yellowcard is in the library; this track is not, under any artist.
scrobble_of("Yellowcard", "A Single She Never Bought"),
],
)
def count(extra):
return store.connection.execute(
"SELECT COUNT(*) FROM scrobble_key k WHERE k.track_id IS NULL"
" AND EXISTS (SELECT 1 FROM lidarr_artist a WHERE a.norm_name = k.norm_artist)"
f" {extra}"
).fetchone()[0]
assert count("") == 2
assert count("AND EXISTS (SELECT 1 FROM lidarr_track t WHERE t.norm_title = k.norm_track)") == 1
def test_the_report_survives_the_attribution_split(tmp_path):
store = indexed(tmp_path, [scrobble_of("Yellowcard", "Hells Bells")])
music_curator.report(store, NOW)