11 Commits
Author SHA1 Message Date
lyrathorpe 08f099aaa9 chore(release): v0.3.1 2026-08-24 16:42:06 +00:00
lyrathorpe 719a6372ad Merge pull request 'fix: distinguish a title collision from an attribution miss, and index for it' (#8) from fix/title-collisions-and-index into main
Build and publish container / build (push) Successful in 4m54s
Reviewed-on: #8
2026-08-24 17:37:12 +01:00
Emma Thorpe 61ce751ac5 fix: distinguish a title collision from an attribution miss, and index for it
Build and publish container / build (pull_request) Successful in 5m16s
Two faults in the split added last change, both visible in the first real run.

It reported 700 pairs as attribution disagreements on the strength of the
library holding the same title under a different artist. The examples show what
that actually caught: "Everyday" matched Def Leppard, "Kaleidoscope" matched
Chappell Roan, "Fight for Your Right" matched Motley Crue. Different songs that
happen to share a name. Across fifty thousand tracks that is not an edge case,
it is the common case, and presenting it as a matcher failure argues for exactly
the title-only matching tier that would produce this rubbish on purpose.

The signal for a real attribution miss is narrower: the library's own title
credits the artist the play is filed under, as in "Voodoo People (Pendulum
Remix)" against a scrobble credited to Pendulum. The normalised title has that
suffix stripped -- which is what let the two meet in the first place -- so the
raw title is searched for the name. Collisions are now counted and named
separately, as what they are.

The same query also took forty-seven seconds. lidarr_track was indexed on
(norm_artist, norm_title), which a lookup by title alone cannot use because its
leading column is the artist, so every unmatched key scanned all eighty-four
thousand tracks. Add the index on the title by itself; the query plan changes
from an automatic partial index to a covering one.
2026-08-24 17:35:32 +01:00
lyrathorpe d1c10d32d9 Merge pull request 'ci: build the image once instead of twice' (#7) from ci/one-build-not-two into main
Reviewed-on: #7
2026-08-24 17:31:47 +01:00
lyrathorpe acd042ad8d chore(release): v0.3.0 2026-08-24 16:28:00 +00:00
Emma Thorpe 45d99ff039 ci: build the image once instead of twice
Build and publish container / build (pull_request) Successful in 5m40s
A pull request took roughly eleven minutes to go green, and the log shows one
CACHED line in the whole run. The image was being built twice, in full.

The test stage is built by the runner's docker daemon. The runtime stage was
then built by docker/build-push-action, which runs under a buildx builder that
setup-buildx-action creates in its own container with its own cache. The two
share nothing, so the second build pulled the base image again, ran pip install
again, and exported the layers again -- about four and a half minutes, plus
another thirty-five seconds to boot buildkit. The comment above the test step
claimed those layers were shared, which is what made this look reasonable.

buildx earns that overhead when producing several architectures. This produces
linux/amd64 only, by an explicit decision recorded in the workflow, so it earns
nothing here. Use plain docker build against the same daemon that ran the
tests, and push with docker push. The runtime stage is a strict prefix of the
test stage, so every layer is a cache hit: measured at 1.3 seconds locally
against roughly four and a half minutes in CI.

The remaining time is the runner itself, which is slow in absolute terms --
pytest takes four seconds locally and a hundred and two in CI. That is not
something the workflow can fix.
2026-08-24 17:24:56 +01:00
lyrathorpe 54e19992e4 Merge pull request 'feat: split unmatched listening by whether the library holds the title' (#6) from diag/split-misses-by-ownership into main
Build and publish container / build (push) Successful in 9m7s
Reviewed-on: #6
2026-08-24 17:19:09 +01:00
Emma Thorpe 2886c02a2e feat: split unmatched listening by whether the library holds the title
Build and publish container / build (pull_request) Successful in 8m41s
"Unmatched by an artist the library holds" was presented as the matcher's
misses. On real data it is not: owning one album by an artist says nothing about
owning a particular single of theirs, and most of that figure turned out to be
drum and bass tracks streamed but never bought.

Split it in two. A title the library holds under some other artist is an
attribution disagreement -- a remixer credited as the artist, a guest billed as
one -- and is a genuine miss worth fixing; the report now names the artist the
library files it under, which is the information needed to judge it. A title the
library does not hold at all under any artist was never bought, and no
improvement to matching will conjure it.

The distinction matters beyond presentation. That figure is the gate on the
cull, and a gate computed from a number that overstates the failure rate blocks
work that is actually safe to do.
2026-08-24 17:15:56 +01:00
lyrathorpe 56a06a6cd5 chore(release): v0.2.4 2026-08-24 16:09:38 +00:00
lyrathorpe a3d0689c0a Merge pull request 'fix: match bracketed guest credits and hyphenated version suffixes' (#5) from fix/normalise-bracketed-credits-and-dash-suffixes into main
Build and publish container / build (push) Successful in 11m42s
Reviewed-on: #5
2026-08-24 16:58:05 +01:00
Emma Thorpe 7588fee302 fix: match bracketed guest credits and hyphenated version suffixes
Build and publish container / build (pull_request) Successful in 10m45s
Two normalisation faults, both found by running a real coverage report's
unmatched list back through the normaliser. Between them they account for five
of the fifteen worst misses by play count.

The guest-credit pattern required whitespace immediately before the word, so it
caught "Yellowcard feat. Tay Jardine" but missed "Self vs Self (feat. In
Flames)" -- and the bracketed form is the more common of the two. An opening
bracket is now allowed in that position.

The trailing-version pattern matched the suffix as a run of non-hyphens, which
cannot cross a hyphen inside the suffix itself: "Gold Dust - Shy FX Re-Edit" and
"Back To Your Roots - Friction & K-Tee Remix" both survived untouched. Matched
lazily instead.

Four version words are added for how drum and bass marks its variants: vip,
bootleg, rework, extended. They only apply inside a bracket or after a trailing
dash, so the exposure is small, and ordinary titles carrying those words --
Editors, Mixed Emotions, Radio Ga Ga, Live and Let Die -- are pinned as tests
against exactly that.

The report also gains the figures that explain why the MBID tier contributes so
little. Two thirds of scrobbles carry a recording id, and only a twentieth of
them join on one: MusicBrainz holds a separate recording per release, and the
two sides rarely choose the same one. Counting the pairs that carried an id and
matched on name anyway measures that disagreement directly, and settles that the
weakness is not a bug in the join.
2026-08-24 16:56:13 +01:00
5 changed files with 310 additions and 39 deletions
+29 -19
View File
@@ -45,7 +45,8 @@ jobs:
# The suite runs inside the image, against the interpreter that ships, # The suite runs inside the image, against the interpreter that ships,
# rather than against whatever the runner happens to provide. A failing # rather than against whatever the runner happens to provide. A failing
# test fails the build. Layers are shared with the push build below. # test fails the build. The runtime stage below is built from the same
# daemon afterwards, so its layers are already in cache.
- name: Run the test suite inside the image - name: Run the test suite inside the image
run: docker build --target test -t music-curator:test . run: docker build --target test -t music-curator:test .
@@ -124,9 +125,6 @@ jobs:
echo "release=${release}" >> "$GITHUB_OUTPUT" echo "release=${release}" >> "$GITHUB_OUTPUT"
echo "Computed bump=${bump}, release=${release}, base=${base}" echo "Computed bump=${bump}, release=${release}, base=${base}"
- name: Set up Buildx
uses: docker/setup-buildx-action@d7f5e7f509e45cec5c76c4d5afdd7de93d0b3df5 # v4
- name: Log in to the Gitea container registry - name: Log in to the Gitea container registry
if: github.event_name != 'pull_request' if: github.event_name != 'pull_request'
uses: docker/login-action@650006c6eb7dba73a995cc03b0b2d7f5ca915bee # v4 uses: docker/login-action@650006c6eb7dba73a995cc03b0b2d7f5ca915bee # v4
@@ -135,21 +133,33 @@ jobs:
username: ${{ github.repository_owner }} username: ${{ github.repository_owner }}
password: ${{ secrets.PACKAGES_TOKEN }} password: ${{ secrets.PACKAGES_TOKEN }}
- name: Build and push # Plain `docker build` rather than buildx. buildx boots its own buildkit
uses: docker/build-push-action@f9f3042f7e2789586610d6e8b85c8f03e5195baf # v7 # in a container with a cache of its own, so it shared nothing with the
with: # test build above and rebuilt the image from the base image up -- two
context: . # full builds per run. It earns that cost when building for several
# Without this the last stage in the Dockerfile -- the test stage -- # platforms; this only ever targets the amd64 NAS, so it does not.
# would be what gets published. #
target: runtime # `--target runtime` is a strict prefix of the test stage, so every layer
# The NAS is the only host this runs on. Building arm64 as well would # is already in the daemon's cache and this resolves in seconds.
# mean emulating it under QEMU for no consumer. - name: Build the runtime image
platforms: linux/amd64 run: |
push: ${{ github.event_name != 'pull_request' }} set -euo pipefail
tags: ${{ steps.version.outputs.tags }} tags=()
labels: | while IFS= read -r tag; do
org.opencontainers.image.source=${{ github.server_url }}/${{ github.repository }} [ -n "$tag" ] && tags+=(-t "$tag")
org.opencontainers.image.revision=${{ github.sha }} done <<< "${{ steps.version.outputs.tags }}"
docker build --target runtime \
--label "org.opencontainers.image.source=${GITHUB_SERVER_URL}/${GITHUB_REPOSITORY}" \
--label "org.opencontainers.image.revision=${GITHUB_SHA}" \
"${tags[@]}" .
- name: Push
if: github.event_name != 'pull_request'
run: |
set -euo pipefail
while IFS= read -r tag; do
[ -n "$tag" ] && docker push "$tag"
done <<< "${{ steps.version.outputs.tags }}"
# Record the release: write the computed version into pyproject.toml, then # Record the release: write the computed version into pyproject.toml, then
# commit and tag it, so the packaging metadata always matches the release # commit and tag it, so the packaging metadata always matches the release
+24 -5
View File
@@ -29,7 +29,14 @@ Two tiers, and no third.
| `name` | Normalised artist and title | Everything the first tier could not carry | | `name` | Normalised artist and title | Everything the first tier could not carry |
| `none` | — | Recorded as a miss, never guessed at | | `none` | — | Recorded as a miss, never guessed at |
The normalisation is the load-bearing part, because the two sides disagree in The name tier does almost all of the work. A recording MBID is exact when it
lands, but MusicBrainz holds a separate recording per release, and Last.fm and
Lidarr rarely pick the same one: on a real library, two thirds of scrobbles
carry a recording id and barely a twentieth of them join on it. The report
counts how many carried an id and matched on name anyway, which is the measure
of that disagreement.
So the normalisation is the load-bearing part, because the two sides disagree in
predictable ways. It folds case and accents, drops guest credits (`Yellowcard predictable ways. It folds case and accents, drops guest credits (`Yellowcard
feat. Tay Jardine` against a tag of `Yellowcard`), strips a trailing feat. Tay Jardine` against a tag of `Yellowcard`), strips a trailing
version suffix (`(Remastered 2011)`, `- Live`), expands `&`, and removes a version suffix (`(Remastered 2011)`, `- Live`), expands `&`, and removes a
@@ -102,10 +109,22 @@ matcher. The line to watch is:
unmatched by an artist the library holds: N pairs, M plays unmatched by an artist the library holds: N pairs, M plays
``` ```
That is a track that was played, sitting beside a file it should have matched. That is a track that was played by an artist the library holds. It is then
Those are the matcher's real misses, and every one is a candidate for being split three ways, because owning an artist is a weak proxy for owning a track
wrongly called cold in stage four. The report lists the worst fifteen by play and a shared title is a weak proxy for a shared song:
count so they can be eyeballed.
- **the library's own title credits the scrobbled artist** — `Voodoo People
(Pendulum Remix)` against a play credited to Pendulum. Same song, filed
under the original artist. These are the genuine misses.
- **the same title under an unrelated artist** — a collision, not a miss.
Across fifty thousand tracks these are constant: `Everyday` is Rusko and
also Def Leppard, `Kaleidoscope` is Delta Heavy and also Chappell Roan.
Matching on title alone would be far worse than missing them, which is why
there is no such tier.
- **the title is nowhere in the library** — never bought.
Only the first is worth chasing. Counting all three as matcher failures
overstates the problem and would over-block the cull.
## How the ingest works ## How the ingest works
+112 -8
View File
@@ -141,6 +141,11 @@ CREATE TABLE IF NOT EXISTS lidarr_track (
); );
CREATE INDEX IF NOT EXISTS lidarr_track_recording ON lidarr_track (recording_mbid); CREATE INDEX IF NOT EXISTS lidarr_track_recording ON lidarr_track (recording_mbid);
CREATE INDEX IF NOT EXISTS lidarr_track_norm ON lidarr_track (norm_artist, norm_title); CREATE INDEX IF NOT EXISTS lidarr_track_norm ON lidarr_track (norm_artist, norm_title);
-- Separate from the composite above, which a lookup by title alone cannot use:
-- its leading column is the artist. The report searches by title on its own, and
-- without this it scans every track for every unmatched key -- forty-seven
-- seconds on a library of eighty-four thousand.
CREATE INDEX IF NOT EXISTS lidarr_track_title ON lidarr_track (norm_title);
CREATE INDEX IF NOT EXISTS lidarr_track_album ON lidarr_track (album_id); CREATE INDEX IF NOT EXISTS lidarr_track_album ON lidarr_track (album_id);
-- One row per distinct thing listened to, with the verdict on whether it could -- One row per distinct thing listened to, with the verdict on whether it could
@@ -187,18 +192,28 @@ VERSION_WORDS = (
"anniversary", "anniversary",
"reissue", "reissue",
"instrumental", "instrumental",
# Drum and bass and its neighbours mark versions their own way.
"vip",
"bootleg",
"rework",
"extended",
) )
_VERSIONS = "|".join(VERSION_WORDS) _VERSIONS = "|".join(VERSION_WORDS)
BRACKETED_VERSION = re.compile( BRACKETED_VERSION = re.compile(
rf"\s*[\(\[][^\)\]]*\b(?:{_VERSIONS})\b[^\)\]]*[\)\]]\s*$", re.IGNORECASE rf"\s*[\(\[][^\)\]]*\b(?:{_VERSIONS})\b[^\)\]]*[\)\]]\s*$", re.IGNORECASE
) )
TRAILING_VERSION = re.compile(rf"\s+-\s+[^-]*\b(?:{_VERSIONS})\b.*$", re.IGNORECASE) # The suffix is matched lazily rather than as a run of non-hyphens, because the
# thing being stripped frequently contains hyphens of its own -- "Gold Dust -
# Shy FX Re-Edit", "Back To Your Roots - Friction & K-Tee Remix".
TRAILING_VERSION = re.compile(rf"\s+-\s+.*?\b(?:{_VERSIONS})\b.*$", re.IGNORECASE)
# Last.fm routinely carries the guest credit in the artist field where the file # Last.fm routinely carries the guest credit where the file tag holds only the
# tag holds only the primary artist -- "Yellowcard feat. Tay Jardine" against a # primary artist -- "Yellowcard feat. Tay Jardine" against a tag of
# tag of "Yellowcard". `with` is deliberately absent: it appears in far too many # "Yellowcard", or "Self vs Self (feat. In Flames)" against "Self vs Self". The
# real titles to cut on sight. # opening bracket has to be allowed for: requiring whitespace immediately before
GUEST_CREDIT = re.compile(r"\s+(?:feat|ft|featuring)\b.*$", re.IGNORECASE) # the word misses every bracketed credit, which is most of them. `with` is
# deliberately absent -- it appears in far too many real titles to cut on sight.
GUEST_CREDIT = re.compile(r"[\s(\[]+(?:feat|ft|featuring)\b.*$", re.IGNORECASE)
LEADING_ARTICLE = re.compile(r"^the\s+") LEADING_ARTICLE = re.compile(r"^the\s+")
# Deleted rather than spaced, so "Don't" and "Dont" agree. Every other mark # Deleted rather than spaced, so "Don't" and "Dont" agree. Every other mark
@@ -1078,18 +1093,107 @@ def coverage_report(store):
row["plays"], row["plays"],
) )
# Why the mbid tier performs the way it does. A recording id on both sides
# that still fails to join means the two disagree about which recording the
# song is -- MusicBrainz holds a separate recording per release, and Last.fm
# and Lidarr need not have picked the same one. That is not a fault to fix
# in the matcher; it is the reason the name tier has to carry the load.
library_with_mbid = store.scalar(
"SELECT COUNT(*) FROM lidarr_track WHERE recording_mbid IS NOT NULL"
)
logger.info(
"library tracks carrying a recording MBID: %d of %d (%.1f%%)",
library_with_mbid,
tracks,
100 * library_with_mbid / tracks,
)
disagreed = store.connection.execute(
"SELECT COUNT(*) AS pairs, COALESCE(SUM(plays), 0) AS plays FROM scrobble_key"
" WHERE track_mbid IS NOT NULL AND method = 'name'"
).fetchone()
logger.info(
"carried a recording MBID, joined on name instead: %d pairs, %d plays"
" -- both sides know the song, they disagree on which recording it is",
disagreed["pairs"],
disagreed["plays"],
)
suspect = store.connection.execute( suspect = store.connection.execute(
"SELECT COUNT(*) AS pairs, COALESCE(SUM(plays), 0) AS plays FROM scrobble_key k" "SELECT COUNT(*) AS pairs, COALESCE(SUM(plays), 0) AS plays FROM scrobble_key k"
" WHERE k.track_id IS NULL" " WHERE k.track_id IS NULL"
" AND EXISTS (SELECT 1 FROM lidarr_artist a WHERE a.norm_name = k.norm_artist)" " AND EXISTS (SELECT 1 FROM lidarr_artist a WHERE a.norm_name = k.norm_artist)"
).fetchone() ).fetchone()
logger.info( logger.info(
"unmatched by an artist the library holds: %d pairs, %d plays -- these are the" "unmatched by an artist the library holds: %d pairs, %d plays",
" matcher's misses, not music you do not own",
suspect["pairs"], suspect["pairs"],
suspect["plays"], suspect["plays"],
) )
# Owning an artist is a weak proxy for owning a track, so that figure alone
# overstates the matcher's failings. Split it -- but not on the title alone.
# Across fifty thousand tracks, titles collide constantly: "Everyday" is
# Rusko and also Def Leppard, "Kaleidoscope" is Delta Heavy and also
# Chappell Roan. Matching those would be worse than missing them.
#
# The signal for a genuine attribution miss is that the library's own title
# credits the artist the scrobble is filed under -- "Voodoo People (Pendulum
# Remix)" against a play credited to Pendulum. The normalised title has that
# suffix stripped, which is exactly what let them meet, so the raw one has to
# be searched for the name.
attribution = store.connection.execute(
"SELECT COUNT(*) AS pairs, COALESCE(SUM(plays), 0) AS plays FROM scrobble_key k"
" WHERE k.track_id IS NULL"
" AND EXISTS (SELECT 1 FROM lidarr_artist a WHERE a.norm_name = k.norm_artist)"
" AND EXISTS (SELECT 1 FROM lidarr_track t"
" WHERE t.norm_title = k.norm_track"
" AND instr(lower(t.title), lower(k.artist)) > 0)"
).fetchone()
collision = store.connection.execute(
"SELECT COUNT(*) AS pairs, COALESCE(SUM(plays), 0) AS plays FROM scrobble_key k"
" WHERE k.track_id IS NULL"
" AND EXISTS (SELECT 1 FROM lidarr_artist a WHERE a.norm_name = k.norm_artist)"
" AND EXISTS (SELECT 1 FROM lidarr_track t WHERE t.norm_title = k.norm_track)"
).fetchone()
logger.info(
" the library's title credits the scrobbled artist: %d pairs, %d plays"
" -- remixes and guest spots, and the genuine misses",
attribution["pairs"],
attribution["plays"],
)
logger.info(
" same title under an unrelated artist: %d pairs, %d plays"
" -- title collisions, not misses; matching these would be a mistake",
collision["pairs"] - attribution["pairs"],
collision["plays"] - attribution["plays"],
)
logger.info(
" the rest, %d pairs, %d plays: you own the artist but not the track",
suspect["pairs"] - collision["pairs"],
suspect["plays"] - collision["plays"],
)
mismatched = store.connection.execute(
"SELECT k.artist, k.track, k.plays,"
" (SELECT t.title FROM lidarr_track t"
" WHERE t.norm_title = k.norm_track"
" AND instr(lower(t.title), lower(k.artist)) > 0 LIMIT 1) AS library_title"
" FROM scrobble_key k"
" WHERE k.track_id IS NULL"
" AND EXISTS (SELECT 1 FROM lidarr_artist a WHERE a.norm_name = k.norm_artist)"
" AND EXISTS (SELECT 1 FROM lidarr_track t"
" WHERE t.norm_title = k.norm_track"
" AND instr(lower(t.title), lower(k.artist)) > 0)"
" ORDER BY k.plays DESC, k.artist LIMIT 10"
).fetchall()
for position, row in enumerate(mismatched, start=1):
logger.info(
" attribution %2d: %-45s %4d plays, library has %r",
position,
f"{row['artist']} - {row['track']}"[:45],
row["plays"],
row["library_title"],
)
with_files = store.scalar("SELECT COUNT(*) FROM lidarr_track WHERE has_file = 1") with_files = store.scalar("SELECT COUNT(*) FROM lidarr_track WHERE has_file = 1")
played = store.scalar( played = store.scalar(
"SELECT COUNT(DISTINCT k.track_id) FROM scrobble_key k" "SELECT COUNT(DISTINCT k.track_id) FROM scrobble_key k"
+1 -1
View File
@@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta"
[project] [project]
name = "music-curator" name = "music-curator"
version = "0.2.3" version = "0.3.1"
description = "Ingest a Last.fm listening history and curate a music library from it" description = "Ingest a Last.fm listening history and curate a music library from it"
readme = "README.md" readme = "README.md"
requires-python = ">=3.11" requires-python = ">=3.11"
+144 -6
View File
@@ -55,8 +55,22 @@ LIBRARY = [
{ {
"name": "The Prodigy", "name": "The Prodigy",
"albums": [ "albums": [
{"title": "The Fat of the Land", "tracks": [{"title": "Breathe (Remastered)"}]} {
"title": "The Fat of the Land",
"tracks": [
{"title": "Breathe (Remastered)"},
# Credits its remixer in the title, which is the only signal
# separating a real attribution miss from a title collision.
{"title": "Voodoo People (Pendulum Remix)"},
], ],
}
],
},
# Held by the library in its own right, which is what puts its scrobbles
# inside the "artist the library holds" filter at all.
{
"name": "Pendulum",
"albums": [{"title": "Immersion", "tracks": [{"title": "Watercolour"}]}],
}, },
] ]
@@ -393,7 +407,7 @@ def test_keys_carry_the_play_count_and_the_span(tmp_path):
def test_re_indexing_drops_what_lidarr_no_longer_has(tmp_path): def test_re_indexing_drops_what_lidarr_no_longer_has(tmp_path):
"""The index is Lidarr's mirror, not an accumulation of everything ever seen.""" """The index is Lidarr's mirror, not an accumulation of everything ever seen."""
store = indexed(tmp_path, [scrobble_of("AC/DC", "Hells Bells")]) store = indexed(tmp_path, [scrobble_of("AC/DC", "Hells Bells")])
assert store.scalar("SELECT COUNT(*) FROM lidarr_artist") == 3 assert store.scalar("SELECT COUNT(*) FROM lidarr_artist") == 4
music_curator.index_library( music_curator.index_library(
music_curator.Lidarr("http://lidarr", "key", transport=FakeLidarr(LIBRARY[:1])), store music_curator.Lidarr("http://lidarr", "key", transport=FakeLidarr(LIBRARY[:1])), store
@@ -512,7 +526,7 @@ def test_albums_come_from_the_unfiltered_endpoint(tmp_path):
album_calls = [query for path, query in api.calls if path == "album"] album_calls = [query for path, query in api.calls if path == "album"]
assert album_calls == [{}] assert album_calls == [{}]
assert store.scalar("SELECT COUNT(*) FROM lidarr_album") == 3 assert store.scalar("SELECT COUNT(*) FROM lidarr_album") == 4
def test_a_bad_album_falls_back_to_asking_per_artist(tmp_path): def test_a_bad_album_falls_back_to_asking_per_artist(tmp_path):
@@ -526,7 +540,7 @@ def test_a_bad_album_falls_back_to_asking_per_artist(tmp_path):
assert store.scalar("SELECT COUNT(*) FROM lidarr_album WHERE artist_id = 2") == 0 assert store.scalar("SELECT COUNT(*) FROM lidarr_album WHERE artist_id = 2") == 0
# The other two artists keep their albums. # The other two artists keep their albums.
assert store.scalar("SELECT COUNT(*) FROM lidarr_album") == 2 assert store.scalar("SELECT COUNT(*) FROM lidarr_album") == 3
assert store.get_state("index_albums_skipped") == "1" assert store.get_state("index_albums_skipped") == "1"
@@ -549,10 +563,10 @@ def test_an_artist_lidarr_cannot_serve_does_not_kill_the_index(tmp_path):
music_curator.index_library(music_curator.Lidarr("http://lidarr", "key", transport=api), store) music_curator.index_library(music_curator.Lidarr("http://lidarr", "key", transport=api), store)
assert store.scalar("SELECT COUNT(*) FROM lidarr_artist") == 3 assert store.scalar("SELECT COUNT(*) FROM lidarr_artist") == 4
assert store.scalar("SELECT COUNT(*) FROM lidarr_track WHERE artist_id = 2") == 0 assert store.scalar("SELECT COUNT(*) FROM lidarr_track WHERE artist_id = 2") == 0
# The other two artists are indexed in full. # The other two artists are indexed in full.
assert store.scalar("SELECT COUNT(*) FROM lidarr_track") == 3 assert store.scalar("SELECT COUNT(*) FROM lidarr_track") == 5
assert store.get_state("index_skipped") == "1" assert store.get_state("index_skipped") == "1"
@@ -687,3 +701,127 @@ def test_lidarr_talks_to_a_real_server_through_keep_alive(http_server):
client.transport.close() client.transport.close()
assert server.connections == 1 assert server.connections == 1
# Titles taken verbatim from a real coverage report's unmatched list. Each one
# was a genuine miss before the normaliser handled it.
@pytest.mark.parametrize(
("scrobbled", "tagged"),
[
("Self vs Self (feat. In Flames)", "Self vs Self"),
("Grime Battle of Hastings (feat. The Town Crier)", "Grime Battle of Hastings"),
("Gold Dust - Shy FX Re-Edit", "Gold Dust"),
("Back To Your Roots - Friction & K-Tee Remix", "Back To Your Roots"),
("Constellations - Forza Horizon 3 VIP", "Constellations"),
("Everyday (Netsky Remix)", "Everyday"),
("Voodoo People [Pendulum Remix] [Live At Brixton Academy]", "Voodoo People"),
],
)
def test_real_unmatched_titles_now_agree_with_their_tags(scrobbled, tagged):
assert music_curator.normalise(scrobbled) == music_curator.normalise(tagged)
@pytest.mark.parametrize(
"title",
[
"Dancing with Myself",
"(Don't Fear) The Reaper",
"Live and Let Die",
"Radio Ga Ga",
"Editors",
"Mixed Emotions",
"Vipassana",
],
)
def test_the_version_words_do_not_eat_ordinary_titles(title):
"""Every one of these contains a version word and must survive intact."""
assert music_curator.normalise(title) == music_curator.normalise(title.lower())
assert len(music_curator.normalise(title).split()) == len(title.split())
def test_a_bracketed_guest_credit_matches_the_bare_tag(tmp_path):
"""Whitespace-then-feat misses the bracketed form, which is most of them."""
store = indexed(tmp_path, [scrobble_of("Yellowcard", "Here I Am Alive (feat. Someone)")])
method, track_id = verdict(store, "Yellowcard", "Here I Am Alive (feat. Someone)")
assert method == "name"
assert track_id is not None
def test_misses_are_split_by_whether_the_library_holds_the_title(tmp_path):
"""Owning an artist is a weak proxy for owning a track. Counting both as
matcher failures overstates the problem and would over-block the cull."""
store = indexed(
tmp_path,
[
# The library holds "Hells Bells", but under AC/DC, not Yellowcard:
# an attribution disagreement, and a real miss.
scrobble_of("Yellowcard", "Hells Bells"),
# Yellowcard is in the library; this track is not, under any artist.
scrobble_of("Yellowcard", "A Single She Never Bought"),
],
)
def count(extra):
return store.connection.execute(
"SELECT COUNT(*) FROM scrobble_key k WHERE k.track_id IS NULL"
" AND EXISTS (SELECT 1 FROM lidarr_artist a WHERE a.norm_name = k.norm_artist)"
f" {extra}"
).fetchone()[0]
assert count("") == 2
assert count("AND EXISTS (SELECT 1 FROM lidarr_track t WHERE t.norm_title = k.norm_track)") == 1
def test_the_report_survives_the_attribution_split(tmp_path):
store = indexed(tmp_path, [scrobble_of("Yellowcard", "Hells Bells")])
music_curator.report(store, NOW)
def attribution_pairs(store):
"""Unmatched pairs where the library's own title credits the scrobbled artist."""
return [
row["artist"]
for row in store.connection.execute(
"SELECT k.artist FROM scrobble_key k WHERE k.track_id IS NULL"
" AND EXISTS (SELECT 1 FROM lidarr_artist a WHERE a.norm_name = k.norm_artist)"
" AND EXISTS (SELECT 1 FROM lidarr_track t"
" WHERE t.norm_title = k.norm_track"
" AND instr(lower(t.title), lower(k.artist)) > 0)"
)
]
def test_a_shared_title_is_not_an_attribution_miss(tmp_path):
"""Across fifty thousand tracks, titles collide constantly: "Everyday" is
Rusko and also Def Leppard. Matching those would be worse than missing."""
store = indexed(tmp_path, [scrobble_of("Yellowcard", "Hells Bells")])
# The library holds "Hells Bells", by AC/DC, and its title says nothing
# about Yellowcard. A collision, not a miss.
assert verdict(store, "Yellowcard", "Hells Bells") == ("none", None)
assert attribution_pairs(store) == []
def test_a_remix_credited_in_the_library_title_is_an_attribution_miss(tmp_path):
"""The library has "Voodoo People (Pendulum Remix)" under The Prodigy; the
scrobble credits Pendulum. Same song, different filing."""
store = indexed(tmp_path, [scrobble_of("Pendulum", "Voodoo People")])
assert attribution_pairs(store) == ["Pendulum"]
def test_the_title_index_is_used_for_the_report_lookup(tmp_path):
"""Without it the report scans every track for every unmatched key: forty-
seven seconds on a real library."""
store = indexed(tmp_path, [])
plan = "\n".join(
row[-1]
for row in store.connection.execute(
"EXPLAIN QUERY PLAN SELECT 1 FROM scrobble_key k WHERE k.track_id IS NULL"
" AND EXISTS (SELECT 1 FROM lidarr_track t WHERE t.norm_title = k.norm_track)"
)
)
assert "lidarr_track_title" in plan, plan