5 Commits
Author SHA1 Message Date
lyrathorpe d1c10d32d9 Merge pull request 'ci: build the image once instead of twice' (#7) from ci/one-build-not-two into main
Reviewed-on: #7
2026-08-24 17:31:47 +01:00
lyrathorpe acd042ad8d chore(release): v0.3.0 2026-08-24 16:28:00 +00:00
Emma Thorpe 45d99ff039 ci: build the image once instead of twice
Build and publish container / build (pull_request) Successful in 5m40s
A pull request took roughly eleven minutes to go green, and the log shows one
CACHED line in the whole run. The image was being built twice, in full.

The test stage is built by the runner's docker daemon. The runtime stage was
then built by docker/build-push-action, which runs under a buildx builder that
setup-buildx-action creates in its own container with its own cache. The two
share nothing, so the second build pulled the base image again, ran pip install
again, and exported the layers again -- about four and a half minutes, plus
another thirty-five seconds to boot buildkit. The comment above the test step
claimed those layers were shared, which is what made this look reasonable.

buildx earns that overhead when producing several architectures. This produces
linux/amd64 only, by an explicit decision recorded in the workflow, so it earns
nothing here. Use plain docker build against the same daemon that ran the
tests, and push with docker push. The runtime stage is a strict prefix of the
test stage, so every layer is a cache hit: measured at 1.3 seconds locally
against roughly four and a half minutes in CI.

The remaining time is the runner itself, which is slow in absolute terms --
pytest takes four seconds locally and a hundred and two in CI. That is not
something the workflow can fix.
2026-08-24 17:24:56 +01:00
lyrathorpe 54e19992e4 Merge pull request 'feat: split unmatched listening by whether the library holds the title' (#6) from diag/split-misses-by-ownership into main
Build and publish container / build (push) Successful in 9m7s
Reviewed-on: #6
2026-08-24 17:19:09 +01:00
Emma Thorpe 2886c02a2e feat: split unmatched listening by whether the library holds the title
Build and publish container / build (pull_request) Successful in 8m41s
"Unmatched by an artist the library holds" was presented as the matcher's
misses. On real data it is not: owning one album by an artist says nothing about
owning a particular single of theirs, and most of that figure turned out to be
drum and bass tracks streamed but never bought.

Split it in two. A title the library holds under some other artist is an
attribution disagreement -- a remixer credited as the artist, a guest billed as
one -- and is a genuine miss worth fixing; the report now names the artist the
library files it under, which is the information needed to judge it. A title the
library does not hold at all under any artist was never bought, and no
improvement to matching will conjure it.

The distinction matters beyond presentation. That figure is the gate on the
cull, and a gate computed from a number that overstates the failure rate blocks
work that is actually safe to do.
2026-08-24 17:15:56 +01:00
5 changed files with 118 additions and 26 deletions
+29 -19
View File
@@ -45,7 +45,8 @@ jobs:
# The suite runs inside the image, against the interpreter that ships,
# rather than against whatever the runner happens to provide. A failing
# test fails the build. Layers are shared with the push build below.
# test fails the build. The runtime stage below is built from the same
# daemon afterwards, so its layers are already in cache.
- name: Run the test suite inside the image
run: docker build --target test -t music-curator:test .
@@ -124,9 +125,6 @@ jobs:
echo "release=${release}" >> "$GITHUB_OUTPUT"
echo "Computed bump=${bump}, release=${release}, base=${base}"
- name: Set up Buildx
uses: docker/setup-buildx-action@d7f5e7f509e45cec5c76c4d5afdd7de93d0b3df5 # v4
- name: Log in to the Gitea container registry
if: github.event_name != 'pull_request'
uses: docker/login-action@650006c6eb7dba73a995cc03b0b2d7f5ca915bee # v4
@@ -135,21 +133,33 @@ jobs:
username: ${{ github.repository_owner }}
password: ${{ secrets.PACKAGES_TOKEN }}
- name: Build and push
uses: docker/build-push-action@f9f3042f7e2789586610d6e8b85c8f03e5195baf # v7
with:
context: .
# Without this the last stage in the Dockerfile -- the test stage --
# would be what gets published.
target: runtime
# The NAS is the only host this runs on. Building arm64 as well would
# mean emulating it under QEMU for no consumer.
platforms: linux/amd64
push: ${{ github.event_name != 'pull_request' }}
tags: ${{ steps.version.outputs.tags }}
labels: |
org.opencontainers.image.source=${{ github.server_url }}/${{ github.repository }}
org.opencontainers.image.revision=${{ github.sha }}
# Plain `docker build` rather than buildx. buildx boots its own buildkit
# in a container with a cache of its own, so it shared nothing with the
# test build above and rebuilt the image from the base image up -- two
# full builds per run. It earns that cost when building for several
# platforms; this only ever targets the amd64 NAS, so it does not.
#
# `--target runtime` is a strict prefix of the test stage, so every layer
# is already in the daemon's cache and this resolves in seconds.
- name: Build the runtime image
run: |
set -euo pipefail
tags=()
while IFS= read -r tag; do
[ -n "$tag" ] && tags+=(-t "$tag")
done <<< "${{ steps.version.outputs.tags }}"
docker build --target runtime \
--label "org.opencontainers.image.source=${GITHUB_SERVER_URL}/${GITHUB_REPOSITORY}" \
--label "org.opencontainers.image.revision=${GITHUB_SHA}" \
"${tags[@]}" .
- name: Push
if: github.event_name != 'pull_request'
run: |
set -euo pipefail
while IFS= read -r tag; do
[ -n "$tag" ] && docker push "$tag"
done <<< "${{ steps.version.outputs.tags }}"
# Record the release: write the computed version into pyproject.toml, then
# commit and tag it, so the packaging metadata always matches the release
+12 -4
View File
@@ -109,10 +109,18 @@ matcher. The line to watch is:
unmatched by an artist the library holds: N pairs, M plays
```
That is a track that was played, sitting beside a file it should have matched.
Those are the matcher's real misses, and every one is a candidate for being
wrongly called cold in stage four. The report lists the worst fifteen by play
count so they can be eyeballed.
That is a track that was played by an artist the library holds. It is then
split again, because owning an artist is a weak proxy for owning a track:
- **the title exists under another artist** — an attribution disagreement, a
remixer or a guest billed as the artist. These are the genuine misses, and
each is a candidate for being wrongly called cold in stage four. The report
names the artist the library files them under.
- **the title is nowhere in the library** — never bought. No amount of matching
conjures a file that does not exist.
Counting both as matcher failures overstates the problem and would over-block
the cull. The report lists the worst of each by play count.
## How the ingest works
+45 -2
View File
@@ -1119,12 +1119,55 @@ def coverage_report(store):
" AND EXISTS (SELECT 1 FROM lidarr_artist a WHERE a.norm_name = k.norm_artist)"
).fetchone()
logger.info(
"unmatched by an artist the library holds: %d pairs, %d plays -- these are the"
" matcher's misses, not music you do not own",
"unmatched by an artist the library holds: %d pairs, %d plays",
suspect["pairs"],
suspect["plays"],
)
# Owning an artist is a weak proxy for owning a track, so that figure alone
# overstates the matcher's failings. Split it. A title the library holds
# under some other artist is an attribution disagreement -- a remixer
# credited as the artist, a guest billed as one -- and is a real miss. A
# title the library does not hold at all was simply never bought, and no
# amount of matching will conjure it.
attribution = store.connection.execute(
"SELECT COUNT(*) AS pairs, COALESCE(SUM(plays), 0) AS plays FROM scrobble_key k"
" WHERE k.track_id IS NULL"
" AND EXISTS (SELECT 1 FROM lidarr_artist a WHERE a.norm_name = k.norm_artist)"
" AND EXISTS (SELECT 1 FROM lidarr_track t WHERE t.norm_title = k.norm_track)"
).fetchone()
logger.info(
" of those, the title exists under another artist: %d pairs, %d plays"
" -- attribution disagreements, and the genuine misses",
attribution["pairs"],
attribution["plays"],
)
logger.info(
" the rest, %d pairs, %d plays: you own the artist but not the track",
suspect["pairs"] - attribution["pairs"],
suspect["plays"] - attribution["plays"],
)
mismatched = store.connection.execute(
"SELECT k.artist, k.track, k.plays,"
" (SELECT a.name FROM lidarr_track t"
" JOIN lidarr_artist a ON a.id = t.artist_id"
" WHERE t.norm_title = k.norm_track LIMIT 1) AS filed_under"
" FROM scrobble_key k"
" WHERE k.track_id IS NULL"
" AND EXISTS (SELECT 1 FROM lidarr_artist a WHERE a.norm_name = k.norm_artist)"
" AND EXISTS (SELECT 1 FROM lidarr_track t WHERE t.norm_title = k.norm_track)"
" ORDER BY k.plays DESC, k.artist LIMIT 10"
).fetchall()
for position, row in enumerate(mismatched, start=1):
logger.info(
" attribution %2d: %-45s %4d plays, filed under %s",
position,
f"{row['artist']} - {row['track']}"[:45],
row["plays"],
row["filed_under"],
)
with_files = store.scalar("SELECT COUNT(*) FROM lidarr_track WHERE has_file = 1")
played = store.scalar(
"SELECT COUNT(DISTINCT k.track_id) FROM scrobble_key k"
+1 -1
View File
@@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta"
[project]
name = "music-curator"
version = "0.2.4"
version = "0.3.0"
description = "Ingest a Last.fm listening history and curate a music library from it"
readme = "README.md"
requires-python = ">=3.11"
+31
View File
@@ -732,3 +732,34 @@ def test_a_bracketed_guest_credit_matches_the_bare_tag(tmp_path):
method, track_id = verdict(store, "Yellowcard", "Here I Am Alive (feat. Someone)")
assert method == "name"
assert track_id is not None
def test_misses_are_split_by_whether_the_library_holds_the_title(tmp_path):
"""Owning an artist is a weak proxy for owning a track. Counting both as
matcher failures overstates the problem and would over-block the cull."""
store = indexed(
tmp_path,
[
# The library holds "Hells Bells", but under AC/DC, not Yellowcard:
# an attribution disagreement, and a real miss.
scrobble_of("Yellowcard", "Hells Bells"),
# Yellowcard is in the library; this track is not, under any artist.
scrobble_of("Yellowcard", "A Single She Never Bought"),
],
)
def count(extra):
return store.connection.execute(
"SELECT COUNT(*) FROM scrobble_key k WHERE k.track_id IS NULL"
" AND EXISTS (SELECT 1 FROM lidarr_artist a WHERE a.norm_name = k.norm_artist)"
f" {extra}"
).fetchone()[0]
assert count("") == 2
assert count("AND EXISTS (SELECT 1 FROM lidarr_track t WHERE t.norm_title = k.norm_track)") == 1
def test_the_report_survives_the_attribution_split(tmp_path):
store = indexed(tmp_path, [scrobble_of("Yellowcard", "Hells Bells")])
music_curator.report(store, NOW)