Files
music-curator/README.md
T
Emma Thorpe 3e78f8ebd4
Build and publish container / build (pull_request) Successful in 8m1s
fix: index albums through the endpoint Lidarr does not throw from
Indexing fetched albums one artist at a time, and `GET /api/v1/album?artistId=`
is Lidarr's unguarded path. It maps straight from the album service with no
hydration: the mapper then dereferences model.Images and model.SecondaryTypes
without a null check, follows model.Artist?.Value where only the first link is
guarded, and selects the monitored release with SingleOrDefault, which throws
outright when an album has two of them. Any of those is a 500 that aborts the
whole index.

The unfiltered `GET /api/v1/album` builds its own artist and release lookups and
skips an album whose metadata is missing rather than dereferencing it. Use that
instead, once, and group by artistId locally. It is the defensive path and it
costs N fewer requests.

Tracks and files have no unfiltered endpoint -- Lidarr rejects a call with no
filter at all -- so those stay per artist. A failure on one artist now skips
that artist rather than ending the run, but the count is recorded in the store
and the coverage report leads with it: a missing artist makes their played music
look unplayed, which is precisely the error that costs music later, so an
incomplete index must not be culled against.

Errors now carry the request URL and whatever the server put in the body. The
original report of this failure was "album: HTTP 500", which points at the URL
and the credentials -- neither of which was at fault.
2026-08-24 14:05:21 +01:00

8.3 KiB

music-curator

Build a local record of what actually gets listened to, from Last.fm.

Companion to music-mirror, which keeps an MP3 copy of a lossless library for an iPod. This one answers the question that mirror cannot: which of it is worth carrying, and which of it has not been played in years.

This is stage two. It ingests the scrobble history, indexes the library from Lidarr, and matches one to the other. There are no playlists yet, and it writes nothing back — every Lidarr call is a GET. See "Where this is going" below.

What it does today

  • Pulls the full Last.fm scrobble history into SQLite, then keeps it current.
  • Tracks loved tracks separately, as the protected set for later stages.
  • Indexes every artist, album and track Lidarr knows about, with file paths and the date each file landed.
  • Ties the two together and reports how well it managed.

Matching

Two tiers, and no third.

Tier Key Notes
mbid MusicBrainz recording id Exact. Last.fm's per-scrobble mbid against Lidarr's ForeignRecordingId
name Normalised artist and title Everything the first tier could not carry
none Recorded as a miss, never guessed at

The normalisation is the load-bearing part, because the two sides disagree in predictable ways. It folds case and accents, drops guest credits (Yellowcard feat. Tay Jardine against a tag of Yellowcard), strips a trailing version suffix ((Remastered 2011), - Live), expands &, and removes a leading The. Punctuation gets two different rules that pull against each other and are both required: apostrophes are deleted, so Don't meets Dont, while every other mark becomes a space, so AC/DC, AC-DC and AC DC all meet as well.

It leans towards collapsing too much. A false match makes something look played; a missed match makes something look abandoned. Only one of those deletes music.

Indexing quirks

Albums are fetched from the unfiltered GET /api/v1/album, not one call per artist. ?artistId= is Lidarr's unguarded path: it dereferences the album's artist metadata with no null check and picks the monitored release with SingleOrDefault, so it returns a 500 for an album with broken metadata or two monitored releases. The unfiltered endpoint skips such albums instead, and costs N fewer requests.

Tracks and files have no unfiltered endpoint — Lidarr rejects a call with no filter — so they stay per artist. If one artist cannot be served, that artist is skipped and the run continues, but the count is recorded and the coverage report says so loudly. A missing artist makes their played music look cold, so an incomplete index must never be culled against.

Reading the coverage report

Matched against unmatched is the wrong comparison — most unmatched listening is music that was never in the library, which says nothing at all about the matcher. The line to watch is:

unmatched by an artist the library holds: N pairs, M plays

That is a track that was played, sitting beside a file it should have matched. Those are the matcher's real misses, and every one is a candidate for being wrongly called cold in stage four. The report lists the worst fifteen by play count so they can be eyeballed.

How the ingest works

Two halves, both taking their bounds from the database rather than from a saved cursor, so an interrupted run resumes from what it actually has.

Half Window Purpose
Catch-up newest heldnow New scrobbles since the last pass
Backfill start → oldest held Walks towards the beginning of the history

The backfill asks repeatedly for the newest page of everything at or before a cursor, and moves the cursor to the oldest scrobble that came back. When a whole page shares a single second the cursor cannot move without stepping over the rest of that second, so it takes the next page of the same window instead.

Both windows are bounded at both ends, so paging cannot shift under the fetch while new scrobbles arrive mid-run.

Three details of the API that are easy to get wrong, all handled:

  • The currently-playing track is prepended to the first page with no timestamp at all. Stored once, it would come back on every pass forever.
  • A lone result is returned as a bare object, not a one-item list.
  • MBIDs are empty strings rather than absent when unknown, and an empty string looks like a usable join key right up until it silently matches everything.

Scrobbles have no identifier, so the primary key is timestamp plus artist plus track. Two plays of the same track in the same second collapse into one; they are genuinely indistinguishable, and a surrogate key would make re-ingest non-idempotent, which is a far worse trade.

Retries cover error 29 (rate limit) and the backend failures, 8, 11 and 16, with an exponential backoff. An invalid or suspended key fails immediately rather than retrying four more times to reach the same conclusion.

Usage

music-curator                 # one pass, then report
music-curator --interval 6h   # keep running
music-curator --report-only   # report on the store, fetch nothing
Option Environment variable Default Meaning
--user MUSIC_CURATOR_LASTFM_USER Last.fm username to read
--api-key MUSIC_CURATOR_LASTFM_API_KEY Last.fm API key
--db MUSIC_CURATOR_DB /data/curator.db Path to the SQLite store
--interval MUSIC_CURATOR_INTERVAL unset Repeat forever, e.g. 45m, 6h, 1d
--request-delay MUSIC_CURATOR_REQUEST_DELAY 0.25 Seconds between API requests
--backfill-limit MUSIC_CURATOR_BACKFILL_LIMIT 0 Cap backfill requests per pass; 0 for no cap
--lidarr-url MUSIC_CURATOR_LIDARR_URL unset Lidarr base URL, e.g. http://lidarr:8686
--lidarr-api-key MUSIC_CURATOR_LIDARR_API_KEY unset Lidarr API key
--skip-index off Match against the index already held
--report-only off Report without fetching

A Last.fm API key is all that is needed. None of the endpoints used here authenticate a user, so there is no shared secret, no session key and no signing.

The first pass over a long history is thousands of requests at 200 scrobbles each. --backfill-limit spreads that over several passes if you would rather not do it in one.

Running it on TrueNAS Scale

compose.yaml is a Custom App definition. Adjust the host path and the user: to match your pool, set the API key through the TrueNAS UI rather than in the file, then add it as a custom app. The image is published to this Gitea's registry on every release:

code.emmathe.dev/lyrathorpe/music-curator:latest

Give the store its own dataset. It is derived data and can be rebuilt from Last.fm, but rebuilding means downloading the whole history again.

Tests

docker build --target test .   # what CI runs
pytest                         # needs pytest on PATH

The suite runs against fake transports for both services. The Last.fm one reproduces its paging, its from/to semantics and its awkward response shapes; the Lidarr one serves a canned library split across the same four endpoints the indexer calls, so the stitching is exercised rather than stubbed. No network, no credentials, no rate limit. On a Nix machine:

nix shell nixpkgs#python3Packages.pytest -c pytest

Where this is going

Stage Status
Last.fm ingest and store done
Lidarr index and the scrobble-to-track matcher done
M3U playlists written into the mirror next
Cold-music report, unmonitoring what is not played last

The cull will unmonitor cold albums in Lidarr and tag their artists. It will never delete files: Lidarr's AlbumResource has no tags at all, so tagging only works per artist, and an artist is only tagged when every one of their albums qualifies. It will also refuse to run at all if the matcher's coverage is poor, because an unmatched track is not an unplayed track.