Files
music-curator/README.md
T
Emma Thorpe 18f05d3d55
Build and publish container / build (push) Failing after 2m16s
feat: ingest a Last.fm scrobble history into a local store
First stage of a curation tool for the music library that music-mirror
mirrors. Before anything can build playlists or decide what has gone cold,
there has to be a local, queryable record of what is actually played; an API
call per question does not scale to a library-sized analysis.

Ingest is in two halves. A catch-up fetches everything scrobbled since the
newest scrobble held, and a backfill walks the history backwards until it runs
out. Both take their bounds from the database rather than from a saved cursor,
so an interrupted run resumes from what it actually has, and both windows are
bounded at each end so paging cannot shift under the fetch while new scrobbles
arrive mid-run.

Scrobbles carry no identifier, so the primary key is timestamp, artist and
track. Two plays of one track in the same second collapse into a single row:
they are indistinguishable in the data, and a surrogate key would make
re-ingest non-idempotent, which is the worse trade.

Three API behaviours are handled explicitly because each fails silently:
the currently-playing track arrives with no timestamp and would be re-ingested
on every pass; a lone result is returned as a bare object rather than a
one-item list; and MBIDs are empty strings rather than absent when unknown,
which would later look like a usable join key.

Retries cover the rate limit and the transient backend errors with an
exponential backoff. An invalid or suspended key fails immediately.

The report exists to surface one number before the next stage is built: the
share of scrobbles carrying a MusicBrainz recording id. Lidarr exposes the same
identifier per track, so those can be joined exactly and the rest must go
through name matching. That percentage bounds how far the matcher can be
trusted.

No runtime dependencies, and the tests run against a fake transport that
reproduces the service's paging and response shapes, so they need neither
network nor credentials.
2026-08-24 12:03:08 +01:00

132 lines
5.8 KiB
Markdown

# music-curator
Build a local record of what actually gets listened to, from Last.fm.
Companion to [music-mirror](https://code.emmathe.dev/lyrathorpe/music-mirror),
which keeps an MP3 copy of a lossless library for an iPod. This one answers the
question that mirror cannot: which of it is worth carrying, and which of it has
not been played in years.
**This is stage one.** It ingests the scrobble history and nothing else. There
are no playlists yet, and nothing touches Lidarr or the music library. See
"Where this is going" below.
## What it does today
- Pulls the full Last.fm scrobble history into SQLite, then keeps it current.
- Tracks loved tracks separately, as the protected set for later stages.
- Reports what it holds, including the number that matters most: how many
scrobbles carry a MusicBrainz recording id.
That last figure decides the next stage. Lidarr exposes a `ForeignRecordingId`
on every track, which is the same identifier, so scrobbles carrying one can be
joined to the library exactly. The rest have to go through name matching, which
is where a curation tool goes wrong and starts recommending the deletion of
music you love. Measure the join rate before trusting the verdict.
## How the ingest works
Two halves, both taking their bounds from the database rather than from a saved
cursor, so an interrupted run resumes from what it actually has.
| Half | Window | Purpose |
| --------- | ---------------------- | ---------------------------------------- |
| Catch-up | `newest held``now` | New scrobbles since the last pass |
| Backfill | start → `oldest held` | Walks towards the beginning of the history |
The backfill asks repeatedly for the newest page of everything at or before a
cursor, and moves the cursor to the oldest scrobble that came back. When a whole
page shares a single second the cursor cannot move without stepping over the
rest of that second, so it takes the next page of the same window instead.
Both windows are bounded at both ends, so paging cannot shift under the fetch
while new scrobbles arrive mid-run.
Three details of the API that are easy to get wrong, all handled:
- The **currently-playing** track is prepended to the first page with no
timestamp at all. Stored once, it would come back on every pass forever.
- A **lone result** is returned as a bare object, not a one-item list.
- **MBIDs are empty strings** rather than absent when unknown, and an empty
string looks like a usable join key right up until it silently matches
everything.
Scrobbles have no identifier, so the primary key is timestamp plus artist plus
track. Two plays of the same track in the same second collapse into one; they
are genuinely indistinguishable, and a surrogate key would make re-ingest
non-idempotent, which is a far worse trade.
Retries cover error 29 (rate limit) and the backend failures, 8, 11 and 16,
with an exponential backoff. An invalid or suspended key fails immediately
rather than retrying four more times to reach the same conclusion.
## Usage
```sh
music-curator # one pass, then report
music-curator --interval 6h # keep running
music-curator --report-only # report on the store, fetch nothing
```
| Option | Environment variable | Default | Meaning |
| ------------------ | ----------------------------- | ------------------ | ------------------------------------------- |
| `--user` | `MUSIC_CURATOR_LASTFM_USER` | — | Last.fm username to read |
| `--api-key` | `MUSIC_CURATOR_LASTFM_API_KEY` | — | Last.fm API key |
| `--db` | `MUSIC_CURATOR_DB` | `/data/curator.db` | Path to the SQLite store |
| `--interval` | `MUSIC_CURATOR_INTERVAL` | unset | Repeat forever, e.g. `45m`, `6h`, `1d` |
| `--request-delay` | `MUSIC_CURATOR_REQUEST_DELAY` | `0.25` | Seconds between API requests |
| `--backfill-limit` | `MUSIC_CURATOR_BACKFILL_LIMIT` | `0` | Cap backfill requests per pass; 0 for no cap |
| `--report-only` | — | off | Report without fetching |
A [Last.fm API key](https://www.last.fm/api/account/create) is all that is
needed. None of the endpoints used here authenticate a user, so there is no
shared secret, no session key and no signing.
The first pass over a long history is thousands of requests at 200 scrobbles
each. `--backfill-limit` spreads that over several passes if you would rather
not do it in one.
## Running it on TrueNAS Scale
`compose.yaml` is a Custom App definition. Adjust the host path and the `user:`
to match your pool, set the API key through the TrueNAS UI rather than in the
file, then add it as a custom app. The image is published to this Gitea's
registry on every release:
```
code.emmathe.dev/lyrathorpe/music-curator:latest
```
Give the store its own dataset. It is derived data and can be rebuilt from
Last.fm, but rebuilding means downloading the whole history again.
## Tests
```sh
docker build --target test . # what CI runs
pytest # needs pytest on PATH
```
The suite runs against a fake transport that reproduces the real service's
paging, its `from`/`to` semantics and its awkward response shapes. No network,
no credentials, no rate limit. On a Nix machine:
```sh
nix shell nixpkgs#python3Packages.pytest -c pytest
```
## Where this is going
| Stage | Status |
| ------------------------------------------------ | ------------ |
| Last.fm ingest and store | done |
| Lidarr index and the scrobble-to-track matcher | next |
| M3U playlists written into the mirror | after that |
| Cold-music report, unmonitoring what is not played | last |
The cull will unmonitor cold albums in Lidarr and tag their artists. It will
never delete files: Lidarr's `AlbumResource` has no tags at all, so tagging
only works per artist, and an artist is only tagged when every one of their
albums qualifies. It will also refuse to run at all if the matcher's coverage
is poor, because an unmatched track is not an unplayed track.