Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
7b6f6e0016 | ||
|
|
0ae2630a79 | ||
|
|
a7d16ca0b2 | ||
|
|
8c4da6e14e | ||
|
|
41f6b290d8 | ||
|
|
fbce764dc5 | ||
|
|
71c7115507 | ||
|
|
40dfce8a4c | ||
|
|
ddd126cc0d | ||
|
|
97b27f8daa | ||
|
|
88b96b8159 | ||
|
|
15e5ee5aea | ||
|
|
05cca3508c | ||
|
|
cda3d8463b | ||
|
|
08f099aaa9 | ||
|
|
719a6372ad | ||
|
|
61ce751ac5 | ||
|
|
d1c10d32d9 | ||
|
|
acd042ad8d | ||
|
|
45d99ff039 | ||
|
|
54e19992e4 | ||
|
|
2886c02a2e | ||
|
|
56a06a6cd5 | ||
|
|
a3d0689c0a | ||
|
|
7588fee302 | ||
|
|
cd66559b55 | ||
|
|
fef082a783 | ||
|
|
edebecc8ea | ||
|
|
75ed26a411 | ||
|
|
aa2bd59320 | ||
|
|
997627f4fe | ||
|
|
e6fa030d9d | ||
|
|
3ac9f84ad7 | ||
|
|
3e78f8ebd4 |
@@ -45,7 +45,8 @@ jobs:
|
|||||||
|
|
||||||
# The suite runs inside the image, against the interpreter that ships,
|
# The suite runs inside the image, against the interpreter that ships,
|
||||||
# rather than against whatever the runner happens to provide. A failing
|
# rather than against whatever the runner happens to provide. A failing
|
||||||
# test fails the build. Layers are shared with the push build below.
|
# test fails the build. The runtime stage below is built from the same
|
||||||
|
# daemon afterwards, so its layers are already in cache.
|
||||||
- name: Run the test suite inside the image
|
- name: Run the test suite inside the image
|
||||||
run: docker build --target test -t music-curator:test .
|
run: docker build --target test -t music-curator:test .
|
||||||
|
|
||||||
@@ -124,9 +125,6 @@ jobs:
|
|||||||
echo "release=${release}" >> "$GITHUB_OUTPUT"
|
echo "release=${release}" >> "$GITHUB_OUTPUT"
|
||||||
echo "Computed bump=${bump}, release=${release}, base=${base}"
|
echo "Computed bump=${bump}, release=${release}, base=${base}"
|
||||||
|
|
||||||
- name: Set up Buildx
|
|
||||||
uses: docker/setup-buildx-action@d7f5e7f509e45cec5c76c4d5afdd7de93d0b3df5 # v4
|
|
||||||
|
|
||||||
- name: Log in to the Gitea container registry
|
- name: Log in to the Gitea container registry
|
||||||
if: github.event_name != 'pull_request'
|
if: github.event_name != 'pull_request'
|
||||||
uses: docker/login-action@650006c6eb7dba73a995cc03b0b2d7f5ca915bee # v4
|
uses: docker/login-action@650006c6eb7dba73a995cc03b0b2d7f5ca915bee # v4
|
||||||
@@ -135,21 +133,33 @@ jobs:
|
|||||||
username: ${{ github.repository_owner }}
|
username: ${{ github.repository_owner }}
|
||||||
password: ${{ secrets.PACKAGES_TOKEN }}
|
password: ${{ secrets.PACKAGES_TOKEN }}
|
||||||
|
|
||||||
- name: Build and push
|
# Plain `docker build` rather than buildx. buildx boots its own buildkit
|
||||||
uses: docker/build-push-action@f9f3042f7e2789586610d6e8b85c8f03e5195baf # v7
|
# in a container with a cache of its own, so it shared nothing with the
|
||||||
with:
|
# test build above and rebuilt the image from the base image up -- two
|
||||||
context: .
|
# full builds per run. It earns that cost when building for several
|
||||||
# Without this the last stage in the Dockerfile -- the test stage --
|
# platforms; this only ever targets the amd64 NAS, so it does not.
|
||||||
# would be what gets published.
|
#
|
||||||
target: runtime
|
# `--target runtime` is a strict prefix of the test stage, so every layer
|
||||||
# The NAS is the only host this runs on. Building arm64 as well would
|
# is already in the daemon's cache and this resolves in seconds.
|
||||||
# mean emulating it under QEMU for no consumer.
|
- name: Build the runtime image
|
||||||
platforms: linux/amd64
|
run: |
|
||||||
push: ${{ github.event_name != 'pull_request' }}
|
set -euo pipefail
|
||||||
tags: ${{ steps.version.outputs.tags }}
|
tags=()
|
||||||
labels: |
|
while IFS= read -r tag; do
|
||||||
org.opencontainers.image.source=${{ github.server_url }}/${{ github.repository }}
|
[ -n "$tag" ] && tags+=(-t "$tag")
|
||||||
org.opencontainers.image.revision=${{ github.sha }}
|
done <<< "${{ steps.version.outputs.tags }}"
|
||||||
|
docker build --target runtime \
|
||||||
|
--label "org.opencontainers.image.source=${GITHUB_SERVER_URL}/${GITHUB_REPOSITORY}" \
|
||||||
|
--label "org.opencontainers.image.revision=${GITHUB_SHA}" \
|
||||||
|
"${tags[@]}" .
|
||||||
|
|
||||||
|
- name: Push
|
||||||
|
if: github.event_name != 'pull_request'
|
||||||
|
run: |
|
||||||
|
set -euo pipefail
|
||||||
|
while IFS= read -r tag; do
|
||||||
|
[ -n "$tag" ] && docker push "$tag"
|
||||||
|
done <<< "${{ steps.version.outputs.tags }}"
|
||||||
|
|
||||||
# Record the release: write the computed version into pyproject.toml, then
|
# Record the release: write the computed version into pyproject.toml, then
|
||||||
# commit and tag it, so the packaging metadata always matches the release
|
# commit and tag it, so the packaging metadata always matches the release
|
||||||
|
|||||||
@@ -7,9 +7,11 @@ which keeps an MP3 copy of a lossless library for an iPod. This one answers the
|
|||||||
question that mirror cannot: which of it is worth carrying, and which of it has
|
question that mirror cannot: which of it is worth carrying, and which of it has
|
||||||
not been played in years.
|
not been played in years.
|
||||||
|
|
||||||
**This is stage two.** It ingests the scrobble history, indexes the library from
|
**This is stage three.** It ingests the scrobble history, indexes the library
|
||||||
Lidarr, and matches one to the other. There are no playlists yet, and it writes
|
from Lidarr, matches one to the other, and writes playlists into the mirror —
|
||||||
nothing back — every Lidarr call is a `GET`. See "Where this is going" below.
|
by listening history and by mood.
|
||||||
|
Nothing is written back to Lidarr — every call there is a `GET`. See "Where this
|
||||||
|
is going" below.
|
||||||
|
|
||||||
## What it does today
|
## What it does today
|
||||||
|
|
||||||
@@ -18,6 +20,7 @@ nothing back — every Lidarr call is a `GET`. See "Where this is going" below.
|
|||||||
- Indexes every artist, album and track Lidarr knows about, with file paths and
|
- Indexes every artist, album and track Lidarr knows about, with file paths and
|
||||||
the date each file landed.
|
the date each file landed.
|
||||||
- Ties the two together and reports how well it managed.
|
- Ties the two together and reports how well it managed.
|
||||||
|
- Writes M3U playlists into the mirror, from the listening history.
|
||||||
|
|
||||||
## Matching
|
## Matching
|
||||||
|
|
||||||
@@ -29,7 +32,14 @@ Two tiers, and no third.
|
|||||||
| `name` | Normalised artist and title | Everything the first tier could not carry |
|
| `name` | Normalised artist and title | Everything the first tier could not carry |
|
||||||
| `none` | — | Recorded as a miss, never guessed at |
|
| `none` | — | Recorded as a miss, never guessed at |
|
||||||
|
|
||||||
The normalisation is the load-bearing part, because the two sides disagree in
|
The name tier does almost all of the work. A recording MBID is exact when it
|
||||||
|
lands, but MusicBrainz holds a separate recording per release, and Last.fm and
|
||||||
|
Lidarr rarely pick the same one: on a real library, two thirds of scrobbles
|
||||||
|
carry a recording id and barely a twentieth of them join on it. The report
|
||||||
|
counts how many carried an id and matched on name anyway, which is the measure
|
||||||
|
of that disagreement.
|
||||||
|
|
||||||
|
So the normalisation is the load-bearing part, because the two sides disagree in
|
||||||
predictable ways. It folds case and accents, drops guest credits (`Yellowcard
|
predictable ways. It folds case and accents, drops guest credits (`Yellowcard
|
||||||
feat. Tay Jardine` against a tag of `Yellowcard`), strips a trailing
|
feat. Tay Jardine` against a tag of `Yellowcard`), strips a trailing
|
||||||
version suffix (`(Remastered 2011)`, `- Live`), expands `&`, and removes a
|
version suffix (`(Remastered 2011)`, `- Live`), expands `&`, and removes a
|
||||||
@@ -42,6 +52,56 @@ It leans towards collapsing too much. A false match makes something look
|
|||||||
played; a missed match makes something look abandoned. Only one of those
|
played; a missed match makes something look abandoned. Only one of those
|
||||||
deletes music.
|
deletes music.
|
||||||
|
|
||||||
|
### Indexing quirks
|
||||||
|
|
||||||
|
Albums are fetched from the **unfiltered** `GET /api/v1/album` first: one
|
||||||
|
request, and the only path that skips albums whose artist metadata is missing
|
||||||
|
rather than dereferencing it.
|
||||||
|
|
||||||
|
That is not enough on its own. Every album endpoint maps through a resource
|
||||||
|
that picks the release with `SingleOrDefault(x => x.Monitored)`, which throws
|
||||||
|
for an album with **two monitored releases** and takes the whole response with
|
||||||
|
it:
|
||||||
|
|
||||||
|
```
|
||||||
|
HTTP 500: Sequence contains more than one element
|
||||||
|
```
|
||||||
|
|
||||||
|
When the bulk call dies that way, the indexer falls back to one request per
|
||||||
|
artist. It cannot avoid the exception, but it confines it to whichever artist
|
||||||
|
owns the offending album and names them in the log — which is the only
|
||||||
|
practical way to find it in a large library. Open that artist in Lidarr and
|
||||||
|
check the Releases tab of each album: exactly one release may be monitored.
|
||||||
|
|
||||||
|
Losing an artist's albums does not cost their tracks, which come from a
|
||||||
|
different endpoint with a different mapper, so matching is unaffected. A cull
|
||||||
|
would not be, and the report says so.
|
||||||
|
|
||||||
|
### Talking to Lidarr
|
||||||
|
|
||||||
|
Indexing is two requests per artist, and more when the album fallback fires. On
|
||||||
|
a large library that is thousands of requests in a few minutes. `urllib` opens a
|
||||||
|
new TCP connection and performs a new DNS lookup for every one of them, which is
|
||||||
|
enough to exhaust a container's resolver and produce `[Errno -3] Try again` on
|
||||||
|
everything at once. The client therefore holds one connection open per host and
|
||||||
|
resolves once.
|
||||||
|
|
||||||
|
Transient failures — a dropped connection, a resolver hiccup, `429`, `502`,
|
||||||
|
`503`, `504` — are retried with a backoff. An HTTP `500` is not: it is an
|
||||||
|
unhandled exception inside Lidarr's own serialisation and will be raised again
|
||||||
|
identically. That distinction also decides whether a failure is worth
|
||||||
|
investigating; a library-wide outage is not probed artist by artist, because
|
||||||
|
doing so multiplies the load that caused it.
|
||||||
|
|
||||||
|
A local address is preferable to a public hostname here. It removes DNS, the
|
||||||
|
reverse proxy and its timeouts from a path that needs none of them.
|
||||||
|
|
||||||
|
Tracks and files have no unfiltered endpoint — Lidarr rejects a call with no
|
||||||
|
filter — so they stay per artist. If one artist cannot be served, that artist is
|
||||||
|
skipped and the run continues, but the count is recorded and the coverage report
|
||||||
|
says so loudly. A missing artist makes their played music look cold, so an
|
||||||
|
incomplete index must never be culled against.
|
||||||
|
|
||||||
### Reading the coverage report
|
### Reading the coverage report
|
||||||
|
|
||||||
Matched against unmatched is the wrong comparison — most unmatched listening is
|
Matched against unmatched is the wrong comparison — most unmatched listening is
|
||||||
@@ -52,10 +112,169 @@ matcher. The line to watch is:
|
|||||||
unmatched by an artist the library holds: N pairs, M plays
|
unmatched by an artist the library holds: N pairs, M plays
|
||||||
```
|
```
|
||||||
|
|
||||||
That is a track that was played, sitting beside a file it should have matched.
|
That is a track that was played by an artist the library holds. It is then
|
||||||
Those are the matcher's real misses, and every one is a candidate for being
|
split three ways, because owning an artist is a weak proxy for owning a track
|
||||||
wrongly called cold in stage four. The report lists the worst fifteen by play
|
and a shared title is a weak proxy for a shared song:
|
||||||
count so they can be eyeballed.
|
|
||||||
|
- **the library's own title credits the scrobbled artist** — `Voodoo People
|
||||||
|
(Pendulum Remix)` against a play credited to Pendulum. Same song, filed
|
||||||
|
under the original artist. These are the genuine misses.
|
||||||
|
- **the same title under an unrelated artist** — a collision, not a miss.
|
||||||
|
Across fifty thousand tracks these are constant: `Everyday` is Rusko and
|
||||||
|
also Def Leppard, `Kaleidoscope` is Delta Heavy and also Chappell Roan.
|
||||||
|
Matching on title alone would be far worse than missing them, which is why
|
||||||
|
there is no such tier.
|
||||||
|
- **the title is nowhere in the library** — never bought.
|
||||||
|
|
||||||
|
Only the first is worth chasing. Counting all three as matcher failures
|
||||||
|
overstates the problem and would over-block the cull.
|
||||||
|
|
||||||
|
## Playlists
|
||||||
|
|
||||||
|
Written into `<mirror>/_playlists/` as extended M3U, rebuilt every pass. Six
|
||||||
|
rules, capped at `--playlist-limit` tracks each:
|
||||||
|
|
||||||
|
| Playlist | Rule |
|
||||||
|
| -------------------- | --------------------------------------------------------- |
|
||||||
|
| `heavy-rotation` | Most played over the last twelve months |
|
||||||
|
| `all-time` | Most played ever |
|
||||||
|
| `neglected` | Played heavily once, silent for twelve months |
|
||||||
|
| `deep-cuts` | Never played, from albums whose other tracks you play constantly |
|
||||||
|
| `unheard-favourites` | Never played, by the artists you play most |
|
||||||
|
| `unheard` | Never played, anywhere in the library |
|
||||||
|
|
||||||
|
Ninety days was the obvious window for "recent" and is the wrong one: on a real
|
||||||
|
history it holds a few hundred plays spread thinly across a twenty-thousand
|
||||||
|
track rotation, so nothing ranks meaningfully. Twelve months does.
|
||||||
|
|
||||||
|
The two `unheard` playlists rotate **weekly**, not per pass. A pass runs every
|
||||||
|
few hours, and a playlist that reorders itself each time is one that has to be
|
||||||
|
re-imported each time — the Music app imports a snapshot of a file, it does not
|
||||||
|
track it.
|
||||||
|
|
||||||
|
### Moods
|
||||||
|
|
||||||
|
A second set of playlists selects by **Last.fm's crowd tags** rather than by
|
||||||
|
listening history. MusicBrainz genres arrive free with the Lidarr index and are
|
||||||
|
no use for this: they are sparse and formal, and will not tell you a record is
|
||||||
|
screamo or synthwave. People typing tags will.
|
||||||
|
|
||||||
|
Tags are fetched once per artist — one `artist.getTopTags` call each — and
|
||||||
|
refreshed every ninety days. `--tag-limit` spreads the first sweep over several
|
||||||
|
passes.
|
||||||
|
|
||||||
|
Artists are looked up **by name**, not by MusicBrainz id, despite Lidarr having
|
||||||
|
an id for every one of them. Last.fm's mbid index is stale and partial: it
|
||||||
|
answers "the artist you supplied could not be found" for Devo, Escape the Fate,
|
||||||
|
Blasterjaxx and a few hundred others whose pages plainly exist and carry exactly
|
||||||
|
the tags wanted. Its name index is the one its own site runs on. The id is kept
|
||||||
|
only as a fallback, for a name Lidarr spells differently.
|
||||||
|
|
||||||
|
An artist neither key resolves is recorded as fetched with no tags, so the next
|
||||||
|
pass does not spend a request on it again. A genuine failure — a rate limit, a
|
||||||
|
bad key — is *not* recorded, so that one is retried.
|
||||||
|
|
||||||
|
The built-in moods are built from the **tag distribution of this library**,
|
||||||
|
measured, rather than from a general taxonomy:
|
||||||
|
|
||||||
|
| Mood | Selected on |
|
||||||
|
| --------------- | -------------------------------------------------------- |
|
||||||
|
| `drum-and-bass` | drum and bass and its six spellings, liquid funk, neurofunk, jungle, techstep, hospital records |
|
||||||
|
| `bass` | dubstep, brostep, grime, trip-hop, big beat |
|
||||||
|
| `dance` | house and its variants, trance, techno, electro, rave |
|
||||||
|
| `pop-punk` | pop punk, punk, emo, emocore, easycore, power pop |
|
||||||
|
| `screamo` | screamo, post-hardcore, metalcore, melodic hardcore, trancecore |
|
||||||
|
| `heavy-metal` | heavy metal, thrash, speed, power, death, prog, NWOBHM |
|
||||||
|
| `hair-metal` | hair metal, glam metal, glam rock, arena rock, AOR — 1975-1994 |
|
||||||
|
| `nu-metal` | nu metal, alternative metal, rapcore, industrial |
|
||||||
|
| `classic-rock` | classic rock, prog, psychedelic, blues rock, 70s, 60s |
|
||||||
|
| `80s-synths` | 80s, new wave, synth pop, electropop, post-punk — 1975-1992, rock excluded |
|
||||||
|
| `indie` | indie, indie rock, indie pop, britpop, singer-songwriter |
|
||||||
|
|
||||||
|
Measuring first mattered. `synthwave`, `edm`, `big room` and `hardstyle` are
|
||||||
|
plausible tags that carry **nothing at all** here, while `techstep`, `easycore`
|
||||||
|
and `hospital records` carry real weight. Guessing produces the first list.
|
||||||
|
|
||||||
|
Three kinds of tag are never used, and there is a test enforcing it:
|
||||||
|
|
||||||
|
- **Nationality** — `american` alone spans 241 artists. A passport is not a
|
||||||
|
sound.
|
||||||
|
- **`rock` and `electronic`** — 340 and 275 artists, most of the library. A
|
||||||
|
mood that matches everything is not a mood.
|
||||||
|
- **Artist names** — Last.fm's most popular tag for an artist is frequently
|
||||||
|
their own name. `green day`, `paramore` and `queen` are single-artist
|
||||||
|
playlists waiting to happen.
|
||||||
|
|
||||||
|
### Exclusions
|
||||||
|
|
||||||
|
A mood may also list `exclude`. An excluded tag drops the artist outright rather
|
||||||
|
than docking their score, and it exists because `80s-synths` cannot be written
|
||||||
|
any other way.
|
||||||
|
|
||||||
|
`80s` is the eleventh most-played tag here, and it sits on Def Leppard and Bon
|
||||||
|
Jovi exactly as heavily as on Eurythmics. Weighting cannot separate them,
|
||||||
|
because the tag it would weight is the one they share. What does separate them
|
||||||
|
is that the stadium rock also carries `hard rock` and `hair metal`, and the
|
||||||
|
synth acts do not.
|
||||||
|
|
||||||
|
Checked against live Last.fm pages, since the tag census only sees artists
|
||||||
|
already in the library:
|
||||||
|
|
||||||
|
| Artist | Tags |
|
||||||
|
| --- | --- |
|
||||||
|
| Eurythmics | `80s`, `new wave`, `pop`, `female vocalists`, `synth pop` |
|
||||||
|
| Frankie Goes to Hollywood | `80s`, `new wave`, `pop`, `british`, `dance` |
|
||||||
|
| Depeche Mode | `electronic`, `synthpop`, `new wave`, `80s`, `synth pop` |
|
||||||
|
| Duran Duran | `new wave`, `80s`, `pop`, `synth pop`, `rock` |
|
||||||
|
|
||||||
|
Four of Eurythmics' five tags are ones no mood may use. Three of the four
|
||||||
|
artists spell it **`synth pop`** with a space; only one spells it `synthpop`.
|
||||||
|
Guessing one spelling would have missed most of the canon.
|
||||||
|
|
||||||
|
An artist qualifies when their tag weights inside a mood sum to at least 30 out
|
||||||
|
of Last.fm's 0-100 scale. One low-weight tag is not a genre, it is somebody's
|
||||||
|
stray opinion.
|
||||||
|
|
||||||
|
`years` filters on the album's release date, which is what separates eighties
|
||||||
|
synth records from everything else a synthpop tag drags in.
|
||||||
|
|
||||||
|
`--vibes` replaces the whole set with a JSON file of the same shape, so a new
|
||||||
|
mood does not need a new release:
|
||||||
|
|
||||||
|
```json
|
||||||
|
[{ "name": "shoegaze", "tags": ["shoegaze", "dream pop"], "min_score": 40 }]
|
||||||
|
```
|
||||||
|
|
||||||
|
Names are validated when the file is read, not when the file is written. A bad
|
||||||
|
one would otherwise surface as a playlist created somewhere unintended.
|
||||||
|
|
||||||
|
### Paths
|
||||||
|
|
||||||
|
Lidarr knows where the lossless source is; the playlists have to point at the
|
||||||
|
MP3s music-mirror made from it. The mapping strips a library root from Lidarr's
|
||||||
|
track paths and re-roots them under the mirror, with the suffix changed.
|
||||||
|
|
||||||
|
`--library-root` is derived from the common parent of the indexed artist folders
|
||||||
|
when unset, so it agrees with Lidarr by construction rather than by being kept
|
||||||
|
in step by hand. Override it if that guess is wrong.
|
||||||
|
|
||||||
|
Entries are written **relative to the playlist file**, so one playlist works
|
||||||
|
from the NAS, from a Mac over SMB, and from Linux, without rewriting.
|
||||||
|
|
||||||
|
Each playlist is given the **owner and group of the mirror** it is written
|
||||||
|
into. The image runs as root by default so that a bind mount of any ownership
|
||||||
|
stays writable, and the cost of that is output owned by root — which the account
|
||||||
|
serving the share cannot read, group bit or no group bit, because the group is
|
||||||
|
also root. Copying the mirror's own ownership avoids having to be told what it
|
||||||
|
should be, and does nothing when the two already agree.
|
||||||
|
|
||||||
|
A track is only listed once its mirror file has been confirmed to exist. Lidarr
|
||||||
|
holding the FLAC says nothing about whether the MP3 has been encoded yet. If a
|
||||||
|
large number are missing, the run says so — that is what a wrong `--library-root`
|
||||||
|
or `--mirror` looks like, since the paths then map to nothing at all.
|
||||||
|
|
||||||
|
`_playlists/` survives music-mirror's prune: it only deletes `*.mp3`, and its
|
||||||
|
empty-directory sweep skips a directory holding M3Us.
|
||||||
|
|
||||||
## How the ingest works
|
## How the ingest works
|
||||||
|
|
||||||
@@ -111,6 +330,11 @@ music-curator --report-only # report on the store, fetch nothing
|
|||||||
| `--backfill-limit` | `MUSIC_CURATOR_BACKFILL_LIMIT` | `0` | Cap backfill requests per pass; 0 for no cap |
|
| `--backfill-limit` | `MUSIC_CURATOR_BACKFILL_LIMIT` | `0` | Cap backfill requests per pass; 0 for no cap |
|
||||||
| `--lidarr-url` | `MUSIC_CURATOR_LIDARR_URL` | unset | Lidarr base URL, e.g. `http://lidarr:8686` |
|
| `--lidarr-url` | `MUSIC_CURATOR_LIDARR_URL` | unset | Lidarr base URL, e.g. `http://lidarr:8686` |
|
||||||
| `--lidarr-api-key` | `MUSIC_CURATOR_LIDARR_API_KEY` | unset | Lidarr API key |
|
| `--lidarr-api-key` | `MUSIC_CURATOR_LIDARR_API_KEY` | unset | Lidarr API key |
|
||||||
|
| `--mirror` | `MUSIC_CURATOR_MIRROR` | unset | Root of the MP3 mirror; playlists go here |
|
||||||
|
| `--library-root` | `MUSIC_CURATOR_LIBRARY_ROOT` | derived | Prefix to strip from Lidarr's paths |
|
||||||
|
| `--playlist-limit` | `MUSIC_CURATOR_PLAYLIST_LIMIT` | `100` | Most tracks in any one playlist |
|
||||||
|
| `--vibes` | `MUSIC_CURATOR_VIBES` | built-in | JSON file of mood definitions |
|
||||||
|
| `--tag-limit` | `MUSIC_CURATOR_TAG_LIMIT` | `0` | Cap artist tag lookups per pass |
|
||||||
| `--skip-index` | — | off | Match against the index already held |
|
| `--skip-index` | — | off | Match against the index already held |
|
||||||
| `--report-only` | — | off | Report without fetching |
|
| `--report-only` | — | off | Report without fetching |
|
||||||
|
|
||||||
@@ -159,7 +383,8 @@ nix shell nixpkgs#python3Packages.pytest -c pytest
|
|||||||
| ------------------------------------------------ | ------------ |
|
| ------------------------------------------------ | ------------ |
|
||||||
| Last.fm ingest and store | done |
|
| Last.fm ingest and store | done |
|
||||||
| Lidarr index and the scrobble-to-track matcher | done |
|
| Lidarr index and the scrobble-to-track matcher | done |
|
||||||
| M3U playlists written into the mirror | next |
|
| M3U playlists from the listening history | done |
|
||||||
|
| Genre and mood playlists from Last.fm tags | done |
|
||||||
| Cold-music report, unmonitoring what is not played | last |
|
| Cold-music report, unmonitoring what is not played | last |
|
||||||
|
|
||||||
The cull will unmonitor cold albums in Lidarr and tag their artists. It will
|
The cull will unmonitor cold albums in Lidarr and tag their artists. It will
|
||||||
|
|||||||
@@ -30,5 +30,11 @@ services:
|
|||||||
# Cap the backfill at this many requests per pass. Unlimited by default,
|
# Cap the backfill at this many requests per pass. Unlimited by default,
|
||||||
# which finishes a long history in one go.
|
# which finishes a long history in one go.
|
||||||
# MUSIC_CURATOR_BACKFILL_LIMIT: "0"
|
# MUSIC_CURATOR_BACKFILL_LIMIT: "0"
|
||||||
|
# The MP3 mirror music-mirror maintains. Playlists are written into
|
||||||
|
# _playlists/ inside it; leave unset to skip them.
|
||||||
|
MUSIC_CURATOR_MIRROR: /mirror
|
||||||
|
# Most tracks in any one playlist.
|
||||||
|
# MUSIC_CURATOR_PLAYLIST_LIMIT: "100"
|
||||||
volumes:
|
volumes:
|
||||||
- /mnt/tank/apps/music-curator:/data
|
- /mnt/tank/apps/music-curator:/data
|
||||||
|
- /mnt/tank/media/music-mp3:/mirror
|
||||||
|
|||||||
+1077
-33
File diff suppressed because it is too large
Load Diff
+1
-1
@@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta"
|
|||||||
|
|
||||||
[project]
|
[project]
|
||||||
name = "music-curator"
|
name = "music-curator"
|
||||||
version = "0.2.0"
|
version = "0.6.0"
|
||||||
description = "Ingest a Last.fm listening history and curate a music library from it"
|
description = "Ingest a Last.fm listening history and curate a music library from it"
|
||||||
readme = "README.md"
|
readme = "README.md"
|
||||||
requires-python = ">=3.11"
|
requires-python = ">=3.11"
|
||||||
|
|||||||
+100
-3
@@ -1,6 +1,10 @@
|
|||||||
|
import http.server
|
||||||
|
import io
|
||||||
import json
|
import json
|
||||||
import os
|
import os
|
||||||
import sys
|
import sys
|
||||||
|
import threading
|
||||||
|
import urllib.error
|
||||||
import urllib.parse
|
import urllib.parse
|
||||||
|
|
||||||
import pytest
|
import pytest
|
||||||
@@ -52,7 +56,9 @@ class FakeLastfm:
|
|||||||
is prepended to the first page with no `date`.
|
is prepended to the first page with no `date`.
|
||||||
"""
|
"""
|
||||||
|
|
||||||
def __init__(self, tracks=(), loved=(), nowplaying=None, outcomes=()):
|
def __init__(self, tracks=(), loved=(), nowplaying=None, outcomes=(), tags=None):
|
||||||
|
# Keyed by mbid or by artist name, whichever the caller asked with.
|
||||||
|
self.tags = tags or {}
|
||||||
self.tracks = sorted(tracks, key=lambda track: int(track["date"]["uts"]), reverse=True)
|
self.tracks = sorted(tracks, key=lambda track: int(track["date"]["uts"]), reverse=True)
|
||||||
self.loved = list(loved)
|
self.loved = list(loved)
|
||||||
self.nowplaying = nowplaying
|
self.nowplaying = nowplaying
|
||||||
@@ -78,6 +84,15 @@ class FakeLastfm:
|
|||||||
return json.dumps(self._recent(query))
|
return json.dumps(self._recent(query))
|
||||||
if method == "user.getlovedtracks":
|
if method == "user.getlovedtracks":
|
||||||
return json.dumps(self._loved(query))
|
return json.dumps(self._loved(query))
|
||||||
|
if method == "artist.gettoptags":
|
||||||
|
key = query.get("mbid") or query.get("artist", "")
|
||||||
|
if key not in self.tags:
|
||||||
|
# What the real service says for a key it cannot resolve, which
|
||||||
|
# for mbids is a great many artists whose pages plainly exist.
|
||||||
|
return json.dumps(
|
||||||
|
{"error": 6, "message": "The artist you supplied could not be found"}
|
||||||
|
)
|
||||||
|
return json.dumps({"toptags": {"tag": self.tags[key]}})
|
||||||
raise AssertionError(f"unexpected method {method}")
|
raise AssertionError(f"unexpected method {method}")
|
||||||
|
|
||||||
def _recent(self, query):
|
def _recent(self, query):
|
||||||
@@ -128,7 +143,10 @@ class FakeLidarr:
|
|||||||
exercise the same stitching the real thing does.
|
exercise the same stitching the real thing does.
|
||||||
"""
|
"""
|
||||||
|
|
||||||
def __init__(self, artists=()):
|
def __init__(self, artists=(), fail=()):
|
||||||
|
# (path, artistId) pairs the fake refuses to serve, standing in for the
|
||||||
|
# 500s Lidarr returns on data it cannot hydrate.
|
||||||
|
self.fail = set(fail)
|
||||||
self.artists, self.albums, self.tracks, self.files = [], [], [], []
|
self.artists, self.albums, self.tracks, self.files = [], [], [], []
|
||||||
for artist_index, entry in enumerate(artists, start=1):
|
for artist_index, entry in enumerate(artists, start=1):
|
||||||
artist_id = artist_index
|
artist_id = artist_index
|
||||||
@@ -190,10 +208,39 @@ class FakeLidarr:
|
|||||||
}
|
}
|
||||||
self.calls.append((path, query))
|
self.calls.append((path, query))
|
||||||
|
|
||||||
|
artist_id = int(query.get("artistId", 0))
|
||||||
|
if (path, artist_id) in self.fail:
|
||||||
|
raise urllib.error.HTTPError(
|
||||||
|
url, 500, "Internal Server Error", {}, io.BytesIO(b'{"message": "boom"}')
|
||||||
|
)
|
||||||
|
|
||||||
|
album_ids = query.get("albumIds")
|
||||||
|
if path == "album" and album_ids:
|
||||||
|
album_id = int(album_ids)
|
||||||
|
if ("albumid", album_id) in self.fail:
|
||||||
|
raise urllib.error.HTTPError(
|
||||||
|
url,
|
||||||
|
500,
|
||||||
|
"Internal Server Error",
|
||||||
|
{},
|
||||||
|
io.BytesIO(b'{"message": "Sequence contains more than one element"}'),
|
||||||
|
)
|
||||||
|
return json.dumps([row for row in self.albums if row["id"] == album_id])
|
||||||
|
|
||||||
if path == "artist":
|
if path == "artist":
|
||||||
return json.dumps(self.artists)
|
return json.dumps(self.artists)
|
||||||
artist_id = int(query.get("artistId", 0))
|
|
||||||
source = {"album": self.albums, "track": self.tracks, "trackfile": self.files}[path]
|
source = {"album": self.albums, "track": self.tracks, "trackfile": self.files}[path]
|
||||||
|
if path == "album" and not artist_id:
|
||||||
|
# Lidarr's unfiltered album endpoint returns the lot.
|
||||||
|
return json.dumps(source)
|
||||||
|
if not artist_id:
|
||||||
|
raise urllib.error.HTTPError(
|
||||||
|
url,
|
||||||
|
400,
|
||||||
|
"Bad Request",
|
||||||
|
{},
|
||||||
|
io.BytesIO(b'{"message": "artistId must be provided"}'),
|
||||||
|
)
|
||||||
return json.dumps([row for row in source if row["artistId"] == artist_id])
|
return json.dumps([row for row in source if row["artistId"] == artist_id])
|
||||||
|
|
||||||
|
|
||||||
@@ -207,3 +254,53 @@ def now_playing():
|
|||||||
"mbid": "",
|
"mbid": "",
|
||||||
"@attr": {"nowplaying": "true"},
|
"@attr": {"nowplaying": "true"},
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|
||||||
|
class _CountingServer(http.server.ThreadingHTTPServer):
|
||||||
|
"""Counts accepted connections, which is what connection reuse is about."""
|
||||||
|
|
||||||
|
daemon_threads = True
|
||||||
|
|
||||||
|
def __init__(self, *args, **kwargs):
|
||||||
|
self.connections = 0
|
||||||
|
super().__init__(*args, **kwargs)
|
||||||
|
|
||||||
|
def process_request(self, request, client_address):
|
||||||
|
self.connections += 1
|
||||||
|
super().process_request(request, client_address)
|
||||||
|
|
||||||
|
|
||||||
|
class _Handler(http.server.BaseHTTPRequestHandler):
|
||||||
|
# Without HTTP/1.1 the server closes after every response and no client
|
||||||
|
# could reuse anything, which would make the test prove nothing.
|
||||||
|
protocol_version = "HTTP/1.1"
|
||||||
|
|
||||||
|
def do_GET(self):
|
||||||
|
if self.path.startswith("/boom"):
|
||||||
|
body, status = b'{"message": "boom"}', 500
|
||||||
|
else:
|
||||||
|
body = json.dumps(
|
||||||
|
{"path": self.path, "key": self.headers.get("X-Api-Key")}
|
||||||
|
).encode()
|
||||||
|
status = 200
|
||||||
|
self.send_response(status)
|
||||||
|
self.send_header("Content-Type", "application/json")
|
||||||
|
self.send_header("Content-Length", str(len(body)))
|
||||||
|
self.end_headers()
|
||||||
|
self.wfile.write(body)
|
||||||
|
|
||||||
|
def log_message(self, *args):
|
||||||
|
pass
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.fixture
|
||||||
|
def http_server():
|
||||||
|
"""A real local HTTP server, for the one component that talks sockets."""
|
||||||
|
server = _CountingServer(("127.0.0.1", 0), _Handler)
|
||||||
|
thread = threading.Thread(target=server.serve_forever, daemon=True)
|
||||||
|
thread.start()
|
||||||
|
try:
|
||||||
|
yield server, f"http://127.0.0.1:{server.server_port}"
|
||||||
|
finally:
|
||||||
|
server.shutdown()
|
||||||
|
server.server_close()
|
||||||
|
|||||||
+844
-2
@@ -1,4 +1,8 @@
|
|||||||
|
import json
|
||||||
|
import os
|
||||||
|
import stat
|
||||||
import urllib.error
|
import urllib.error
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
import pytest
|
import pytest
|
||||||
from conftest import FakeLastfm, FakeLidarr, make_loved, make_tracks
|
from conftest import FakeLastfm, FakeLidarr, make_loved, make_tracks
|
||||||
@@ -54,9 +58,23 @@ LIBRARY = [
|
|||||||
{
|
{
|
||||||
"name": "The Prodigy",
|
"name": "The Prodigy",
|
||||||
"albums": [
|
"albums": [
|
||||||
{"title": "The Fat of the Land", "tracks": [{"title": "Breathe (Remastered)"}]}
|
{
|
||||||
|
"title": "The Fat of the Land",
|
||||||
|
"tracks": [
|
||||||
|
{"title": "Breathe (Remastered)"},
|
||||||
|
# Credits its remixer in the title, which is the only signal
|
||||||
|
# separating a real attribution miss from a title collision.
|
||||||
|
{"title": "Voodoo People (Pendulum Remix)"},
|
||||||
|
],
|
||||||
|
}
|
||||||
],
|
],
|
||||||
},
|
},
|
||||||
|
# Held by the library in its own right, which is what puts its scrobbles
|
||||||
|
# inside the "artist the library holds" filter at all.
|
||||||
|
{
|
||||||
|
"name": "Pendulum",
|
||||||
|
"albums": [{"title": "Immersion", "tracks": [{"title": "Watercolour"}]}],
|
||||||
|
},
|
||||||
]
|
]
|
||||||
|
|
||||||
|
|
||||||
@@ -392,7 +410,7 @@ def test_keys_carry_the_play_count_and_the_span(tmp_path):
|
|||||||
def test_re_indexing_drops_what_lidarr_no_longer_has(tmp_path):
|
def test_re_indexing_drops_what_lidarr_no_longer_has(tmp_path):
|
||||||
"""The index is Lidarr's mirror, not an accumulation of everything ever seen."""
|
"""The index is Lidarr's mirror, not an accumulation of everything ever seen."""
|
||||||
store = indexed(tmp_path, [scrobble_of("AC/DC", "Hells Bells")])
|
store = indexed(tmp_path, [scrobble_of("AC/DC", "Hells Bells")])
|
||||||
assert store.scalar("SELECT COUNT(*) FROM lidarr_artist") == 3
|
assert store.scalar("SELECT COUNT(*) FROM lidarr_artist") == 4
|
||||||
|
|
||||||
music_curator.index_library(
|
music_curator.index_library(
|
||||||
music_curator.Lidarr("http://lidarr", "key", transport=FakeLidarr(LIBRARY[:1])), store
|
music_curator.Lidarr("http://lidarr", "key", transport=FakeLidarr(LIBRARY[:1])), store
|
||||||
@@ -500,3 +518,827 @@ def test_matching_is_redone_when_new_scrobbles_arrive(tmp_path):
|
|||||||
music_curator.run_once(client_for(api), store, "lyra", NOW, 0)
|
music_curator.run_once(client_for(api), store, "lyra", NOW, 0)
|
||||||
|
|
||||||
assert verdict(store, "Yellowcard", "Transmission Home")[0] == "name"
|
assert verdict(store, "Yellowcard", "Transmission Home")[0] == "name"
|
||||||
|
|
||||||
|
|
||||||
|
def test_albums_come_from_the_unfiltered_endpoint(tmp_path):
|
||||||
|
"""One request, and the only path that skips albums it cannot hydrate."""
|
||||||
|
api = FakeLidarr(LIBRARY)
|
||||||
|
store = store_at(tmp_path)
|
||||||
|
|
||||||
|
music_curator.index_library(music_curator.Lidarr("http://lidarr", "key", transport=api), store)
|
||||||
|
|
||||||
|
album_calls = [query for path, query in api.calls if path == "album"]
|
||||||
|
assert album_calls == [{}]
|
||||||
|
assert store.scalar("SELECT COUNT(*) FROM lidarr_album") == 4
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_bad_album_falls_back_to_asking_per_artist(tmp_path):
|
||||||
|
"""An album with two monitored releases throws in the resource mapper, so
|
||||||
|
the bulk call dies wholesale. Per artist, only its owner is lost."""
|
||||||
|
# The bulk call fails because artist 2 owns the offending album.
|
||||||
|
api = FakeLidarr(LIBRARY, fail=[("album", 0), ("album", 2)])
|
||||||
|
store = store_at(tmp_path)
|
||||||
|
|
||||||
|
music_curator.index_library(music_curator.Lidarr("http://lidarr", "key", transport=api), store)
|
||||||
|
|
||||||
|
assert store.scalar("SELECT COUNT(*) FROM lidarr_album WHERE artist_id = 2") == 0
|
||||||
|
# The other two artists keep their albums.
|
||||||
|
assert store.scalar("SELECT COUNT(*) FROM lidarr_album") == 3
|
||||||
|
assert store.get_state("index_albums_skipped") == "1"
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_bad_album_does_not_cost_that_artist_their_tracks(tmp_path):
|
||||||
|
"""Tracks come from a different endpoint with a different mapper, so
|
||||||
|
matching survives an album Lidarr cannot serialise."""
|
||||||
|
api = FakeLidarr(LIBRARY, fail=[("album", 0), ("album", 2)])
|
||||||
|
store = store_at(tmp_path)
|
||||||
|
|
||||||
|
music_curator.index_library(music_curator.Lidarr("http://lidarr", "key", transport=api), store)
|
||||||
|
|
||||||
|
assert store.scalar("SELECT COUNT(*) FROM lidarr_track WHERE artist_id = 2") == 1
|
||||||
|
assert store.get_state("index_skipped") == "0"
|
||||||
|
|
||||||
|
|
||||||
|
def test_an_artist_lidarr_cannot_serve_does_not_kill_the_index(tmp_path):
|
||||||
|
# Artist 2 is AC/DC in LIBRARY; its track lookup fails.
|
||||||
|
api = FakeLidarr(LIBRARY, fail=[("track", 2)])
|
||||||
|
store = store_at(tmp_path)
|
||||||
|
|
||||||
|
music_curator.index_library(music_curator.Lidarr("http://lidarr", "key", transport=api), store)
|
||||||
|
|
||||||
|
assert store.scalar("SELECT COUNT(*) FROM lidarr_artist") == 4
|
||||||
|
assert store.scalar("SELECT COUNT(*) FROM lidarr_track WHERE artist_id = 2") == 0
|
||||||
|
# The other two artists are indexed in full.
|
||||||
|
assert store.scalar("SELECT COUNT(*) FROM lidarr_track") == 5
|
||||||
|
assert store.get_state("index_skipped") == "1"
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_skipped_artist_is_recorded_so_the_report_can_disown_the_numbers(tmp_path):
|
||||||
|
"""An incomplete index makes played music look cold. It has to be loud."""
|
||||||
|
api = FakeLidarr(LIBRARY, fail=[("trackfile", 1)])
|
||||||
|
store = store_at(tmp_path)
|
||||||
|
|
||||||
|
music_curator.index_library(music_curator.Lidarr("http://lidarr", "key", transport=api), store)
|
||||||
|
music_curator.match_library(store)
|
||||||
|
|
||||||
|
assert store.get_state("index_skipped") == "1"
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_clean_index_records_no_skips(tmp_path):
|
||||||
|
store = indexed(tmp_path, [])
|
||||||
|
|
||||||
|
assert store.get_state("index_skipped") == "0"
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_lidarr_error_carries_the_url_and_what_the_server_said():
|
||||||
|
"""A bare status code sends you looking at the wrong thing entirely."""
|
||||||
|
api = FakeLidarr(LIBRARY, fail=[("album", 0)])
|
||||||
|
client = music_curator.Lidarr("http://lidarr:8686", "key", transport=api)
|
||||||
|
|
||||||
|
with pytest.raises(music_curator.LidarrError) as raised:
|
||||||
|
client.get("album")
|
||||||
|
|
||||||
|
message = str(raised.value)
|
||||||
|
assert "http://lidarr:8686/api/v1/album" in message
|
||||||
|
assert "HTTP 500" in message
|
||||||
|
assert "boom" in message
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_offending_album_is_named_not_just_its_artist(tmp_path, caplog):
|
||||||
|
"""An artist's whole discography is too much to click through by hand."""
|
||||||
|
# AC/DC is artist 2; its only album is id 201, holding "Hells Bells".
|
||||||
|
api = FakeLidarr(LIBRARY, fail=[("album", 0), ("album", 2), ("albumid", 201)])
|
||||||
|
store = store_at(tmp_path)
|
||||||
|
|
||||||
|
with caplog.at_level("WARNING"):
|
||||||
|
music_curator.index_library(
|
||||||
|
music_curator.Lidarr("http://lidarr", "key", transport=api), store
|
||||||
|
)
|
||||||
|
|
||||||
|
assert "album id 201" in caplog.text
|
||||||
|
assert "Hells Bells" in caplog.text
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_transient_failure_is_not_probed_album_by_album(tmp_path, caplog):
|
||||||
|
"""When the resolver is the problem every artist fails, and probing each of
|
||||||
|
them multiplies the load that caused it."""
|
||||||
|
|
||||||
|
def unresolvable(url, timeout=None, headers=None):
|
||||||
|
if "artistId" in url:
|
||||||
|
raise urllib.error.URLError("[Errno -3] Try again")
|
||||||
|
return json.dumps(FakeLidarr(LIBRARY).artists) if url.endswith("artist") else "[]"
|
||||||
|
|
||||||
|
store = store_at(tmp_path)
|
||||||
|
client = music_curator.Lidarr("http://lidarr", "key", transport=unresolvable, backoff=0)
|
||||||
|
|
||||||
|
with caplog.at_level("WARNING"):
|
||||||
|
music_curator.index_library(client, store)
|
||||||
|
|
||||||
|
assert "Try again" in caplog.text
|
||||||
|
assert "cannot serialise" not in caplog.text
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_transient_failure_is_retried():
|
||||||
|
attempts = []
|
||||||
|
|
||||||
|
def flaky(url, timeout=None, headers=None):
|
||||||
|
attempts.append(url)
|
||||||
|
if len(attempts) < 3:
|
||||||
|
raise urllib.error.URLError("[Errno -3] Try again")
|
||||||
|
return "[]"
|
||||||
|
|
||||||
|
client = music_curator.Lidarr("http://lidarr", "key", transport=flaky, backoff=0)
|
||||||
|
|
||||||
|
assert client.get("artist") == []
|
||||||
|
assert len(attempts) == 3
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_lidarr_500_is_not_retried():
|
||||||
|
"""It is an exception inside Lidarr's serialisation, not a busy server."""
|
||||||
|
api = FakeLidarr(LIBRARY, fail=[("album", 0)])
|
||||||
|
client = music_curator.Lidarr("http://lidarr", "key", transport=api, backoff=0)
|
||||||
|
|
||||||
|
with pytest.raises(music_curator.LidarrError) as raised:
|
||||||
|
client.get("album")
|
||||||
|
|
||||||
|
assert raised.value.transient is False
|
||||||
|
assert len(api.calls) == 1
|
||||||
|
|
||||||
|
|
||||||
|
def test_keep_alive_uses_one_connection_for_many_requests(http_server):
|
||||||
|
"""The point of the whole class: one DNS lookup and one socket, not N."""
|
||||||
|
server, base = http_server
|
||||||
|
transport = music_curator.KeepAlive()
|
||||||
|
|
||||||
|
try:
|
||||||
|
for index in range(5):
|
||||||
|
body = transport(f"{base}/api/v1/artist?n={index}", headers={"X-Api-Key": "key"})
|
||||||
|
assert json.loads(body)["key"] == "key"
|
||||||
|
finally:
|
||||||
|
transport.close()
|
||||||
|
|
||||||
|
assert server.connections == 1
|
||||||
|
|
||||||
|
|
||||||
|
def test_keep_alive_maps_an_error_status_onto_httperror(http_server):
|
||||||
|
_, base = http_server
|
||||||
|
transport = music_curator.KeepAlive()
|
||||||
|
|
||||||
|
try:
|
||||||
|
with pytest.raises(urllib.error.HTTPError) as raised:
|
||||||
|
transport(f"{base}/boom", headers={"X-Api-Key": "key"})
|
||||||
|
assert raised.value.code == 500
|
||||||
|
assert music_curator.error_detail(raised.value) == "boom"
|
||||||
|
finally:
|
||||||
|
transport.close()
|
||||||
|
|
||||||
|
|
||||||
|
def test_lidarr_talks_to_a_real_server_through_keep_alive(http_server):
|
||||||
|
server, base = http_server
|
||||||
|
client = music_curator.Lidarr(base, "secret")
|
||||||
|
|
||||||
|
try:
|
||||||
|
assert client.get("artist", {"x": 1})["key"] == "secret"
|
||||||
|
assert client.get("album")["path"] == "/api/v1/album"
|
||||||
|
finally:
|
||||||
|
client.transport.close()
|
||||||
|
|
||||||
|
assert server.connections == 1
|
||||||
|
|
||||||
|
|
||||||
|
# Titles taken verbatim from a real coverage report's unmatched list. Each one
|
||||||
|
# was a genuine miss before the normaliser handled it.
|
||||||
|
@pytest.mark.parametrize(
|
||||||
|
("scrobbled", "tagged"),
|
||||||
|
[
|
||||||
|
("Self vs Self (feat. In Flames)", "Self vs Self"),
|
||||||
|
("Grime Battle of Hastings (feat. The Town Crier)", "Grime Battle of Hastings"),
|
||||||
|
("Gold Dust - Shy FX Re-Edit", "Gold Dust"),
|
||||||
|
("Back To Your Roots - Friction & K-Tee Remix", "Back To Your Roots"),
|
||||||
|
("Constellations - Forza Horizon 3 VIP", "Constellations"),
|
||||||
|
("Everyday (Netsky Remix)", "Everyday"),
|
||||||
|
("Voodoo People [Pendulum Remix] [Live At Brixton Academy]", "Voodoo People"),
|
||||||
|
],
|
||||||
|
)
|
||||||
|
def test_real_unmatched_titles_now_agree_with_their_tags(scrobbled, tagged):
|
||||||
|
assert music_curator.normalise(scrobbled) == music_curator.normalise(tagged)
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize(
|
||||||
|
"title",
|
||||||
|
[
|
||||||
|
"Dancing with Myself",
|
||||||
|
"(Don't Fear) The Reaper",
|
||||||
|
"Live and Let Die",
|
||||||
|
"Radio Ga Ga",
|
||||||
|
"Editors",
|
||||||
|
"Mixed Emotions",
|
||||||
|
"Vipassana",
|
||||||
|
],
|
||||||
|
)
|
||||||
|
def test_the_version_words_do_not_eat_ordinary_titles(title):
|
||||||
|
"""Every one of these contains a version word and must survive intact."""
|
||||||
|
assert music_curator.normalise(title) == music_curator.normalise(title.lower())
|
||||||
|
assert len(music_curator.normalise(title).split()) == len(title.split())
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_bracketed_guest_credit_matches_the_bare_tag(tmp_path):
|
||||||
|
"""Whitespace-then-feat misses the bracketed form, which is most of them."""
|
||||||
|
store = indexed(tmp_path, [scrobble_of("Yellowcard", "Here I Am Alive (feat. Someone)")])
|
||||||
|
|
||||||
|
method, track_id = verdict(store, "Yellowcard", "Here I Am Alive (feat. Someone)")
|
||||||
|
assert method == "name"
|
||||||
|
assert track_id is not None
|
||||||
|
|
||||||
|
|
||||||
|
def test_misses_are_split_by_whether_the_library_holds_the_title(tmp_path):
|
||||||
|
"""Owning an artist is a weak proxy for owning a track. Counting both as
|
||||||
|
matcher failures overstates the problem and would over-block the cull."""
|
||||||
|
store = indexed(
|
||||||
|
tmp_path,
|
||||||
|
[
|
||||||
|
# The library holds "Hells Bells", but under AC/DC, not Yellowcard:
|
||||||
|
# an attribution disagreement, and a real miss.
|
||||||
|
scrobble_of("Yellowcard", "Hells Bells"),
|
||||||
|
# Yellowcard is in the library; this track is not, under any artist.
|
||||||
|
scrobble_of("Yellowcard", "A Single She Never Bought"),
|
||||||
|
],
|
||||||
|
)
|
||||||
|
|
||||||
|
def count(extra):
|
||||||
|
return store.connection.execute(
|
||||||
|
"SELECT COUNT(*) FROM scrobble_key k WHERE k.track_id IS NULL"
|
||||||
|
" AND EXISTS (SELECT 1 FROM lidarr_artist a WHERE a.norm_name = k.norm_artist)"
|
||||||
|
f" {extra}"
|
||||||
|
).fetchone()[0]
|
||||||
|
|
||||||
|
assert count("") == 2
|
||||||
|
assert count("AND EXISTS (SELECT 1 FROM lidarr_track t WHERE t.norm_title = k.norm_track)") == 1
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_report_survives_the_attribution_split(tmp_path):
|
||||||
|
store = indexed(tmp_path, [scrobble_of("Yellowcard", "Hells Bells")])
|
||||||
|
|
||||||
|
music_curator.report(store, NOW)
|
||||||
|
|
||||||
|
|
||||||
|
def attribution_pairs(store):
|
||||||
|
"""Unmatched pairs where the library's own title credits the scrobbled artist."""
|
||||||
|
return [
|
||||||
|
row["artist"]
|
||||||
|
for row in store.connection.execute(
|
||||||
|
"SELECT k.artist FROM scrobble_key k WHERE k.track_id IS NULL"
|
||||||
|
" AND EXISTS (SELECT 1 FROM lidarr_artist a WHERE a.norm_name = k.norm_artist)"
|
||||||
|
" AND EXISTS (SELECT 1 FROM lidarr_track t"
|
||||||
|
" WHERE t.norm_title = k.norm_track"
|
||||||
|
" AND instr(lower(t.title), lower(k.artist)) > 0)"
|
||||||
|
)
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_shared_title_is_not_an_attribution_miss(tmp_path):
|
||||||
|
"""Across fifty thousand tracks, titles collide constantly: "Everyday" is
|
||||||
|
Rusko and also Def Leppard. Matching those would be worse than missing."""
|
||||||
|
store = indexed(tmp_path, [scrobble_of("Yellowcard", "Hells Bells")])
|
||||||
|
|
||||||
|
# The library holds "Hells Bells", by AC/DC, and its title says nothing
|
||||||
|
# about Yellowcard. A collision, not a miss.
|
||||||
|
assert verdict(store, "Yellowcard", "Hells Bells") == ("none", None)
|
||||||
|
assert attribution_pairs(store) == []
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_remix_credited_in_the_library_title_is_an_attribution_miss(tmp_path):
|
||||||
|
"""The library has "Voodoo People (Pendulum Remix)" under The Prodigy; the
|
||||||
|
scrobble credits Pendulum. Same song, different filing."""
|
||||||
|
store = indexed(tmp_path, [scrobble_of("Pendulum", "Voodoo People")])
|
||||||
|
|
||||||
|
assert attribution_pairs(store) == ["Pendulum"]
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_title_index_is_used_for_the_report_lookup(tmp_path):
|
||||||
|
"""Without it the report scans every track for every unmatched key: forty-
|
||||||
|
seven seconds on a real library."""
|
||||||
|
store = indexed(tmp_path, [])
|
||||||
|
|
||||||
|
plan = "\n".join(
|
||||||
|
row[-1]
|
||||||
|
for row in store.connection.execute(
|
||||||
|
"EXPLAIN QUERY PLAN SELECT 1 FROM scrobble_key k WHERE k.track_id IS NULL"
|
||||||
|
" AND EXISTS (SELECT 1 FROM lidarr_track t WHERE t.norm_title = k.norm_track)"
|
||||||
|
)
|
||||||
|
)
|
||||||
|
assert "lidarr_track_title" in plan, plan
|
||||||
|
|
||||||
|
|
||||||
|
def playlist_library(tmp_path):
|
||||||
|
"""A library on disk, with mirror MP3s beside the Lidarr source paths."""
|
||||||
|
source = tmp_path / "music"
|
||||||
|
mirror = tmp_path / "mirror"
|
||||||
|
library = [
|
||||||
|
{
|
||||||
|
"name": "Played Band",
|
||||||
|
"albums": [
|
||||||
|
{
|
||||||
|
"title": "Known",
|
||||||
|
"tracks": [{"title": "Hit"}, {"title": "Album Track"}],
|
||||||
|
}
|
||||||
|
],
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "Silent Band",
|
||||||
|
"albums": [{"title": "Unknown", "tracks": [{"title": "Never Heard"}]}],
|
||||||
|
},
|
||||||
|
]
|
||||||
|
api = FakeLidarr(library)
|
||||||
|
# FakeLidarr invents /music/<artist>/<album>/<title>.flac; put the mirror
|
||||||
|
# MP3s at the paths music-mirror would have produced from those.
|
||||||
|
for handle in api.files:
|
||||||
|
relative = Path(handle["path"]).relative_to("/music")
|
||||||
|
handle["path"] = str(source / relative)
|
||||||
|
mp3 = (mirror / relative).with_suffix(".mp3")
|
||||||
|
mp3.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
mp3.write_bytes(b"not really an mp3")
|
||||||
|
for artist in api.artists:
|
||||||
|
artist["path"] = str(source / Path(artist["path"]).name)
|
||||||
|
return api, source, mirror
|
||||||
|
|
||||||
|
|
||||||
|
def test_playlists_are_written_into_the_mirror(tmp_path):
|
||||||
|
api, source, mirror = playlist_library(tmp_path)
|
||||||
|
store = store_at(tmp_path)
|
||||||
|
ingest(store, FakeLastfm([scrobble_of("Played Band", "Hit")] * 1))
|
||||||
|
music_curator.index_library(
|
||||||
|
music_curator.Lidarr("http://lidarr", "key", transport=api), store
|
||||||
|
)
|
||||||
|
music_curator.match_library(store)
|
||||||
|
|
||||||
|
music_curator.build_playlists(store, mirror, str(source), 100, NOW)
|
||||||
|
|
||||||
|
written = sorted(p.name for p in (mirror / "_playlists").glob("*.m3u"))
|
||||||
|
assert written == [
|
||||||
|
"all-time.m3u",
|
||||||
|
"deep-cuts.m3u",
|
||||||
|
"heavy-rotation.m3u",
|
||||||
|
"neglected.m3u",
|
||||||
|
"unheard-favourites.m3u",
|
||||||
|
"unheard.m3u",
|
||||||
|
]
|
||||||
|
|
||||||
|
played = (mirror / "_playlists" / "all-time.m3u").read_text().splitlines()
|
||||||
|
assert played[0] == "#EXTM3U"
|
||||||
|
assert played[1].startswith("#EXTINF:")
|
||||||
|
assert "Played Band - Hit" in played[1]
|
||||||
|
# Relative to the playlist file, so the same file works from any mount.
|
||||||
|
assert played[2] == "../Played Band/Known/Hit.mp3"
|
||||||
|
assert (mirror / "_playlists" / "all-time.m3u").parent.joinpath(played[2]).resolve().is_file()
|
||||||
|
|
||||||
|
|
||||||
|
def test_an_unplayed_track_lands_in_the_unheard_playlist(tmp_path):
|
||||||
|
api, source, mirror = playlist_library(tmp_path)
|
||||||
|
store = store_at(tmp_path)
|
||||||
|
ingest(store, FakeLastfm([scrobble_of("Played Band", "Hit")]))
|
||||||
|
music_curator.index_library(
|
||||||
|
music_curator.Lidarr("http://lidarr", "key", transport=api), store
|
||||||
|
)
|
||||||
|
music_curator.match_library(store)
|
||||||
|
|
||||||
|
music_curator.build_playlists(store, mirror, str(source), 100, NOW)
|
||||||
|
|
||||||
|
unheard = (mirror / "_playlists" / "unheard.m3u").read_text()
|
||||||
|
assert "Never Heard" in unheard
|
||||||
|
assert "Album Track" in unheard
|
||||||
|
# The one thing that was played must not be in it.
|
||||||
|
assert "Played Band - Hit\n" not in unheard
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_track_with_no_mirror_file_is_left_out(tmp_path):
|
||||||
|
"""The index knows what Lidarr holds; whether music-mirror has encoded the
|
||||||
|
MP3 yet is a different question, and is checked rather than assumed."""
|
||||||
|
api, source, mirror = playlist_library(tmp_path)
|
||||||
|
for mp3 in mirror.rglob("*.mp3"):
|
||||||
|
mp3.unlink()
|
||||||
|
store = store_at(tmp_path)
|
||||||
|
ingest(store, FakeLastfm([scrobble_of("Played Band", "Hit")]))
|
||||||
|
music_curator.index_library(
|
||||||
|
music_curator.Lidarr("http://lidarr", "key", transport=api), store
|
||||||
|
)
|
||||||
|
music_curator.match_library(store)
|
||||||
|
|
||||||
|
total = music_curator.build_playlists(store, mirror, str(source), 100, NOW)
|
||||||
|
|
||||||
|
assert total == 0
|
||||||
|
assert (mirror / "_playlists" / "all-time.m3u").read_text() == "#EXTM3U\n"
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_library_root_is_derived_from_the_artist_folders(tmp_path):
|
||||||
|
api, source, mirror = playlist_library(tmp_path)
|
||||||
|
store = store_at(tmp_path)
|
||||||
|
music_curator.index_library(
|
||||||
|
music_curator.Lidarr("http://lidarr", "key", transport=api), store
|
||||||
|
)
|
||||||
|
|
||||||
|
assert music_curator.library_root_of(store) == str(source)
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_playlist_limit_is_honoured(tmp_path):
|
||||||
|
api, source, mirror = playlist_library(tmp_path)
|
||||||
|
store = store_at(tmp_path)
|
||||||
|
music_curator.index_library(
|
||||||
|
music_curator.Lidarr("http://lidarr", "key", transport=api), store
|
||||||
|
)
|
||||||
|
music_curator.match_library(store)
|
||||||
|
|
||||||
|
music_curator.build_playlists(store, mirror, str(source), 1, NOW)
|
||||||
|
|
||||||
|
unheard = (mirror / "_playlists" / "unheard.m3u").read_text().splitlines()
|
||||||
|
assert len([line for line in unheard if line.startswith("#EXTINF")]) == 1
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_rotation_moves_weekly_not_every_pass(tmp_path):
|
||||||
|
"""A playlist that reorders on every pass is one that has to be re-imported
|
||||||
|
on every pass; the Music app imports a snapshot, it does not track a file."""
|
||||||
|
api, source, mirror = playlist_library(tmp_path)
|
||||||
|
store = store_at(tmp_path)
|
||||||
|
music_curator.index_library(
|
||||||
|
music_curator.Lidarr("http://lidarr", "key", transport=api), store
|
||||||
|
)
|
||||||
|
music_curator.match_library(store)
|
||||||
|
|
||||||
|
def unheard_at(when):
|
||||||
|
music_curator.build_playlists(store, mirror, str(source), 100, when)
|
||||||
|
return (mirror / "_playlists" / "unheard.m3u").read_text()
|
||||||
|
|
||||||
|
same_week = unheard_at(NOW), unheard_at(NOW + 3600)
|
||||||
|
assert same_week[0] == same_week[1]
|
||||||
|
|
||||||
|
|
||||||
|
def test_playlists_are_group_readable(tmp_path):
|
||||||
|
api, source, mirror = playlist_library(tmp_path)
|
||||||
|
store = store_at(tmp_path)
|
||||||
|
music_curator.index_library(
|
||||||
|
music_curator.Lidarr("http://lidarr", "key", transport=api), store
|
||||||
|
)
|
||||||
|
music_curator.match_library(store)
|
||||||
|
|
||||||
|
music_curator.build_playlists(store, mirror, str(source), 100, NOW)
|
||||||
|
|
||||||
|
for playlist in (mirror / "_playlists").glob("*.m3u"):
|
||||||
|
assert playlist.stat().st_mode & stat.S_IRGRP, playlist
|
||||||
|
|
||||||
|
|
||||||
|
def test_tag_weights_fall_back_to_rank_when_no_count_is_sent():
|
||||||
|
"""The documented sample carries only a name and a URL; a live response also
|
||||||
|
carries a count. Neither may be relied on alone."""
|
||||||
|
with_count = music_curator.parse_tags(
|
||||||
|
{"toptags": {"tag": [{"name": "Screamo", "count": 100}, {"name": "emo", "count": 40}]}}
|
||||||
|
)
|
||||||
|
assert with_count == [("screamo", 100), ("emo", 40)]
|
||||||
|
|
||||||
|
without = music_curator.parse_tags(
|
||||||
|
{"toptags": {"tag": [{"name": "screamo"}, {"name": "emo"}]}}
|
||||||
|
)
|
||||||
|
assert [tag for tag, _ in without] == ["screamo", "emo"]
|
||||||
|
assert without[0][1] > without[1][1]
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_lone_tag_is_not_a_list():
|
||||||
|
assert music_curator.parse_tags({"toptags": {"tag": {"name": "dnb", "count": 90}}}) == [
|
||||||
|
("dnb", 90)
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
def test_an_artist_lastfm_cannot_answer_for_is_not_asked_again(tmp_path):
|
||||||
|
"""Recording the fetch even when it returns nothing is what stops a pass
|
||||||
|
spending a request per unknown artist, forever."""
|
||||||
|
api, source, mirror = playlist_library(tmp_path)
|
||||||
|
store = store_at(tmp_path)
|
||||||
|
music_curator.index_library(
|
||||||
|
music_curator.Lidarr("http://lidarr", "key", transport=api), store
|
||||||
|
)
|
||||||
|
lastfm = FakeLastfm(tags={})
|
||||||
|
|
||||||
|
assert music_curator.sync_tags(client_for(lastfm), store, NOW) == 2
|
||||||
|
before = len(lastfm.calls)
|
||||||
|
|
||||||
|
assert music_curator.sync_tags(client_for(lastfm), store, NOW) == 0
|
||||||
|
assert len(lastfm.calls) == before
|
||||||
|
|
||||||
|
|
||||||
|
def test_tags_are_refetched_once_they_go_stale(tmp_path):
|
||||||
|
api, source, mirror = playlist_library(tmp_path)
|
||||||
|
store = store_at(tmp_path)
|
||||||
|
music_curator.index_library(
|
||||||
|
music_curator.Lidarr("http://lidarr", "key", transport=api), store
|
||||||
|
)
|
||||||
|
lastfm = FakeLastfm(tags={})
|
||||||
|
music_curator.sync_tags(client_for(lastfm), store, NOW)
|
||||||
|
|
||||||
|
later = NOW + music_curator.TAG_REFRESH_SECONDS + 1
|
||||||
|
assert music_curator.sync_tags(client_for(lastfm), store, later) == 2
|
||||||
|
|
||||||
|
|
||||||
|
def tagged_store(tmp_path, tags):
|
||||||
|
api, source, mirror = playlist_library(tmp_path)
|
||||||
|
store = store_at(tmp_path)
|
||||||
|
music_curator.index_library(
|
||||||
|
music_curator.Lidarr("http://lidarr", "key", transport=api), store
|
||||||
|
)
|
||||||
|
music_curator.sync_tags(client_for(FakeLastfm(tags=tags)), store, NOW)
|
||||||
|
return store, source, mirror
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_vibe_selects_by_tag(tmp_path):
|
||||||
|
store, source, mirror = tagged_store(
|
||||||
|
tmp_path,
|
||||||
|
{
|
||||||
|
"Played Band": [{"name": "screamo", "count": 100}],
|
||||||
|
"Silent Band": [{"name": "classic rock", "count": 100}],
|
||||||
|
},
|
||||||
|
)
|
||||||
|
vibes = [{"name": "screamo", "tags": ["screamo"]}]
|
||||||
|
|
||||||
|
music_curator.build_vibe_playlists(store, vibes, mirror, str(source), 100, NOW)
|
||||||
|
|
||||||
|
written = (mirror / "_playlists" / "screamo.m3u").read_text()
|
||||||
|
assert "Played Band" in written
|
||||||
|
assert "Silent Band" not in written
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_weakly_tagged_artist_is_below_the_threshold(tmp_path):
|
||||||
|
"""A single low-weight tag is not a genre, it is somebody's stray opinion."""
|
||||||
|
store, source, mirror = tagged_store(
|
||||||
|
tmp_path, {"Played Band": [{"name": "screamo", "count": 3}]}
|
||||||
|
)
|
||||||
|
vibes = [{"name": "screamo", "tags": ["screamo"]}]
|
||||||
|
|
||||||
|
music_curator.build_vibe_playlists(store, vibes, mirror, str(source), 100, NOW)
|
||||||
|
|
||||||
|
assert (mirror / "_playlists" / "screamo.m3u").read_text() == "#EXTM3U\n"
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_vibe_can_be_restricted_by_release_year(tmp_path):
|
||||||
|
"""What separates eighties synth records from everything else a synthpop
|
||||||
|
tag drags in."""
|
||||||
|
store, source, mirror = tagged_store(
|
||||||
|
tmp_path, {"Played Band": [{"name": "synthpop", "count": 100}]}
|
||||||
|
)
|
||||||
|
inside = [{"name": "eighties", "tags": ["synthpop"], "years": [1975, 1992]}]
|
||||||
|
outside = [{"name": "nineties", "tags": ["synthpop"], "years": [1993, 1999]}]
|
||||||
|
|
||||||
|
# FakeLidarr dates every album 2019, so neither window should catch it.
|
||||||
|
music_curator.build_vibe_playlists(store, inside, mirror, str(source), 100, NOW)
|
||||||
|
music_curator.build_vibe_playlists(store, outside, mirror, str(source), 100, NOW)
|
||||||
|
assert (mirror / "_playlists" / "eighties.m3u").read_text() == "#EXTM3U\n"
|
||||||
|
assert (mirror / "_playlists" / "nineties.m3u").read_text() == "#EXTM3U\n"
|
||||||
|
|
||||||
|
modern = [{"name": "modern", "tags": ["synthpop"], "years": [2000, 2030]}]
|
||||||
|
music_curator.build_vibe_playlists(store, modern, mirror, str(source), 100, NOW)
|
||||||
|
assert "Played Band" in (mirror / "_playlists" / "modern.m3u").read_text()
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_built_in_vibes_are_all_usable_filenames():
|
||||||
|
for vibe in music_curator.DEFAULT_VIBES:
|
||||||
|
assert music_curator.SAFE_VIBE_NAME.fullmatch(vibe["name"]), vibe["name"]
|
||||||
|
assert vibe["tags"]
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_vibes_file_replaces_the_built_in_set(tmp_path):
|
||||||
|
path = tmp_path / "vibes.json"
|
||||||
|
path.write_text(json.dumps([{"name": "mine", "tags": ["shoegaze"]}]))
|
||||||
|
|
||||||
|
assert music_curator.load_vibes(str(path)) == ({"name": "mine", "tags": ["shoegaze"]},)
|
||||||
|
assert music_curator.load_vibes(None) is music_curator.DEFAULT_VIBES
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize(
|
||||||
|
"content",
|
||||||
|
[
|
||||||
|
'{"not": "a list"}',
|
||||||
|
'[{"tags": ["x"]}]',
|
||||||
|
'[{"name": "../escape", "tags": ["x"]}]',
|
||||||
|
'[{"name": "ok"}]',
|
||||||
|
"not json at all",
|
||||||
|
],
|
||||||
|
)
|
||||||
|
def test_a_bad_vibes_file_is_refused_up_front(tmp_path, content):
|
||||||
|
"""A bad name would otherwise surface as a file written somewhere
|
||||||
|
unintended, which is a poor way to learn about a typo."""
|
||||||
|
path = tmp_path / "vibes.json"
|
||||||
|
path.write_text(content)
|
||||||
|
|
||||||
|
with pytest.raises(ValueError):
|
||||||
|
music_curator.load_vibes(str(path))
|
||||||
|
|
||||||
|
|
||||||
|
def test_tags_are_looked_up_by_name_not_by_mbid(tmp_path):
|
||||||
|
"""Last.fm's mbid index is stale: it cannot find Devo or Escape the Fate by
|
||||||
|
one, though their pages plainly exist. Its name index can."""
|
||||||
|
api, source, mirror = playlist_library(tmp_path)
|
||||||
|
store = store_at(tmp_path)
|
||||||
|
music_curator.index_library(
|
||||||
|
music_curator.Lidarr("http://lidarr", "key", transport=api), store
|
||||||
|
)
|
||||||
|
lastfm = FakeLastfm(tags={"Played Band": [{"name": "screamo", "count": 90}]})
|
||||||
|
|
||||||
|
music_curator.sync_tags(client_for(lastfm), store, NOW)
|
||||||
|
|
||||||
|
asked = [call for call in lastfm.calls if call.get("method") == "artist.getTopTags"]
|
||||||
|
# The name is always tried first; the mbid only appears as a fallback for
|
||||||
|
# the artist that the name could not resolve.
|
||||||
|
assert "artist" in asked[0]
|
||||||
|
assert [call for call in asked if call.get("artist") == "Played Band"]
|
||||||
|
assert not [call for call in asked if call.get("mbid") == "artist-mbid-1"]
|
||||||
|
assert store.scalar("SELECT COUNT(*) FROM artist_tag WHERE tag = 'screamo'") == 1
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_mbid_is_tried_when_the_name_is_not_found(tmp_path):
|
||||||
|
"""Kept only for a name Lidarr spells differently to Last.fm."""
|
||||||
|
api, source, mirror = playlist_library(tmp_path)
|
||||||
|
store = store_at(tmp_path)
|
||||||
|
music_curator.index_library(
|
||||||
|
music_curator.Lidarr("http://lidarr", "key", transport=api), store
|
||||||
|
)
|
||||||
|
lastfm = FakeLastfm(tags={"artist-mbid-1": [{"name": "dnb", "count": 80}]})
|
||||||
|
|
||||||
|
music_curator.sync_tags(client_for(lastfm), store, NOW)
|
||||||
|
|
||||||
|
assert store.scalar("SELECT COUNT(*) FROM artist_tag WHERE tag = 'dnb'") == 1
|
||||||
|
asked = [call for call in lastfm.calls if call.get("method") == "artist.getTopTags"]
|
||||||
|
assert any("mbid" in call for call in asked)
|
||||||
|
|
||||||
|
|
||||||
|
def test_an_artist_neither_key_resolves_is_recorded_and_not_retried(tmp_path):
|
||||||
|
"""Otherwise every pass spends a request on it again, for ever."""
|
||||||
|
api, source, mirror = playlist_library(tmp_path)
|
||||||
|
store = store_at(tmp_path)
|
||||||
|
music_curator.index_library(
|
||||||
|
music_curator.Lidarr("http://lidarr", "key", transport=api), store
|
||||||
|
)
|
||||||
|
lastfm = FakeLastfm(tags={})
|
||||||
|
|
||||||
|
assert music_curator.sync_tags(client_for(lastfm), store, NOW) == 2
|
||||||
|
spent = len(lastfm.calls)
|
||||||
|
|
||||||
|
assert music_curator.sync_tags(client_for(lastfm), store, NOW) == 0
|
||||||
|
assert len(lastfm.calls) == spent
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_real_failure_is_not_recorded_so_the_next_pass_retries(tmp_path):
|
||||||
|
"""A rate limit is not the same as an artist not existing."""
|
||||||
|
api, source, mirror = playlist_library(tmp_path)
|
||||||
|
store = store_at(tmp_path)
|
||||||
|
music_curator.index_library(
|
||||||
|
music_curator.Lidarr("http://lidarr", "key", transport=api), store
|
||||||
|
)
|
||||||
|
lastfm = FakeLastfm(
|
||||||
|
tags={"Played Band": [], "Silent Band": []},
|
||||||
|
outcomes=[{"error": 10, "message": "Invalid API key"}],
|
||||||
|
)
|
||||||
|
|
||||||
|
music_curator.sync_tags(client_for(lastfm), store, NOW)
|
||||||
|
|
||||||
|
# One artist failed hard and must still be pending.
|
||||||
|
assert store.scalar("SELECT COUNT(*) FROM artist_tag_fetched") == 1
|
||||||
|
|
||||||
|
|
||||||
|
# Tags that carry real weight in a real library, against tags that describe a
|
||||||
|
# passport, span most of the collection, or are somebody's artist name.
|
||||||
|
USELESS_TAGS = {
|
||||||
|
"american", "british", "australian", "canadian", "swedish", "dutch", "german",
|
||||||
|
"scottish", "english", "uk", "usa", "canada",
|
||||||
|
"rock", "electronic", "pop", "alternative", "metal ", "all", "heavy",
|
||||||
|
"female vocalists", "male vocalists", "female vocalist",
|
||||||
|
"my top songs", "cover", "covers", "not emo",
|
||||||
|
"green day", "paramore", "queen", "bon jovi", "shinedown", "aerosmith",
|
||||||
|
"journey", "fleetwood mac",
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def test_no_mood_selects_on_a_useless_tag():
|
||||||
|
"""Nationality is not a sound; `rock` and `electronic` span most of the
|
||||||
|
library; and Last.fm's top tag for an artist is often their own name."""
|
||||||
|
for vibe in music_curator.DEFAULT_VIBES:
|
||||||
|
overlap = {tag.casefold() for tag in vibe["tags"]} & USELESS_TAGS
|
||||||
|
assert not overlap, f"{vibe['name']} selects on {overlap}"
|
||||||
|
|
||||||
|
|
||||||
|
def test_mood_names_are_unique():
|
||||||
|
names = [vibe["name"] for vibe in music_curator.DEFAULT_VIBES]
|
||||||
|
assert len(names) == len(set(names))
|
||||||
|
|
||||||
|
|
||||||
|
ROCK_TAGS = {
|
||||||
|
"classic rock", "hard rock", "blues rock", "southern rock", "arena rock",
|
||||||
|
"glam rock", "hair metal", "glam metal", "heavy metal", "metal", "art rock",
|
||||||
|
"psychedelic rock", "progressive rock", "rock and roll", "rock n roll",
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_decade_tag_in_a_non_rock_mood_must_exclude_the_rock():
|
||||||
|
"""`80s` sits on Def Leppard and Bon Jovi as heavily as on Eurythmics, so a
|
||||||
|
mood that reaches for a decade without wanting rock has to say so. A mood
|
||||||
|
that does want it -- classic-rock reaching for 70s -- is exempt."""
|
||||||
|
for vibe in music_curator.DEFAULT_VIBES:
|
||||||
|
tags = {tag.casefold() for tag in vibe["tags"]}
|
||||||
|
decades = {"60s", "70s", "80s", "90s"} & tags
|
||||||
|
if not decades or tags & ROCK_TAGS:
|
||||||
|
continue
|
||||||
|
excluded = {tag.casefold() for tag in vibe.get("exclude", [])}
|
||||||
|
missing = {"hard rock", "hair metal"} - excluded
|
||||||
|
assert not missing, f"{vibe['name']} selects on {decades} without excluding {missing}"
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_eighties_mood_covers_the_canon():
|
||||||
|
"""Depeche Mode, Duran Duran, Eurythmics and Frankie Goes to Hollywood --
|
||||||
|
checked against their live Last.fm tags. Three of the four carry "synth pop"
|
||||||
|
with a space; only one carries it without."""
|
||||||
|
synths = next(v for v in music_curator.DEFAULT_VIBES if v["name"] == "80s-synths")
|
||||||
|
tags = {t.casefold() for t in synths["tags"]}
|
||||||
|
for artist_tags in (
|
||||||
|
{"80s", "new wave", "pop", "female vocalists", "synth pop"}, # Eurythmics
|
||||||
|
{"80s", "new wave", "pop", "british", "dance"}, # Frankie
|
||||||
|
{"electronic", "synthpop", "new wave", "80s", "synth pop"}, # Depeche Mode
|
||||||
|
{"new wave", "80s", "pop", "synth pop", "rock"}, # Duran Duran
|
||||||
|
):
|
||||||
|
assert tags & artist_tags, artist_tags
|
||||||
|
assert synths["years"] == [1975, 1992]
|
||||||
|
|
||||||
|
|
||||||
|
def test_an_excluded_tag_drops_the_artist(tmp_path):
|
||||||
|
"""Weighting cannot separate the eighties synth acts from the eighties
|
||||||
|
stadium rock, because the tag they would be weighted on is the one they
|
||||||
|
share."""
|
||||||
|
store, source, mirror = tagged_store(
|
||||||
|
tmp_path,
|
||||||
|
{
|
||||||
|
"Played Band": [{"name": "80s", "count": 100}, {"name": "hard rock", "count": 90}],
|
||||||
|
"Silent Band": [{"name": "80s", "count": 100}, {"name": "synth pop", "count": 90}],
|
||||||
|
},
|
||||||
|
)
|
||||||
|
vibes = [{"name": "eighties", "tags": ["80s", "synth pop"], "exclude": ["hard rock"]}]
|
||||||
|
|
||||||
|
music_curator.build_vibe_playlists(store, vibes, mirror, str(source), 100, NOW)
|
||||||
|
|
||||||
|
written = (mirror / "_playlists" / "eighties.m3u").read_text()
|
||||||
|
assert "Silent Band" in written
|
||||||
|
assert "Played Band" not in written
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_vibe_cannot_both_select_and_exclude_a_tag(tmp_path):
|
||||||
|
path = tmp_path / "vibes.json"
|
||||||
|
path.write_text(json.dumps([{"name": "x", "tags": ["80s"], "exclude": ["80s"]}]))
|
||||||
|
|
||||||
|
with pytest.raises(ValueError, match="selects on and excludes"):
|
||||||
|
music_curator.load_vibes(str(path))
|
||||||
|
|
||||||
|
|
||||||
|
def test_matching_ownership_is_a_no_op_when_it_already_agrees(tmp_path):
|
||||||
|
target = tmp_path / "file"
|
||||||
|
target.write_text("x")
|
||||||
|
|
||||||
|
assert music_curator.match_ownership(target, tmp_path) is False
|
||||||
|
|
||||||
|
|
||||||
|
def test_ownership_failure_is_tolerated(tmp_path):
|
||||||
|
"""Not permitted unless running as root -- which is exactly the case where
|
||||||
|
the ownership is already whatever the caller runs as."""
|
||||||
|
target = tmp_path / "file"
|
||||||
|
target.write_text("x")
|
||||||
|
|
||||||
|
# uid 0 from a non-root test process: refused, and must not raise.
|
||||||
|
assert music_curator.set_ownership(target, 0, 0) is False
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_playlist_is_chowned_to_match_the_mirror(tmp_path, monkeypatch):
|
||||||
|
"""The image runs as root, so its output is root-owned, and a root-owned
|
||||||
|
playlist in an apps-owned mirror is unreadable to whatever serves it."""
|
||||||
|
api, source, mirror = playlist_library(tmp_path)
|
||||||
|
store = store_at(tmp_path)
|
||||||
|
music_curator.index_library(
|
||||||
|
music_curator.Lidarr("http://lidarr", "key", transport=api), store
|
||||||
|
)
|
||||||
|
music_curator.match_library(store)
|
||||||
|
|
||||||
|
attempted = []
|
||||||
|
real_stat = music_curator.Path.stat
|
||||||
|
|
||||||
|
def pretend_mirror_is_owned_by_568(self, *args, **kwargs):
|
||||||
|
info = real_stat(self, *args, **kwargs)
|
||||||
|
if self == mirror:
|
||||||
|
return os.stat_result(
|
||||||
|
(info.st_mode, info.st_ino, info.st_dev, info.st_nlink, 568, 568,
|
||||||
|
info.st_size, int(info.st_atime), int(info.st_mtime), int(info.st_ctime))
|
||||||
|
)
|
||||||
|
return info
|
||||||
|
|
||||||
|
monkeypatch.setattr(music_curator.Path, "stat", pretend_mirror_is_owned_by_568)
|
||||||
|
monkeypatch.setattr(
|
||||||
|
music_curator.os, "chown", lambda p, u, g: attempted.append((str(p), u, g))
|
||||||
|
)
|
||||||
|
|
||||||
|
music_curator.build_playlists(store, mirror, str(source), 100, NOW)
|
||||||
|
|
||||||
|
assert attempted, "no ownership was applied"
|
||||||
|
assert all(tuple(owner) == (568, 568) for _, *owner in attempted)
|
||||||
|
# The temporary file, before the rename, never the finished playlist.
|
||||||
|
assert all(path.endswith(".part") or path.endswith("_playlists") for path, *_ in attempted)
|
||||||
|
|||||||
Reference in New Issue
Block a user