Compare commits
16
Commits
v0.2.3
..
97b27f8daa
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
97b27f8daa | ||
|
|
88b96b8159 | ||
|
|
15e5ee5aea | ||
|
|
05cca3508c | ||
|
|
cda3d8463b | ||
|
|
08f099aaa9 | ||
|
|
719a6372ad | ||
|
|
61ce751ac5 | ||
|
|
d1c10d32d9 | ||
|
|
acd042ad8d | ||
|
|
45d99ff039 | ||
|
|
54e19992e4 | ||
|
|
2886c02a2e | ||
|
|
56a06a6cd5 | ||
|
|
a3d0689c0a | ||
|
|
7588fee302 |
@@ -45,7 +45,8 @@ jobs:
|
|||||||
|
|
||||||
# The suite runs inside the image, against the interpreter that ships,
|
# The suite runs inside the image, against the interpreter that ships,
|
||||||
# rather than against whatever the runner happens to provide. A failing
|
# rather than against whatever the runner happens to provide. A failing
|
||||||
# test fails the build. Layers are shared with the push build below.
|
# test fails the build. The runtime stage below is built from the same
|
||||||
|
# daemon afterwards, so its layers are already in cache.
|
||||||
- name: Run the test suite inside the image
|
- name: Run the test suite inside the image
|
||||||
run: docker build --target test -t music-curator:test .
|
run: docker build --target test -t music-curator:test .
|
||||||
|
|
||||||
@@ -124,9 +125,6 @@ jobs:
|
|||||||
echo "release=${release}" >> "$GITHUB_OUTPUT"
|
echo "release=${release}" >> "$GITHUB_OUTPUT"
|
||||||
echo "Computed bump=${bump}, release=${release}, base=${base}"
|
echo "Computed bump=${bump}, release=${release}, base=${base}"
|
||||||
|
|
||||||
- name: Set up Buildx
|
|
||||||
uses: docker/setup-buildx-action@d7f5e7f509e45cec5c76c4d5afdd7de93d0b3df5 # v4
|
|
||||||
|
|
||||||
- name: Log in to the Gitea container registry
|
- name: Log in to the Gitea container registry
|
||||||
if: github.event_name != 'pull_request'
|
if: github.event_name != 'pull_request'
|
||||||
uses: docker/login-action@650006c6eb7dba73a995cc03b0b2d7f5ca915bee # v4
|
uses: docker/login-action@650006c6eb7dba73a995cc03b0b2d7f5ca915bee # v4
|
||||||
@@ -135,21 +133,33 @@ jobs:
|
|||||||
username: ${{ github.repository_owner }}
|
username: ${{ github.repository_owner }}
|
||||||
password: ${{ secrets.PACKAGES_TOKEN }}
|
password: ${{ secrets.PACKAGES_TOKEN }}
|
||||||
|
|
||||||
- name: Build and push
|
# Plain `docker build` rather than buildx. buildx boots its own buildkit
|
||||||
uses: docker/build-push-action@f9f3042f7e2789586610d6e8b85c8f03e5195baf # v7
|
# in a container with a cache of its own, so it shared nothing with the
|
||||||
with:
|
# test build above and rebuilt the image from the base image up -- two
|
||||||
context: .
|
# full builds per run. It earns that cost when building for several
|
||||||
# Without this the last stage in the Dockerfile -- the test stage --
|
# platforms; this only ever targets the amd64 NAS, so it does not.
|
||||||
# would be what gets published.
|
#
|
||||||
target: runtime
|
# `--target runtime` is a strict prefix of the test stage, so every layer
|
||||||
# The NAS is the only host this runs on. Building arm64 as well would
|
# is already in the daemon's cache and this resolves in seconds.
|
||||||
# mean emulating it under QEMU for no consumer.
|
- name: Build the runtime image
|
||||||
platforms: linux/amd64
|
run: |
|
||||||
push: ${{ github.event_name != 'pull_request' }}
|
set -euo pipefail
|
||||||
tags: ${{ steps.version.outputs.tags }}
|
tags=()
|
||||||
labels: |
|
while IFS= read -r tag; do
|
||||||
org.opencontainers.image.source=${{ github.server_url }}/${{ github.repository }}
|
[ -n "$tag" ] && tags+=(-t "$tag")
|
||||||
org.opencontainers.image.revision=${{ github.sha }}
|
done <<< "${{ steps.version.outputs.tags }}"
|
||||||
|
docker build --target runtime \
|
||||||
|
--label "org.opencontainers.image.source=${GITHUB_SERVER_URL}/${GITHUB_REPOSITORY}" \
|
||||||
|
--label "org.opencontainers.image.revision=${GITHUB_SHA}" \
|
||||||
|
"${tags[@]}" .
|
||||||
|
|
||||||
|
- name: Push
|
||||||
|
if: github.event_name != 'pull_request'
|
||||||
|
run: |
|
||||||
|
set -euo pipefail
|
||||||
|
while IFS= read -r tag; do
|
||||||
|
[ -n "$tag" ] && docker push "$tag"
|
||||||
|
done <<< "${{ steps.version.outputs.tags }}"
|
||||||
|
|
||||||
# Record the release: write the computed version into pyproject.toml, then
|
# Record the release: write the computed version into pyproject.toml, then
|
||||||
# commit and tag it, so the packaging metadata always matches the release
|
# commit and tag it, so the packaging metadata always matches the release
|
||||||
|
|||||||
@@ -7,9 +7,11 @@ which keeps an MP3 copy of a lossless library for an iPod. This one answers the
|
|||||||
question that mirror cannot: which of it is worth carrying, and which of it has
|
question that mirror cannot: which of it is worth carrying, and which of it has
|
||||||
not been played in years.
|
not been played in years.
|
||||||
|
|
||||||
**This is stage two.** It ingests the scrobble history, indexes the library from
|
**This is stage three.** It ingests the scrobble history, indexes the library
|
||||||
Lidarr, and matches one to the other. There are no playlists yet, and it writes
|
from Lidarr, matches one to the other, and writes playlists into the mirror —
|
||||||
nothing back — every Lidarr call is a `GET`. See "Where this is going" below.
|
by listening history and by mood.
|
||||||
|
Nothing is written back to Lidarr — every call there is a `GET`. See "Where this
|
||||||
|
is going" below.
|
||||||
|
|
||||||
## What it does today
|
## What it does today
|
||||||
|
|
||||||
@@ -18,6 +20,7 @@ nothing back — every Lidarr call is a `GET`. See "Where this is going" below.
|
|||||||
- Indexes every artist, album and track Lidarr knows about, with file paths and
|
- Indexes every artist, album and track Lidarr knows about, with file paths and
|
||||||
the date each file landed.
|
the date each file landed.
|
||||||
- Ties the two together and reports how well it managed.
|
- Ties the two together and reports how well it managed.
|
||||||
|
- Writes M3U playlists into the mirror, from the listening history.
|
||||||
|
|
||||||
## Matching
|
## Matching
|
||||||
|
|
||||||
@@ -29,7 +32,14 @@ Two tiers, and no third.
|
|||||||
| `name` | Normalised artist and title | Everything the first tier could not carry |
|
| `name` | Normalised artist and title | Everything the first tier could not carry |
|
||||||
| `none` | — | Recorded as a miss, never guessed at |
|
| `none` | — | Recorded as a miss, never guessed at |
|
||||||
|
|
||||||
The normalisation is the load-bearing part, because the two sides disagree in
|
The name tier does almost all of the work. A recording MBID is exact when it
|
||||||
|
lands, but MusicBrainz holds a separate recording per release, and Last.fm and
|
||||||
|
Lidarr rarely pick the same one: on a real library, two thirds of scrobbles
|
||||||
|
carry a recording id and barely a twentieth of them join on it. The report
|
||||||
|
counts how many carried an id and matched on name anyway, which is the measure
|
||||||
|
of that disagreement.
|
||||||
|
|
||||||
|
So the normalisation is the load-bearing part, because the two sides disagree in
|
||||||
predictable ways. It folds case and accents, drops guest credits (`Yellowcard
|
predictable ways. It folds case and accents, drops guest credits (`Yellowcard
|
||||||
feat. Tay Jardine` against a tag of `Yellowcard`), strips a trailing
|
feat. Tay Jardine` against a tag of `Yellowcard`), strips a trailing
|
||||||
version suffix (`(Remastered 2011)`, `- Live`), expands `&`, and removes a
|
version suffix (`(Remastered 2011)`, `- Live`), expands `&`, and removes a
|
||||||
@@ -102,10 +112,108 @@ matcher. The line to watch is:
|
|||||||
unmatched by an artist the library holds: N pairs, M plays
|
unmatched by an artist the library holds: N pairs, M plays
|
||||||
```
|
```
|
||||||
|
|
||||||
That is a track that was played, sitting beside a file it should have matched.
|
That is a track that was played by an artist the library holds. It is then
|
||||||
Those are the matcher's real misses, and every one is a candidate for being
|
split three ways, because owning an artist is a weak proxy for owning a track
|
||||||
wrongly called cold in stage four. The report lists the worst fifteen by play
|
and a shared title is a weak proxy for a shared song:
|
||||||
count so they can be eyeballed.
|
|
||||||
|
- **the library's own title credits the scrobbled artist** — `Voodoo People
|
||||||
|
(Pendulum Remix)` against a play credited to Pendulum. Same song, filed
|
||||||
|
under the original artist. These are the genuine misses.
|
||||||
|
- **the same title under an unrelated artist** — a collision, not a miss.
|
||||||
|
Across fifty thousand tracks these are constant: `Everyday` is Rusko and
|
||||||
|
also Def Leppard, `Kaleidoscope` is Delta Heavy and also Chappell Roan.
|
||||||
|
Matching on title alone would be far worse than missing them, which is why
|
||||||
|
there is no such tier.
|
||||||
|
- **the title is nowhere in the library** — never bought.
|
||||||
|
|
||||||
|
Only the first is worth chasing. Counting all three as matcher failures
|
||||||
|
overstates the problem and would over-block the cull.
|
||||||
|
|
||||||
|
## Playlists
|
||||||
|
|
||||||
|
Written into `<mirror>/_playlists/` as extended M3U, rebuilt every pass. Six
|
||||||
|
rules, capped at `--playlist-limit` tracks each:
|
||||||
|
|
||||||
|
| Playlist | Rule |
|
||||||
|
| -------------------- | --------------------------------------------------------- |
|
||||||
|
| `heavy-rotation` | Most played over the last twelve months |
|
||||||
|
| `all-time` | Most played ever |
|
||||||
|
| `neglected` | Played heavily once, silent for twelve months |
|
||||||
|
| `deep-cuts` | Never played, from albums whose other tracks you play constantly |
|
||||||
|
| `unheard-favourites` | Never played, by the artists you play most |
|
||||||
|
| `unheard` | Never played, anywhere in the library |
|
||||||
|
|
||||||
|
Ninety days was the obvious window for "recent" and is the wrong one: on a real
|
||||||
|
history it holds a few hundred plays spread thinly across a twenty-thousand
|
||||||
|
track rotation, so nothing ranks meaningfully. Twelve months does.
|
||||||
|
|
||||||
|
The two `unheard` playlists rotate **weekly**, not per pass. A pass runs every
|
||||||
|
few hours, and a playlist that reorders itself each time is one that has to be
|
||||||
|
re-imported each time — the Music app imports a snapshot of a file, it does not
|
||||||
|
track it.
|
||||||
|
|
||||||
|
### Moods
|
||||||
|
|
||||||
|
A second set of playlists selects by **Last.fm's crowd tags** rather than by
|
||||||
|
listening history. MusicBrainz genres arrive free with the Lidarr index and are
|
||||||
|
no use for this: they are sparse and formal, and will not tell you a record is
|
||||||
|
screamo or synthwave. People typing tags will.
|
||||||
|
|
||||||
|
Tags are fetched once per artist — one `artist.getTopTags` call each, by
|
||||||
|
MusicBrainz id where Lidarr has one — and refreshed every ninety days. An
|
||||||
|
artist Last.fm has never heard of is recorded as fetched with no tags, so it is
|
||||||
|
not asked about again on every pass. `--tag-limit` spreads the first sweep over
|
||||||
|
several passes.
|
||||||
|
|
||||||
|
The built-in moods are chosen for this library rather than as a general
|
||||||
|
taxonomy:
|
||||||
|
|
||||||
|
| Mood | Selected on |
|
||||||
|
| ------------------ | ------------------------------------------------- |
|
||||||
|
| `80s-synths` | synthpop, new wave, synthwave — released 1975-1992 |
|
||||||
|
| `high-energy-rock` | hard rock, punk, pop punk, alternative |
|
||||||
|
| `screamo` | screamo, post-hardcore, metalcore, emo |
|
||||||
|
| `drum-and-bass` | drum and bass, liquid funk, neurofunk, jungle |
|
||||||
|
| `dance` | house, big room, hardstyle, trance, dubstep |
|
||||||
|
| `classic-rock` | classic rock, prog, psychedelic, blues rock |
|
||||||
|
|
||||||
|
An artist qualifies when their tag weights inside a mood sum to at least 30 out
|
||||||
|
of Last.fm's 0-100 scale. One low-weight tag is not a genre, it is somebody's
|
||||||
|
stray opinion.
|
||||||
|
|
||||||
|
`years` filters on the album's release date, which is what separates eighties
|
||||||
|
synth records from everything else a synthpop tag drags in.
|
||||||
|
|
||||||
|
`--vibes` replaces the whole set with a JSON file of the same shape, so a new
|
||||||
|
mood does not need a new release:
|
||||||
|
|
||||||
|
```json
|
||||||
|
[{ "name": "shoegaze", "tags": ["shoegaze", "dream pop"], "min_score": 40 }]
|
||||||
|
```
|
||||||
|
|
||||||
|
Names are validated when the file is read, not when the file is written. A bad
|
||||||
|
one would otherwise surface as a playlist created somewhere unintended.
|
||||||
|
|
||||||
|
### Paths
|
||||||
|
|
||||||
|
Lidarr knows where the lossless source is; the playlists have to point at the
|
||||||
|
MP3s music-mirror made from it. The mapping strips a library root from Lidarr's
|
||||||
|
track paths and re-roots them under the mirror, with the suffix changed.
|
||||||
|
|
||||||
|
`--library-root` is derived from the common parent of the indexed artist folders
|
||||||
|
when unset, so it agrees with Lidarr by construction rather than by being kept
|
||||||
|
in step by hand. Override it if that guess is wrong.
|
||||||
|
|
||||||
|
Entries are written **relative to the playlist file**, so one playlist works
|
||||||
|
from the NAS, from a Mac over SMB, and from Linux, without rewriting.
|
||||||
|
|
||||||
|
A track is only listed once its mirror file has been confirmed to exist. Lidarr
|
||||||
|
holding the FLAC says nothing about whether the MP3 has been encoded yet. If a
|
||||||
|
large number are missing, the run says so — that is what a wrong `--library-root`
|
||||||
|
or `--mirror` looks like, since the paths then map to nothing at all.
|
||||||
|
|
||||||
|
`_playlists/` survives music-mirror's prune: it only deletes `*.mp3`, and its
|
||||||
|
empty-directory sweep skips a directory holding M3Us.
|
||||||
|
|
||||||
## How the ingest works
|
## How the ingest works
|
||||||
|
|
||||||
@@ -161,6 +269,11 @@ music-curator --report-only # report on the store, fetch nothing
|
|||||||
| `--backfill-limit` | `MUSIC_CURATOR_BACKFILL_LIMIT` | `0` | Cap backfill requests per pass; 0 for no cap |
|
| `--backfill-limit` | `MUSIC_CURATOR_BACKFILL_LIMIT` | `0` | Cap backfill requests per pass; 0 for no cap |
|
||||||
| `--lidarr-url` | `MUSIC_CURATOR_LIDARR_URL` | unset | Lidarr base URL, e.g. `http://lidarr:8686` |
|
| `--lidarr-url` | `MUSIC_CURATOR_LIDARR_URL` | unset | Lidarr base URL, e.g. `http://lidarr:8686` |
|
||||||
| `--lidarr-api-key` | `MUSIC_CURATOR_LIDARR_API_KEY` | unset | Lidarr API key |
|
| `--lidarr-api-key` | `MUSIC_CURATOR_LIDARR_API_KEY` | unset | Lidarr API key |
|
||||||
|
| `--mirror` | `MUSIC_CURATOR_MIRROR` | unset | Root of the MP3 mirror; playlists go here |
|
||||||
|
| `--library-root` | `MUSIC_CURATOR_LIBRARY_ROOT` | derived | Prefix to strip from Lidarr's paths |
|
||||||
|
| `--playlist-limit` | `MUSIC_CURATOR_PLAYLIST_LIMIT` | `100` | Most tracks in any one playlist |
|
||||||
|
| `--vibes` | `MUSIC_CURATOR_VIBES` | built-in | JSON file of mood definitions |
|
||||||
|
| `--tag-limit` | `MUSIC_CURATOR_TAG_LIMIT` | `0` | Cap artist tag lookups per pass |
|
||||||
| `--skip-index` | — | off | Match against the index already held |
|
| `--skip-index` | — | off | Match against the index already held |
|
||||||
| `--report-only` | — | off | Report without fetching |
|
| `--report-only` | — | off | Report without fetching |
|
||||||
|
|
||||||
@@ -209,7 +322,8 @@ nix shell nixpkgs#python3Packages.pytest -c pytest
|
|||||||
| ------------------------------------------------ | ------------ |
|
| ------------------------------------------------ | ------------ |
|
||||||
| Last.fm ingest and store | done |
|
| Last.fm ingest and store | done |
|
||||||
| Lidarr index and the scrobble-to-track matcher | done |
|
| Lidarr index and the scrobble-to-track matcher | done |
|
||||||
| M3U playlists written into the mirror | next |
|
| M3U playlists from the listening history | done |
|
||||||
|
| Genre and mood playlists from Last.fm tags | done |
|
||||||
| Cold-music report, unmonitoring what is not played | last |
|
| Cold-music report, unmonitoring what is not played | last |
|
||||||
|
|
||||||
The cull will unmonitor cold albums in Lidarr and tag their artists. It will
|
The cull will unmonitor cold albums in Lidarr and tag their artists. It will
|
||||||
|
|||||||
@@ -30,5 +30,11 @@ services:
|
|||||||
# Cap the backfill at this many requests per pass. Unlimited by default,
|
# Cap the backfill at this many requests per pass. Unlimited by default,
|
||||||
# which finishes a long history in one go.
|
# which finishes a long history in one go.
|
||||||
# MUSIC_CURATOR_BACKFILL_LIMIT: "0"
|
# MUSIC_CURATOR_BACKFILL_LIMIT: "0"
|
||||||
|
# The MP3 mirror music-mirror maintains. Playlists are written into
|
||||||
|
# _playlists/ inside it; leave unset to skip them.
|
||||||
|
MUSIC_CURATOR_MIRROR: /mirror
|
||||||
|
# Most tracks in any one playlist.
|
||||||
|
# MUSIC_CURATOR_PLAYLIST_LIMIT: "100"
|
||||||
volumes:
|
volumes:
|
||||||
- /mnt/tank/apps/music-curator:/data
|
- /mnt/tank/apps/music-curator:/data
|
||||||
|
- /mnt/tank/media/music-mp3:/mirror
|
||||||
|
|||||||
+644
-10
@@ -26,6 +26,7 @@ import re
|
|||||||
import signal
|
import signal
|
||||||
import sqlite3
|
import sqlite3
|
||||||
import sys
|
import sys
|
||||||
|
import tempfile
|
||||||
import time
|
import time
|
||||||
import unicodedata
|
import unicodedata
|
||||||
import urllib.error
|
import urllib.error
|
||||||
@@ -62,6 +63,28 @@ REQUEST_DELAY_SECONDS = 0.25
|
|||||||
BACKOFF_SECONDS = 2.0
|
BACKOFF_SECONDS = 2.0
|
||||||
BACKOFF_CEILING_SECONDS = 60.0
|
BACKOFF_CEILING_SECONDS = 60.0
|
||||||
|
|
||||||
|
# What music-mirror names its output, and where playlists are written inside the
|
||||||
|
# mirror. music-mirror's prune only deletes `*.mp3` and only removes directories
|
||||||
|
# it finds empty, so a directory of M3Us survives it untouched.
|
||||||
|
MIRROR_SUFFIX = ".mp3"
|
||||||
|
PLAYLIST_DIRECTORY = "_playlists"
|
||||||
|
|
||||||
|
# Tags are re-fetched this often. They move slowly, and the first pass over a
|
||||||
|
# library already costs one request per artist.
|
||||||
|
TAG_REFRESH_SECONDS = 90 * 86400
|
||||||
|
|
||||||
|
# How much of an artist's tag weight has to fall inside a vibe before their
|
||||||
|
# tracks are eligible for it. Weights are Last.fm's own 0-100 popularity, summed
|
||||||
|
# across whichever of the vibe's tags the artist carries.
|
||||||
|
VIBE_MIN_SCORE = 30
|
||||||
|
|
||||||
|
# A vibe writes a file named after itself, so the name has to be a filename.
|
||||||
|
SAFE_VIBE_NAME = re.compile(r"^[a-z0-9][a-z0-9-]*$")
|
||||||
|
|
||||||
|
# Group-readable, for the same reason music-mirror sets it: the mirror is read
|
||||||
|
# back by whatever serves it, and a playlist nobody can read is not a playlist.
|
||||||
|
GROUP_READ = 0o040
|
||||||
|
|
||||||
SCHEMA_VERSION = "2"
|
SCHEMA_VERSION = "2"
|
||||||
|
|
||||||
# Versions this build upgrades in place. Everything added since version 1 is a
|
# Versions this build upgrades in place. Everything added since version 1 is a
|
||||||
@@ -141,6 +164,11 @@ CREATE TABLE IF NOT EXISTS lidarr_track (
|
|||||||
);
|
);
|
||||||
CREATE INDEX IF NOT EXISTS lidarr_track_recording ON lidarr_track (recording_mbid);
|
CREATE INDEX IF NOT EXISTS lidarr_track_recording ON lidarr_track (recording_mbid);
|
||||||
CREATE INDEX IF NOT EXISTS lidarr_track_norm ON lidarr_track (norm_artist, norm_title);
|
CREATE INDEX IF NOT EXISTS lidarr_track_norm ON lidarr_track (norm_artist, norm_title);
|
||||||
|
-- Separate from the composite above, which a lookup by title alone cannot use:
|
||||||
|
-- its leading column is the artist. The report searches by title on its own, and
|
||||||
|
-- without this it scans every track for every unmatched key -- forty-seven
|
||||||
|
-- seconds on a library of eighty-four thousand.
|
||||||
|
CREATE INDEX IF NOT EXISTS lidarr_track_title ON lidarr_track (norm_title);
|
||||||
CREATE INDEX IF NOT EXISTS lidarr_track_album ON lidarr_track (album_id);
|
CREATE INDEX IF NOT EXISTS lidarr_track_album ON lidarr_track (album_id);
|
||||||
|
|
||||||
-- One row per distinct thing listened to, with the verdict on whether it could
|
-- One row per distinct thing listened to, with the verdict on whether it could
|
||||||
@@ -161,6 +189,24 @@ CREATE TABLE IF NOT EXISTS scrobble_key (
|
|||||||
PRIMARY KEY (artist, track)
|
PRIMARY KEY (artist, track)
|
||||||
);
|
);
|
||||||
CREATE INDEX IF NOT EXISTS scrobble_key_track ON scrobble_key (track_id);
|
CREATE INDEX IF NOT EXISTS scrobble_key_track ON scrobble_key (track_id);
|
||||||
|
|
||||||
|
-- Last.fm's crowd tags for each library artist. MusicBrainz genres come free
|
||||||
|
-- with the Lidarr index and are useless for this: they are sparse and formal,
|
||||||
|
-- and will not tell you a record is screamo or synthwave. People typing tags
|
||||||
|
-- will. Keyed on the normalised name rather than the Lidarr id so it survives
|
||||||
|
-- an artist being removed and re-added there.
|
||||||
|
CREATE TABLE IF NOT EXISTS artist_tag (
|
||||||
|
norm_artist TEXT NOT NULL,
|
||||||
|
tag TEXT NOT NULL,
|
||||||
|
weight INTEGER NOT NULL,
|
||||||
|
PRIMARY KEY (norm_artist, tag)
|
||||||
|
);
|
||||||
|
CREATE INDEX IF NOT EXISTS artist_tag_tag ON artist_tag (tag);
|
||||||
|
|
||||||
|
CREATE TABLE IF NOT EXISTS artist_tag_fetched (
|
||||||
|
norm_artist TEXT PRIMARY KEY,
|
||||||
|
fetched_at INTEGER NOT NULL
|
||||||
|
);
|
||||||
"""
|
"""
|
||||||
|
|
||||||
|
|
||||||
@@ -187,18 +233,28 @@ VERSION_WORDS = (
|
|||||||
"anniversary",
|
"anniversary",
|
||||||
"reissue",
|
"reissue",
|
||||||
"instrumental",
|
"instrumental",
|
||||||
|
# Drum and bass and its neighbours mark versions their own way.
|
||||||
|
"vip",
|
||||||
|
"bootleg",
|
||||||
|
"rework",
|
||||||
|
"extended",
|
||||||
)
|
)
|
||||||
_VERSIONS = "|".join(VERSION_WORDS)
|
_VERSIONS = "|".join(VERSION_WORDS)
|
||||||
BRACKETED_VERSION = re.compile(
|
BRACKETED_VERSION = re.compile(
|
||||||
rf"\s*[\(\[][^\)\]]*\b(?:{_VERSIONS})\b[^\)\]]*[\)\]]\s*$", re.IGNORECASE
|
rf"\s*[\(\[][^\)\]]*\b(?:{_VERSIONS})\b[^\)\]]*[\)\]]\s*$", re.IGNORECASE
|
||||||
)
|
)
|
||||||
TRAILING_VERSION = re.compile(rf"\s+-\s+[^-]*\b(?:{_VERSIONS})\b.*$", re.IGNORECASE)
|
# The suffix is matched lazily rather than as a run of non-hyphens, because the
|
||||||
|
# thing being stripped frequently contains hyphens of its own -- "Gold Dust -
|
||||||
|
# Shy FX Re-Edit", "Back To Your Roots - Friction & K-Tee Remix".
|
||||||
|
TRAILING_VERSION = re.compile(rf"\s+-\s+.*?\b(?:{_VERSIONS})\b.*$", re.IGNORECASE)
|
||||||
|
|
||||||
# Last.fm routinely carries the guest credit in the artist field where the file
|
# Last.fm routinely carries the guest credit where the file tag holds only the
|
||||||
# tag holds only the primary artist -- "Yellowcard feat. Tay Jardine" against a
|
# primary artist -- "Yellowcard feat. Tay Jardine" against a tag of
|
||||||
# tag of "Yellowcard". `with` is deliberately absent: it appears in far too many
|
# "Yellowcard", or "Self vs Self (feat. In Flames)" against "Self vs Self". The
|
||||||
# real titles to cut on sight.
|
# opening bracket has to be allowed for: requiring whitespace immediately before
|
||||||
GUEST_CREDIT = re.compile(r"\s+(?:feat|ft|featuring)\b.*$", re.IGNORECASE)
|
# the word misses every bracketed credit, which is most of them. `with` is
|
||||||
|
# deliberately absent -- it appears in far too many real titles to cut on sight.
|
||||||
|
GUEST_CREDIT = re.compile(r"[\s(\[]+(?:feat|ft|featuring)\b.*$", re.IGNORECASE)
|
||||||
|
|
||||||
LEADING_ARTICLE = re.compile(r"^the\s+")
|
LEADING_ARTICLE = re.compile(r"^the\s+")
|
||||||
# Deleted rather than spaced, so "Don't" and "Dont" agree. Every other mark
|
# Deleted rather than spaced, so "Don't" and "Dont" agree. Every other mark
|
||||||
@@ -520,6 +576,24 @@ class Store:
|
|||||||
tracks,
|
tracks,
|
||||||
)
|
)
|
||||||
|
|
||||||
|
def replace_tags(self, norm_artist, pairs, now):
|
||||||
|
"""Replace one artist's tags, and record that they were fetched.
|
||||||
|
|
||||||
|
The fetch is recorded even when nothing came back, so an artist Last.fm
|
||||||
|
has never heard of is not asked about again on every pass.
|
||||||
|
"""
|
||||||
|
with self.connection:
|
||||||
|
self.connection.execute("DELETE FROM artist_tag WHERE norm_artist = ?", (norm_artist,))
|
||||||
|
self.connection.executemany(
|
||||||
|
"INSERT OR IGNORE INTO artist_tag (norm_artist, tag, weight) VALUES (?, ?, ?)",
|
||||||
|
[(norm_artist, tag, weight) for tag, weight in pairs],
|
||||||
|
)
|
||||||
|
self.connection.execute(
|
||||||
|
"INSERT INTO artist_tag_fetched (norm_artist, fetched_at) VALUES (?, ?)"
|
||||||
|
" ON CONFLICT (norm_artist) DO UPDATE SET fetched_at = excluded.fetched_at",
|
||||||
|
(norm_artist, now),
|
||||||
|
)
|
||||||
|
|
||||||
def rebuild_keys(self):
|
def rebuild_keys(self):
|
||||||
"""Collapse the scrobble history into one row per distinct track.
|
"""Collapse the scrobble history into one row per distinct track.
|
||||||
|
|
||||||
@@ -1031,6 +1105,419 @@ def format_time(uts):
|
|||||||
return datetime.fromtimestamp(uts, tz=timezone.utc).strftime("%Y-%m-%d %H:%M")
|
return datetime.fromtimestamp(uts, tz=timezone.utc).strftime("%Y-%m-%d %H:%M")
|
||||||
|
|
||||||
|
|
||||||
|
# Playlists written into the mirror. Each is a rule over the listening history
|
||||||
|
# and the library index; none of them look at the audio.
|
||||||
|
#
|
||||||
|
# The rotating ones are shuffled by week rather than by pass. A pass runs every
|
||||||
|
# few hours, and a playlist that reorders itself every time is one that has to
|
||||||
|
# be re-imported every time -- the Music app does not track a file, it imports a
|
||||||
|
# snapshot of one.
|
||||||
|
ROTATION_PERIOD_SECONDS = 7 * 86400
|
||||||
|
SHUFFLE_MULTIPLIER = 2654435761
|
||||||
|
SHUFFLE_MODULUS = 104729
|
||||||
|
|
||||||
|
# Common table expressions the rules share. `played` is every library track that
|
||||||
|
# has ever been matched to a scrobble, with its count and the last time it was
|
||||||
|
# heard; `recent` is the same restricted to a window.
|
||||||
|
PLAYED_CTE = """
|
||||||
|
WITH played AS (
|
||||||
|
SELECT sk.track_id AS track_id, COUNT(*) AS plays, MAX(s.uts) AS last_uts
|
||||||
|
FROM scrobble s
|
||||||
|
JOIN scrobble_key sk ON sk.artist = s.artist AND sk.track = s.track
|
||||||
|
WHERE sk.track_id IS NOT NULL
|
||||||
|
GROUP BY sk.track_id
|
||||||
|
)
|
||||||
|
"""
|
||||||
|
RECENT_CTE = """
|
||||||
|
WITH recent AS (
|
||||||
|
SELECT sk.track_id AS track_id, COUNT(*) AS plays
|
||||||
|
FROM scrobble s
|
||||||
|
JOIN scrobble_key sk ON sk.artist = s.artist AND sk.track = s.track
|
||||||
|
WHERE sk.track_id IS NOT NULL AND s.uts >= :year_ago
|
||||||
|
GROUP BY sk.track_id
|
||||||
|
)
|
||||||
|
"""
|
||||||
|
SELECT_TRACK = """
|
||||||
|
SELECT t.id AS id, a.name AS artist, t.title AS title,
|
||||||
|
t.duration AS duration, t.path AS path
|
||||||
|
FROM lidarr_track t
|
||||||
|
JOIN lidarr_artist a ON a.id = t.artist_id
|
||||||
|
"""
|
||||||
|
PLAYABLE = " t.has_file = 1 AND t.path IS NOT NULL AND t.path <> '' "
|
||||||
|
SHUFFLE = f" ((t.id * {SHUFFLE_MULTIPLIER}) + :week) % {SHUFFLE_MODULUS} "
|
||||||
|
|
||||||
|
PLAYLISTS = (
|
||||||
|
(
|
||||||
|
"heavy-rotation",
|
||||||
|
"most played over the last twelve months",
|
||||||
|
RECENT_CTE + SELECT_TRACK + f"""
|
||||||
|
JOIN recent r ON r.track_id = t.id
|
||||||
|
WHERE {PLAYABLE}
|
||||||
|
ORDER BY r.plays DESC, a.name, t.title
|
||||||
|
LIMIT :limit
|
||||||
|
""",
|
||||||
|
),
|
||||||
|
(
|
||||||
|
"all-time",
|
||||||
|
"most played ever",
|
||||||
|
PLAYED_CTE + SELECT_TRACK + f"""
|
||||||
|
JOIN played p ON p.track_id = t.id
|
||||||
|
WHERE {PLAYABLE}
|
||||||
|
ORDER BY p.plays DESC, a.name, t.title
|
||||||
|
LIMIT :limit
|
||||||
|
""",
|
||||||
|
),
|
||||||
|
(
|
||||||
|
"neglected",
|
||||||
|
"played heavily once, silent for twelve months",
|
||||||
|
PLAYED_CTE + SELECT_TRACK + f"""
|
||||||
|
JOIN played p ON p.track_id = t.id
|
||||||
|
WHERE {PLAYABLE} AND p.last_uts < :year_ago
|
||||||
|
ORDER BY p.plays DESC, a.name, t.title
|
||||||
|
LIMIT :limit
|
||||||
|
""",
|
||||||
|
),
|
||||||
|
(
|
||||||
|
"deep-cuts",
|
||||||
|
"never played, from albums whose other tracks you play constantly",
|
||||||
|
PLAYED_CTE + """,
|
||||||
|
album_plays AS (
|
||||||
|
SELECT t.album_id AS album_id, SUM(p.plays) AS plays
|
||||||
|
FROM played p JOIN lidarr_track t ON t.id = p.track_id
|
||||||
|
GROUP BY t.album_id
|
||||||
|
)
|
||||||
|
""" + SELECT_TRACK + f"""
|
||||||
|
JOIN album_plays ap ON ap.album_id = t.album_id
|
||||||
|
LEFT JOIN played p ON p.track_id = t.id
|
||||||
|
WHERE {PLAYABLE} AND p.track_id IS NULL
|
||||||
|
ORDER BY ap.plays DESC, t.id
|
||||||
|
LIMIT :limit
|
||||||
|
""",
|
||||||
|
),
|
||||||
|
(
|
||||||
|
"unheard-favourites",
|
||||||
|
"never played, by the artists you play most; rotates weekly",
|
||||||
|
PLAYED_CTE + """,
|
||||||
|
artist_plays AS (
|
||||||
|
SELECT t.artist_id AS artist_id, SUM(p.plays) AS plays
|
||||||
|
FROM played p JOIN lidarr_track t ON t.id = p.track_id
|
||||||
|
GROUP BY t.artist_id
|
||||||
|
)
|
||||||
|
""" + SELECT_TRACK + f"""
|
||||||
|
LEFT JOIN played p ON p.track_id = t.id
|
||||||
|
WHERE {PLAYABLE} AND p.track_id IS NULL
|
||||||
|
AND t.artist_id IN (SELECT artist_id FROM artist_plays
|
||||||
|
ORDER BY plays DESC LIMIT 50)
|
||||||
|
ORDER BY {SHUFFLE}
|
||||||
|
LIMIT :limit
|
||||||
|
""",
|
||||||
|
),
|
||||||
|
(
|
||||||
|
"unheard",
|
||||||
|
"never played, anywhere in the library; rotates weekly",
|
||||||
|
PLAYED_CTE + SELECT_TRACK + f"""
|
||||||
|
LEFT JOIN played p ON p.track_id = t.id
|
||||||
|
WHERE {PLAYABLE} AND p.track_id IS NULL
|
||||||
|
ORDER BY {SHUFFLE}
|
||||||
|
LIMIT :limit
|
||||||
|
""",
|
||||||
|
),
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
# Moods, as sets of Last.fm tags. Chosen for this library -- drum and bass,
|
||||||
|
# punk and its descendants, big-room dance, classic rock -- rather than as a
|
||||||
|
# general taxonomy. `--vibes` replaces the lot with a JSON file of the same
|
||||||
|
# shape, so a new one does not need a new release.
|
||||||
|
#
|
||||||
|
# `years` filters on the album's release date, which is what separates eighties
|
||||||
|
# synth records from everything a synthpop tag would otherwise drag in.
|
||||||
|
DEFAULT_VIBES = (
|
||||||
|
{
|
||||||
|
"name": "80s-synths",
|
||||||
|
"tags": [
|
||||||
|
"synthpop", "synth pop", "synth-pop", "new wave", "synthwave",
|
||||||
|
"new romantic", "electropop", "80s", "1980s",
|
||||||
|
],
|
||||||
|
"years": [1975, 1992],
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "high-energy-rock",
|
||||||
|
"tags": [
|
||||||
|
"hard rock", "punk rock", "pop punk", "punk", "alternative rock",
|
||||||
|
"rock", "garage rock", "skate punk",
|
||||||
|
],
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "screamo",
|
||||||
|
"tags": [
|
||||||
|
"screamo", "post-hardcore", "metalcore", "emo", "hardcore",
|
||||||
|
"melodic hardcore", "emocore",
|
||||||
|
],
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "drum-and-bass",
|
||||||
|
"tags": [
|
||||||
|
"drum and bass", "drum n bass", "dnb", "liquid funk", "neurofunk",
|
||||||
|
"jungle", "breakbeat",
|
||||||
|
],
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "dance",
|
||||||
|
"tags": [
|
||||||
|
"electro house", "house", "big room", "electronic dance music",
|
||||||
|
"edm", "hardstyle", "trance", "dubstep", "electro",
|
||||||
|
],
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "classic-rock",
|
||||||
|
"tags": [
|
||||||
|
"classic rock", "progressive rock", "psychedelic rock",
|
||||||
|
"blues rock", "70s", "60s",
|
||||||
|
],
|
||||||
|
},
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def load_vibes(path):
|
||||||
|
"""Return the vibe definitions, from a file when one is given.
|
||||||
|
|
||||||
|
Validated up front rather than at write time: a bad name would otherwise
|
||||||
|
surface as a file created somewhere unintended, which is a poor way to find
|
||||||
|
out about a typo.
|
||||||
|
"""
|
||||||
|
if not path:
|
||||||
|
return DEFAULT_VIBES
|
||||||
|
try:
|
||||||
|
loaded = json.loads(Path(path).read_text(encoding="utf-8"))
|
||||||
|
except (OSError, json.JSONDecodeError) as error:
|
||||||
|
raise ValueError(f"cannot read vibes from {path}: {error}") from error
|
||||||
|
|
||||||
|
if not isinstance(loaded, list):
|
||||||
|
raise ValueError(f"{path}: expected a list of vibes")
|
||||||
|
for vibe in loaded:
|
||||||
|
name = vibe.get("name") if isinstance(vibe, dict) else None
|
||||||
|
if not name or not SAFE_VIBE_NAME.fullmatch(str(name)):
|
||||||
|
raise ValueError(f"{path}: {name!r} is not a usable vibe name (a-z, 0-9, -)")
|
||||||
|
if not vibe.get("tags"):
|
||||||
|
raise ValueError(f"{path}: vibe {name!r} lists no tags")
|
||||||
|
return tuple(loaded)
|
||||||
|
|
||||||
|
|
||||||
|
def parse_tags(payload):
|
||||||
|
"""Return (tag, weight) pairs from an artist.getTopTags response.
|
||||||
|
|
||||||
|
The documented sample carries only a name and a URL per tag; a live response
|
||||||
|
also carries a 0-100 `count`. Rather than depend on which, the count is used
|
||||||
|
when present and the documented ordering -- by popularity -- stands in for it
|
||||||
|
when it is not.
|
||||||
|
"""
|
||||||
|
block = payload.get("toptags") or {}
|
||||||
|
pairs = []
|
||||||
|
for position, entry in enumerate(as_list(block.get("tag"))):
|
||||||
|
tag = (entry.get("name") or "").strip().casefold()
|
||||||
|
if not tag:
|
||||||
|
continue
|
||||||
|
weight = int(entry.get("count") or 0) or max(1, 100 - position * 5)
|
||||||
|
pairs.append((tag, weight))
|
||||||
|
return pairs
|
||||||
|
|
||||||
|
|
||||||
|
def sync_tags(client, store, now, limit=0):
|
||||||
|
"""Fetch crowd tags for library artists that have none, or stale ones.
|
||||||
|
|
||||||
|
One request per artist, once, and then only for artists newly added. An
|
||||||
|
artist Last.fm has never heard of is recorded as fetched with no tags, so it
|
||||||
|
is not asked about again every pass.
|
||||||
|
"""
|
||||||
|
stale = store.connection.execute(
|
||||||
|
"SELECT a.norm_name AS norm_name, a.name AS name, a.mbid AS mbid"
|
||||||
|
" FROM lidarr_artist a"
|
||||||
|
" LEFT JOIN artist_tag_fetched f ON f.norm_artist = a.norm_name"
|
||||||
|
" WHERE a.norm_name <> ''"
|
||||||
|
" AND (f.fetched_at IS NULL OR f.fetched_at < :cutoff)"
|
||||||
|
" ORDER BY a.name",
|
||||||
|
{"cutoff": now - TAG_REFRESH_SECONDS},
|
||||||
|
).fetchall()
|
||||||
|
if not stale:
|
||||||
|
return 0
|
||||||
|
|
||||||
|
if limit > 0 and len(stale) > limit:
|
||||||
|
logger.info("tagging %d of %d artists this pass; the rest follow next", limit, len(stale))
|
||||||
|
stale = stale[:limit]
|
||||||
|
else:
|
||||||
|
logger.info("fetching tags for %d artists", len(stale))
|
||||||
|
|
||||||
|
tagged = 0
|
||||||
|
for artist in stale:
|
||||||
|
query = {"mbid": artist["mbid"]} if artist["mbid"] else {"artist": artist["name"]}
|
||||||
|
try:
|
||||||
|
payload = client.call("artist.getTopTags", {**query, "autocorrect": 1})
|
||||||
|
except LastfmError as error:
|
||||||
|
# One artist Last.fm cannot answer for is not worth losing the pass.
|
||||||
|
logger.warning("no tags for %s: %s", artist["name"], error)
|
||||||
|
continue
|
||||||
|
store.replace_tags(artist["norm_name"], parse_tags(payload), now)
|
||||||
|
tagged += 1
|
||||||
|
|
||||||
|
logger.info("tagged %d artists", tagged)
|
||||||
|
return tagged
|
||||||
|
|
||||||
|
|
||||||
|
def build_vibe_playlists(store, vibes, mirror_root, library_root, limit, now):
|
||||||
|
"""Write one playlist per mood. Returns how many tracks were listed."""
|
||||||
|
if not store.scalar("SELECT COUNT(*) FROM artist_tag"):
|
||||||
|
return 0
|
||||||
|
|
||||||
|
directory = Path(mirror_root) / PLAYLIST_DIRECTORY
|
||||||
|
week = now // ROTATION_PERIOD_SECONDS
|
||||||
|
total = 0
|
||||||
|
|
||||||
|
for vibe in vibes:
|
||||||
|
tags = [str(tag).strip().casefold() for tag in vibe["tags"]]
|
||||||
|
placeholders = ",".join("?" * len(tags))
|
||||||
|
years = vibe.get("years")
|
||||||
|
parameters = [*tags, vibe.get("min_score", VIBE_MIN_SCORE)]
|
||||||
|
year_clause = ""
|
||||||
|
if years:
|
||||||
|
year_clause = (
|
||||||
|
" AND CAST(substr(al.release_date, 1, 4) AS INTEGER) BETWEEN ? AND ?"
|
||||||
|
)
|
||||||
|
parameters += [int(years[0]), int(years[1])]
|
||||||
|
parameters += [week, limit]
|
||||||
|
|
||||||
|
sql = f"""
|
||||||
|
WITH vibe AS (
|
||||||
|
SELECT norm_artist, SUM(weight) AS score
|
||||||
|
FROM artist_tag
|
||||||
|
WHERE tag IN ({placeholders})
|
||||||
|
GROUP BY norm_artist
|
||||||
|
HAVING SUM(weight) >= ?
|
||||||
|
)
|
||||||
|
SELECT t.id AS id, a.name AS artist, t.title AS title,
|
||||||
|
t.duration AS duration, t.path AS path
|
||||||
|
FROM lidarr_track t
|
||||||
|
JOIN lidarr_artist a ON a.id = t.artist_id
|
||||||
|
JOIN vibe v ON v.norm_artist = a.norm_name
|
||||||
|
LEFT JOIN lidarr_album al ON al.id = t.album_id
|
||||||
|
WHERE {PLAYABLE}{year_clause}
|
||||||
|
ORDER BY ((t.id * {SHUFFLE_MULTIPLIER}) + ?) % {SHUFFLE_MODULUS}
|
||||||
|
LIMIT ?
|
||||||
|
"""
|
||||||
|
|
||||||
|
entries = []
|
||||||
|
for row in store.connection.execute(sql, parameters):
|
||||||
|
mirror = mirror_path_for(row["path"], library_root, mirror_root)
|
||||||
|
if mirror is None or not mirror.is_file():
|
||||||
|
continue
|
||||||
|
entries.append({**dict(row), "mirror": mirror})
|
||||||
|
|
||||||
|
write_playlist(directory / f"{vibe['name']}.m3u", entries)
|
||||||
|
total += len(entries)
|
||||||
|
logger.info("playlist %-20s %4d tracks -- by tag", vibe["name"], len(entries))
|
||||||
|
|
||||||
|
return total
|
||||||
|
|
||||||
|
|
||||||
|
def library_root_of(store):
|
||||||
|
"""Return the directory Lidarr's artist folders sit under.
|
||||||
|
|
||||||
|
Derived rather than configured, because it has to agree with what Lidarr
|
||||||
|
reports and no one wants to keep a second copy of that in step by hand.
|
||||||
|
"""
|
||||||
|
paths = [
|
||||||
|
row["path"]
|
||||||
|
for row in store.connection.execute(
|
||||||
|
"SELECT path FROM lidarr_artist WHERE path IS NOT NULL AND path <> ''"
|
||||||
|
)
|
||||||
|
]
|
||||||
|
if not paths:
|
||||||
|
return None
|
||||||
|
if len(paths) == 1:
|
||||||
|
return str(Path(paths[0]).parent)
|
||||||
|
try:
|
||||||
|
return os.path.commonpath(paths)
|
||||||
|
except ValueError:
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
def mirror_path_for(source, library_root, mirror_root):
|
||||||
|
"""Return where music-mirror would have put the MP3 for a source file."""
|
||||||
|
try:
|
||||||
|
relative = Path(source).relative_to(library_root)
|
||||||
|
except ValueError:
|
||||||
|
return None
|
||||||
|
return (Path(mirror_root) / relative).with_suffix(MIRROR_SUFFIX)
|
||||||
|
|
||||||
|
|
||||||
|
def write_playlist(path, entries):
|
||||||
|
"""Write one extended M3U, atomically.
|
||||||
|
|
||||||
|
Paths are relative to the playlist file, so the same playlist works from the
|
||||||
|
NAS, from a Mac over SMB and from Linux without rewriting.
|
||||||
|
"""
|
||||||
|
lines = ["#EXTM3U"]
|
||||||
|
for entry in entries:
|
||||||
|
seconds = round((entry["duration"] or 0) / 1000)
|
||||||
|
lines.append(f"#EXTINF:{seconds},{entry['artist']} - {entry['title']}")
|
||||||
|
lines.append(os.path.relpath(entry["mirror"], path.parent))
|
||||||
|
|
||||||
|
path.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
handle, temporary = tempfile.mkstemp(dir=path.parent, suffix=".m3u.part")
|
||||||
|
os.close(handle)
|
||||||
|
temporary = Path(temporary)
|
||||||
|
try:
|
||||||
|
temporary.write_text("\n".join(lines) + "\n", encoding="utf-8")
|
||||||
|
# The mirror is read back by something else; see music-mirror, which had
|
||||||
|
# to learn this the hard way.
|
||||||
|
mode = temporary.stat().st_mode
|
||||||
|
if not mode & GROUP_READ:
|
||||||
|
temporary.chmod(mode | GROUP_READ)
|
||||||
|
os.replace(temporary, path)
|
||||||
|
finally:
|
||||||
|
temporary.unlink(missing_ok=True)
|
||||||
|
|
||||||
|
|
||||||
|
def build_playlists(store, mirror_root, library_root, limit, now):
|
||||||
|
"""Write every playlist into the mirror. Returns how many tracks were listed.
|
||||||
|
|
||||||
|
A track is only listed once its mirror file has been confirmed to exist. The
|
||||||
|
index knows what Lidarr holds, which is the lossless source; whether the MP3
|
||||||
|
beside it has been encoded yet is music-mirror's business and is checked
|
||||||
|
rather than assumed.
|
||||||
|
"""
|
||||||
|
directory = Path(mirror_root) / PLAYLIST_DIRECTORY
|
||||||
|
parameters = {
|
||||||
|
"limit": limit,
|
||||||
|
"year_ago": now - 365 * 86400,
|
||||||
|
"week": now // ROTATION_PERIOD_SECONDS,
|
||||||
|
}
|
||||||
|
|
||||||
|
total = 0
|
||||||
|
missing = 0
|
||||||
|
for name, description, sql in PLAYLISTS:
|
||||||
|
entries = []
|
||||||
|
for row in store.connection.execute(sql, parameters):
|
||||||
|
mirror = mirror_path_for(row["path"], library_root, mirror_root)
|
||||||
|
if mirror is None or not mirror.is_file():
|
||||||
|
missing += 1
|
||||||
|
continue
|
||||||
|
entries.append({**dict(row), "mirror": mirror})
|
||||||
|
write_playlist(directory / f"{name}.m3u", entries)
|
||||||
|
total += len(entries)
|
||||||
|
logger.info("playlist %-20s %4d tracks -- %s", name, len(entries), description)
|
||||||
|
|
||||||
|
if missing:
|
||||||
|
logger.warning(
|
||||||
|
"%d selected tracks had no file in the mirror and were left out."
|
||||||
|
" A handful means music-mirror has not encoded them yet; a large"
|
||||||
|
" number means --library-root or --mirror is pointing at the wrong"
|
||||||
|
" place, since the paths are then being mapped to nothing.",
|
||||||
|
missing,
|
||||||
|
)
|
||||||
|
return total
|
||||||
|
|
||||||
|
|
||||||
def coverage_report(store):
|
def coverage_report(store):
|
||||||
"""Log how much of the listening history could be tied to the library.
|
"""Log how much of the listening history could be tied to the library.
|
||||||
|
|
||||||
@@ -1078,18 +1565,107 @@ def coverage_report(store):
|
|||||||
row["plays"],
|
row["plays"],
|
||||||
)
|
)
|
||||||
|
|
||||||
|
# Why the mbid tier performs the way it does. A recording id on both sides
|
||||||
|
# that still fails to join means the two disagree about which recording the
|
||||||
|
# song is -- MusicBrainz holds a separate recording per release, and Last.fm
|
||||||
|
# and Lidarr need not have picked the same one. That is not a fault to fix
|
||||||
|
# in the matcher; it is the reason the name tier has to carry the load.
|
||||||
|
library_with_mbid = store.scalar(
|
||||||
|
"SELECT COUNT(*) FROM lidarr_track WHERE recording_mbid IS NOT NULL"
|
||||||
|
)
|
||||||
|
logger.info(
|
||||||
|
"library tracks carrying a recording MBID: %d of %d (%.1f%%)",
|
||||||
|
library_with_mbid,
|
||||||
|
tracks,
|
||||||
|
100 * library_with_mbid / tracks,
|
||||||
|
)
|
||||||
|
disagreed = store.connection.execute(
|
||||||
|
"SELECT COUNT(*) AS pairs, COALESCE(SUM(plays), 0) AS plays FROM scrobble_key"
|
||||||
|
" WHERE track_mbid IS NOT NULL AND method = 'name'"
|
||||||
|
).fetchone()
|
||||||
|
logger.info(
|
||||||
|
"carried a recording MBID, joined on name instead: %d pairs, %d plays"
|
||||||
|
" -- both sides know the song, they disagree on which recording it is",
|
||||||
|
disagreed["pairs"],
|
||||||
|
disagreed["plays"],
|
||||||
|
)
|
||||||
|
|
||||||
suspect = store.connection.execute(
|
suspect = store.connection.execute(
|
||||||
"SELECT COUNT(*) AS pairs, COALESCE(SUM(plays), 0) AS plays FROM scrobble_key k"
|
"SELECT COUNT(*) AS pairs, COALESCE(SUM(plays), 0) AS plays FROM scrobble_key k"
|
||||||
" WHERE k.track_id IS NULL"
|
" WHERE k.track_id IS NULL"
|
||||||
" AND EXISTS (SELECT 1 FROM lidarr_artist a WHERE a.norm_name = k.norm_artist)"
|
" AND EXISTS (SELECT 1 FROM lidarr_artist a WHERE a.norm_name = k.norm_artist)"
|
||||||
).fetchone()
|
).fetchone()
|
||||||
logger.info(
|
logger.info(
|
||||||
"unmatched by an artist the library holds: %d pairs, %d plays -- these are the"
|
"unmatched by an artist the library holds: %d pairs, %d plays",
|
||||||
" matcher's misses, not music you do not own",
|
|
||||||
suspect["pairs"],
|
suspect["pairs"],
|
||||||
suspect["plays"],
|
suspect["plays"],
|
||||||
)
|
)
|
||||||
|
|
||||||
|
# Owning an artist is a weak proxy for owning a track, so that figure alone
|
||||||
|
# overstates the matcher's failings. Split it -- but not on the title alone.
|
||||||
|
# Across fifty thousand tracks, titles collide constantly: "Everyday" is
|
||||||
|
# Rusko and also Def Leppard, "Kaleidoscope" is Delta Heavy and also
|
||||||
|
# Chappell Roan. Matching those would be worse than missing them.
|
||||||
|
#
|
||||||
|
# The signal for a genuine attribution miss is that the library's own title
|
||||||
|
# credits the artist the scrobble is filed under -- "Voodoo People (Pendulum
|
||||||
|
# Remix)" against a play credited to Pendulum. The normalised title has that
|
||||||
|
# suffix stripped, which is exactly what let them meet, so the raw one has to
|
||||||
|
# be searched for the name.
|
||||||
|
attribution = store.connection.execute(
|
||||||
|
"SELECT COUNT(*) AS pairs, COALESCE(SUM(plays), 0) AS plays FROM scrobble_key k"
|
||||||
|
" WHERE k.track_id IS NULL"
|
||||||
|
" AND EXISTS (SELECT 1 FROM lidarr_artist a WHERE a.norm_name = k.norm_artist)"
|
||||||
|
" AND EXISTS (SELECT 1 FROM lidarr_track t"
|
||||||
|
" WHERE t.norm_title = k.norm_track"
|
||||||
|
" AND instr(lower(t.title), lower(k.artist)) > 0)"
|
||||||
|
).fetchone()
|
||||||
|
collision = store.connection.execute(
|
||||||
|
"SELECT COUNT(*) AS pairs, COALESCE(SUM(plays), 0) AS plays FROM scrobble_key k"
|
||||||
|
" WHERE k.track_id IS NULL"
|
||||||
|
" AND EXISTS (SELECT 1 FROM lidarr_artist a WHERE a.norm_name = k.norm_artist)"
|
||||||
|
" AND EXISTS (SELECT 1 FROM lidarr_track t WHERE t.norm_title = k.norm_track)"
|
||||||
|
).fetchone()
|
||||||
|
logger.info(
|
||||||
|
" the library's title credits the scrobbled artist: %d pairs, %d plays"
|
||||||
|
" -- remixes and guest spots, and the genuine misses",
|
||||||
|
attribution["pairs"],
|
||||||
|
attribution["plays"],
|
||||||
|
)
|
||||||
|
logger.info(
|
||||||
|
" same title under an unrelated artist: %d pairs, %d plays"
|
||||||
|
" -- title collisions, not misses; matching these would be a mistake",
|
||||||
|
collision["pairs"] - attribution["pairs"],
|
||||||
|
collision["plays"] - attribution["plays"],
|
||||||
|
)
|
||||||
|
logger.info(
|
||||||
|
" the rest, %d pairs, %d plays: you own the artist but not the track",
|
||||||
|
suspect["pairs"] - collision["pairs"],
|
||||||
|
suspect["plays"] - collision["plays"],
|
||||||
|
)
|
||||||
|
|
||||||
|
mismatched = store.connection.execute(
|
||||||
|
"SELECT k.artist, k.track, k.plays,"
|
||||||
|
" (SELECT t.title FROM lidarr_track t"
|
||||||
|
" WHERE t.norm_title = k.norm_track"
|
||||||
|
" AND instr(lower(t.title), lower(k.artist)) > 0 LIMIT 1) AS library_title"
|
||||||
|
" FROM scrobble_key k"
|
||||||
|
" WHERE k.track_id IS NULL"
|
||||||
|
" AND EXISTS (SELECT 1 FROM lidarr_artist a WHERE a.norm_name = k.norm_artist)"
|
||||||
|
" AND EXISTS (SELECT 1 FROM lidarr_track t"
|
||||||
|
" WHERE t.norm_title = k.norm_track"
|
||||||
|
" AND instr(lower(t.title), lower(k.artist)) > 0)"
|
||||||
|
" ORDER BY k.plays DESC, k.artist LIMIT 10"
|
||||||
|
).fetchall()
|
||||||
|
for position, row in enumerate(mismatched, start=1):
|
||||||
|
logger.info(
|
||||||
|
" attribution %2d: %-45s %4d plays, library has %r",
|
||||||
|
position,
|
||||||
|
f"{row['artist']} - {row['track']}"[:45],
|
||||||
|
row["plays"],
|
||||||
|
row["library_title"],
|
||||||
|
)
|
||||||
|
|
||||||
with_files = store.scalar("SELECT COUNT(*) FROM lidarr_track WHERE has_file = 1")
|
with_files = store.scalar("SELECT COUNT(*) FROM lidarr_track WHERE has_file = 1")
|
||||||
played = store.scalar(
|
played = store.scalar(
|
||||||
"SELECT COUNT(DISTINCT k.track_id) FROM scrobble_key k"
|
"SELECT COUNT(DISTINCT k.track_id) FROM scrobble_key k"
|
||||||
@@ -1158,7 +1734,10 @@ def report(store, now):
|
|||||||
logger.info(" top track 90d %2d: %-50s %d", position, label[:50], row["plays"])
|
logger.info(" top track 90d %2d: %-50s %d", position, label[:50], row["plays"])
|
||||||
|
|
||||||
|
|
||||||
def run_once(client, store, user, now, backfill_limit, lidarr=None):
|
def run_once(
|
||||||
|
client, store, user, now, backfill_limit, lidarr=None, mirror=None,
|
||||||
|
library_root=None, playlist_limit=100, vibes=DEFAULT_VIBES, tag_limit=0,
|
||||||
|
):
|
||||||
"""Run a single pass. Returns the number of scrobbles added."""
|
"""Run a single pass. Returns the number of scrobbles added."""
|
||||||
started = time.monotonic()
|
started = time.monotonic()
|
||||||
added = catch_up(client, store, user, now)
|
added = catch_up(client, store, user, now)
|
||||||
@@ -1171,6 +1750,17 @@ def run_once(client, store, user, now, backfill_limit, lidarr=None):
|
|||||||
# arrived, and they need a verdict too.
|
# arrived, and they need a verdict too.
|
||||||
if store.scalar("SELECT COUNT(*) FROM lidarr_track"):
|
if store.scalar("SELECT COUNT(*) FROM lidarr_track"):
|
||||||
match_library(store)
|
match_library(store)
|
||||||
|
if mirror is not None:
|
||||||
|
root = library_root or library_root_of(store)
|
||||||
|
if root is None:
|
||||||
|
logger.warning(
|
||||||
|
"cannot work out where Lidarr's music lives, so no playlists were"
|
||||||
|
" written; set --library-root"
|
||||||
|
)
|
||||||
|
else:
|
||||||
|
build_playlists(store, mirror, root, playlist_limit, now)
|
||||||
|
sync_tags(client, store, now, limit=tag_limit)
|
||||||
|
build_vibe_playlists(store, vibes, mirror, root, playlist_limit, now)
|
||||||
|
|
||||||
logger.info(
|
logger.info(
|
||||||
"pass complete in %.1fs: %d scrobbles added, %d loved",
|
"pass complete in %.1fs: %d scrobbles added, %d loved",
|
||||||
@@ -1243,6 +1833,37 @@ def build_parser():
|
|||||||
default=os.getenv("MUSIC_CURATOR_LIDARR_API_KEY"),
|
default=os.getenv("MUSIC_CURATOR_LIDARR_API_KEY"),
|
||||||
help="Lidarr API key (env MUSIC_CURATOR_LIDARR_API_KEY)",
|
help="Lidarr API key (env MUSIC_CURATOR_LIDARR_API_KEY)",
|
||||||
)
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--mirror",
|
||||||
|
default=os.getenv("MUSIC_CURATOR_MIRROR"),
|
||||||
|
help="root of the MP3 mirror; playlists are written into it"
|
||||||
|
" (env MUSIC_CURATOR_MIRROR)",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--library-root",
|
||||||
|
default=os.getenv("MUSIC_CURATOR_LIBRARY_ROOT"),
|
||||||
|
help="prefix to strip from Lidarr's track paths when mapping them into the"
|
||||||
|
" mirror; derived from the indexed artist folders when unset"
|
||||||
|
" (env MUSIC_CURATOR_LIBRARY_ROOT)",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--playlist-limit",
|
||||||
|
type=int,
|
||||||
|
default=int(os.getenv("MUSIC_CURATOR_PLAYLIST_LIMIT", "100")),
|
||||||
|
help="most tracks to put in any one playlist (env MUSIC_CURATOR_PLAYLIST_LIMIT)",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--vibes",
|
||||||
|
default=os.getenv("MUSIC_CURATOR_VIBES"),
|
||||||
|
help="JSON file of mood definitions, replacing the built-in set"
|
||||||
|
" (env MUSIC_CURATOR_VIBES)",
|
||||||
|
)
|
||||||
|
parser.add_argument(
|
||||||
|
"--tag-limit",
|
||||||
|
type=int,
|
||||||
|
default=int(os.getenv("MUSIC_CURATOR_TAG_LIMIT", "0")),
|
||||||
|
help="cap artist tag lookups per pass; 0 for no cap (env MUSIC_CURATOR_TAG_LIMIT)",
|
||||||
|
)
|
||||||
parser.add_argument(
|
parser.add_argument(
|
||||||
"--skip-index",
|
"--skip-index",
|
||||||
action="store_true",
|
action="store_true",
|
||||||
@@ -1267,6 +1888,7 @@ def main(argv=None, clock=time.time):
|
|||||||
|
|
||||||
try:
|
try:
|
||||||
interval = parse_interval(args.interval) if args.interval else None
|
interval = parse_interval(args.interval) if args.interval else None
|
||||||
|
vibes = load_vibes(args.vibes)
|
||||||
except ValueError as error:
|
except ValueError as error:
|
||||||
logger.error("%s", error)
|
logger.error("%s", error)
|
||||||
return 2
|
return 2
|
||||||
@@ -1307,7 +1929,19 @@ def main(argv=None, clock=time.time):
|
|||||||
while True:
|
while True:
|
||||||
now = int(clock())
|
now = int(clock())
|
||||||
try:
|
try:
|
||||||
run_once(client, store, args.user, now, args.backfill_limit, lidarr)
|
run_once(
|
||||||
|
client,
|
||||||
|
store,
|
||||||
|
args.user,
|
||||||
|
now,
|
||||||
|
args.backfill_limit,
|
||||||
|
lidarr,
|
||||||
|
args.mirror,
|
||||||
|
args.library_root,
|
||||||
|
args.playlist_limit,
|
||||||
|
vibes,
|
||||||
|
args.tag_limit,
|
||||||
|
)
|
||||||
except (LastfmError, LidarrError) as error:
|
except (LastfmError, LidarrError) as error:
|
||||||
logger.error("%s", error)
|
logger.error("%s", error)
|
||||||
if interval is None:
|
if interval is None:
|
||||||
|
|||||||
+1
-1
@@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta"
|
|||||||
|
|
||||||
[project]
|
[project]
|
||||||
name = "music-curator"
|
name = "music-curator"
|
||||||
version = "0.2.3"
|
version = "0.4.0"
|
||||||
description = "Ingest a Last.fm listening history and curate a music library from it"
|
description = "Ingest a Last.fm listening history and curate a music library from it"
|
||||||
readme = "README.md"
|
readme = "README.md"
|
||||||
requires-python = ">=3.11"
|
requires-python = ">=3.11"
|
||||||
|
|||||||
+6
-1
@@ -56,7 +56,9 @@ class FakeLastfm:
|
|||||||
is prepended to the first page with no `date`.
|
is prepended to the first page with no `date`.
|
||||||
"""
|
"""
|
||||||
|
|
||||||
def __init__(self, tracks=(), loved=(), nowplaying=None, outcomes=()):
|
def __init__(self, tracks=(), loved=(), nowplaying=None, outcomes=(), tags=None):
|
||||||
|
# Keyed by mbid or by artist name, whichever the caller asked with.
|
||||||
|
self.tags = tags or {}
|
||||||
self.tracks = sorted(tracks, key=lambda track: int(track["date"]["uts"]), reverse=True)
|
self.tracks = sorted(tracks, key=lambda track: int(track["date"]["uts"]), reverse=True)
|
||||||
self.loved = list(loved)
|
self.loved = list(loved)
|
||||||
self.nowplaying = nowplaying
|
self.nowplaying = nowplaying
|
||||||
@@ -82,6 +84,9 @@ class FakeLastfm:
|
|||||||
return json.dumps(self._recent(query))
|
return json.dumps(self._recent(query))
|
||||||
if method == "user.getlovedtracks":
|
if method == "user.getlovedtracks":
|
||||||
return json.dumps(self._loved(query))
|
return json.dumps(self._loved(query))
|
||||||
|
if method == "artist.gettoptags":
|
||||||
|
key = query.get("mbid") or query.get("artist", "")
|
||||||
|
return json.dumps({"toptags": {"tag": self.tags.get(key, [])}})
|
||||||
raise AssertionError(f"unexpected method {method}")
|
raise AssertionError(f"unexpected method {method}")
|
||||||
|
|
||||||
def _recent(self, query):
|
def _recent(self, query):
|
||||||
|
|||||||
+446
-6
@@ -1,5 +1,7 @@
|
|||||||
import json
|
import json
|
||||||
|
import stat
|
||||||
import urllib.error
|
import urllib.error
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
import pytest
|
import pytest
|
||||||
from conftest import FakeLastfm, FakeLidarr, make_loved, make_tracks
|
from conftest import FakeLastfm, FakeLidarr, make_loved, make_tracks
|
||||||
@@ -55,9 +57,23 @@ LIBRARY = [
|
|||||||
{
|
{
|
||||||
"name": "The Prodigy",
|
"name": "The Prodigy",
|
||||||
"albums": [
|
"albums": [
|
||||||
{"title": "The Fat of the Land", "tracks": [{"title": "Breathe (Remastered)"}]}
|
{
|
||||||
|
"title": "The Fat of the Land",
|
||||||
|
"tracks": [
|
||||||
|
{"title": "Breathe (Remastered)"},
|
||||||
|
# Credits its remixer in the title, which is the only signal
|
||||||
|
# separating a real attribution miss from a title collision.
|
||||||
|
{"title": "Voodoo People (Pendulum Remix)"},
|
||||||
|
],
|
||||||
|
}
|
||||||
],
|
],
|
||||||
},
|
},
|
||||||
|
# Held by the library in its own right, which is what puts its scrobbles
|
||||||
|
# inside the "artist the library holds" filter at all.
|
||||||
|
{
|
||||||
|
"name": "Pendulum",
|
||||||
|
"albums": [{"title": "Immersion", "tracks": [{"title": "Watercolour"}]}],
|
||||||
|
},
|
||||||
]
|
]
|
||||||
|
|
||||||
|
|
||||||
@@ -393,7 +409,7 @@ def test_keys_carry_the_play_count_and_the_span(tmp_path):
|
|||||||
def test_re_indexing_drops_what_lidarr_no_longer_has(tmp_path):
|
def test_re_indexing_drops_what_lidarr_no_longer_has(tmp_path):
|
||||||
"""The index is Lidarr's mirror, not an accumulation of everything ever seen."""
|
"""The index is Lidarr's mirror, not an accumulation of everything ever seen."""
|
||||||
store = indexed(tmp_path, [scrobble_of("AC/DC", "Hells Bells")])
|
store = indexed(tmp_path, [scrobble_of("AC/DC", "Hells Bells")])
|
||||||
assert store.scalar("SELECT COUNT(*) FROM lidarr_artist") == 3
|
assert store.scalar("SELECT COUNT(*) FROM lidarr_artist") == 4
|
||||||
|
|
||||||
music_curator.index_library(
|
music_curator.index_library(
|
||||||
music_curator.Lidarr("http://lidarr", "key", transport=FakeLidarr(LIBRARY[:1])), store
|
music_curator.Lidarr("http://lidarr", "key", transport=FakeLidarr(LIBRARY[:1])), store
|
||||||
@@ -512,7 +528,7 @@ def test_albums_come_from_the_unfiltered_endpoint(tmp_path):
|
|||||||
|
|
||||||
album_calls = [query for path, query in api.calls if path == "album"]
|
album_calls = [query for path, query in api.calls if path == "album"]
|
||||||
assert album_calls == [{}]
|
assert album_calls == [{}]
|
||||||
assert store.scalar("SELECT COUNT(*) FROM lidarr_album") == 3
|
assert store.scalar("SELECT COUNT(*) FROM lidarr_album") == 4
|
||||||
|
|
||||||
|
|
||||||
def test_a_bad_album_falls_back_to_asking_per_artist(tmp_path):
|
def test_a_bad_album_falls_back_to_asking_per_artist(tmp_path):
|
||||||
@@ -526,7 +542,7 @@ def test_a_bad_album_falls_back_to_asking_per_artist(tmp_path):
|
|||||||
|
|
||||||
assert store.scalar("SELECT COUNT(*) FROM lidarr_album WHERE artist_id = 2") == 0
|
assert store.scalar("SELECT COUNT(*) FROM lidarr_album WHERE artist_id = 2") == 0
|
||||||
# The other two artists keep their albums.
|
# The other two artists keep their albums.
|
||||||
assert store.scalar("SELECT COUNT(*) FROM lidarr_album") == 2
|
assert store.scalar("SELECT COUNT(*) FROM lidarr_album") == 3
|
||||||
assert store.get_state("index_albums_skipped") == "1"
|
assert store.get_state("index_albums_skipped") == "1"
|
||||||
|
|
||||||
|
|
||||||
@@ -549,10 +565,10 @@ def test_an_artist_lidarr_cannot_serve_does_not_kill_the_index(tmp_path):
|
|||||||
|
|
||||||
music_curator.index_library(music_curator.Lidarr("http://lidarr", "key", transport=api), store)
|
music_curator.index_library(music_curator.Lidarr("http://lidarr", "key", transport=api), store)
|
||||||
|
|
||||||
assert store.scalar("SELECT COUNT(*) FROM lidarr_artist") == 3
|
assert store.scalar("SELECT COUNT(*) FROM lidarr_artist") == 4
|
||||||
assert store.scalar("SELECT COUNT(*) FROM lidarr_track WHERE artist_id = 2") == 0
|
assert store.scalar("SELECT COUNT(*) FROM lidarr_track WHERE artist_id = 2") == 0
|
||||||
# The other two artists are indexed in full.
|
# The other two artists are indexed in full.
|
||||||
assert store.scalar("SELECT COUNT(*) FROM lidarr_track") == 3
|
assert store.scalar("SELECT COUNT(*) FROM lidarr_track") == 5
|
||||||
assert store.get_state("index_skipped") == "1"
|
assert store.get_state("index_skipped") == "1"
|
||||||
|
|
||||||
|
|
||||||
@@ -687,3 +703,427 @@ def test_lidarr_talks_to_a_real_server_through_keep_alive(http_server):
|
|||||||
client.transport.close()
|
client.transport.close()
|
||||||
|
|
||||||
assert server.connections == 1
|
assert server.connections == 1
|
||||||
|
|
||||||
|
|
||||||
|
# Titles taken verbatim from a real coverage report's unmatched list. Each one
|
||||||
|
# was a genuine miss before the normaliser handled it.
|
||||||
|
@pytest.mark.parametrize(
|
||||||
|
("scrobbled", "tagged"),
|
||||||
|
[
|
||||||
|
("Self vs Self (feat. In Flames)", "Self vs Self"),
|
||||||
|
("Grime Battle of Hastings (feat. The Town Crier)", "Grime Battle of Hastings"),
|
||||||
|
("Gold Dust - Shy FX Re-Edit", "Gold Dust"),
|
||||||
|
("Back To Your Roots - Friction & K-Tee Remix", "Back To Your Roots"),
|
||||||
|
("Constellations - Forza Horizon 3 VIP", "Constellations"),
|
||||||
|
("Everyday (Netsky Remix)", "Everyday"),
|
||||||
|
("Voodoo People [Pendulum Remix] [Live At Brixton Academy]", "Voodoo People"),
|
||||||
|
],
|
||||||
|
)
|
||||||
|
def test_real_unmatched_titles_now_agree_with_their_tags(scrobbled, tagged):
|
||||||
|
assert music_curator.normalise(scrobbled) == music_curator.normalise(tagged)
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize(
|
||||||
|
"title",
|
||||||
|
[
|
||||||
|
"Dancing with Myself",
|
||||||
|
"(Don't Fear) The Reaper",
|
||||||
|
"Live and Let Die",
|
||||||
|
"Radio Ga Ga",
|
||||||
|
"Editors",
|
||||||
|
"Mixed Emotions",
|
||||||
|
"Vipassana",
|
||||||
|
],
|
||||||
|
)
|
||||||
|
def test_the_version_words_do_not_eat_ordinary_titles(title):
|
||||||
|
"""Every one of these contains a version word and must survive intact."""
|
||||||
|
assert music_curator.normalise(title) == music_curator.normalise(title.lower())
|
||||||
|
assert len(music_curator.normalise(title).split()) == len(title.split())
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_bracketed_guest_credit_matches_the_bare_tag(tmp_path):
|
||||||
|
"""Whitespace-then-feat misses the bracketed form, which is most of them."""
|
||||||
|
store = indexed(tmp_path, [scrobble_of("Yellowcard", "Here I Am Alive (feat. Someone)")])
|
||||||
|
|
||||||
|
method, track_id = verdict(store, "Yellowcard", "Here I Am Alive (feat. Someone)")
|
||||||
|
assert method == "name"
|
||||||
|
assert track_id is not None
|
||||||
|
|
||||||
|
|
||||||
|
def test_misses_are_split_by_whether_the_library_holds_the_title(tmp_path):
|
||||||
|
"""Owning an artist is a weak proxy for owning a track. Counting both as
|
||||||
|
matcher failures overstates the problem and would over-block the cull."""
|
||||||
|
store = indexed(
|
||||||
|
tmp_path,
|
||||||
|
[
|
||||||
|
# The library holds "Hells Bells", but under AC/DC, not Yellowcard:
|
||||||
|
# an attribution disagreement, and a real miss.
|
||||||
|
scrobble_of("Yellowcard", "Hells Bells"),
|
||||||
|
# Yellowcard is in the library; this track is not, under any artist.
|
||||||
|
scrobble_of("Yellowcard", "A Single She Never Bought"),
|
||||||
|
],
|
||||||
|
)
|
||||||
|
|
||||||
|
def count(extra):
|
||||||
|
return store.connection.execute(
|
||||||
|
"SELECT COUNT(*) FROM scrobble_key k WHERE k.track_id IS NULL"
|
||||||
|
" AND EXISTS (SELECT 1 FROM lidarr_artist a WHERE a.norm_name = k.norm_artist)"
|
||||||
|
f" {extra}"
|
||||||
|
).fetchone()[0]
|
||||||
|
|
||||||
|
assert count("") == 2
|
||||||
|
assert count("AND EXISTS (SELECT 1 FROM lidarr_track t WHERE t.norm_title = k.norm_track)") == 1
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_report_survives_the_attribution_split(tmp_path):
|
||||||
|
store = indexed(tmp_path, [scrobble_of("Yellowcard", "Hells Bells")])
|
||||||
|
|
||||||
|
music_curator.report(store, NOW)
|
||||||
|
|
||||||
|
|
||||||
|
def attribution_pairs(store):
|
||||||
|
"""Unmatched pairs where the library's own title credits the scrobbled artist."""
|
||||||
|
return [
|
||||||
|
row["artist"]
|
||||||
|
for row in store.connection.execute(
|
||||||
|
"SELECT k.artist FROM scrobble_key k WHERE k.track_id IS NULL"
|
||||||
|
" AND EXISTS (SELECT 1 FROM lidarr_artist a WHERE a.norm_name = k.norm_artist)"
|
||||||
|
" AND EXISTS (SELECT 1 FROM lidarr_track t"
|
||||||
|
" WHERE t.norm_title = k.norm_track"
|
||||||
|
" AND instr(lower(t.title), lower(k.artist)) > 0)"
|
||||||
|
)
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_shared_title_is_not_an_attribution_miss(tmp_path):
|
||||||
|
"""Across fifty thousand tracks, titles collide constantly: "Everyday" is
|
||||||
|
Rusko and also Def Leppard. Matching those would be worse than missing."""
|
||||||
|
store = indexed(tmp_path, [scrobble_of("Yellowcard", "Hells Bells")])
|
||||||
|
|
||||||
|
# The library holds "Hells Bells", by AC/DC, and its title says nothing
|
||||||
|
# about Yellowcard. A collision, not a miss.
|
||||||
|
assert verdict(store, "Yellowcard", "Hells Bells") == ("none", None)
|
||||||
|
assert attribution_pairs(store) == []
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_remix_credited_in_the_library_title_is_an_attribution_miss(tmp_path):
|
||||||
|
"""The library has "Voodoo People (Pendulum Remix)" under The Prodigy; the
|
||||||
|
scrobble credits Pendulum. Same song, different filing."""
|
||||||
|
store = indexed(tmp_path, [scrobble_of("Pendulum", "Voodoo People")])
|
||||||
|
|
||||||
|
assert attribution_pairs(store) == ["Pendulum"]
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_title_index_is_used_for_the_report_lookup(tmp_path):
|
||||||
|
"""Without it the report scans every track for every unmatched key: forty-
|
||||||
|
seven seconds on a real library."""
|
||||||
|
store = indexed(tmp_path, [])
|
||||||
|
|
||||||
|
plan = "\n".join(
|
||||||
|
row[-1]
|
||||||
|
for row in store.connection.execute(
|
||||||
|
"EXPLAIN QUERY PLAN SELECT 1 FROM scrobble_key k WHERE k.track_id IS NULL"
|
||||||
|
" AND EXISTS (SELECT 1 FROM lidarr_track t WHERE t.norm_title = k.norm_track)"
|
||||||
|
)
|
||||||
|
)
|
||||||
|
assert "lidarr_track_title" in plan, plan
|
||||||
|
|
||||||
|
|
||||||
|
def playlist_library(tmp_path):
|
||||||
|
"""A library on disk, with mirror MP3s beside the Lidarr source paths."""
|
||||||
|
source = tmp_path / "music"
|
||||||
|
mirror = tmp_path / "mirror"
|
||||||
|
library = [
|
||||||
|
{
|
||||||
|
"name": "Played Band",
|
||||||
|
"albums": [
|
||||||
|
{
|
||||||
|
"title": "Known",
|
||||||
|
"tracks": [{"title": "Hit"}, {"title": "Album Track"}],
|
||||||
|
}
|
||||||
|
],
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"name": "Silent Band",
|
||||||
|
"albums": [{"title": "Unknown", "tracks": [{"title": "Never Heard"}]}],
|
||||||
|
},
|
||||||
|
]
|
||||||
|
api = FakeLidarr(library)
|
||||||
|
# FakeLidarr invents /music/<artist>/<album>/<title>.flac; put the mirror
|
||||||
|
# MP3s at the paths music-mirror would have produced from those.
|
||||||
|
for handle in api.files:
|
||||||
|
relative = Path(handle["path"]).relative_to("/music")
|
||||||
|
handle["path"] = str(source / relative)
|
||||||
|
mp3 = (mirror / relative).with_suffix(".mp3")
|
||||||
|
mp3.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
mp3.write_bytes(b"not really an mp3")
|
||||||
|
for artist in api.artists:
|
||||||
|
artist["path"] = str(source / Path(artist["path"]).name)
|
||||||
|
return api, source, mirror
|
||||||
|
|
||||||
|
|
||||||
|
def test_playlists_are_written_into_the_mirror(tmp_path):
|
||||||
|
api, source, mirror = playlist_library(tmp_path)
|
||||||
|
store = store_at(tmp_path)
|
||||||
|
ingest(store, FakeLastfm([scrobble_of("Played Band", "Hit")] * 1))
|
||||||
|
music_curator.index_library(
|
||||||
|
music_curator.Lidarr("http://lidarr", "key", transport=api), store
|
||||||
|
)
|
||||||
|
music_curator.match_library(store)
|
||||||
|
|
||||||
|
music_curator.build_playlists(store, mirror, str(source), 100, NOW)
|
||||||
|
|
||||||
|
written = sorted(p.name for p in (mirror / "_playlists").glob("*.m3u"))
|
||||||
|
assert written == [
|
||||||
|
"all-time.m3u",
|
||||||
|
"deep-cuts.m3u",
|
||||||
|
"heavy-rotation.m3u",
|
||||||
|
"neglected.m3u",
|
||||||
|
"unheard-favourites.m3u",
|
||||||
|
"unheard.m3u",
|
||||||
|
]
|
||||||
|
|
||||||
|
played = (mirror / "_playlists" / "all-time.m3u").read_text().splitlines()
|
||||||
|
assert played[0] == "#EXTM3U"
|
||||||
|
assert played[1].startswith("#EXTINF:")
|
||||||
|
assert "Played Band - Hit" in played[1]
|
||||||
|
# Relative to the playlist file, so the same file works from any mount.
|
||||||
|
assert played[2] == "../Played Band/Known/Hit.mp3"
|
||||||
|
assert (mirror / "_playlists" / "all-time.m3u").parent.joinpath(played[2]).resolve().is_file()
|
||||||
|
|
||||||
|
|
||||||
|
def test_an_unplayed_track_lands_in_the_unheard_playlist(tmp_path):
|
||||||
|
api, source, mirror = playlist_library(tmp_path)
|
||||||
|
store = store_at(tmp_path)
|
||||||
|
ingest(store, FakeLastfm([scrobble_of("Played Band", "Hit")]))
|
||||||
|
music_curator.index_library(
|
||||||
|
music_curator.Lidarr("http://lidarr", "key", transport=api), store
|
||||||
|
)
|
||||||
|
music_curator.match_library(store)
|
||||||
|
|
||||||
|
music_curator.build_playlists(store, mirror, str(source), 100, NOW)
|
||||||
|
|
||||||
|
unheard = (mirror / "_playlists" / "unheard.m3u").read_text()
|
||||||
|
assert "Never Heard" in unheard
|
||||||
|
assert "Album Track" in unheard
|
||||||
|
# The one thing that was played must not be in it.
|
||||||
|
assert "Played Band - Hit\n" not in unheard
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_track_with_no_mirror_file_is_left_out(tmp_path):
|
||||||
|
"""The index knows what Lidarr holds; whether music-mirror has encoded the
|
||||||
|
MP3 yet is a different question, and is checked rather than assumed."""
|
||||||
|
api, source, mirror = playlist_library(tmp_path)
|
||||||
|
for mp3 in mirror.rglob("*.mp3"):
|
||||||
|
mp3.unlink()
|
||||||
|
store = store_at(tmp_path)
|
||||||
|
ingest(store, FakeLastfm([scrobble_of("Played Band", "Hit")]))
|
||||||
|
music_curator.index_library(
|
||||||
|
music_curator.Lidarr("http://lidarr", "key", transport=api), store
|
||||||
|
)
|
||||||
|
music_curator.match_library(store)
|
||||||
|
|
||||||
|
total = music_curator.build_playlists(store, mirror, str(source), 100, NOW)
|
||||||
|
|
||||||
|
assert total == 0
|
||||||
|
assert (mirror / "_playlists" / "all-time.m3u").read_text() == "#EXTM3U\n"
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_library_root_is_derived_from_the_artist_folders(tmp_path):
|
||||||
|
api, source, mirror = playlist_library(tmp_path)
|
||||||
|
store = store_at(tmp_path)
|
||||||
|
music_curator.index_library(
|
||||||
|
music_curator.Lidarr("http://lidarr", "key", transport=api), store
|
||||||
|
)
|
||||||
|
|
||||||
|
assert music_curator.library_root_of(store) == str(source)
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_playlist_limit_is_honoured(tmp_path):
|
||||||
|
api, source, mirror = playlist_library(tmp_path)
|
||||||
|
store = store_at(tmp_path)
|
||||||
|
music_curator.index_library(
|
||||||
|
music_curator.Lidarr("http://lidarr", "key", transport=api), store
|
||||||
|
)
|
||||||
|
music_curator.match_library(store)
|
||||||
|
|
||||||
|
music_curator.build_playlists(store, mirror, str(source), 1, NOW)
|
||||||
|
|
||||||
|
unheard = (mirror / "_playlists" / "unheard.m3u").read_text().splitlines()
|
||||||
|
assert len([line for line in unheard if line.startswith("#EXTINF")]) == 1
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_rotation_moves_weekly_not_every_pass(tmp_path):
|
||||||
|
"""A playlist that reorders on every pass is one that has to be re-imported
|
||||||
|
on every pass; the Music app imports a snapshot, it does not track a file."""
|
||||||
|
api, source, mirror = playlist_library(tmp_path)
|
||||||
|
store = store_at(tmp_path)
|
||||||
|
music_curator.index_library(
|
||||||
|
music_curator.Lidarr("http://lidarr", "key", transport=api), store
|
||||||
|
)
|
||||||
|
music_curator.match_library(store)
|
||||||
|
|
||||||
|
def unheard_at(when):
|
||||||
|
music_curator.build_playlists(store, mirror, str(source), 100, when)
|
||||||
|
return (mirror / "_playlists" / "unheard.m3u").read_text()
|
||||||
|
|
||||||
|
same_week = unheard_at(NOW), unheard_at(NOW + 3600)
|
||||||
|
assert same_week[0] == same_week[1]
|
||||||
|
|
||||||
|
|
||||||
|
def test_playlists_are_group_readable(tmp_path):
|
||||||
|
api, source, mirror = playlist_library(tmp_path)
|
||||||
|
store = store_at(tmp_path)
|
||||||
|
music_curator.index_library(
|
||||||
|
music_curator.Lidarr("http://lidarr", "key", transport=api), store
|
||||||
|
)
|
||||||
|
music_curator.match_library(store)
|
||||||
|
|
||||||
|
music_curator.build_playlists(store, mirror, str(source), 100, NOW)
|
||||||
|
|
||||||
|
for playlist in (mirror / "_playlists").glob("*.m3u"):
|
||||||
|
assert playlist.stat().st_mode & stat.S_IRGRP, playlist
|
||||||
|
|
||||||
|
|
||||||
|
def test_tag_weights_fall_back_to_rank_when_no_count_is_sent():
|
||||||
|
"""The documented sample carries only a name and a URL; a live response also
|
||||||
|
carries a count. Neither may be relied on alone."""
|
||||||
|
with_count = music_curator.parse_tags(
|
||||||
|
{"toptags": {"tag": [{"name": "Screamo", "count": 100}, {"name": "emo", "count": 40}]}}
|
||||||
|
)
|
||||||
|
assert with_count == [("screamo", 100), ("emo", 40)]
|
||||||
|
|
||||||
|
without = music_curator.parse_tags(
|
||||||
|
{"toptags": {"tag": [{"name": "screamo"}, {"name": "emo"}]}}
|
||||||
|
)
|
||||||
|
assert [tag for tag, _ in without] == ["screamo", "emo"]
|
||||||
|
assert without[0][1] > without[1][1]
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_lone_tag_is_not_a_list():
|
||||||
|
assert music_curator.parse_tags({"toptags": {"tag": {"name": "dnb", "count": 90}}}) == [
|
||||||
|
("dnb", 90)
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
def test_an_artist_lastfm_cannot_answer_for_is_not_asked_again(tmp_path):
|
||||||
|
"""Recording the fetch even when it returns nothing is what stops a pass
|
||||||
|
spending a request per unknown artist, forever."""
|
||||||
|
api, source, mirror = playlist_library(tmp_path)
|
||||||
|
store = store_at(tmp_path)
|
||||||
|
music_curator.index_library(
|
||||||
|
music_curator.Lidarr("http://lidarr", "key", transport=api), store
|
||||||
|
)
|
||||||
|
lastfm = FakeLastfm(tags={})
|
||||||
|
|
||||||
|
assert music_curator.sync_tags(client_for(lastfm), store, NOW) == 2
|
||||||
|
before = len(lastfm.calls)
|
||||||
|
|
||||||
|
assert music_curator.sync_tags(client_for(lastfm), store, NOW) == 0
|
||||||
|
assert len(lastfm.calls) == before
|
||||||
|
|
||||||
|
|
||||||
|
def test_tags_are_refetched_once_they_go_stale(tmp_path):
|
||||||
|
api, source, mirror = playlist_library(tmp_path)
|
||||||
|
store = store_at(tmp_path)
|
||||||
|
music_curator.index_library(
|
||||||
|
music_curator.Lidarr("http://lidarr", "key", transport=api), store
|
||||||
|
)
|
||||||
|
lastfm = FakeLastfm(tags={})
|
||||||
|
music_curator.sync_tags(client_for(lastfm), store, NOW)
|
||||||
|
|
||||||
|
later = NOW + music_curator.TAG_REFRESH_SECONDS + 1
|
||||||
|
assert music_curator.sync_tags(client_for(lastfm), store, later) == 2
|
||||||
|
|
||||||
|
|
||||||
|
def tagged_store(tmp_path, tags):
|
||||||
|
api, source, mirror = playlist_library(tmp_path)
|
||||||
|
store = store_at(tmp_path)
|
||||||
|
music_curator.index_library(
|
||||||
|
music_curator.Lidarr("http://lidarr", "key", transport=api), store
|
||||||
|
)
|
||||||
|
music_curator.sync_tags(client_for(FakeLastfm(tags=tags)), store, NOW)
|
||||||
|
return store, source, mirror
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_vibe_selects_by_tag(tmp_path):
|
||||||
|
store, source, mirror = tagged_store(
|
||||||
|
tmp_path,
|
||||||
|
{
|
||||||
|
"artist-mbid-1": [{"name": "screamo", "count": 100}],
|
||||||
|
"artist-mbid-2": [{"name": "classic rock", "count": 100}],
|
||||||
|
},
|
||||||
|
)
|
||||||
|
vibes = [{"name": "screamo", "tags": ["screamo"]}]
|
||||||
|
|
||||||
|
music_curator.build_vibe_playlists(store, vibes, mirror, str(source), 100, NOW)
|
||||||
|
|
||||||
|
written = (mirror / "_playlists" / "screamo.m3u").read_text()
|
||||||
|
assert "Played Band" in written
|
||||||
|
assert "Silent Band" not in written
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_weakly_tagged_artist_is_below_the_threshold(tmp_path):
|
||||||
|
"""A single low-weight tag is not a genre, it is somebody's stray opinion."""
|
||||||
|
store, source, mirror = tagged_store(
|
||||||
|
tmp_path, {"artist-mbid-1": [{"name": "screamo", "count": 3}]}
|
||||||
|
)
|
||||||
|
vibes = [{"name": "screamo", "tags": ["screamo"]}]
|
||||||
|
|
||||||
|
music_curator.build_vibe_playlists(store, vibes, mirror, str(source), 100, NOW)
|
||||||
|
|
||||||
|
assert (mirror / "_playlists" / "screamo.m3u").read_text() == "#EXTM3U\n"
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_vibe_can_be_restricted_by_release_year(tmp_path):
|
||||||
|
"""What separates eighties synth records from everything else a synthpop
|
||||||
|
tag drags in."""
|
||||||
|
store, source, mirror = tagged_store(
|
||||||
|
tmp_path, {"artist-mbid-1": [{"name": "synthpop", "count": 100}]}
|
||||||
|
)
|
||||||
|
inside = [{"name": "eighties", "tags": ["synthpop"], "years": [1975, 1992]}]
|
||||||
|
outside = [{"name": "nineties", "tags": ["synthpop"], "years": [1993, 1999]}]
|
||||||
|
|
||||||
|
# FakeLidarr dates every album 2019, so neither window should catch it.
|
||||||
|
music_curator.build_vibe_playlists(store, inside, mirror, str(source), 100, NOW)
|
||||||
|
music_curator.build_vibe_playlists(store, outside, mirror, str(source), 100, NOW)
|
||||||
|
assert (mirror / "_playlists" / "eighties.m3u").read_text() == "#EXTM3U\n"
|
||||||
|
assert (mirror / "_playlists" / "nineties.m3u").read_text() == "#EXTM3U\n"
|
||||||
|
|
||||||
|
modern = [{"name": "modern", "tags": ["synthpop"], "years": [2000, 2030]}]
|
||||||
|
music_curator.build_vibe_playlists(store, modern, mirror, str(source), 100, NOW)
|
||||||
|
assert "Played Band" in (mirror / "_playlists" / "modern.m3u").read_text()
|
||||||
|
|
||||||
|
|
||||||
|
def test_the_built_in_vibes_are_all_usable_filenames():
|
||||||
|
for vibe in music_curator.DEFAULT_VIBES:
|
||||||
|
assert music_curator.SAFE_VIBE_NAME.fullmatch(vibe["name"]), vibe["name"]
|
||||||
|
assert vibe["tags"]
|
||||||
|
|
||||||
|
|
||||||
|
def test_a_vibes_file_replaces_the_built_in_set(tmp_path):
|
||||||
|
path = tmp_path / "vibes.json"
|
||||||
|
path.write_text(json.dumps([{"name": "mine", "tags": ["shoegaze"]}]))
|
||||||
|
|
||||||
|
assert music_curator.load_vibes(str(path)) == ({"name": "mine", "tags": ["shoegaze"]},)
|
||||||
|
assert music_curator.load_vibes(None) is music_curator.DEFAULT_VIBES
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.mark.parametrize(
|
||||||
|
"content",
|
||||||
|
[
|
||||||
|
'{"not": "a list"}',
|
||||||
|
'[{"tags": ["x"]}]',
|
||||||
|
'[{"name": "../escape", "tags": ["x"]}]',
|
||||||
|
'[{"name": "ok"}]',
|
||||||
|
"not json at all",
|
||||||
|
],
|
||||||
|
)
|
||||||
|
def test_a_bad_vibes_file_is_refused_up_front(tmp_path, content):
|
||||||
|
"""A bad name would otherwise surface as a file written somewhere
|
||||||
|
unintended, which is a poor way to learn about a typo."""
|
||||||
|
path = tmp_path / "vibes.json"
|
||||||
|
path.write_text(content)
|
||||||
|
|
||||||
|
with pytest.raises(ValueError):
|
||||||
|
music_curator.load_vibes(str(path))
|
||||||
|
|||||||
Reference in New Issue
Block a user