feat: mirror a lossless library to MP3 for iPod sync (#1)
Build and publish container / build (push) Failing after 51s

## What

A path-for-path MP3 mirror of a lossless library. FLAC in, MP3 out, same relative layout, tags and cover art carried across. Already-MP3 sources are copied rather than re-encoded. Mirror files whose source has gone are deleted, along with any directory they emptied. The source library is never written to — mounted read-only in the compose file, and never opened for writing in code.

| Situation                        | Action                               |
| -------------------------------- | ------------------------------------ |
| No mirror file                   | encode                               |
| Source modified since the mirror | re-encode (a Lidarr quality upgrade) |
| Mirror up to date                | skip                                 |
| Source is already MP3            | copy verbatim                        |
| Source gone                      | delete, prune empty dirs             |

## Design notes

- **No database.** Freshness is mtime: an encode is stamped with its source's mtime, so a file is stale exactly when the two differ. Lidarr owns the library; a second tool with its own index would only fall out of step with it. This is also why it is not beets.
- **Atomic writes.** Encode to a temp file, rename into place. An interrupted run cannot leave a truncated MP3 that the next run treats as finished.
- **A lock file** in the mirror root stops two passes overlapping.
- **Refuses a mirror inside the source tree**, which would otherwise recurse.
- `--subdir` never prunes: a partial pass cannot distinguish an orphan from a file outside its scope.

## Shipping

- Python package with a `music-mirror` console script, no runtime dependencies beyond ffmpeg.
- `Dockerfile` plus `compose.yaml` as a TrueNAS Scale Custom App: source dataset read-only, mirror dataset writable, `MUSIC_MIRROR_INTERVAL=6h`.
- CI runs the tests **inside the image**, against the ffmpeg that ships, on every pull request, and on merge publishes multi-arch (amd64 + arm64) to this Gitea's registry using the `PACKAGES_SECRET` repository secret. Versioning follows the same conventional-commit scheme as `legacy-email-proxy`, including writing the released version back into `pyproject.toml` so the packaging metadata cannot drift behind the tag.

## Behaviour under Lidarr

| Lidarr does this                        | The mirror does this                                                             |
| --------------------------------------- | -------------------------------------------------------------------------------- |
| Replaces a file with a better rip        | Re-encodes in place; same path in, same path out, so no duplicate                 |
| Upgrades MP3 to FLAC                     | Both map to the same `.mp3` mirror path, so the old one is overwritten            |
| Renames a track, album or artist folder  | Old path pruned, new path encoded — correct, though it re-encodes rather than moving |
| Deletes an album or artist               | Every orphaned mirror file goes, and the directories they emptied with them        |

Three defects were found and fixed while writing those tests, each verified to
fail against the previous code:

- Pruning probed the source tree for a mirror file's original name, lowercase
  extensions only. A `.FLAC` source was never found, so its mirror file was
  deleted as an orphan and rebuilt on the next pass, for ever. Pruning now works
  from the set of paths the pass actually accounted for.
- Two sources could claim one mirror path — `01 Song.flac` beside a leftover
  `01 Song.mp3`. Both encoded to the same destination and each pass found the
  loser stale. The better format now wins, ties break on path.
- The source mtime was read after encoding rather than before, so a file still
  being written when the pass reached it could be stamped current while holding
  truncated audio.

## Why there is no flake

The deployment target is a container. A flake here would sit on no path between
the source and the NAS, and `nix flake check` would test against nixpkgs' ffmpeg
while the artefact ships Debian's — precisely the layer these tests exercise. It
was removed in favour of running the suite inside the image. The sibling
`legacy-email-proxy` keeps its flake because there the flake *is* the deployment
mechanism.

## Verification

- 21 tests, all real ffmpeg round-trips: layout, tag survival, MP3 output, skip-when-current, re-encode-on-change, orphan pruning, empty-dir removal, `--no-prune`, MP3 passthrough, `--dry-run`, `--subdir`, external cover art embedding, both refusal paths, and the Lidarr lifecycle cases below.
- The suite also runs in the multi-stage Docker `test` stage: 21 passed. The published `runtime` stage carries neither the tests nor pytest, verified by inspecting the image.
- Container built and run locally against a sample library: correct output path, Cyrillic tags intact.
- The release step was extracted from the workflow and run against a scratch repository: it commits and tags when the version changes, and skips the commit but still tags when `pyproject.toml` already carries it.
- **Not tested against the real library** — the first run there should be `--dry-run`.

## Follow-ups, not in this PR

- Lidarr imports are picked up on the next scheduled pass rather than instantly. A webhook trigger is the obvious next step if six hours feels slow.
- `PACKAGES_SECRET` must exist as a repository secret before the first merge, or the login step fails.

---------

Co-authored-by: Emma Thorpe <emma.thorpe@citrix.com>
Reviewed-on: #1
This commit was merged in pull request #1.
This commit is contained in:
2026-08-21 15:30:58 +01:00
co-authored by Emma Thorpe
parent 38f545ac32
commit d4ccff3b75
11 changed files with 1346 additions and 1 deletions
+8
View File
@@ -0,0 +1,8 @@
.git
.gitignore
result
result-*
__pycache__
*.pyc
.pytest_cache
compose.yaml
+193
View File
@@ -0,0 +1,193 @@
name: Build and publish container
on:
# On merge to main, only build/release when image-affecting files change;
# CI-config and docs changes do not produce a new image. pyproject.toml is
# deliberately absent: the release step below commits to it, and that commit
# must not start another run.
push:
branches: [main]
paths:
- "Dockerfile"
- ".dockerignore"
- "music_mirror.py"
# Pull requests always run (tests and the image build are the checks); no
# path filter.
pull_request:
branches: [main]
workflow_dispatch:
# A newer run cancels an older in-flight run in the same group (keyed by ref),
# so a fresh merge to main supersedes the previous build and only the latest
# release is produced. Each pull request likewise supersedes only its own
# earlier runs.
concurrency:
group: ${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: true
defaults:
run:
shell: bash
jobs:
build:
runs-on: ubuntu-latest
permissions:
contents: write
packages: write
steps:
- name: Checkout
uses: actions/checkout@df4cb1c069e1874edd31b4311f1884172cec0e10 # v6
with:
# Full history and tags are required to derive the next version
# from the conventional-commit messages since the last release.
fetch-depth: 0
# The suite runs inside the image, against the ffmpeg that ships, rather
# than against whatever the runner happens to provide. A failing test
# fails the build. Layers are shared with the push build below.
- name: Run the test suite inside the image
run: docker build --target test -t music-mirror:test .
- name: Determine registry host
run: echo "REGISTRY=${GITHUB_SERVER_URL#*://}" >> "$GITHUB_ENV"
# Derive the release version from conventional commits since the last
# v* tag: feat -> minor, fix/perf -> patch, ! or BREAKING CHANGE -> major.
# Anything else (chore, ci, docs, build) produces no release; those builds
# are published under a sha-<short> tag only.
- name: Compute version and image tags
id: version
run: |
set -euo pipefail
image="${REGISTRY}/${GITHUB_REPOSITORY,,}"
last_tag="$(git tag --list 'v*' --sort=-v:refname | head -n1 || true)"
if [ -n "$last_tag" ]; then
range="${last_tag}..HEAD"
base="${last_tag#v}"
else
range=""
base="0.0.0"
fi
subjects="$(git log ${range} --format='%s')"
bodies="$(git log ${range} --format='%B')"
bump="none"
if printf '%s\n' "$bodies" | grep -qiE 'BREAKING[ -]CHANGE' \
|| printf '%s\n' "$subjects" | grep -qE '^[a-z]+([(][^)]*[)])?!:'; then
bump="major"
elif printf '%s\n' "$subjects" | grep -qE '^feat([(][^)]*[)])?:'; then
bump="minor"
elif printf '%s\n' "$subjects" | grep -qE '^(fix|perf)([(][^)]*[)])?:'; then
bump="patch"
fi
major="${base%%.*}"
rest="${base#*.}"
minor="${rest%%.*}"
patch="${rest##*.}"
release="false"
if [ "${GITHUB_EVENT_NAME}" = "push" ] && [ "$bump" != "none" ]; then
release="true"
case "$bump" in
major) major=$((major + 1)); minor=0; patch=0 ;;
minor) minor=$((minor + 1)); patch=0 ;;
patch) patch=$((patch + 1)) ;;
esac
version="${major}.${minor}.${patch}"
{
echo "tags<<__EOT__"
echo "${image}:${version}"
echo "${image}:${major}.${minor}"
echo "${image}:${major}"
echo "${image}:latest"
echo "__EOT__"
} >> "$GITHUB_OUTPUT"
echo "version=${version}" >> "$GITHUB_OUTPUT"
else
short="$(git rev-parse --short HEAD)"
{
echo "tags<<__EOT__"
echo "${image}:sha-${short}"
echo "__EOT__"
} >> "$GITHUB_OUTPUT"
fi
echo "release=${release}" >> "$GITHUB_OUTPUT"
echo "Computed bump=${bump}, release=${release}, base=${base}"
- name: Set up Buildx
uses: docker/setup-buildx-action@d7f5e7f509e45cec5c76c4d5afdd7de93d0b3df5 # v4
- name: Log in to the Gitea container registry
if: github.event_name != 'pull_request'
uses: docker/login-action@650006c6eb7dba73a995cc03b0b2d7f5ca915bee # v4
with:
registry: ${{ env.REGISTRY }}
username: ${{ github.repository_owner }}
password: ${{ secrets.PACKAGES_SECRET }}
- name: Build and push
uses: docker/build-push-action@f9f3042f7e2789586610d6e8b85c8f03e5195baf # v7
with:
context: .
# Without this the last stage in the Dockerfile -- the test stage --
# would be what gets published.
target: runtime
# The NAS is the only host this runs on. Building arm64 as well would
# mean emulating it under QEMU for no consumer.
platforms: linux/amd64
push: ${{ github.event_name != 'pull_request' }}
tags: ${{ steps.version.outputs.tags }}
labels: |
org.opencontainers.image.source=${{ github.server_url }}/${{ github.repository }}
org.opencontainers.image.revision=${{ github.sha }}
# Record the release: write the computed version into pyproject.toml, then
# commit and tag it, so the packaging metadata always matches the release
# instead of drifting behind it. The version is derived from commit
# messages and only known here, after the build, so it cannot be set by
# hand in the pull request that causes the release.
- name: Record and tag the release
if: steps.version.outputs.release == 'true'
env:
VERSION: ${{ steps.version.outputs.version }}
run: |
set -euo pipefail
python3 - "$VERSION" <<'PY'
import pathlib
import re
import sys
version = sys.argv[1]
path = pathlib.Path("pyproject.toml")
text = path.read_text()
text, count = re.subn(
r'(?m)^version = ".*"$', f'version = "{version}"', text, count=1
)
if count != 1:
raise SystemExit("no version line found in pyproject.toml")
path.write_text(text)
PY
git config user.name "${{ github.actor }}"
git config user.email "${{ github.actor }}@users.noreply.${REGISTRY}"
git add pyproject.toml
# The file may already carry this version, in which case there is
# nothing to commit and `git commit` would fail the job.
if git diff --cached --quiet; then
echo "pyproject.toml is already at ${VERSION}"
else
git commit -m "chore(release): v${VERSION}"
# Push the branch before the tag. If main has moved on and this push
# is rejected, the job fails without having left a tag pointing at a
# commit that is not on main.
git push origin "HEAD:${GITHUB_REF_NAME}"
fi
git tag -a "v${VERSION}" -m "v${VERSION}"
git push origin "v${VERSION}"
+7
View File
@@ -0,0 +1,7 @@
result
result-*
__pycache__/
*.pyc
.venv/
.pytest_cache/
*.egg-info/
+31
View File
@@ -0,0 +1,31 @@
# Alpine rather than Debian slim purely because of ffmpeg's dependencies.
# Debian's ffmpeg depends on libavdevice, which can render to X11, Wayland and
# OpenGL, so it drags in Mesa and with it LLVM (127 MB) and Z3 (27 MB) to
# encode an MP3 on a headless NAS. Alpine's build brings none of that: 576 MB
# down to 189 MB.
FROM python:3.13-alpine AS runtime
ENV PYTHONUNBUFFERED=1
# ffmpeg does the encoding; the application itself has no Python dependencies.
RUN apk add --no-cache ffmpeg
WORKDIR /app
COPY pyproject.toml README.md ./
COPY music_mirror.py ./
RUN pip install --no-cache-dir . \
&& rm -rf build music_mirror.egg-info
# Runs as root by default so a bind-mounted dataset of any ownership is
# writable. Override with `user:` in compose to run as the dataset's owner.
ENTRYPOINT ["music-mirror"]
# Test stage: the suite runs against this image's own ffmpeg, which is the one
# that ships. Build it with `--target test`; a failing test fails the build.
# The published image is the `runtime` stage above and carries none of this.
FROM runtime AS test
RUN pip install --no-cache-dir pytest
COPY pytest.ini ./
COPY tests ./tests
RUN python -m pytest
+141 -1
View File
@@ -1,3 +1,143 @@
# music-mirror
Maintain a lossy MP3 mirror of a lossless music library
Maintain a lossy MP3 mirror of a lossless music library.
Walks a source library and reproduces it, path for path, as MP3 in a separate
tree. Tags and cover art are carried across; sources that are already MP3 are
copied rather than re-encoded; mirror files whose source has been deleted are
removed. **The source library is never written to** — it is mounted read-only
in the supplied compose file, and nothing in the code opens it for writing.
The intended use is an iPod. Apple's Music app cannot read FLAC at all, so a
converted copy has to exist somewhere; this keeps that copy next to the
library on a NAS instead of on a laptop, and keeps it current without a human
remembering to do anything.
## How it decides what to do
| Situation | Action |
| -------------------------------- | ---------------------------------------- |
| No mirror file | encode |
| Source modified since the mirror | re-encode (a Lidarr quality upgrade) |
| Mirror up to date | skip |
| Source is already MP3 | copy verbatim |
| Source gone | delete the mirror file, prune empty dirs |
Freshness is modification time: an encoded file is stamped with its source's
mtime, so a file is stale exactly when the two differ. There is no database to
fall out of step with the library, which matters when something else — Lidarr,
in this case — is the thing that owns and reorganises it.
### What that means for Lidarr
| Lidarr does this | The mirror does this |
| --------------------------------------- | -------------------------------------------------------------------------------- |
| Replaces a file with a better rip | Re-encodes in place. Same path in, same path out, so no duplicate |
| Upgrades MP3 to FLAC | Both map to the same `.mp3` mirror path, so the old one is overwritten |
| Renames a track, album or artist folder | Old path pruned, new path encoded. Correct, but it re-encodes rather than moving |
| Deletes an album or artist | Every orphaned mirror file is deleted and the emptied directories go too |
Pruning is driven by what the pass actually found, not by guessing source
filenames from mirror ones: a `.FLAC` source would not be found by a search for
`.flac`, and the mirror file would be deleted and rebuilt on alternate passes
for ever.
If two sources want the same mirror path — a `01 Song.flac` next to a leftover
`01 Song.mp3`, which is what an interrupted upgrade leaves — the better format
wins, ties break on path, and the loser is logged. Without that rule both
encode to the same destination and every pass finds one of them stale.
The mtime is read _before_ encoding rather than after. A file still being
written when the pass reaches it would otherwise be stamped with its final
mtime while holding truncated audio, and never be revisited.
Encodes are written to a temporary file and renamed into place, so an
interrupted run cannot leave a truncated MP3 that the next run mistakes for
finished work. A lock file in the mirror root stops two passes overlapping.
## Usage
```sh
music-mirror --source /music --mirror /music-mp3 # one pass
music-mirror --source /music --mirror /music-mp3 --interval 6h # keep running
music-mirror --source /music --mirror /music-mp3 --dry-run # report only
music-mirror --source /music --mirror /music-mp3 --subdir "Artist/Album"
```
| Option | Environment variable | Default | Meaning |
| ------------ | ----------------------- | --------- | ---------------------------------------------- |
| `--source` | `MUSIC_MIRROR_SOURCE` | — | Root of the lossless library, read-only |
| `--mirror` | `MUSIC_MIRROR_MIRROR` | — | Root of the MP3 mirror |
| `--quality` | `MUSIC_MIRROR_QUALITY` | `V0` | LAME VBR level `V0``V9`, or kbps e.g. `256` |
| `--jobs` | `MUSIC_MIRROR_JOBS` | CPU count | Concurrent encodes |
| `--interval` | `MUSIC_MIRROR_INTERVAL` | unset | Repeat forever, e.g. `45m`, `6h`, `1d` |
| `--subdir` | — | unset | Limit the pass to one directory; skips pruning |
| `--no-prune` | — | off | Keep mirror files whose source has gone |
| `--dry-run` | — | off | Report what would change, write nothing |
`--subdir` never prunes: a partial pass cannot tell an orphan from a file
outside its own scope.
Requires `ffmpeg` and `ffprobe` on `PATH`. The container image provides both.
## Running it on TrueNAS Scale
`compose.yaml` is a Custom App definition. Adjust the two host paths and the
`user:` to match your pool, then add it as a custom app. The image is published
to this Gitea's registry on every release:
```
code.emmathe.dev/lyrathorpe/music-mirror:latest
```
Tags are `latest`, the full version, and the truncated `major.minor` and
`major` forms; builds that are not releases are published as `sha-<short>`.
Point the mirror at its own dataset rather than a directory inside the music
dataset — it is derived data, so it wants its own snapshot policy, its own
quota, and its own SMB share. The tool refuses to run with a mirror inside the
source tree.
New Lidarr imports are picked up on the next pass. With `MUSIC_MIRROR_INTERVAL`
at `6h` that is the worst case; run `--subdir` by hand if you want an album
immediately.
## Tests
```sh
docker build --target test . # what CI runs
pytest # needs ffmpeg and pytest on PATH
```
The suite runs real ffmpeg encodes rather than mocking them. The interesting
failures are in what ffmpeg actually does with tags, cover art and container
formats, and a mock cannot fail that way — which is also why CI runs the tests
_inside the image_, against the ffmpeg that ships, rather than against whatever
the build runner provides. The published image is the `runtime` stage and
carries neither the tests nor pytest.
Run them directly instead if you prefer; they skip when ffmpeg is absent. On a
Nix machine:
```sh
nix shell nixpkgs#python3Packages.pytest nixpkgs#ffmpeg -c pytest
```
## Getting the result onto an iPod
The mirror is just a directory of MP3s, so any client will do:
- **macOS.** Add the mirror's SMB share to the Music app with _Copy files to
Music Media folder_ and _Keep Media folder organised_ both **off**. The Mac
then stores a library database and nothing else. Keep the share mounted at a
stable path — if it is missing when Music opens, every track shows `!`.
- **Linux.** Rhythmbox links `libgpod` and handles iPod sync. An iPod Video
(5th generation) predates the models whose database has to be signed, so no
firmware-hash trickery is needed.
Neither client transcodes at sync time; they copy finished MP3s.
Two device-side details worth knowing: the iPod reads cover art from the file's
tags and ignores `folder.jpg`, which is why art is embedded here; and volume
levelling on the device uses iTunes' Soundcheck tag, not ReplayGain, so
ReplayGain tags in the source are not carried over as such.
+27
View File
@@ -0,0 +1,27 @@
# TrueNAS Scale "Custom App" definition.
#
# Adjust the two host paths and the user to match your pool. The source is
# mounted read-only on purpose: nothing here should ever be able to write to
# the lossless library.
services:
music-mirror:
image: code.emmathe.dev/lyrathorpe/music-mirror:latest
container_name: music-mirror
restart: unless-stopped
# The dataset owner, so the mirror is not written as root. `id apps` or the
# ownership of the mirror dataset will tell you the right numbers.
user: "568:568"
environment:
MUSIC_MIRROR_SOURCE: /music
MUSIC_MIRROR_MIRROR: /mirror
# LAME V0 averages about 245 kbps. Use 256 for constant bitrate instead.
MUSIC_MIRROR_QUALITY: V0
# How long to wait between passes. Lidarr imports are picked up on the
# next one.
MUSIC_MIRROR_INTERVAL: 6h
# Concurrent encodes; defaults to the CPU count. Lower it to leave the
# NAS responsive during the first full pass.
# MUSIC_MIRROR_JOBS: "4"
volumes:
- /mnt/tank/media/music:/music:ro
- /mnt/tank/media/music-mp3:/mirror
+509
View File
@@ -0,0 +1,509 @@
"""Maintain a lossy MP3 mirror of a lossless music library.
Walks a source library and reproduces it, path for path, as MP3 in a separate
directory tree: FLAC in, MP3 out, same relative layout, tags and cover art
carried across. Sources that are already MP3 are copied rather than re-encoded.
The mirror is derived state. It is only ever written to, never read as a
source of truth, so it can be deleted and rebuilt at any time. Nothing here
writes to the source library.
Staleness is tracked by modification time: an encoded file is given its
source's mtime, so a file is out of date exactly when the two differ. That
makes runs idempotent without a database to keep in step.
"""
import argparse
import concurrent.futures
import fcntl
import logging
import os
import re
import shutil
import signal
import subprocess
import sys
import tempfile
import time
from dataclasses import dataclass
from pathlib import Path
logger = logging.getLogger("music-mirror")
# Sources that are transcoded. Anything ffmpeg can decode works; this list
# decides what the walker picks up in the first place.
SOURCE_EXTENSIONS = {
".flac",
".wav",
".aif",
".aiff",
".ape",
".wv",
".m4a",
".alac",
".ogg",
".opus",
".wma",
}
# Already MP3: copied through. Re-encoding lossy audio to lossy audio costs
# quality for nothing.
COPY_EXTENSIONS = {".mp3"}
# Best first. Used only to settle which source wins when two of them want the
# same mirror path; see plan().
SOURCE_PRIORITY = [
".flac",
".wav",
".aif",
".aiff",
".ape",
".wv",
".alac",
".m4a",
".ogg",
".opus",
".wma",
".mp3",
]
# Looked for in the source directory when a file has no embedded picture.
COVER_NAMES = ("cover.jpg", "folder.jpg", "front.jpg", "cover.png", "folder.png")
# Files the mirror is allowed to contain, and therefore allowed to delete.
MIRROR_SUFFIX = ".mp3"
# Filesystems disagree about mtime precision; SMB in particular rounds.
MTIME_TOLERANCE_SECONDS = 2
@dataclass
class Result:
"""Outcome of processing one file."""
action: str # encoded | copied | skipped | failed
path: Path
error: str = ""
def parse_quality(quality):
"""Return the ffmpeg arguments for a quality setting.
Accepts LAME VBR levels (``V0``..``V9``) or a constant bitrate in kbps
(``256``). VBR is the better trade at a given average bitrate; CBR is for
when a fixed size matters more.
"""
text = str(quality).strip().lower()
if re.fullmatch(r"v[0-9]", text):
return ["-q:a", text[1:]]
if re.fullmatch(r"[0-9]{2,3}", text):
return ["-b:a", f"{text}k"]
raise ValueError(f"unrecognised quality {quality!r}: expected V0-V9 or a bitrate like 256")
def parse_interval(interval):
"""Return seconds for an interval such as ``30m``, ``6h`` or ``90``."""
text = str(interval).strip().lower()
match = re.fullmatch(r"([0-9]+)([smhd]?)", text)
if not match:
raise ValueError(f"unrecognised interval {interval!r}: expected e.g. 45m, 6h, 1d")
value = int(match.group(1))
return value * {"": 1, "s": 1, "m": 60, "h": 3600, "d": 86400}[match.group(2)]
def mirror_path_for(source, source_root, mirror_root):
"""Return the mirror path corresponding to a source file."""
return (mirror_root / source.relative_to(source_root)).with_suffix(MIRROR_SUFFIX)
def is_current(source, mirror):
"""Return whether the mirror file is up to date with its source."""
if not mirror.exists():
return False
return abs(source.stat().st_mtime - mirror.stat().st_mtime) <= MTIME_TOLERANCE_SECONDS
def find_cover(directory):
"""Return an external cover image for a directory, if one is present."""
for name in COVER_NAMES:
candidate = directory / name
if candidate.is_file():
return candidate
return None
def has_embedded_picture(source):
"""Return whether the source carries its own cover art."""
try:
probe = subprocess.run(
[
"ffprobe",
"-v",
"error",
"-select_streams",
"v",
"-show_entries",
"stream=index",
"-of",
"csv=p=0",
str(source),
],
capture_output=True,
text=True,
check=True,
)
except (subprocess.CalledProcessError, FileNotFoundError):
return False
return bool(probe.stdout.strip())
def build_command(source, destination, quality_args, cover):
"""Return the ffmpeg command that encodes one file."""
command = ["ffmpeg", "-nostdin", "-hide_banner", "-loglevel", "error", "-y", "-i", str(source)]
if cover is not None:
command += ["-i", str(cover), "-map", "1:v:0"]
else:
# Optional: the source may have no picture stream at all.
command += ["-map", "0:v:0?"]
command += [
"-map",
"0:a:0",
"-map_metadata",
"0",
"-c:a",
"libmp3lame",
*quality_args,
"-c:v",
"copy",
"-disposition:v",
"attached_pic",
# ID3v2.3 is the widest-compatibility tag version, and what the iPod
# firmware is happiest with; the v1 tag costs 128 bytes.
"-id3v2_version",
"3",
"-write_id3v1",
"1",
# Stated rather than inferred: the destination is a temporary file
# whose suffix ffmpeg would not recognise.
"-f",
"mp3",
str(destination),
]
return command
def encode(source, mirror, quality_args, dry_run):
"""Encode one source file into the mirror, atomically."""
if dry_run:
logger.info("would encode %s", source)
return Result("encoded", mirror)
mirror.parent.mkdir(parents=True, exist_ok=True)
cover = None if has_embedded_picture(source) else find_cover(source.parent)
# Read the source's mtime before encoding, not after. If the file is still
# being written -- a Lidarr import landing mid-pass -- stamping the mirror
# with the later mtime would mark truncated output as current. Stamping the
# earlier one leaves the two mismatched, so the next pass re-encodes it.
stat = source.stat()
# Encode to a temporary file in the destination directory and rename it
# into place, so an interrupted run cannot leave a truncated MP3 that the
# next run would treat as complete.
handle, temporary = tempfile.mkstemp(dir=mirror.parent, suffix=".mp3.part")
os.close(handle)
temporary = Path(temporary)
try:
command = build_command(source, temporary, quality_args, cover)
completed = subprocess.run(command, capture_output=True, text=True)
if completed.returncode != 0:
lines = completed.stderr.strip().splitlines()
return Result("failed", source, lines[-1] if lines else "ffmpeg failed")
os.utime(temporary, (stat.st_atime, stat.st_mtime))
os.replace(temporary, mirror)
except Exception as error: # noqa: BLE001 - reported per file, run continues
return Result("failed", source, str(error))
finally:
temporary.unlink(missing_ok=True)
logger.info("encoded %s", source)
return Result("encoded", mirror)
def copy(source, mirror, dry_run):
"""Copy an already-MP3 source into the mirror."""
if dry_run:
logger.info("would copy %s", source)
return Result("copied", mirror)
mirror.parent.mkdir(parents=True, exist_ok=True)
try:
shutil.copy2(source, mirror)
except OSError as error:
return Result("failed", source, str(error))
logger.info("copied %s", source)
return Result("copied", mirror)
def process(source, mirror, quality_args, dry_run):
"""Bring one source file's mirror entry up to date."""
if is_current(source, mirror):
return Result("skipped", mirror)
if source.suffix.lower() in COPY_EXTENSIONS:
return copy(source, mirror, dry_run)
return encode(source, mirror, quality_args, dry_run)
def find_sources(root):
"""Yield every audio file under a root, in a stable order."""
extensions = SOURCE_EXTENSIONS | COPY_EXTENSIONS
for path in sorted(root.rglob("*")):
if path.is_file() and path.suffix.lower() in extensions:
yield path
def plan(scan_root, source_root, mirror_root):
"""Map each mirror path to the one source that should produce it.
Two sources can want the same mirror path -- `01 Song.flac` alongside a
leftover `01 Song.mp3`, which is what an interrupted Lidarr upgrade leaves
behind. Without a decision here both would encode to the same destination,
each pass would find the loser stale, and the mirror would be rewritten
forever. Preferring the highest-quality source, ties broken by path, makes
the outcome stable and predictable instead.
"""
chosen = {}
for source in find_sources(scan_root):
mirror = mirror_path_for(source, source_root, mirror_root)
rival = chosen.get(mirror)
if rival is None:
chosen[mirror] = source
continue
winner, loser = sorted((source, rival), key=source_rank)
logger.warning("%s and %s both map to %s; using %s", rival, source, mirror, winner)
chosen[mirror] = winner
return chosen
def source_rank(source):
"""Sort key preferring better source formats, then a stable path order."""
suffix = source.suffix.lower()
position = SOURCE_PRIORITY.index(suffix) if suffix in SOURCE_PRIORITY else len(SOURCE_PRIORITY)
return (position, str(source))
def prune(mirror_root, expected, dry_run):
"""Delete mirror files this pass did not account for, and empty dirs.
Driven by the set of paths the pass expects to exist rather than by
probing the source tree for names, which would disagree with it over
letter case and over any extension the walker does not collect.
"""
removed = 0
for mirror in sorted(mirror_root.rglob(f"*{MIRROR_SUFFIX}")):
if mirror in expected:
continue
removed += 1
if dry_run:
logger.info("would remove orphan %s", mirror)
continue
logger.info("removing orphan %s", mirror)
mirror.unlink(missing_ok=True)
if not dry_run:
# Deepest first, so a directory emptied by the loop above is caught.
for directory in sorted(mirror_root.rglob("*"), reverse=True):
if directory.is_dir() and not any(directory.iterdir()):
directory.rmdir()
return removed
def run_once(scan_root, source_root, mirror_root, quality_args, jobs, dry_run, do_prune):
"""Run a single pass. Returns the number of failures.
`scan_root` is what gets walked and `source_root` is what mirror paths are
computed against; they differ only for a partial pass over one directory.
"""
started = time.monotonic()
counts = {"encoded": 0, "copied": 0, "skipped": 0, "failed": 0}
failures = []
work = plan(scan_root, source_root, mirror_root)
with concurrent.futures.ThreadPoolExecutor(max_workers=jobs) as pool:
futures = [
pool.submit(process, source, mirror, quality_args, dry_run)
for mirror, source in work.items()
]
for future in concurrent.futures.as_completed(futures):
result = future.result()
counts[result.action] += 1
if result.action == "failed":
failures.append(result)
removed = prune(mirror_root, set(work), dry_run) if do_prune else 0
for failure in failures:
logger.error("failed: %s: %s", failure.path, failure.error)
logger.info(
"pass complete in %.1fs: %d encoded, %d copied, %d up to date, %d removed, %d failed",
time.monotonic() - started,
counts["encoded"],
counts["copied"],
counts["skipped"],
removed,
counts["failed"],
)
return counts["failed"]
def acquire_lock(mirror_root):
"""Take an exclusive lock so two passes cannot run over one mirror."""
mirror_root.mkdir(parents=True, exist_ok=True)
handle = open(mirror_root / ".music-mirror.lock", "w") # noqa: SIM115 - held for the process
try:
fcntl.flock(handle, fcntl.LOCK_EX | fcntl.LOCK_NB)
except OSError:
handle.close()
return None
return handle
def build_parser():
"""Return the argument parser. Every option also reads an env var, so the
container can be configured without a command line."""
parser = argparse.ArgumentParser(
prog="music-mirror",
description="Maintain a lossy MP3 mirror of a lossless music library.",
)
parser.add_argument(
"--source",
default=os.getenv("MUSIC_MIRROR_SOURCE"),
help="root of the lossless library; never written to (env MUSIC_MIRROR_SOURCE)",
)
parser.add_argument(
"--mirror",
default=os.getenv("MUSIC_MIRROR_MIRROR"),
help="root of the MP3 mirror (env MUSIC_MIRROR_MIRROR)",
)
parser.add_argument(
"--quality",
default=os.getenv("MUSIC_MIRROR_QUALITY", "V0"),
help="LAME VBR level (V0-V9) or a CBR bitrate in kbps (env MUSIC_MIRROR_QUALITY)",
)
parser.add_argument(
"--jobs",
type=int,
default=int(os.getenv("MUSIC_MIRROR_JOBS", "0")) or (os.cpu_count() or 4),
help="concurrent encodes (env MUSIC_MIRROR_JOBS; default: CPU count)",
)
parser.add_argument(
"--interval",
default=os.getenv("MUSIC_MIRROR_INTERVAL"),
help="repeat forever, waiting this long between passes, e.g. 6h (env MUSIC_MIRROR_INTERVAL)",
)
parser.add_argument(
"--subdir",
default=None,
help="limit the pass to one directory below --source; skips pruning",
)
parser.add_argument(
"--no-prune",
action="store_true",
help="keep mirror files whose source has been deleted",
)
parser.add_argument(
"--dry-run",
action="store_true",
help="report what would change without touching the mirror",
)
return parser
def main(argv=None):
"""Entry point. Returns a process exit code."""
logging.basicConfig(format="%(asctime)s %(levelname)s %(message)s", level=logging.INFO)
args = build_parser().parse_args(argv)
if not args.source or not args.mirror:
logger.error("both --source and --mirror are required")
return 2
source_root = Path(args.source).resolve()
mirror_root = Path(args.mirror).resolve()
if not source_root.is_dir():
logger.error("source %s is not a directory", source_root)
return 2
if mirror_root == source_root or mirror_root.is_relative_to(source_root):
logger.error("mirror %s must not sit inside the source library", mirror_root)
return 2
try:
quality_args = parse_quality(args.quality)
interval = parse_interval(args.interval) if args.interval else None
except ValueError as error:
logger.error("%s", error)
return 2
scan_root = source_root
do_prune = not args.no_prune
if args.subdir:
scan_root = (source_root / args.subdir).resolve()
if not scan_root.is_relative_to(source_root) or not scan_root.is_dir():
logger.error("--subdir %s is not a directory below the source", args.subdir)
return 2
# A partial pass cannot tell an orphan from a file outside its scope.
do_prune = False
lock = acquire_lock(mirror_root)
if lock is None:
logger.error("another pass is already running over %s", mirror_root)
return 3
stopping = False
def stop(signum, _frame):
nonlocal stopping
stopping = True
logger.info("signal %d received; finishing the current pass", signum)
signal.signal(signal.SIGTERM, stop)
signal.signal(signal.SIGINT, stop)
try:
while True:
failures = run_once(
scan_root,
source_root,
mirror_root,
quality_args,
args.jobs,
args.dry_run,
do_prune,
)
if interval is None or stopping:
return 1 if failures else 0
logger.info("sleeping %ds", interval)
for _ in range(interval):
if stopping:
return 1 if failures else 0
time.sleep(1)
finally:
lock.close()
def run():
"""Console-script entry point."""
sys.exit(main())
if __name__ == "__main__":
run()
+22
View File
@@ -0,0 +1,22 @@
[build-system]
requires = ["setuptools>=77"]
build-backend = "setuptools.build_meta"
[project]
name = "music-mirror"
version = "0.1.0"
description = "Maintain a lossy MP3 mirror of a lossless music library"
readme = "README.md"
requires-python = ">=3.11"
# No runtime Python dependencies: the work is done by ffmpeg, which must be on
# PATH.
dependencies = []
[project.scripts]
music-mirror = "music_mirror:run"
[project.urls]
Homepage = "https://code.emmathe.dev/lyrathorpe/music-mirror"
[tool.setuptools]
py-modules = ["music_mirror"]
+5
View File
@@ -0,0 +1,5 @@
[pytest]
minversion = 7.0
testpaths = tests
python_files = test_*.py
addopts = -q
+80
View File
@@ -0,0 +1,80 @@
import os
import shutil
import subprocess
import sys
import pytest
# Ensure the project root is on sys.path when running tests.
ROOT = os.path.dirname(os.path.dirname(__file__))
if ROOT not in sys.path:
sys.path.insert(0, ROOT)
@pytest.fixture(scope="session", autouse=True)
def require_ffmpeg():
"""The tests exercise real encodes; there is little point faking them."""
for tool in ("ffmpeg", "ffprobe"):
if shutil.which(tool) is None:
pytest.skip(f"{tool} is not on PATH", allow_module_level=True)
@pytest.fixture
def make_flac():
"""Return a factory writing a short tagged FLAC file."""
def factory(path, title="Test Title", artist="Test Artist", album="Test Album", seconds=1):
path.parent.mkdir(parents=True, exist_ok=True)
subprocess.run(
[
"ffmpeg",
"-nostdin",
"-hide_banner",
"-loglevel",
"error",
"-y",
"-f",
"lavfi",
"-i",
f"sine=frequency=440:duration={seconds}",
"-metadata",
f"title={title}",
"-metadata",
f"artist={artist}",
"-metadata",
f"album={album}",
"-metadata",
"track=3",
str(path),
],
check=True,
capture_output=True,
)
return path
return factory
@pytest.fixture
def probe_tag():
"""Return a helper reading a single metadata tag from a file."""
def reader(path, tag):
completed = subprocess.run(
[
"ffprobe",
"-v",
"error",
"-show_entries",
f"format_tags={tag}",
"-of",
"default=noprint_wrappers=1:nokey=1",
str(path),
],
check=True,
capture_output=True,
text=True,
)
return completed.stdout.strip()
return reader
+323
View File
@@ -0,0 +1,323 @@
import os
import shutil
import subprocess
import time
import pytest
import music_mirror
def run(source, mirror, *extra):
return music_mirror.main(["--source", str(source), "--mirror", str(mirror), *extra])
def test_parse_quality_accepts_vbr_and_cbr():
assert music_mirror.parse_quality("V0") == ["-q:a", "0"]
assert music_mirror.parse_quality("v2") == ["-q:a", "2"]
assert music_mirror.parse_quality("256") == ["-b:a", "256k"]
def test_parse_quality_rejects_nonsense():
with pytest.raises(ValueError):
music_mirror.parse_quality("best")
def test_parse_interval_units():
assert music_mirror.parse_interval("90") == 90
assert music_mirror.parse_interval("30m") == 1800
assert music_mirror.parse_interval("6h") == 21600
assert music_mirror.parse_interval("1d") == 86400
def test_encodes_and_preserves_layout_and_tags(tmp_path, make_flac, probe_tag):
source = tmp_path / "src"
mirror = tmp_path / "dst"
make_flac(source / "Test Artist" / "Test Album" / "03 Song.flac")
assert run(source, mirror) == 0
output = mirror / "Test Artist" / "Test Album" / "03 Song.mp3"
assert output.is_file()
assert probe_tag(output, "title") == "Test Title"
assert probe_tag(output, "album") == "Test Album"
def test_output_is_mp3(tmp_path, make_flac):
source = tmp_path / "src"
mirror = tmp_path / "dst"
make_flac(source / "a.flac")
run(source, mirror)
codec = subprocess.run(
[
"ffprobe",
"-v",
"error",
"-select_streams",
"a:0",
"-show_entries",
"stream=codec_name",
"-of",
"csv=p=0",
str(mirror / "a.mp3"),
],
check=True,
capture_output=True,
text=True,
).stdout.strip()
assert codec == "mp3"
def test_second_pass_skips_unchanged_files(tmp_path, make_flac):
source = tmp_path / "src"
mirror = tmp_path / "dst"
make_flac(source / "a.flac")
run(source, mirror)
first = (mirror / "a.mp3").stat().st_mtime_ns
run(source, mirror)
assert (mirror / "a.mp3").stat().st_mtime_ns == first
def test_changed_source_is_re_encoded(tmp_path, make_flac):
source = tmp_path / "src"
mirror = tmp_path / "dst"
track = make_flac(source / "a.flac")
run(source, mirror)
before = (mirror / "a.mp3").stat().st_mtime
# A replaced file with a newer mtime is what a Lidarr quality upgrade
# looks like on disk.
later = time.time() + 120
os.utime(track, (later, later))
run(source, mirror)
assert (mirror / "a.mp3").stat().st_mtime > before
def test_deleted_source_is_pruned(tmp_path, make_flac):
source = tmp_path / "src"
mirror = tmp_path / "dst"
make_flac(source / "Album" / "a.flac")
make_flac(source / "Album" / "b.flac")
run(source, mirror)
(source / "Album" / "b.flac").unlink()
run(source, mirror)
assert (mirror / "Album" / "a.mp3").is_file()
assert not (mirror / "Album" / "b.mp3").exists()
def test_emptied_directory_is_removed(tmp_path, make_flac):
source = tmp_path / "src"
mirror = tmp_path / "dst"
make_flac(source / "Gone" / "a.flac")
run(source, mirror)
(source / "Gone" / "a.flac").unlink()
run(source, mirror)
assert not (mirror / "Gone").exists()
def test_no_prune_keeps_orphans(tmp_path, make_flac):
source = tmp_path / "src"
mirror = tmp_path / "dst"
make_flac(source / "a.flac")
run(source, mirror)
(source / "a.flac").unlink()
run(source, mirror, "--no-prune")
assert (mirror / "a.mp3").is_file()
def test_existing_mp3_is_copied_not_re_encoded(tmp_path, make_flac):
source = tmp_path / "src"
mirror = tmp_path / "dst"
flac = make_flac(source / "a.flac")
subprocess.run(
["ffmpeg", "-loglevel", "error", "-y", "-i", str(flac), str(source / "b.mp3")],
check=True,
capture_output=True,
)
flac.unlink()
run(source, mirror)
assert (mirror / "b.mp3").read_bytes() == (source / "b.mp3").read_bytes()
def test_format_upgrade_replaces_rather_than_duplicating(tmp_path, make_flac):
"""Lidarr replacing an MP3 with a FLAC must not leave two mirror files."""
source = tmp_path / "src"
mirror = tmp_path / "dst"
flac = make_flac(source / "Album" / "01 Song.flac")
subprocess.run(
["ffmpeg", "-loglevel", "error", "-y", "-i", str(flac), str(source / "Album" / "01 Song.mp3")],
check=True,
capture_output=True,
)
flac.unlink()
run(source, mirror)
assert sorted(p.name for p in (mirror / "Album").iterdir()) == ["01 Song.mp3"]
# The upgrade: the MP3 goes, a FLAC arrives at the same stem.
(source / "Album" / "01 Song.mp3").unlink()
make_flac(source / "Album" / "01 Song.flac", title="Upgraded")
run(source, mirror)
assert sorted(p.name for p in (mirror / "Album").iterdir()) == ["01 Song.mp3"]
def test_renamed_album_leaves_nothing_behind(tmp_path, make_flac):
"""A Lidarr rename is a delete plus an add; the old tree must not linger."""
source = tmp_path / "src"
mirror = tmp_path / "dst"
make_flac(source / "Artist" / "Album (2019)" / "01 Song.flac")
run(source, mirror)
(source / "Artist" / "Album (2019)").rename(source / "Artist" / "Album (2020)")
run(source, mirror)
assert (mirror / "Artist" / "Album (2020)" / "01 Song.mp3").is_file()
assert not (mirror / "Artist" / "Album (2019)").exists()
def test_competing_sources_pick_the_lossless_one_and_stay_stable(tmp_path, make_flac, probe_tag):
"""Both a FLAC and an MP3 at one stem: the FLAC wins, and stays won."""
source = tmp_path / "src"
mirror = tmp_path / "dst"
make_flac(source / "a.flac", title="From FLAC")
other = make_flac(tmp_path / "scratch" / "other.flac", title="From MP3")
subprocess.run(
["ffmpeg", "-loglevel", "error", "-y", "-i", str(other), str(source / "a.mp3")],
check=True,
capture_output=True,
)
# Age the loser well beyond the mtime tolerance, so "is it current?" gives
# a definite answer for it rather than one that depends on the clock.
older = time.time() - 3600
os.utime(source / "a.mp3", (older, older))
run(source, mirror)
assert probe_tag(mirror / "a.mp3", "title") == "From FLAC"
# The loser must not make the mirror look stale on the next pass, or every
# run would re-encode for ever.
first = (mirror / "a.mp3").stat().st_mtime_ns
run(source, mirror)
assert (mirror / "a.mp3").stat().st_mtime_ns == first
def test_uppercase_extension_is_not_pruned_and_re_encoded(tmp_path, make_flac):
"""A .FLAC source must not be treated as an orphan on the next pass."""
source = tmp_path / "src"
mirror = tmp_path / "dst"
made = make_flac(source / "a.flac")
made.rename(source / "a.FLAC")
run(source, mirror)
first = (mirror / "a.mp3").stat().st_mtime_ns
run(source, mirror)
assert (mirror / "a.mp3").is_file()
assert (mirror / "a.mp3").stat().st_mtime_ns == first
def test_whole_library_deleted_empties_the_mirror(tmp_path, make_flac):
source = tmp_path / "src"
mirror = tmp_path / "dst"
make_flac(source / "A" / "one.flac")
make_flac(source / "B" / "two.flac")
run(source, mirror)
shutil.rmtree(source / "A")
shutil.rmtree(source / "B")
run(source, mirror)
assert list(mirror.rglob("*.mp3")) == []
assert not (mirror / "A").exists()
assert not (mirror / "B").exists()
def test_dry_run_writes_nothing(tmp_path, make_flac):
source = tmp_path / "src"
mirror = tmp_path / "dst"
make_flac(source / "a.flac")
assert run(source, mirror, "--dry-run") == 0
assert not (mirror / "a.mp3").exists()
def test_subdir_limits_the_pass_and_does_not_prune(tmp_path, make_flac):
source = tmp_path / "src"
mirror = tmp_path / "dst"
make_flac(source / "One" / "a.flac")
make_flac(source / "Two" / "b.flac")
assert run(source, mirror, "--subdir", "One") == 0
assert (mirror / "One" / "a.mp3").is_file()
# Outside the requested directory: neither encoded nor treated as an orphan.
assert not (mirror / "Two").exists()
def test_external_cover_is_embedded(tmp_path, make_flac):
source = tmp_path / "src"
mirror = tmp_path / "dst"
make_flac(source / "Album" / "a.flac")
subprocess.run(
[
"ffmpeg",
"-loglevel",
"error",
"-y",
"-f",
"lavfi",
"-i",
"color=c=red:s=64x64:d=1",
"-frames:v",
"1",
str(source / "Album" / "cover.jpg"),
],
check=True,
capture_output=True,
)
run(source, mirror)
streams = subprocess.run(
[
"ffprobe",
"-v",
"error",
"-select_streams",
"v",
"-show_entries",
"stream=codec_name",
"-of",
"csv=p=0",
str(mirror / "Album" / "a.mp3"),
],
check=True,
capture_output=True,
text=True,
).stdout.strip()
assert streams # a picture stream is present
def test_mirror_inside_source_is_refused(tmp_path, make_flac):
source = tmp_path / "src"
make_flac(source / "a.flac")
assert run(source, source / "mp3") == 2
def test_missing_source_is_refused(tmp_path):
assert run(tmp_path / "nope", tmp_path / "dst") == 2