fix: hold one connection to Lidarr open, and retry what deserves retrying
Build and publish container / build (pull_request) Successful in 6m0s
Build and publish container / build (pull_request) Successful in 6m0s
Indexing makes two requests per artist, and more when the album fallback fires. urllib opens a new TCP connection and performs a new DNS lookup for every one of them, so a large library becomes thousands of lookups inside a few minutes. That is enough to exhaust a container's resolver, and the result is "[Errno -3] Try again" on every artist at once -- a failure caused entirely by how the requests were made rather than by anything wrong with Lidarr. Add a transport that keeps one connection open per host, so the name is resolved once and the socket is reused. It retries once on a connection the server has already closed, since a stale keep-alive announces itself only on use. Retries were previously declined on the grounds that Lidarr is on the same LAN. That is not a safe assumption -- it may sit behind a public hostname and a reverse proxy -- and a transient failure currently costs an artist their entire entry for that pass. Transient failures are now retried with a backoff. HTTP 500 is deliberately excluded: it is an exception inside Lidarr's serialisation, not a busy server, and three attempts only delay finding that out. The same distinction gates the album probe added alongside this. Naming the offending album costs one request per album of that artist, which is worth it for a deterministic fault and actively harmful during a network-wide one, where every artist fails and probing each of them multiplies the load responsible. The keep-alive transport is tested against a real local HTTP server rather than a fake, because connection reuse and status mapping are exactly the properties a fake would assume rather than demonstrate.
This commit is contained in:
@@ -67,6 +67,25 @@ Losing an artist's albums does not cost their tracks, which come from a
|
||||
different endpoint with a different mapper, so matching is unaffected. A cull
|
||||
would not be, and the report says so.
|
||||
|
||||
### Talking to Lidarr
|
||||
|
||||
Indexing is two requests per artist, and more when the album fallback fires. On
|
||||
a large library that is thousands of requests in a few minutes. `urllib` opens a
|
||||
new TCP connection and performs a new DNS lookup for every one of them, which is
|
||||
enough to exhaust a container's resolver and produce `[Errno -3] Try again` on
|
||||
everything at once. The client therefore holds one connection open per host and
|
||||
resolves once.
|
||||
|
||||
Transient failures — a dropped connection, a resolver hiccup, `429`, `502`,
|
||||
`503`, `504` — are retried with a backoff. An HTTP `500` is not: it is an
|
||||
unhandled exception inside Lidarr's own serialisation and will be raised again
|
||||
identically. That distinction also decides whether a failure is worth
|
||||
investigating; a library-wide outage is not probed artist by artist, because
|
||||
doing so multiplies the load that caused it.
|
||||
|
||||
A local address is preferable to a public hostname here. It removes DNS, the
|
||||
reverse proxy and its timeouts from a path that needs none of them.
|
||||
|
||||
Tracks and files have no unfiltered endpoint — Lidarr rejects a call with no
|
||||
filter — so they stay per artist. If one artist cannot be served, that artist is
|
||||
skipped and the run continues, but the count is recorded and the coverage report
|
||||
|
||||
Reference in New Issue
Block a user