Files
bookshelf/docs/HANDOFF.md
T
Spriteandclaude ab83287b5d Move Open Library lookup to the Editions API; stop asking Google Books for its placeholder
Two user-reported shelf-testing symptoms, both a third party answering
misleadingly and the app believing it.

1. Open Library's /api/books?bibkeys=... now 404s for EVERY ISBN, including
   books OL demonstrably still holds, with OL's own x-ol-stats header on the
   response while /isbn/, /search.json, /api/volumes/brief and covers.* all
   serve normally. OL's docs call it the "Legacy Books API" that "may be phased
   out" and it is gone from their API index, so this reads as a retirement
   rather than an outage. The app had been effectively single-sourced on Google
   Books since it broke: every failure the user saw was a book GB lacks.

   Lookup now uses /isbn/{isbn}.json — current, non-legacy, edition-level, and
   the only option of the three that carries a description. /api/volumes/brief
   is a near drop-in for the old response shape and was rejected precisely
   because it is also legacy.

   Its costs, all handled: authors are references, so AuthorNameCache resolves
   and caches them for the process lifetime (books by one author get scanned in
   runs off one shelf); an edition may carry NO authors, in which case they live
   on the work — 9780898707168 on the user's own shelf is exactly this, so
   without the work fallback the move would have silently dropped its author;
   author/work requests are best-effort and can only degrade a record, never
   turn Found into Unavailable.

   404 on this endpoint is authoritative NotFound. The legacy endpoint reported
   a miss as 200 with an empty object, which is why every non-2xx there was a
   failure. Every other non-2xx still is.

2. Google Books answers zoom=2 with a grey "image not available" PNG at HTTP
   200 — not a 404 — for any volume it holds no full preview of. Coil loads it
   as a success, so BookCover's placeholder never fires and the cover pipeline
   uploads Google's placeholder to PocketBase as the book's cover. Measured over
   18 real volumes: 11 placeholders at zoom=2, 0 at zoom=1&w=400. zoom=0/3/6 are
   placeholders too. normalizeCoverUrl now pins zoom=1, adds w=400 and strips
   edge=curl.

SPEC.md's "Book metadata lookup" is rewritten with both rules and the evidence
for them — it was the source of the zoom=2 instruction, and would otherwise be
the reason someone restores it.

Also fixed, because it blocked verification: LibraryViewModelTest never cleared
the view models it built, and LibraryViewModel's eleven WhileSubscribed(5_000)
flows kept running five seconds into later tests, racing resetMain(). It now
cancels each viewModelScope in tearDown.

NOT fixed, reported instead: AddBookViewModel.performSave's in-flight guard is a
check-then-act and two coroutines can both pass it. Unrelated to this change
(that VM has no metadata dependency) and out of scope. See HAZARD #13.

Verified: assembleDebug exit 0; testDebugUnitTest --rerun-tasks 378 tests,
2 skipped, 0 failures (was 355); verifyPaparazziDebug exit 0, no pixels moved;
0 "always 'false'" warnings; no build files touched. LIVE_METADATA=1 live test
ran (not skipped): 9/9 Found with cover art, with the GB key absent, so Open
Library alone answered through the new endpoint.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-20 02:13:01 +00:00

62 KiB
Raw Blame History

Bookshelf — session handoff

Written 2026-09-06 by the Opus orchestrator, after a sprite restart forced a fresh session.

FIRST COMMAND OF A NEW SESSION — is a wave already finished?

Workers run fully detached (hazard #5) and OUTLIVE the orchestrator. A wave can finish while no orchestrator exists. Nothing will "notify" a session that was not running, so completion is recorded ON DISK. Run this before anything else:

cd ~/bookshelf && cat logs/WAVE*-DONE 2>/dev/null; \
  ps -eo args= | awk '$1~/bash$/ && $2~/run-task\.sh$/' | wc -l ; tail -2 logs/*.state
# NOT `pgrep -fc 'run-task\.sh'` — that counts this very shell (HAZARD #6/#10)
  • logs/WAVE<N>-DONE exists -> that wave's workers have STOPPED. Read it, then independently verify (assembleDebug + testDebugUnitTest + git status --porcelain) before accepting anything. Workers self-report optimistically; two of three waves so far over-claimed.
  • no sentinel + pgrep count > 0 -> still running; arm a Monitor and wait.
  • no sentinel + count 0 -> workers were KILLED. Check logs/<name>.json: 0 bytes means killed, not failed (hazard #3). Relaunch with setsid per hazard #5; run-task.sh will RESUME the existing session id rather than restart, so context/quota is preserved.

The sentinel is written by tasks/wave-sentinel.sh, itself launched detached: setsid nohup ./tasks/wave-sentinel.sh WAVE3-DONE E1-shell E2-books </dev/null >/dev/null 2>&1 & An in-process Monitor is only a convenience for a LIVE orchestrator; it dies with the process and caps at 1h. Never rely on it as the record that a wave completed.

Read these first, in order

  1. docs/SPEC.md — the authoritative product/technical contract. Unchanged and still correct. Every worker prompt must point at it. Do not restate it; do not let it drift.
  2. This file — operational state, what's done, what bit us, what's next.

Operating model (the user explicitly asked for this — keep it)

The user is on the $20/mo Pro plan and wants Opus used sparingly.

  • Opus = orchestrator only. Write specs, launch workers, verify results, decide. Do NOT write app code yourself. Do NOT read large files into Opus context.
  • Sonnet = all implementation, via claude -p (NOT the Agent tool — the user asked for claude -p specifically, and it keeps worker output out of the orchestrator's context).
  • Read worker output via logs/<name>.summary / jq -r '.result', never by cat-ing source.
  • Quota status: user reported 62% of the 5h window consumed at ~19:45 on 09-05. Worker A alone cost $2.40 / 85 turns. Budget accordingly; prefer resuming a session over restarting one.

How to launch a worker (BOTH steps — the guard is not optional)

cd ~/bookshelf
setsid nohup ./tasks/run-task.sh <NAME> ./tasks/<NAME>.txt </dev/null >/dev/null 2>&1 &
setsid nohup ./tasks/wave-guard.sh WAVE<N>-DONE <NAME> [<NAME>...] </dev/null >/dev/null 2>&1 &

Without the guard the sprite auto-suspends as soon as the user's console goes idle and the whole wave is lost (hazard #5). The guard holds a /v1/tasks lease, renews it every 15 min, writes logs/WAVE<N>-DONE at the end, and releases the lease so the box can sleep. Check it with sprite-env curl /v1/tasks and cat logs/wave-guard.log. tasks/run-task.sh is quota-aware: on a usage-limit error it sleeps POLL (600s) and resumes the SAME session rather than restarting, up to MAX_WALL (24h), and does not count quota waits against its 3-strike hard-failure budget. It writes: logs/<name>.json (final result), .err, .sid (session id), .state (progress), .summary.

HAZARD — do not repeat: never edit run-task.sh while workers are running. Bash reads scripts by byte offset; swapping the file mid-run makes live workers resume inside unrelated code and can spawn duplicate claude processes that burn quota on finished work. If you must change it, write a NEW file and use that for the next wave.

Environment

  • JDK 21: ~/toolchain/jdk21 (system java is 25 — too new for AGP, do not use it)
  • Android SDK: ~/toolchain/android-sdk (platforms;android-37.0, build-tools;37.0.0, platform-tools)
  • No KVM, no emulator. Verify only via ./gradlew assembleDebug, JVM unit tests, and Paparazzi PNG rendering. Never claim the app was "run".
  • PocketBase v0.40.2 service, bound to 0.0.0.0:8090 and INTERNET-EXPOSED via the sprite proxy at https://bookshelf-dev-b2jqx.sprites.app (the service is registered with --http-port 8090, so the proxy routes to it). This is the live server the user's phone talks to. Restart: sprite-env services restart pocketbase. Logs: /.sprite/logs/services/pocketbase.log. See "Where the server lives" below.
  • Superuser creds: server/.dev-credentials (gitignored).

Where the server lives — ANSWERED 2026-09-12: on this sprite, publicly

This was an open question through waves 1-7 and it now has an answer: for now the server stays here on the sprite and is reachable from the open internet. The user's phone talks to https://bookshelf-dev-b2jqx.sprites.app. Earlier notes in this file and in server/deploy/ said the opposite (localhost-only, unreachable from outside); they were true when written and are now corrected in place. server/deploy/ still documents systemd/Docker/Tailscale for an eventual move to home hardware — that is a future option, not what is running.

Consequences that were not true when the app was designed:

  1. PocketBase's own API rules are now the ONLY thing between this library and the internet. There is no NAT, no Tailscale, no reverse proxy in front of it. The security curls in "Verification standard" below stopped being a formality the day this changed — run them against the PUBLIC hostname, not 127.0.0.1, because localhost cannot tell you what the world can reach.
  2. The sprite auto-suspends when idle and wakes on an incoming HTTP request (that is what --http-port buys). So the server is still "often unreachable" in the sense SPEC's offline-first rule cares about — just for a different reason than residential NAT, and with a cold-start delay on the first request after a pause rather than a hard failure. The size of that delay has not been measured; if a sync or a first login ever looks pathologically slow, measure it before assuming a bug in the app.
  3. The URL is a sprite-scoped hostname. If this sprite is ever renamed or rebuilt the URL changes, and every phone has to be re-pointed at it on the setup screen. That is an argument for moving to Tailscale + real hardware eventually, not a reason to hardcode anything.

Posture re-verified against the PUBLIC hostname, 2026-09-12

Check (anonymous, over the internet) Result
GET /api/collections/books/records 403 {"message":"Authentication required."}
GET /api/collections/shelves/records 403
GET /api/collections/bookcases/records 403
POST /api/collections/users/records (self-registration) 403
GET /api/health 200 (intended — it is a health check and leaks nothing)
GET /api/collections/users/records 403 (was 200 {"items":[]} — fixed same day, below)

The users gap, found and closed on 2026-09-12. users LIST was answering 200 with an empty array instead of 403. No data escaped — there are two real accounts in the DB and an anonymous caller saw neither, because listRule filters the rows out — but the status code was wrong, and it was the exact PocketBase quirk pb_hooks/main.pb.js exists to paper over. That hook listed only bookcases, shelves, books; users was never added, which was defensible while the server was localhost-only and is not now that the collection is world-reachable.

Fixed by adding "users" to the hook's collection list and restarting the service. Re-verified over the internet after the restart: all four collections 403 anonymous, self-registration 403, health 200. Login was re-verified too, because this hook runs on a collection the app authenticates against: POST /api/collections/users/auth-with-password still returns 200 with a token, and an authenticated LIST of all four collections still returns 200. The hook only rejects UNAUTHENTICATED list/search, and auth-with-password is not a list request.

STATE: what is DONE

Wave 1A — server: COMPLETE and verified by the orchestrator (not just self-reported)

server/ contains setup-schema.sh (idempotent), create-user.sh, pb_hooks/main.pb.js, pb_migrations/, deploy/ (systemd unit, Dockerfile, compose, backup.sh, README covering Tailscale vs port-forward+Caddy), README.md, .gitignore.

Independently re-verified on 09-06 after fixing the service:

Check Result
anonymous LIST books/shelves/bookcases 403 / 403 / 403
anonymous self-registration 403
/api/health 200

pb_hooks/main.pb.js is INTENTIONAL, not scope drift. PocketBase's listRule is a row filter, so anonymous LIST would otherwise return 200 [] instead of an error. The hook forces 403. Keep it; it is why the table above passes. It is auto-loaded by the stock binary.

Wave 1B — Android scaffold + design system: COMPLETE, verified by the orchestrator

The pre-restart worker had gotten much further than the last handoff recorded. On 09-06 the orchestrator found everything on disk (theme, 7 shared components, MainActivity, BookshelfApplication, Paparazzi test, all 8 Literata TTFs) and only THREE compile errors, all the same class of trivial import bug — fixed directly by the orchestrator rather than spending a worker session on two-line edits:

  • import androidx.compose.foundation.layout.weight (x2: BookshelfScaffold.kt, the Paparazzi test) — that resolves to the internal RowColumnParentData.weight. weight is a ColumnScope/RowScope member; it needs NO import. Delete the line.
  • SyncStatusBar.kt was missing import androidx.compose.runtime.getValue, so val x by transition.animateFloat(...) had no delegate.
Check Result
./gradlew assembleDebug exit 0 — app-debug.apk, 47MB
./gradlew testDebugUnitTest exit 0
./gradlew recordPaparazziDebug exit 0 — 10 PNGs, light+dark

Snapshots: app/app/src/test/snapshots/images/. The orchestrator eyeballed scaffold-light: warm paper ground, Literata serif title, thin gold hairline rule. Matches the design language.

Repo is now a git repo

git init + baseline commit 8bcd9f7 at the 1B-green point. This is deliberate: it lets the orchestrator verify a wave with git diff --stat / git log instead of reading source files into Opus context, and gives a rollback that isn't a whole-sprite checkpoint restore. Root .gitignore covers build outputs, server/pb_data, .dev-credentials, worker logs.

STATE: what is IN FLIGHT

Wave 2 — C (data layer) + D (metadata/scanning): LAUNCHED 09-06 ~01:59Z, running in parallel

Prompts: tasks/C-data.txt, tasks/D-metadata.txt. Sessions: C-data=c2b92ca5-55d7-49b2-8a89-dc36e3ba4c9f, D-metadata=9a2c0de8-e475-4a43-9302-bc66889ee2bd

Two coordination devices were put in place before launch; keep them for wave 3:

  1. tasks/gw — a flock-serialized gradle wrapper. Both workers share ONE Gradle project dir; concurrent ./gradlew runs clobber each other's outputs. Both prompts forbid ./gradlew and require tasks/gw. Reuse this for every future parallel wave.
  2. Disjoint file ownership, stated as a hard boundary in each prompt. C owns data.local, data.remote, data.repo, data.prefs, AppContainer, BookshelfApplication. D owns data.metadata and ui.scan plumbing. NEITHER may touch app/build.gradle.kts or libs.versions.toml — the orchestrator confirmed every wave-2 dependency is ALREADY declared and wired, so there is no legitimate reason for a worker to edit a build file. D must not wire MetadataRepository into AppContainer (C owns it); D reports the one-line snippet instead, to be applied later.

STATE: what is NOT done

Waves 3-4 — not started. Prompts not yet written.

  • Wave 3 (after C+D land and are verified): E — the six screens (setup, library, detail, scan, locations, settings).
  • Wave 4: F — Paparazzi screenshots for the user to judge the look, release keystore + signed APK, top-level README, end-to-end sync test against the live PocketBase.

Gotchas already paid for — do not rediscover

  1. Migration filename ↔ _migrations desync. Worker A renamed 1788636563_created_books.js to ...564... to fix an alphabetical-replay ordering bug (books sorted before shelves, breaking the relation). Correct for fresh instances, but the dev DB still had the old name recorded applied, so PocketBase tried to re-create books and crash-looped 9 times. Fixed via UPDATE _migrations SET file=.... If you ever rename a migration, update that table too. DB backup: scratchpad data.db.bak.
  2. Rule semantics: in PocketBase "" means PUBLIC, null means superuser-only. Confusing these is exactly how the library would end up world-readable.
  3. claude -p --output-format json writes its log only at exit; a 0-byte .json means the worker is still running or was killed, not that it failed.
  4. System JDK is 25 and will break AGP. Workers must export JAVA_HOME=~/toolchain/jdk21 (run-task.sh already does).

Verification standard (hold workers to this)

Workers self-report optimistically. Before accepting any wave:

  • Re-run the security curls above yourself. The user's stated requirement is that this not be "accessible to everyone in the world"; that check is non-negotiable and cheap.
  • Require assembleDebug + test exit 0, and confirm artifacts exist on disk.
  • Treat "I couldn't get Paparazzi working so I skipped screenshots" as a finding to report to the user, not something to paper over — the user explicitly cares how this looks.

Open questions for the user (not yet asked — deferred, not forgotten)

  • Where the server will actually live (home box vs a sprite) — ANSWERED 2026-09-12: it stays on this sprite, internet-exposed. See "Where the server lives" above.
  • Their two account emails, for create-user.sh. Not needed until the app can log in.

HAZARD #5 — THE SPRITE AUTO-SUSPENDS; detached workers do NOT keep it awake

This is the real cause of the wave-3 loss on 09-06, and an earlier note in this file blamed the wrong thing (it claimed nohup process-group semantics). Correct diagnosis, credit to the user: /.sprite/llm.txt says "When idle, sprites pause automatically" and "Services and sessions keep sprites alive." A detached background process is on NEITHER list. Wave 3 was launched at 03:25 with setsid nohup, the user's console session went away, the sprite went COLD, and all process state was lost. Evidence: at 10:07 uptime read "up 2 min" (boot 10:05:40) while the workers' session files stopped at 03:25 — a machine that stopped and rebooted, not a signalled process.

FIX — hold a sprite task lease for the duration of the wave. Undocumented in /.sprite/docs but live on the API socket: POST /v1/tasks {"name":"","expire":"3600s"} -> holds the sprite HOT GET /v1/tasks -> list active leases DELETE /v1/tasks/ -> release

  • max expire is 3600s (a 2h request is rejected: "exceeds maximum 3600 seconds")
  • re-POSTing a live name returns 409, so RENEWAL = DELETE then POST
  • the lease lives server-side, so it keeps the box up independently of any process; a renewal loop running on the sprite therefore sustains itself

tasks/wave-guard.sh does all of this: acquires the lease, renews every 15 min while workers run, then writes logs/WAVE<N>-DONE and RELEASES the lease so the sprite can suspend instead of idling hot on the user's dime. Launch it detached alongside a wave: setsid nohup ./tasks/wave-guard.sh WAVE3-DONE E1-shell E2-books </dev/null >/dev/null 2>&1 & Still launch workers with setsid (needed so they survive the orchestrator exiting), but understand that alone it does NOT survive a suspend. The lease is what does.

HAZARD #6 — pgrep -f / pkill -f match YOUR OWN shell

Bitten three times in one session, once fatally: pkill -f wave-sentinel.sh killed the orchestrator's own shell (exit 144) because the bash -c command line contained that literal string. Same bug made pgrep -f "claude -p" report a phantom running worker. Use the bracket trick (ps aux | grep "[c]laude -p") or match on argv shape (ps -eo pid,args | grep -E "tasks/(wave-guard|run-task)" | grep -v grep).

HAZARD #7 — never git add -A while workers are running

The orchestrator committed twice (10:09, 10:22) while E1/E2 were actively writing files. git add -A swept half-finished wave-3 SOURCE into commits whose messages said "orchestration tooling". That silently defeats the whole reason this repo exists as git: per-wave git diff --stat verification. It also produced a fake-clean git status, which briefly looked like the workers had produced nothing at all. Rules:

  • While a wave is in flight, commit with an EXPLICIT pathspec only, e.g. git add docs/HANDOFF.md tasks/ && git commit ... — never -A, never ..
  • Do the wave's own git add -A commit only AFTER the sentinel exists and the build and tests have been independently verified.
  • If it happens anyway: git reset --soft <last-good>, git reset, then re-commit in honest slices. Safe here — the repo has no remote and checkpoints exist. Done once already (commits d947aa7/058e864 were rebuilt into dbb0726 + a022a1b).

Wave 3 — COMPLETE, verified by the orchestrator on 09-06

E1 (nav/setup/locations/settings) and E2 (library/detail/scan) both reported SUCCESS; E2 took one quota wait and run-task.sh resumed it correctly.

Check Result
assembleDebug exit 0
testDebugUnitTest exit 0 — 91 tests, 0 failures, 0 errors (was 68)
boundary check clean: neither touched build files or the other's packages
Commit a022a1b (32 files, +3081).

Known gaps carried into wave 4 — do not lose these:

  1. Cover pipeline has NEVER run against a real PocketBase. Download-on-create and multipart upload-on-sync were only exercised against fakes (worker C's own report). This is the single most likely place a real bug is hiding. Wave 4 must do a live round-trip against 127.0.0.1:8090.
  2. Settings shows the PocketBase user id, not the email — AuthRepository/SettingsStore never persist the login email (E1's report). Cosmetic, needs a data-layer change.
  3. No Room foreign keys between books/shelves/bookcases (deliberate, worker C).
  4. No emulator on this box: nothing has ever been run, only compiled and unit-tested. Paparazzi PNGs are the only evidence of how any of it actually looks.

Wave 4 — COMPLETE, verified by the orchestrator on 2026-09-08

This is the last planned wave. All of SPEC's build/verify gates now pass.

Check Result
./tasks/gw assembleDebug exit 0
./tasks/gw testDebugUnitTest exit 0 — 102 tests, 1 skipped, 0 failures, 0 errors
./tasks/gw assembleRelease exit 0
signed APK app/app/build/outputs/apk/release/app-release.apk, 41,777,344 bytes
apksigner verify --print-certs V2 signer CN=Bookshelf, O=Montanaro — the real release key, not the debug cert
secrets app/release-keystore.jks + app/keystore.properties gitignored and NOT committed — re-verified against the staged file list before committing

The skipped test is LiveSyncTest — opt-in, it needs the live PocketBase. It PASSED in wave 4's first half; the cover round-trip gap from wave 3 is closed.

Commits 5455df2 (app work) + 0f47ee9 (orchestration). Pushed to origin/main (ssh://git@git.jfmonty2.com:2022/jfmonty2/bookshelf.git) — the repo now has a remote, so pushing after each commit is the norm.

How wave 4 actually ended — F3 never reported

F3-release wrote all four deliverables between 20:34 and 20:46 on 09-08 and then could not finish. Two separate mechanisms:

  • claude -p terminates background tasks after 600s (Background tasks still running after 600s; terminating. Set CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS=0 to wait indefinitely.). F3 had launched assembleRelease in the background with a Monitor and ended its turn saying it would report when done — the exact failure its own prompt forbade.
  • It then hit the 5h quota and went into a 600s wait loop, so the service supervisor would have kept re-running verification on finished work indefinitely.

The orchestrator ran the verification itself, wrote logs/WAVE4-DONE by hand (noting it was NOT written by service-worker.sh), and stopped the service. If you launch another worker, set CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS=0 in run-task.sh, or tell the worker to run builds in the FOREGROUND. A worker that backgrounds a 10-minute Gradle build cannot ever report on it.

The one thing wave 4 owed and did not deliver

F3 was asked to say candidly what the rendered screens get WRONG against SPEC's design language, now that they can finally be seen. It never reported. Fourteen new PNGs are on disk, unreviewed by anyone. The user explicitly cares how this looks, so this is the top open item — not a build problem, a design-review one: app/app/src/test/snapshots/images/ (setup, detail, scan ×2, locations, settings, library, scaffold — each light + dark).

STATE: what is NOT done (as of 2026-09-08)

  1. Nobody has looked at the screenshots. See above. Highest-value next step.
  2. The app has still never run on a device or emulator (no KVM on this box). Everything is compile + unit-test + Paparazzi evidence only. Installing the signed APK on a real phone is the only way past this, and it needs the user.
  3. R8 is off. Acceptable per F3's prompt, but the release APK is 41.8MB. Turning it on requires proving Room/Retrofit/kotlinx-serialization/ML Kit survive minification.
  4. Still-open user questions: the two account emails for create-user.sh. (Where the server will live was ANSWERED on 2026-09-12 — it stays on this sprite, internet-exposed. See "Where the server lives" near the top.)

First on-device test — 2026-09-09

The user installed the signed APK on a real phone. It runs. This closes the "never been run" gap that waves 1-4 all carried. Six issues came back; five were fixed directly by the orchestrator in commit 356f639 (they were small, and spinning up Sonnet workers for two-line Compose edits costs more than it saves).

The one worth remembering — BookCover branched on painter.state, but in coil3 that is a StateFlow<State>, not a State. Every is AsyncImagePainter.State.X arm was therefore always false, and because a when used as a statement needs no else, it compiled clean and drew NOTHING — no cover, no placeholder, no error icon. Kotlin emitted "Check for instance is always 'false'" as a warning on four consecutive lines and the build stayed green. Grep the build log for always 'false' before accepting a wave; that warning class is a silent-dead-code detector and this build had it for months.

Two more cover defects sat behind it, both verified against the live service:

  • covers.openlibrary.org/b/isbn/{isbn}-L.jpg answers 200 with a 43-byte 1x1 transparent GIF for an edition with no art. Any image loader calls that a successful load. Only ?default=false turns a miss into a 404.
  • OL's DTO synthesized that URL unconditionally, so MetadataMerger's fill-blanks rule could never reach Google Books' thumbnail. The SPEC'd cover fallback was dead code. Cover URLs now come from OL's own cover object.

Also fixed: sync bar clipped by rounded display corners (now owns its navigation-bar inset, wider horizontal padding, moved into Scaffold's bottomBar slot); setup screen's Password field hidden behind the IME (safeDrawingPadding outside verticalScroll, plus Next/Next/Done IME actions and windowSoftInputMode=adjustResize); library card titles reflowed (20sp leading, author gets its own 4dp gap); scan sheet now names the ISBN and says "Searching…" instead of showing a bare spinner.

Check Result
./tasks/gw assembleDebug exit 0
./tasks/gw testDebugUnitTest exit 0 — 106 tests, 1 skipped, 0 failures
./tasks/gw verifyPaparazziDebug exit 0 against re-recorded snapshots
./tasks/gw assembleRelease exit 0 — 41,777,376 bytes
apksigner verify V2 signer CN=Bookshelf, O=Montanaro — real release key

Open, not started: metadata coverage

The user reported 1 of 3 scans resolving, and asked for research, not a change. Findings are in docs/METADATA-SOURCES.md. Headline: Open Library answered 88% of a 60-ISBN sample, keyless Google Books returned 429 on 60 of 60 requests, and both clients collapse every non-200 into null — so a rate-limited lookup reaches the user as "No match found." Recommended order is a free Google Books API key, then distinguishing "couldn't ask" from "not found", then retry/backoff, before adding any new source. Awaiting the user's decision; do not implement unasked.

Wave 5 — G-diagnostics: IN FLIGHT, launched 2026-09-09 10:42Z

Prompt: tasks/G-diagnostics.txt. Session id in logs/G-diagnostics.sid. Lease bookshelf-wave held and VERIFIED by tasks/wave-guard.sh, sentinel logs/WAVE5-DONE. Launched from a console the user was about to disconnect, so the lease is the only thing keeping the sprite hot — check it first if anything looks stalled: sprite-env curl /v1/tasks and tail logs/wave-guard.log.

Scope: make the app distinguish three outcomes it currently conflates — barcode-didn't-decode (silent today), lookup-request-failed (reported as "No match found"), and genuinely-not-found. SPEC's "Book metadata lookup" and "Barcode scanning" sections were rewritten to state the three-way contract (Found / NotFound / Unavailable) BEFORE launch, so the worker implements a spec rather than inventing one. This is the only wave-5 task; there is no parallel worker, so the disjoint-ownership device isn't needed, but the prompt still forbids build files and every package outside data.metadata + ui.scan.

Why this and not more sources: the two books that failed on the phone (9781883937386, 9781883937676) are BOTH fully present in Open Library, with cover art, and the app's own parser handles their real responses — there are regression fixtures and a test proving it. The coverage hypothesis is dead. See docs/METADATA-SOURCES.md § "What actually failed". Do not let a future worker "fix" this by bolting on a third data source.

Verify before accepting (workers self-report optimistically; two of five waves over-claimed): cd ~/bookshelf && ./tasks/gw assembleDebug && ./tasks/gw testDebugUnitTest
&& ./tasks/gw recordPaparazziDebug && git status --porcelain 107 tests pass today; the count must go UP and nothing may regress. Also grep the build log for always 'false' — that warning class silently blanked every book cover in this app for months and is now an explicit item in the worker's prompt. Then eyeball the two new Paparazzi PNGs (LookupFailed sheet, rejected-barcode overlay) — the user cares how this looks and no worker has ever been trusted on that.

Still open, unchanged: the free Google Books API key (keyless returns 429; worth doing on its own merits but no longer the leading theory), R8 still off so the release APK is 41.8MB and too large to send over the file channel (30MB cap), and the two account emails for create-user.sh. (Where the server will live was ANSWERED on 2026-09-12: it stays on this sprite, internet-exposed.)

Wave 5 — G-diagnostics: COMPLETE, verified by the orchestrator 2026-09-09

Commit 93f972b. The app now distinguishes barcode-didn't-decode from lookup-request-failed from genuinely-not-found; see the commit message and docs/METADATA-SOURCES.md.

Check Result
assembleDebug exit 0
testDebugUnitTest exit 0 — 138 tests, 1 skipped, 0 failures (was 107)
verifyPaparazziDebug exit 0
assembleRelease exit 0 — 41,793,760 bytes, V2 signer CN=Bookshelf
boundary check clean — no build files, no forbidden packages
grep "always 'false'" 0 hits on touched files

The worker was honest this time: everything it claimed checked out. Cost $0.28, 6 turns, one quota wait that run-task.sh resumed correctly.

The orchestrator added one thing the worker's brief didn't cover: the manual-ISBN dialog silently discarded an unparseable entry — the same silent failure this wave existed to eliminate, sitting just outside the prompt's scope. It now marks the field in error and disables "Look up" until the checksum passes. Lesson for future prompts: scope a wave by failure class, not by file list, or the instances of the class that live outside the listed files survive.

HAZARD #8 — the wave-guard can die without writing its sentinel

logs/WAVE5-DONE was written BY HAND. The guard renewed at 11:13, the worker succeeded at 11:21, and the guard neither wrote the sentinel nor logged its "guard exiting" trap line — it was killed outright. The lease expired on its own an hour later.

This breaks the first-command heuristic at the top of this file. "no sentinel

  • pgrep count 0 -> workers were KILLED" was WRONG here: the worker had finished successfully. Use these instead, in this order:
    1. ls -l logs/<name>.json — 0 bytes means killed; non-zero means it finished.
    2. tail logs/<name>.state — says SUCCESS / GIVING UP / WALL CLOCK explicitly.
    3. git status --porcelain — is there actually work in the tree? The sentinel is a convenience, not the record of truth. logs/<name>.state is.

Wave 6 — second on-device feedback round: COMPLETE, verified 2026-09-09

The user tested the phone build again and sent eight items. All eight are done. Two Sonnet workers (tasks/H1-screens.txt, tasks/H2-picker.txt) took the UI work; the ORCHESTRATOR did the retry/backoff work itself in data/metadata and AppContainer, because it needed the live measurement below to design it.

Check Result
./tasks/gw assembleDebug exit 0
./tasks/gw testDebugUnitTest exit 0 — 172 tests, 1 skipped, 0 failures (was 138)
./tasks/gw verifyPaparazziDebug exit 0
grep "always 'false'" on a --rerun-tasks rebuild 0 hits
./tasks/gw assembleRelease exit 0 — 41,810,740 bytes, V2 signer CN=Bookshelf, O=Montanaro
boundary check clean — neither worker touched a build file or the other's packages

The ghost bookcase was an inset bug, not a data bug

LocationsScreen's list branch dropped the Scaffold's innerPadding while its empty-state branch applied it, so the FIRST bookcase row rendered underneath the top app bar and was invisible. Every symptom the user described follows from that: invisible first bookcase, no empty state on re-entry (the list was genuinely non-empty), a second bookcase created, both showing in the filter menu. Both records were always real and healthy — the user should delete the spare. Every other screen was checked for the same class of bug; Locations was the only one. Fix: fold innerPadding into the LazyColumn's contentPadding (NOT Modifier.padding, which would clip the scroll area instead of insetting it).

Metadata: measured, not guessed

See docs/METADATA-SOURCES.md § "Measured again 2026-09-09" for the full data. Two things that change how you should think about this app:

  1. Google Books keyless is dead for everyone, permanently. The user's residential-IP test returned a quota error naming project_number:624717413613 — a shared anonymous project, not an IP. The old note in METADATA-SOURCES.md guessing that a residential IP "may well get answers" is now marked CORRECTED in place. Because combine() turns any Failed-with-no-Found into Unavailable, this standing failure meant every Open Library hiccup surfaced as "one or more sources couldn't be reached". The app has been single-sourced all along. The user has deliberately deferred the API key — do not add it unasked.
  2. Our own timeouts were manufacturing failures. 30 live requests: 13% failed, all fast TLS resets (<2.5s); successes had a median of 4.3s but a max of 22.0s, and 2 of 26 successes exceeded the old 12s callTimeout. Timeouts are now 25s/20s/20s. Failures are fast and successes are slow, so a short timeout buys nothing on the failure path and costs real successes on the slow path.

RetryPolicy + withRetry (new, data/metadata/) retry TRANSPORT and SERVER_ERROR only. It deliberately does NOT retry:

  • TIMEOUT — the budget is already spent; retrying could triple the wait.
  • RATE_LIMITED — hammering a quota is how an intermittent block becomes a permanent one, and METADATA-SOURCES.md records that happening to this project's IP. Revisit when the Google Books key lands: a keyed 429 is a per-second limit and does deserve one Retry-After-respecting retry. SourceResult.Failed now carries a FailureKind alongside its human reason, and reason names the specific exception ("tls connection reset, 3 attempts") instead of a generic "network error". That string is now rendered on the scan sheet and is our ONLY diagnostic channel from a real phone. Nothing may parse it.

Known-soft spots in wave 6 — do not mistake these for verified

  1. The Paparazzi "regression" snapshot for the ghost bookcase is a lookalike, not the real screen. LocationsScreenPaparazziTest hand-rolls its own Scaffold+LazyColumn copy because the real LocationsScreen needs an AppContainer (Room + DataStore). The orchestrator verified the REAL fix by reading the diff; the PNG only proves the test's copy is right, and the two can drift — the copy already omits the bottom inset the real screen adds. Splitting a stateless LocationsContent(state, callbacks) out of the screen would make this snapshot genuine. Worth doing before anyone trusts it as regression cover.
  2. The auto-focus calls are unverified. Three dialogs now do LaunchedEffect(Unit) { runCatching { focusRequester.requestFocus() } }. That is the idiomatic form, but there is no emulator here and runCatching means a too-early call fails SILENTLY rather than crashing. If a dialog opens unfocused on the phone, that is why; the fix is to await a frame before requesting.
  3. The shelf picker opens as a bottom sheet stacked on top of the save sheet (H2's own flagged judgement call). It renders correctly in Paparazzi but sheet-over-sheet is awkward on real Android. Watch it on the device.

Worker lessons (both are repeats — the prompts already forbade them)

  • H1 backgrounded a Gradle build and ended its turn, exactly the wave-4 failure, despite an explicit foreground-only instruction AND CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS=0 being set in run-task.sh. Its final message was "I'll wait for this background build to complete." run-task.sh still recorded SUCCESS because the process exited 0. .state saying SUCCESS means the process exited cleanly, NOT that the worker finished its task — read logs/<name>.summary and check that the result is an actual report. H1's work was fine, but nobody verified it except the orchestrator.
  • H2 (108 turns, $4.00) followed the brief closely, ran builds in the foreground, and reported honestly, including flagging its own stacked-sheet judgement call. Cost ratio to H1 ($0.66, 5 turns) is roughly the ratio of work actually done.

Wave 7 — I-gbkey (Google Books API key): COMPLETE, verified by the orchestrator 2026-09-11

Prompt: tasks/I-gbkey.txt. The user asked for the key on 09-11, which lifts the standing "do not add it unasked" instruction recorded in wave 6.

Split of work, deliberately: the ORCHESTRATOR did the build plumbing (app/app/build.gradle.kts: read GOOGLE_BOOKS_API_KEY from local.properties, enable buildConfig, emit BuildConfig.GOOGLE_BOOKS_API_KEY) and verified it green BEFORE launching the worker, because workers on this project are barred from build files. The Sonnet worker did the Kotlin against a tree where the key was already available. Reuse this pattern for anything needing a build-file change.

Check Result
./tasks/gw assembleDebug exit 0
./tasks/gw testDebugUnitTest --rerun-tasks exit 0 — 189 tests, 1 skipped, 0 failures (was 172)
test count source summed from TEST-*.xml, not the console
./tasks/gw verifyPaparazziDebug exit 0 — no pixels moved, as intended
./tasks/gw assembleRelease exit 0 — 41,810,740 bytes
grep "always 'false'" on a --rerun-tasks build 0 hits
boundary check clean — worker touched only data/metadata + the one AppContainer line
key-leak check git grep finds no real key in the tree; tests use test-key-123
live API check HTTP 200 for both previously-failing ISBNs and a control

The worker's report was honest: every claim re-verified, including the test count, and it flagged its own judgement calls (how it reconciled the slightly ambiguous Retry-After cap wording) rather than papering over them. 51 turns, $1.59.

The live check was not ceremony. HttpUrl.Builder percent-encodes the colon, so the app sends q=isbn%3A... where every earlier hand-run test sent q=isbn:.... Offline tests cannot distinguish those. Verified: the API accepts both.

Both books that failed on the phone are in Google Books — so the restored fallback now covers exactly the Open Library TLS-reset failure mode that actually broke those two scans. The app has been effectively single-sourced since it was written and is now genuinely two-sourced. Details in docs/METADATA-SOURCES.md § "The key landed".

HAZARD #9 — run-task.sh reads the WORKER'S OWN OUTPUT for quota strings

run-task.sh's quota detector greps the worker's result blob for usage limit|...|429|too many requests|.... That blob includes .result — the worker's own prose. This wave's task was ABOUT HTTP 429, so the moment the worker finished and wrote a report mentioning 429, the runner declared QUOTA hit (wait #1), slept 600s, and was about to --resume a session that had already SUCCEEDED — which would have burned quota redoing finished work and let a fresh worker turn loose on a completed tree.

Caught it by checking the log rather than trusting the state line: jq '{is_error, subtype, num_turns}' logs/I-gbkey.json said is_error:false, subtype:"success", num_turns:51. The orchestrator killed the runner (PID from ps -o pid,args -p <pid>, per HAZARD #6 — not pkill -f) before the sleep elapsed, and appended a note to logs/<name>.state saying why.

Before believing any QUOTA hit line, check whether the worker actually finished: a non-empty logs/<name>.json with is_error:false means it SUCCEEDED and the runner is about to waste a session.

FIXED 2026-09-12 in run-task.sh itself (safe: no workers were running — the standing rule is only about editing it while a wave is live). Two defences, because either alone would have prevented this:

  1. The detector blob is now stderr plus jq -r 'select(.is_error==true) | .result' — the worker's prose reaches it ONLY when the run actually errored. If the log isn't valid JSON at all (a hard crash), it falls back to the whole log, where there is no prose to be confused by.
  2. The success check moved ABOVE the quota and session-vanished checks. A finished worker is finished regardless of what strings appear in its output.

Regression-checked against the real logs/I-gbkey.json that caused this: the old logic matches the quota pattern, the new logic yields SUCCESS and feeds the detector an empty blob.

Wave 8 — J-crashsafe: launched 2026-09-12 17:50Z — COMPLETE, see "accepted" below

Prompt: tasks/J-crashsafe.txt. Session id in logs/J-crashsafe.sid. Lease bookshelf-wave held by tasks/wave-guard.sh, sentinel logs/WAVE8-DONE.

Why this wave exists. The user installed the wave-7 build and the FIRST barcode scanned crashed the app: ISBN decoded and displayed, spinner ran a few seconds, process died. Never reproduced, on that book or any other. Their theory — that the scan fell through to Google Books and crashed there — is structurally the best fit: keyless GB 429'd every caller on earth, so GoogleBooksClient.classify's 2xx branch, toBookMetadata(), normalizeCoverUrl() and the two-source merge had NEVER EXECUTED in production before 486f6eb. A first crash belongs in a path's first real exercise.

It does not reproduce off-device, and that is now evidence rather than a guess. app/app/src/test/java/org/modg/bookshelf/livemetadata/LiveMetadataLookupTest.kt (commit 93546ed, opt-in on LIVE_METADATA=1) drives the real MetadataRepository.lookup against both real APIs with the real key. Nine lookups — both of the user's previously-failing ISBNs plus a control, three times each — all returned Found with cover art, 254ms to 4.7s, nothing thrown. Static review found no unsafe operation in that path either. Do not let a future worker "fix" this by rewriting the parse/merge code; that code was exercised live and is fine.

What the wave actually fixes is the defect found while looking: nothing on that path is exception-safe. Both clients' fetch catch only IOException, classify catches only SerializationException/IllegalArgumentException, and ScanViewModel.runLookup sits inside viewModelScope.launch { ...collect { } } with no try/catch at all. So any throwable that is not an IOException — platform TLS, an OkHttp internal, memory pressure, an API-level difference, none of which this JVM reproduces — kills the process instead of surfacing as "couldn't be reached". Both clients' KDoc claims "Never throws"; that claim is false today. Plus: the app has NO crash capture whatsoever, which is why a one-time crash left nothing to work from.

Scope, per the user's decision: (1) both clients catch Throwable -> new FailureKind.UNEXPECTED, non-retryable, reason names the exception class only (never its message — that can carry the key); (2) ScanViewModel.runLookup guards the same way, which also covers the Room call in it; (3) a CrashReporter that persists the stack trace and chains to the previous handler; (4) a Diagnostics section in settings to read and SHARE it off the phone. CancellationException must be rethrown, not converted, in every one of those catches — it is normal control flow here (dismissing the sheet cancels the lookup) and is the easiest thing in this wave to get wrong.

Verify before accepting (workers self-report optimistically; several waves have over-claimed): cd ~/bookshelf && ./tasks/gw assembleDebug && ./tasks/gw testDebugUnitTest
&& ./tasks/gw verifyPaparazziDebug && git status --porcelain 190 tests, 2 skipped today (LiveSyncTest + LiveMetadataLookupTest, both opt-in live). Count from the TEST-*.xml files, not the console. The count must go UP. Grep the build log for always 'false'. Then eyeball the new settings Paparazzi PNGs. Check specifically that CancellationException is rethrown in every new catch, and that CrashReporter calls the previously-installed handler — a handler that does not chain leaves the process hung instead of dying.

Not in scope, deliberately: finding the original throwing line. Nobody knows what it was and the prompt says not to hunt for it. If the guard lands and the user ever sees "unexpected: SomeException" on the scan sheet, THAT is when we learn the answer.

Wave 8 — J-crashsafe: COMPLETE, verified by the orchestrator 2026-09-13

Commit 3409726. Worker finished 09-12 18:20Z: 1 attempt, 0 quota waits, 111 turns, $4.02. Its report was honest and every claim below re-checked.

Check Result
./tasks/gw assembleDebug --rerun-tasks exit 0
./tasks/gw testDebugUnitTest exit 0 — 205 tests, 2 skipped, 0 failures (was 190), from TEST-*.xml
./tasks/gw verifyPaparazziDebug exit 0
grep "always 'false'" on the rerun build 0 hits
boundary check clean — data/metadata, ui/scan, ui/settings, new diagnostics/, CrashReporter wiring in AppContainer + BookshelfApplication; no build files

Read, not just trusted: CancellationException is caught and rethrown AHEAD of the Throwable catch in both clients and in ScanViewModel.runLookup (the import is kotlinx.coroutines.CancellationException, the same type). Reasons carry e.javaClass.simpleName only. RetryPolicy does not retry UNEXPECTED. CrashReporter.install() calls previous?.uncaughtException in a finally, so the process still dies even if writing the report throws. The Diagnostics PNGs (empty and populated, light and dark) match the rest of settings. One nit: outlined mahogany buttons on the dark ground are low-contrast, but "Sign out" already looked like that.

Open, flagged by the worker, deliberately out of scope: performSave / performSaveManualEntry make the same unguarded Room calls runLookup used to. A disk failure while SAVING can still crash the process. Same failure class; small follow-up if wanted.

The payoff is on the phone, not here. If the original crash recurs, it now lands either as "unexpected: SomeException" on the scan sheet (caught) or as a stored trace under Settings → Diagnostics → Share (uncaught). Either one finally tells us the class.

HAZARD #10 — the wave-guard held the sprite hot for 17 hours

The guard's loop was while pgrep -f 'run-task\.sh|run-resume\.sh'. The orchestrator launched the wave from ONE bash -c '... setsid nohup ./tasks/run-task.sh ... & ... wave-guard.sh ... & sleep 8; ...', and that shell (re-parented to PID 1) was still alive the next day. Its command line contains run-task.sh, so the guard saw "workers still running" from 18:20Z on 09-12 until 11:16Z on 09-13, renewing the lease every 15 min. This is HAZARD #6 again, inside the guard itself. Worse than hazard #8: #8 fails SAFE (the sprite sleeps); this fails EXPENSIVE, and silently, because the guard log says "workers still running" either way.

Resolved by killing that one stale shell by PID. The guard then exited normally and wrote WAVE8-DONE and released the lease.

Fixed in tasks/wave-guard.sh (safe: nothing was running). workers_running() now matches argv SHAPE: argv[0] is bash, argv[1] IS run-task.sh/run-resume.sh, and argv[2] is one of THIS wave's task names. A bash -c has argv[1] = -c and cannot match. Tested with no worker, with a decoy bash -c containing ./tasks/run-task.sh J-fake, with a real-shaped detached worker (and with a non-matching task name), and after killing it. Only the real worker counts as running.

When checking a wave, compare the guard log with the worker's .state: a workers still running renewal timestamped AFTER .state says SUCCESS means the guard is stuck. It is not a slow worker.

Waves 9 + 10 — K-manual then L-search, CHAINED: launched 2026-09-13

The Wave-8 crash never recurred and the user has parked it. Two user-requested features, run IN SEQUENCE (they share ui/library, ui/nav) with NO orchestrator between them — the user asked for wave 10 to start straight after wave 9 without waiting on them.

  • Wave 9 tasks/K-manual.txt — add a book by hand with no ISBN. New ui/add/ screen + Routes.addBook(title, isbn), entry points on the library FAB/empty state and the scan screen's manual-ISBN dialog. Before this there was NO way to add a book with no barcode (saveManualEntry required an ISBN).
  • Wave 10 tasks/L-search.txt — library search bar also searches Open Library search.json + Google Books volumes?q=, on SUBMIT only; "In your library" on top; deduped across all three sources. User decisions recorded in the prompt: submit-only, "in your library" = owns SOME edition, ISBN filled only when unambiguous (OL work with exactly one distinct ISBN-13, else exactly one GB ISBN-13, else none). Live-verified API fact: OL search omits isbn unless fields= asks for it.

Runner: tasks/wave-chain.sh 9:K-manual 10:L-search (new). Holds the lease itself (wave-guard would release it during the between-wave gate), runs each worker in the foreground, then a MECHANICAL gate (assembleDebug, testDebugUnitTest, verifyPaparazziDebug, XML test count up, no build files, no always 'false'). Red -> one fresh <task>-fix worker -> re-gate -> still red = sentinel says FAILED and the chain STOPS. Green -> git add -A commit (message says "ORCHESTRATOR REVIEW STILL PENDING") + push, then next wave. Logs: logs/wave-chain.log, logs/gate-<task>.log, logs/WAVE9-DONE, logs/WAVE10-DONE, logs/CHAIN-DONE. Baseline: 205 tests, 2 skipped.

The gate is not acceptance. Still owed by the orchestrator for BOTH waves: read each .summary, read the diff, eyeball the new Paparazzi PNGs, check CancellationException is rethrown in every new catch, and check wave 10's dedupe judgement call in its report.

Waves 9 + 10 — REVIEWED 2026-09-16; accepted after a review-fix round

Chain finished 2026-09-15 18:28Z: b600df6 (wave 9, 241 tests), cf82cbc (wave 10, 308). Orchestrator independently re-ran assembleDebug / testDebugUnitTest / verifyPaparazziDebug on both — green — then read the diffs and PNGs. The gate was right that they built; it could not see these, all fixed directly by the orchestrator in the review-fix commit that follows:

  1. "In your library" never showed the merged set. The screen passed the plain text match (uiState.books) instead of OnlineSearchState.Done.inLibrary, so an owned edition reached only via an online ISBN was removed from "Online" and shown NOWHERE; wave 10's VM test asserted inLibrary but nothing rendered it. Now rendered; the shelf filter applies to it and hidden owned matches are counted ("1 more in your library, outside the current filter").
  2. "Edit details" left the save sheet open -> save from the add screen, Back, sheet still up with Save enabled -> duplicate book. LibraryViewModel.takeForEditing() closes it.
  3. rememberShelf inside the save try (add + online save): a failed prefs write after a successful insert showed "Couldn't save" -> retry -> duplicate. Now after, and swallowed.
  4. Per-keystroke search on the main thread: two regexes recompiled per call, every book re-normalized per keystroke (~60ms/keystroke at 2,000 books on this server, measured with jshell). Precompiled regexes, LocalBookMatcher.index rebuilt only when books change, flowOn(Dispatchers.Default).
  5. Add-by-hand duplicate check unguarded (a Room throw from a keystroke would crash). Guarded.
  6. Scan screen save was unguarded (wave 8's known gap): now catch/rethrow-cancellation, in-flight guard, error line. performSaveManualEntry (post-lookup NotFound sheet) is STILL unguarded — its sheet has no error UI; left as is.

Cover preload — design decided WITH the user, 2026-09-16

User-reported: saving had a distinct delay even with the cover already visible on the sheet. Cause: createBook downloaded the cover (fresh OkHttpClient) BEFORE writing the row, and on failure silently saved with no localCoverPath. User rejected "save first, fetch cover in the background" (a kill/crash loses the upload; a failure has nowhere to be shown). Their design:

  • Preload (data/repo/CoverPreload.kt): the download starts when a sheet/screen gets a cover URL — scan Found sheet, online save sheet on open (ONE result, never the whole list — the user asked), add screen when "Edit details" carries a cover. Into cacheDir/cover-preload/.
  • Save waits for it if still running; createBook(coverFile=) COPIES it to filesDir/covers/<id>.jpg (copy, so a failed insert leaves the preload for a retry; the copy is deleted on insert failure).
  • Failure = Option A (user's choice): failure line + Retry on the sheet; Save becomes "Save without cover". A download that fails WHILE Save waits stops the save (the user hasn't seen it yet); one that had already failed when tapped proceeds without the file.
  • Cleanup: preload discarded on skip/dismiss/save/onCleared; BookshelfApplication sweeps files older than process start. NOT Coil's disk cache — public API, but Coil's eviction and keys aren't a contract; the user agreed it wasn't worth depending on.
  • Downloads use metadataHttpClient (never the PocketBase-token client).

Still open after review (not fixed — judgement calls or out of scope)

  • Wave 9 has NO worker report (worker hit its wall clock mid-verification, see hazard #11). Its undocumented calls: small EditNote FAB above Scan; no "More fields…" on the in-scan sheet; imePadding rather than the prompt's safeDrawingPadding — needs an on-device look.
  • Dedupe (wave 10's own analysis, agreed): false merges for same-author colon series ("Star Wars: X"/"Star Wars: Y") and same-title same-surname authors; one bad edge chains via union-find; misses translations, surname spelling variants, author-less books.
  • Retry re-runs BOTH sources (an OL failure spends a Google Books request).
  • Search toolbar still renders as a blank bar in Paparazzi (pre-existing), so no snapshot shows a query. Online result rows are a non-lazy Column (≤ ~40 thumbnails per search; accepted).

HAZARD #11 — the lease did NOT stop a freeze; the wall clock counts frozen time

Chain took the lease 2026-09-13 17:30 (expire 3600s). The sprite froze ~17:45 — BEFORE the first renewal, lease still valid — and stayed frozen ~26.5h until the user connected 09-14 20:21. No reboot (uptime continuous); processes paused (ps start times shifted by the gap). It froze again during K-manual-fix (09-14 20:37 -> 23:43 -> 09-15 11:27 log gaps). Consequences: (a) an unattended chain only progresses while someone is connected — do not promise the user overnight progress; (b) run-task.sh's 24h wall= counts the freeze, so K-manual was killed as "WALL CLOCK EXCEEDED" seconds after waking, mid-test-run, with no report. The 8 red tests the gate then saw were its own unfinished tests. If this recurs, measure elapsed time with the monotonic clock (/proc/uptime), which does not advance while frozen.

Also seen 2026-09-16: Claude Code's background-task runner killed Gradle runs "because the system is running low on memory" even with ~6GB available (8GB box, no swap; Kotlin daemon ~2.1GB, Gradle daemon ~1.2GB). Foreground runs of the same command succeeded.

Wave 11 — M-authexpiry: sync broken by an expired token — ACCEPTED 2026-09-17

User report: every sync on the phone said "Sync failed: HTTP 400". Server and public URL were both fine. PocketBase's request log (pb_data/auxiliary.db, _logs) showed every phone request with auth: "": POST books -> 400 "create rule failure", PATCH books -> 404. Root cause: user tokens lasted PocketBase's default 5 days, the app logged in once and never refreshed, and PocketBase does NOT reject a bad token; it silently serves the request anonymously. Last good sync was 09-13.

The worse half: data loss on the phone. SyncEngine.update* hard-deletes the local row on 404 (SPEC rule). The anonymous PATCHes 404'd on five books that still exist on the server, so they were deleted locally, and the pull cursor was already past them. Server copies intact (kclh1e51t7hmj0z, insggwycbrd6o48, 9tp7rumid3rzsn9, 9plhmyj5s68f58e, mle7o641pn7zzvt). Any unsynced local edits to those five on the phone are lost. Lesson: when you debug a server's 4xx, check the app's 4xx handling for side effects.

Fixes:

  • ab1d294 (orchestrator): migration sets user token duration to 180 days (the user chose it). Verified: fresh login token exp = 180.0 days; refresh also returns 180 days.
  • 009cb9a (worker + chain fix-worker, $4.89 + $1.34): refresh at the top of every sync; 401/403 -> SyncResult.AuthExpired (token cleared, URL/email kept, nothing pushed); library sync bar "Signed out — sign in again to sync" (tap -> setup); settings "Sign in" button; setup pre-fills URL + email; login clears pull cursors; one-time full re-pull (full_repull_2026_09_17_done flag) to restore the five books; cancellation rethrown in sync(). First gate RED: a pre-existing race in LibraryViewModelTest (cancelled search job not joined). The fix worker made the test join it. That was a test fix, not a logic change.

Orchestrator verification: read the full data-layer + VM diff; assembleDebug / testDebugUnitTest (355, 2 skipped, 0 fail; was 337) / verifyPaparazziDebug all exit 0; no "always 'false'"; eyeballed the 4 new PNGs (fine, light+dark). LiveSyncTest run against the PUBLIC URL (PB_URL=https://bookshelf-dev-b2jqx.sprites.app server/live-sync-test.sh): passed, not skipped; server log shows 4 auth-refresh calls, all 200, all authenticated. Anonymous LIST on all four collections still 403.

Recovery on the phone: install the new build. The first sync gets 401 on refresh and shows "Signed out". The user signs in (password only), which clears the cursors, and the next sync re-pulls everything, which brings back the five books. The queued PENDING_CREATE books from the failed syncs push normally. NOT verified on-device yet.

Known soft spots (accepted):

  • The 401/403 "belt and braces" catch mid-sync would NOT fire for a token that dies mid-sync. PocketBase would 404 anonymously, not 401. The window is one sync right after a successful refresh that issued a 180-day token, so it's negligible in practice. It would matter if tokens are ever revoked server-side (password change, tokenKey reset) during a sync.
  • The library sync bar is tappable, but nothing visual marks it as tappable. The label says what to do.
  • SetupViewModel pre-fill is async and could overwrite text the user types in the first few ms. Negligible.
  • One unnecessary full pull on fresh installs (the flag starts false). Harmless.
  • SettingsStore.resetFullRepullDoneForTesting is an internal test-only method.

HAZARD #12 — sprite storage degrades; restart fixes it

2026-09-17 ~01:30Z: testDebugUnitTest ran more than 20 min (normally ~75s). ComponentGalleryPaparazziTest took 763s and BookDaoTest 327s, and git was slow too. Nothing was hung: result XML was still being written, just very slowly. The user restarted the sprite and the same run took 74s. If builds or git crawl, suspect the box, not the code: compare against the gate logs' "BUILD SUCCESSFUL in" times and ask the user to restart.

tasks/wave-chain.sh now reads its commit subject from tasks/<task>.subject (it was hard-coded for waves 9/10).

Wave 12 — Open Library endpoint move + Google Books cover fix: DONE 2026-09-20

Done by the orchestrator directly, not a worker — the user left the choice open ("do whatever you think will yield the best results"). The change is confined to data/metadata, and the live API shapes had already been measured during diagnosis, so a worker would have spent a session rediscovering them. Full write-up in docs/METADATA-SOURCES.md § "The Open Library endpoint move".

Two user-reported symptoms, both third parties answering misleadingly:

  1. /api/books?bibkeys=... now 404s for EVERY ISBN, including books OL still holds. OL's docs call it the "Legacy Books API" that "may be phased out"; it is gone from their API index. Treat as retired, not down. The app had been effectively single-sourced on Google Books since it broke.
  2. Google Books answers zoom=2 with a grey "image not available" PNG at HTTP 200 for volumes with no full preview — 11 of 18 sampled. SPEC told us to force zoom=2.

Changed: OpenLibraryClient now uses /isbn/{isbn}.json (+ /works/ fallback for authors, + /authors/ resolution via the new AuthorNameCache); normalizeCoverUrl pins zoom=1, adds w=400, strips edge=curl. SPEC.md's "Book metadata lookup" section rewritten with both rules and WHY, so nobody restores either. Old-endpoint fixtures deleted; 9 new fixtures captured from the live API.

Three things that would have been silent bugs (all caught pre-ship, all pinned by tests): an edition record can carry NO authors (9780898707168 — the user's own shelf — has them only on the work, so a naive move drops the author); covers uses -1 as a no-cover sentinel; description is a bare string on some records and {type,value} on others. Also: on this endpoint 404 IS authoritative NotFound, where on the old one every non-2xx was a failure.

Check Result
./tasks/gw assembleDebug exit 0
./tasks/gw testDebugUnitTest --rerun-tasks exit 0 — 378 tests, 2 skipped, 0 failures (was 355)
./tasks/gw verifyPaparazziDebug exit 0 — no pixels moved
LIVE_METADATA=1 live test RAN, not skipped — 9/9 Found with cover art, keyless GB, so OL alone answered
grep "always 'false'" 0 hits

HAZARD #13 — two latent test races, surfaced by adding tests

Adding ~20 tests shifted suite timing and made two pre-existing races start flapping. Neither was caused by the metadata change; both flake in whichever test happens to be running when a window expires, NOT in the one at fault. Do not chase the named test.

  1. FIXED. LibraryViewModelTest leaked view models. LibraryViewModel has eleven stateIn(viewModelScope, WhileSubscribed(5_000), ...) flows, so its upstreams keep running five seconds after the last collector. Nothing ever cleared the VM, so they were still touching Dispatchers.Main while the NEXT test's tearDown called resetMain() -> "Dispatchers.Main is used concurrently with setting it". The test now tracks every VM it builds and cancels viewModelScope in tearDown (and guards db.close() with ::db.isInitialized, since a failed setUp otherwise masks the real failure with an UninitializedPropertyAccessException).
  2. NOT FIXED — reported to the user, out of scope. AddBookViewModel.performSave has a check-then-act in-flight guard: if (_formState.value.isSaving) return null and only then update { isSaving = true }. Two coroutines can both pass the check, which is what a second save while one is in flight does not create a second book catches when timing allows. The KDoc right above it claims a double-tap "still can't create two books" — that claim is false today. AddBookViewModel takes no metadata dependency at all, so this is provably unrelated to wave 12. One atomic getAndUpdate fixes it.

Lesson: a green suite on this project is worth one re-run before you trust it, and --rerun-tasks is mandatory — a plain testDebugUnitTest after a stash happily reports BUILD SUCCESSFUL FROM-CACHE without executing a single test. That nearly produced a false "pre-existing, not mine" conclusion here.