Committed by tasks/wave-chain.sh after its mechanical gate passed
(assembleDebug, testDebugUnitTest = 355 tests, verifyPaparazziDebug, no build
files touched, no "always 'false'"). ORCHESTRATOR REVIEW STILL PENDING.
Prompt: tasks/M-authexpiry.txt. Worker: $4.886013000000001, 123 turns.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
wave-chain.sh: commit subject now comes from tasks/<task>.subject instead
of being hard-coded for waves 9/10; drop a stale session trailer.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Committed by tasks/wave-chain.sh after its mechanical gate passed
(assembleDebug, testDebugUnitTest = 241 tests, verifyPaparazziDebug, no build
files touched, no "always 'false'"). ORCHESTRATOR REVIEW STILL PENDING.
Prompt: tasks/K-manual.txt. Worker: $5.051114400000003, 108 turns.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RWoivrRUEmJLFFwbsrGkqQ
The guard's `pgrep -f 'run-task\.sh'` matched the orchestrator's own stale
launching shell, whose command line contains that text, so it kept the sprite
awake for 17 hours after J-crashsafe finished. workers_running() now requires
argv[1] to be the runner script and argv[2] to be one of this wave's tasks.
Also commits the wave-7 sid/sentinel that were never added.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AQb3d68LqWD4EdND3C8jkD
The user's first scan on the wave-7 build crashed and never reproduced. Their
theory fits structurally — the Google Books Found path had never executed in
production before the key landed — but nine live lookups through the real
repository, including both of their ISBNs, threw nothing, so the parse/merge
code is exonerated by evidence rather than by argument. Written down so no
future worker "fixes" code that was measured working.
What the wave fixes is the defect found while looking: the lookup path catches
only IOException and the ViewModel catches nothing, so any other throwable kills
the process instead of reaching the user as "couldn't be reached" — and the app
has no crash capture at all, which is why one crash left no evidence.
Committed with an explicit pathspec: a wave is in flight (HAZARD #7).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PPpdG8VnRfS3KkisR3HUAE
The quota detector grepped `cat "$LOG" "$ERR"` for strings like "429" and "usage
limit". $LOG holds the worker's final JSON, including its .result prose — so a
worker whose TASK was about HTTP 429 finished successfully, wrote a report saying
so, and the runner read its own worker's words, logged "QUOTA hit", slept 600s
and was about to --resume a session that had already succeeded. That would have
burned quota redoing finished work with a fresh worker loose on a completed tree.
Two independent defences, because either alone suffices:
- the blob is now stderr plus the result text ONLY when is_error is true,
falling back to the whole log when it isn't valid JSON (a hard crash writes
no JSON, and has no prose to be confused by).
- the success check runs BEFORE the quota and session-vanished checks. A
finished worker is finished regardless of what strings its output contains.
Regression-checked against the real logs/I-gbkey.json that triggered this: the
old logic matches the quota pattern, the new logic yields SUCCESS.
Recorded as HAZARD #9 in docs/HANDOFF.md.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The user obtained a restricted Google Books key. Keyless requests 429 for every
caller on the internet — all anonymous traffic bills to one shared Google Cloud
project whose daily quota is permanently exhausted — so the documented fallback
has never once answered. Because MetadataRepository.combine turns "a source
failed, none found" into Unavailable, that standing failure meant every Open
Library hiccup reached the user as "couldn't be reached". The app has been
effectively single-sourced since it was written.
Build plumbing reads GOOGLE_BOOKS_API_KEY from local.properties (gitignored),
falling back to the environment and then to empty. A blank key is a supported
state: a fresh clone still builds a working app that falls back to the keyless
endpoint, rather than failing to build.
GoogleBooksClient appends the key only when non-blank, building the URL with
HttpUrl.Builder in a pure requestUrl() so it is testable without a socket. The
key is scrubbed from SourceResult.Failed.reason before that string can reach the
scan sheet — it is rendered to the user and is our only diagnostic channel from a
real phone, and some okhttp/JDK IOExceptions embed the full request URL in their
message. Defensive, not a response to an observed leak.
Resolves the RATE_LIMITED decision parked in RetryPolicy's KDoc: a keyed 429 is
the short per-user rate limit and gets exactly one retry, honouring Retry-After
capped at 2s. A keyless 429 is still the dead daily quota and is still never
retried.
Verified against the live API, not only offline: both ISBNs that failed on the
phone (9781883937386, 9781883937676) plus a control return HTTP 200, in the
percent-encoded URL shape HttpUrl actually produces. Both books are in Google
Books, so the restored fallback now covers precisely the Open Library TLS-reset
failure that broke those scans.
189 unit tests (was 172), 0 failures; Paparazzi unchanged; release APK builds.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01J7WHnTx2Cso4VV245WDAJY
The ghost bookcase was an inset bug, not a data bug. LocationsScreen's list
branch dropped the Scaffold's innerPadding while its empty-state branch applied
it, so the first bookcase row rendered under the top app bar and was invisible.
Both of the user's bookcases were always real; they just could not see the first
one, so they made a second. Every other screen was checked for the same class of
bug — Locations was the only one.
Metadata lookup is now designed against a measurement rather than a guess
(docs/METADATA-SOURCES.md § "Measured again 2026-09-09"):
- Google Books keyless is dead for everyone, permanently. The user's
residential-IP test returned a quota error naming a shared anonymous PROJECT,
not an IP, so the earlier "your phone may well get answers" guess is wrong and
is now marked CORRECTED in place. Because combine() turns any Failed-with-no-
Found into Unavailable, that standing failure meant every Open Library hiccup
surfaced as "one or more sources couldn't be reached". The API key is
deliberately deferred by the user; this commit leaves the source broken.
- Our own timeouts were manufacturing failures. Over 30 live requests, 13%
failed — all fast TLS resets under 2.5s — while successes ran to a median of
4.3s and a max of 22.0s. Two of 26 successes exceeded the old 12s callTimeout,
so ~8% of lookups that were about to work were cancelled and reported as
unreachable. Timeouts are now 25s/20s/20s.
That asymmetry (cheap failures, expensive successes) is what RetryPolicy encodes.
It retries TRANSPORT and SERVER_ERROR with a 250ms/750ms jittered backoff, and
deliberately does not retry TIMEOUT (the budget is already spent) or RATE_LIMITED
(hammering a quota is how an intermittent block becomes a permanent one — this
project's IP has already been refused outright once during research).
SourceResult.Failed now carries a FailureKind alongside its human reason, and the
reason names the specific failure ("tls connection reset, 3 attempts") instead of
a generic "network error". That string was already threaded to the UI and dropped
on the floor; LookupFailedSheet now renders it. It is the only diagnostic channel
we have from a real phone, so nothing may parse it.
Also from the same feedback round:
- Grouped ModalBottomSheet shelf picker, replacing two near-duplicate flat
dropdowns that listed every bookcase x shelf pair. Sections per bookcase,
empty bookcases say so, and the most recently used shelf is pinned on top.
- The recent shelf persists across sessions (SettingsStore.LAST_SHELF_ID) and is
cleared on sign-out. It is offered, never pre-selected: the user weighed that
and chose one tap over the risk of silently mis-shelving a book.
- Locations dialogs and the manual-ISBN dialog auto-focus their first field.
- The library filter menu offers "Add a bookcase to enable filtering" instead of
a lone "All books" that is already the active state and cannot be changed.
- The Locations button is Material Symbols' "shelves" (a bookcase) instead of
Warehouse (a barn). material-icons-extended 1.7.8 has no bookcase glyph.
- The scan sheet drops "you can lower the book" — the ISBN echo already says it.
assembleDebug exit 0; testDebugUnitTest 172 tests, 1 skipped, 0 failures (was
138); verifyPaparazziDebug exit 0; assembleRelease exit 0, signed with the real
release key; zero "always 'false'" warnings on a --rerun-tasks rebuild.
Three soft spots are recorded in docs/HANDOFF.md and are NOT verified: the
ghost-bookcase Paparazzi snapshot renders a lookalike of the screen rather than
the screen, the auto-focus calls swallow their own failure and no emulator exists
here, and the picker opens as a sheet stacked on the save sheet.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LSnVqWdiQNEcPFRq1hGZAi
SPEC now states that a source which could not be reached must never be
reported as a book that does not exist: Found / NotFound / Unavailable,
where NotFound requires every source to have answered authoritatively.
Also records that a rejected barcode must not be silent, and that the
by-ISBN cover URL is not evidence a cover exists.
run-task.sh exports CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS=0, per the wave-4
post-mortem: `claude -p` otherwise kills background tasks at 600s, so a
worker that backgrounds a Gradle build can never report on it. Safe to
edit now — no workers are running (hazard: never edit it while they are).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CDcPottghJXEvfYKqFM7zf
setsid nohup did not survive orchestrator teardown on 09-06 (uptime was
continuous, so this was a teardown kill, not the hazard-#5 suspend).
tasks/service-worker.sh runs a worker under the sprite service supervisor
instead, which both outlives the orchestrator and holds the box awake, making
the task-lease guard redundant. It is sentinel-guarded so a supervisor restart
does not re-run a finished wave, and it stops its own service afterwards so the
sprite can suspend.
logs/WAVE4-DONE records that F3's files were all complete but the worker was
stuck re-running verification it could not finish, so the orchestrator ran the
verification itself and accepted the work. F3's own written report — including
the design critique of the rendered screens it was asked for — was never
produced; that gap is recorded in the sentinel.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014TyzeWmdTqi7U85iYNGy7P
wave-guard.sh post-mortem: lease_hold() did DELETE-then-POST with both results
discarded and logged "lease renewed" unconditionally, so a failed re-POST left
the box with no lease while the log claimed it was protected. That matches the
wave-4 loss exactly (last "renewal" 11:39, workers dead 11:46, reboot 12:55).
Now: every acquire is verified against GET /v1/tasks before it is believed, a
failed acquire retries and is logged as a failure, and the lease is re-checked
every POLL rather than only every RENEW so a lease lost between renewals is
caught in seconds.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016mTs3kQXQsQwonXpEq7aEw
Guard now polls every 30s but renews every 15 min. Coupling them meant a wave that
finished just after a renewal sat undetected for a full interval with the sprite
pinned hot.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The wave-3 loss was caused by sprite auto-suspend, not nohup process-group
semantics. /.sprite/llm.txt: 'When idle, sprites pause automatically. Services and
sessions keep sprites alive.' Detached processes are on neither list, so setsid is
necessary but not sufficient.
tasks/wave-guard.sh holds a /v1/tasks lease (max 3600s, renewal is DELETE+POST since
re-POST returns 409), renews every 15 min while workers run, writes logs/WAVE<N>-DONE,
then releases the lease so the sprite can suspend rather than idle hot.
Also documents hazard #6 (pgrep -f / pkill -f matching the orchestrator's own shell)
and adds the wave-3 worker prompts and shared screen contract.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>