docs: record wave 8 (J-crashsafe) in flight

The user's first scan on the wave-7 build crashed and never reproduced. Their
theory fits structurally — the Google Books Found path had never executed in
production before the key landed — but nine live lookups through the real
repository, including both of their ISBNs, threw nothing, so the parse/merge
code is exonerated by evidence rather than by argument. Written down so no
future worker "fixes" code that was measured working.

What the wave fixes is the defect found while looking: the lookup path catches
only IOException and the ViewModel catches nothing, so any other throwable kills
the process instead of reaching the user as "couldn't be reached" — and the app
has no crash capture at all, which is why one crash left no evidence.

Committed with an explicit pathspec: a wave is in flight (HAZARD #7).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PPpdG8VnRfS3KkisR3HUAE
This commit is contained in:
Spriteandclaude committed 2026-09-12 17:51:13 +00:00
1 parent 0455794603
commit 1295d88d6e
2 files changed
+272

No files matched your search

+55
View File
@@ -627,3 +627,58 @@ because either alone would have prevented this:
Regression-checked against the real `logs/I-gbkey.json` that caused this: the old
logic matches the quota pattern, the new logic yields SUCCESS and feeds the
detector an empty blob.
## Wave 8 — J-crashsafe: IN FLIGHT, launched 2026-09-12 17:50Z
Prompt: `tasks/J-crashsafe.txt`. Session id in `logs/J-crashsafe.sid`. Lease
`bookshelf-wave` held by `tasks/wave-guard.sh`, sentinel `logs/WAVE8-DONE`.
**Why this wave exists.** The user installed the wave-7 build and the FIRST barcode
scanned crashed the app: ISBN decoded and displayed, spinner ran a few seconds, process
died. Never reproduced, on that book or any other. Their theory — that the scan fell
through to Google Books and crashed there — is structurally the best fit: keyless GB
429'd every caller on earth, so `GoogleBooksClient.classify`'s 2xx branch,
`toBookMetadata()`, `normalizeCoverUrl()` and the two-source merge had NEVER EXECUTED
in production before 486f6eb. A first crash belongs in a path's first real exercise.
**It does not reproduce off-device, and that is now evidence rather than a guess.**
`app/app/src/test/java/org/modg/bookshelf/livemetadata/LiveMetadataLookupTest.kt`
(commit 93546ed, opt-in on `LIVE_METADATA=1`) drives the real `MetadataRepository.lookup`
against both real APIs with the real key. Nine lookups — both of the user's
previously-failing ISBNs plus a control, three times each — all returned Found with
cover art, 254ms to 4.7s, nothing thrown. Static review found no unsafe operation in
that path either. **Do not let a future worker "fix" this by rewriting the parse/merge
code; that code was exercised live and is fine.**
**What the wave actually fixes** is the defect found while looking: nothing on that path
is exception-safe. Both clients' `fetch` catch only `IOException`, `classify` catches
only `SerializationException`/`IllegalArgumentException`, and `ScanViewModel.runLookup`
sits inside `viewModelScope.launch { ...collect { } }` with no try/catch at all. So any
throwable that is not an `IOException` — platform TLS, an OkHttp internal, memory
pressure, an API-level difference, none of which this JVM reproduces — kills the
process instead of surfacing as "couldn't be reached". Both clients' KDoc claims "Never
throws"; that claim is false today. Plus: the app has NO crash capture whatsoever,
which is why a one-time crash left nothing to work from.
Scope, per the user's decision: (1) both clients catch `Throwable` -> new
`FailureKind.UNEXPECTED`, non-retryable, reason names the exception class only (never
its message — that can carry the key); (2) `ScanViewModel.runLookup` guards the same
way, which also covers the Room call in it; (3) a `CrashReporter` that persists the
stack trace and chains to the previous handler; (4) a Diagnostics section in settings to
read and SHARE it off the phone. **`CancellationException` must be rethrown, not
converted, in every one of those catches** — it is normal control flow here (dismissing
the sheet cancels the lookup) and is the easiest thing in this wave to get wrong.
**Verify before accepting** (workers self-report optimistically; several waves have
over-claimed):
cd ~/bookshelf && ./tasks/gw assembleDebug && ./tasks/gw testDebugUnitTest \
&& ./tasks/gw verifyPaparazziDebug && git status --porcelain
**190 tests, 2 skipped today** (LiveSyncTest + LiveMetadataLookupTest, both opt-in
live). Count from the `TEST-*.xml` files, not the console. The count must go UP.
Grep the build log for `always 'false'`. Then eyeball the new settings Paparazzi PNGs.
Check specifically that `CancellationException` is rethrown in every new catch, and
that `CrashReporter` calls the previously-installed handler — a handler that does not
chain leaves the process hung instead of dying.
**Not in scope, deliberately:** finding the original throwing line. Nobody knows what it
was and the prompt says not to hunt for it. If the guard lands and the user ever sees
"unexpected: SomeException" on the scan sheet, THAT is when we learn the answer.