Files
bookshelf/docs/METADATA-SOURCES.md
T
Spriteandclaude 486f6ebc48 Use the Google Books API key: the fallback source actually answers now
The user obtained a restricted Google Books key. Keyless requests 429 for every
caller on the internet — all anonymous traffic bills to one shared Google Cloud
project whose daily quota is permanently exhausted — so the documented fallback
has never once answered. Because MetadataRepository.combine turns "a source
failed, none found" into Unavailable, that standing failure meant every Open
Library hiccup reached the user as "couldn't be reached". The app has been
effectively single-sourced since it was written.

Build plumbing reads GOOGLE_BOOKS_API_KEY from local.properties (gitignored),
falling back to the environment and then to empty. A blank key is a supported
state: a fresh clone still builds a working app that falls back to the keyless
endpoint, rather than failing to build.

GoogleBooksClient appends the key only when non-blank, building the URL with
HttpUrl.Builder in a pure requestUrl() so it is testable without a socket. The
key is scrubbed from SourceResult.Failed.reason before that string can reach the
scan sheet — it is rendered to the user and is our only diagnostic channel from a
real phone, and some okhttp/JDK IOExceptions embed the full request URL in their
message. Defensive, not a response to an observed leak.

Resolves the RATE_LIMITED decision parked in RetryPolicy's KDoc: a keyed 429 is
the short per-user rate limit and gets exactly one retry, honouring Retry-After
capped at 2s. A keyless 429 is still the dead daily quota and is still never
retried.

Verified against the live API, not only offline: both ISBNs that failed on the
phone (9781883937386, 9781883937676) plus a control return HTTP 200, in the
percent-encoded URL shape HttpUrl actually produces. Both books are in Google
Books, so the restored fallback now covers precisely the Open Library TLS-reset
failure that broke those scans.

189 unit tests (was 172), 0 failures; Paparazzi unchanged; release APK builds.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01J7WHnTx2Cso4VV245WDAJY
2026-09-11 23:34:43 +00:00

18 KiB
Raw Blame History

Book metadata lookup — why scans miss, and what else we could ask

Research note, 2026-09-09. Written in response to "out of the 3 barcodes I've scanned, only 1 has been discovered properly."

Nothing in here has been implemented. SPEC's two-source design (Open Library primary, Google Books fallback) is unchanged. This is the evidence for deciding whether to change it.

Update, same day — the coverage hypothesis is dead. The user supplied the two ISBNs that failed: 9781883937386 (The Hittite Warrior) and 9781883937676 (Shadow Hawk), both Bethlehem Books. Both are fully present in Open Library — title, author, publisher, page count and cover art — and the app's own parser handles both real responses correctly (regression test: OpenLibraryClientTest.parses the real responses for the two books the app failed to identify). Whatever went wrong on the phone was upstream of the metadata sources. Everything below still holds as background, but do not act on "add a third source" until we know why a book Open Library has did not reach the lookup. See What actually failed.


Short version

Three separate defects were making lookups look far worse than the underlying data actually is, and all three are now fixed (commit 356f639). They are not the same problem as "this book isn't in the database":

  1. A cover that loaded fine still rendered as nothing, so a successful lookup looked like a failed one. That alone could account for the book you did find appearing broken.
  2. Open Library's cover URL was synthesized for every book whether or not art existed, and a missing cover comes back as a 200 with a 43-byte 1×1 transparent GIF — a "successful" load that paints nothing.
  3. Because that synthesized URL was never blank, the merge rule could never fall through to Google Books' thumbnail. The documented fallback was dead code for covers.

What is left is a real coverage question, and there the measurements point at one thing above all others: the Google Books fallback is probably not answering at all. Every keyless request from this machine returned HTTP 429, and the app turns any non-200 into null, which the UI presents as "No match found — enter the details by hand." A rate-limited lookup and a book that genuinely exists nowhere are, right now, indistinguishable to both you and me.

My recommendation is to fix the diagnosis before buying more data. Details in Recommendation.


What I measured

Sample. 60 ISBNs drawn from the Harvard Library catalog — deliberately a third party, so the sample doesn't presuppose the answer by coming from one of the two sources under test. Ten publishers, weighted toward the small Catholic and homeschool presses that a MODG family's shelf actually carries (Ignatius, TAN, Sophia Institute, Bethlehem Books, Baronius) alongside mainstream trade (Penguin, Random House, Scholastic, Crossway, Loyola).

Method. Direct HTTP against each API, one ISBN at a time, 1.2 s apart. A source "hits" only if it returns a usable title.

Results

Source Hit Miss Error Hit rate
Open Library Books API (what the app calls today) 53 4 3 88%
Google Books, keyless (the app's fallback) 0 0 60 × HTTP 429 0%
Harvard LibraryCloud 56 4 0 93% *
Open Library cover art exists for the ISBN 47 13 0 78%

* Harvard is where the sample came from, so its number is inflated by construction. It's here to show the API works and answers by ISBN-13, not as a fair comparison.

Two further observations from the same runs:

  • Concurrency is punished. The same 60 ISBNs run six-at-a-time dropped Open Library from 88% to 70%, entirely through transport errors. The app makes one request per scan, so this doesn't bite in normal use — but it does mean "Open Library missed" in a log is not proof the book is absent. By the end of this research my own IP was refused outright for a while.
  • 17% of successful Open Library lookups have no cover art at all (9 of 53). Even with everything working, roughly one book in six will legitimately show the placeholder. That is a data fact, not a bug, and it's worth knowing before you read a placeholder as a failure.

What this does not tell us

Worth saying plainly, because it bounds how much weight the numbers carry:

  • n = 60, and the sample comes from a research library. It under-represents recent mass-market paperbacks, reprints and print-on-demand editions — which is exactly where Open Library is thinnest. Real shelf coverage is probably below 88%.
  • Every request came from a datacenter IP. The Google Books 429 may partly be this host sharing a quota pool with other tenants; your phone, on a residential or mobile IP, may well get answers. That's precisely why the app needs to be able to tell us which it got.

    CORRECTED 2026-09-09 — this guess was WRONG. The user tested from a residential IP and got the same refusal, naming a shared project quota rather than an IP one. See "Measured again 2026-09-09" at the end of this file. Do not act on the sentence above.

  • I don't know which three ISBNs you scanned. If you still have the books to hand, those three numbers are worth more than another 60 sampled ones.

The options

A. Give Google Books an API key — DONE 2026-09-11, see the end of this file

Free, 1,000 requests/day, no billing account required. Turns the fallback from "silently 429" into a working source. Roughly a dozen lines: a key in local.properties → BuildConfig → &key= on the query.

The key ships inside the APK and can be extracted, so restrict it to the Books API in the Google Cloud console. At our volume, someone stealing it costs us nothing but the quota.

Effort: hours. Cost: free. Likely the single biggest win.

B. Tell the difference between "not found" and "couldn't ask"

Both clients collapse every non-200, timeout and parse failure into null, and MetadataRepository collapses that into "no match", and the UI writes "No match found." A book that's offline, rate-limited, or hit a 500 is reported to you as a book that does not exist.

Distinguishing these gets you a retry button instead of a manual-entry form, and gets me a real answer next time you say "it missed."

Effort: half a day. Cost: free. Do this regardless of what else we choose.

C. Retry with backoff

One retry on 429/5xx, a couple of seconds apart. Standing at a bookshelf, a two-second retry is invisible; a manual-entry form is not.

Effort: an hour. Cost: free.

D. Add a third free source

Only worth doing after A–C, when we can see what's actually still missing. Ranked by what I'd try first:

Source Key? Notes
Open Library search.json No Searches the whole OL index rather than the edition table the Books API reads. Cheapest possible fallback — same service, one more request, no new failure modes.
Harvard LibraryCloud No Demoted. Free, no registration, answers by ISBN-13 (?identifier=<isbn>), verified working — but it returned numFound: 0 for both of the user's real failing books. It is a research library: strong on scholarly and older material, and it simply does not hold small-press children's historical fiction. The 93% in the table above is inflated by construction (the sample came from Harvard) and is misleading in exactly the direction that matters. Not a fit for this shelf.
Library of Congress No The SRU endpoint (port 210) is blocked from here; the loc.gov JSON API responds. Excellent for US imprints. Needs more probing before I'd commit.
K10plus SRU No Free German-led union catalogue, large and international. Cataloguing conventions differ enough that merging would need care.

Dead ends, so nobody re-investigates them: OCLC Classify (retired 2021), Goodreads API (retired 2020), Amazon Product Advertising API (requires an affiliate account with qualifying sales), WorldCat Search (requires OCLC membership — institutional pricing).

E. Pay for ISBNdb

~$15–50/month depending on tier. Genuinely better coverage than anything free, including cover art, and a single clean API. It is also a subscription for a two-person home library, and I'd want proof that A–D leave a real gap before recommending it.

Effort: hours. Cost: $180–600/year.


What actually failed

Both failing ISBNs resolve cleanly against the source the app already uses:

Check 9781883937386 9781883937676
Open Library Books API Hittite warrior, Joanne S. Williamson Shadow hawk, Andre Norton
Cover art yes yes
App's own parser (unit test) parses parses
ISBN-13 checksum valid valid
Harvard not held not held

So the request either never went out, or went out and failed in a way the app reported as "not found". Candidates, in the order I'd chase them:

  1. The barcode never decoded into a valid ISBN-13. ScanCodeFilter drops anything that fails the checksum, silently and with no UI feedback at all — no sheet, no message, nothing. A book whose barcode carries a price add-on, or is worn, or is a UPC-A rather than a Bookland EAN, looks to the user exactly like a camera that isn't working. This is my leading theory, and it fits "wasn't able to scan" better than "wasn't found".
  2. The HTTP request failed. Both clients turn every non-200, timeout and parse error into null, which reaches the user as "No match found". The metadata OkHttpClient is constructed with no call timeout, so a stalled connection hangs on default socket timeouts.
  3. The lookup ran and the sheet was dismissed before it landed. While Loading, the sheet passes an empty onDismissRequest, so it can't be swiped away — but onScanned early-returns whenever a sheet is already showing, so a stuck sheet blocks every subsequent scan.

Nothing here is a data-source problem. Note the sting in (1) and (2): both failure modes are invisible or actively misleading, which is why three scans produced no usable diagnosis.

Recommendation

Revised after the two real ISBNs came in. Adding sources is now the wrong next move: the books that failed are already in the source we query.

Do B first and on its own — make the app say what happened. A scan that decodes nothing should say so on the camera screen; a lookup that fails should offer retry, not a manual-entry form captioned "No match found". Add a call timeout while in there.

Then rescan those two books. The app will tell us which of the three candidates above it is, and that determines everything after it. A (the free Google Books key) and C (retry/backoff) are still worth doing — cheap, and the 429 result is real — but they are no longer the leading theory.

D and E are on hold. Harvard specifically is off the list for this shelf. Paying ISBNdb for coverage we demonstrably already have would be the wrong order.

One thing worth deciding separately: 17% of books legitimately have no cover art anywhere. The placeholder now looks deliberate rather than broken, but if you want covers on everything, that's a different feature — photograph the book, store it as the cover — and not a metadata-source problem at all.

Measured again 2026-09-09, after the user tested from a residential IP

Two things were settled that the first round could only guess at.

Google Books keyless is dead everywhere, not just from this box

The user ran the app's exact Google Books call from their home connection and got:

Quota exceeded for quota metric 'Queries' and limit 'Queries per day'
of service 'books.googleapis.com' for consumer 'project_number:624717413613'

That names a Google Cloud project, not an IP. Every keyless caller on the internet is billed to that one shared anonymous project and its daily quota is exhausted. So:

  • The earlier caveat — "your phone, on a residential or mobile IP, may well get answers" — is WRONG. Delete it from your mental model. It was tested and it is not true.
  • Backoff cannot help. This is a daily quota, not a per-second rate limit.
  • A free API key is the only fix, and it is a complete one: it moves the app into its own project with its own quota (free tier 1,000 req/day).

The knock-on is the part that actually hurt the user. MetadataRepository.combine turns "any source Failed, none Found" into Unavailable. Google Books is a PERMANENT standing failure, so every Open Library hiccup became Unavailable. The app has effectively been single-sourced this whole time while reporting failures as though two sources had been consulted.

The user deferred the API key at the time; they asked for it on 2026-09-11 and it is now implemented and verified. See "The key landed" at the end of this file.

Open Library: 13% failure, and our own timeout was manufacturing more

30 requests, the exact call OpenLibraryClient makes, 1.5s apart, from the sprite:

failure rate 13% (4 of 30)
every failure curl exit 35 — TLS-stage Connection reset by peer
failure latency 0.23s, 0.31s, 0.59s, 2.46s — all fast
success latency median 4.3s, p75 6.0s, p90 9.2s, max 22.0s
successes over the old 12s callTimeout 2 of 26 (8%)

Two conclusions, and they point in opposite directions:

  1. Failures are cheap and transient, so retrying is nearly free. 13% → ~1.7% at two attempts → ~0.2% at three. This is why RetryPolicy exists and why its backoff is milliseconds rather than the conventional seconds.
  2. Successes are slow and long-tailed, and the app's own callTimeout(12s)/connectTimeout(10s) were cutting off roughly 8% of lookups that were about to succeed — then reporting them to the user as "couldn't be reached". The app was generating a meaningful share of its own failures. Timeouts are now 25s/20s/20s, above the 22.0s worst observed success.

The asymmetry is the whole design: a short timeout buys nothing on the failure path (failures return in under 2.5s regardless) and costs real successes on the slow path. That is also why RetryPolicy.isRetryable refuses to repeat a TIMEOUT — a timeout means the budget was already spent, and the evidence says slow requests mostly succeed if you let them finish.

Connection reuse is visible in the data and matters in real use: cold connects ran 2-19s while warm ones ran 0.07s. OkHttp pools connections for 5 minutes, so scanning a box of books in sequence stays on the fast path after the first book.

What still is not known

  • All 30 requests came from this datacenter IP. The user's phone may see a different failure rate. The reason string now shown on the scan sheet (SourceResult.Failed.reason, e.g. "tls connection reset, 3 attempts") is how we find out — it is the only diagnostic channel we have from a real device.
  • The user's own 3-request sample showed 2 failures. That is consistent with 13% (p ~ 5%) but does not confirm it. If their phone reports "3 attempts" often, their network is worse than this one and the retry count deserves revisiting.

The key landed — 2026-09-11

The user obtained a restricted Google Books API key and it is wired in (commit below). This closes option A, which had been the single biggest outstanding win in this file.

Verified live, not just in tests

The key was tested from this machine against the real API before any code was written, and again afterwards in the exact URL shape the app now builds:

ISBN Result
9781883937386 Hittite Warrior HTTP 200, 1 item
9781883937676 Shadow Hawk HTTP 200, 1 item
9780140449136 Crime and Punishment (control) HTTP 200, 1 item

Both of the books that failed on the user's phone are in Google Books. They were always in Open Library too — what actually failed was the TLS-stage connection reset documented above, not coverage. The significance is that the fallback would now cover exactly that failure mode: an Open Library transport error no longer leaves the lookup with nothing to fall back to. The app has been effectively single-sourced since it was written; it is now genuinely two-sourced.

Note the second URL test was not redundant. HttpUrl.Builder percent-encodes the colon, so the app sends q=isbn%3A9781883937386 where every hand-run test in this file sent q=isbn:9781883937386. Offline unit tests cannot tell those apart and the API accepts both — but that is a fact worth having measured rather than assumed, because the failure mode would have been a feature that passes every test and returns nothing on the phone.

What was built

  • The key lives in app/local.properties (gitignored) as GOOGLE_BOOKS_API_KEY, read by app/app/build.gradle.kts into BuildConfig.GOOGLE_BOOKS_API_KEY, falling back to an environment variable of the same name and then to empty. A blank key is a supported state — a fresh clone builds a working app that falls back to the keyless (429ing) endpoint rather than failing to build.
  • GoogleBooksClient takes the key and appends it only when non-blank, with URL construction in a pure requestUrl() so it is testable with no socket.
  • The key is scrubbed out of SourceResult.Failed.reason before it can reach the scan sheet. That string is rendered to the user and is our only diagnostic channel from a real phone; some okhttp/JDK IOExceptions embed the full request URL in their message, so the scrub is defensive rather than a response to an observed leak. Do not remove it on the grounds that nothing currently leaks.
  • The RetryPolicy RATE_LIMITED decision parked in its KDoc is now resolved: a keyed 429 is the short per-user rate limit and gets exactly one retry, honouring Retry-After capped at 2s; a keyless 429 is still the dead daily quota and is still never retried.

What this does NOT fix

The 17% of books with no cover art anywhere is unchanged, and so is Open Library's ~13% transport failure rate. What changes is that a failure of one source is now much more likely to be covered by the other instead of surfacing as Unavailable.