Skip to content

fix(extraction): report pages the OCR cap left out, and show scanned books - #1004

Open
bloosqr wants to merge 2 commits into
Drakonis96:mainfrom
bloosqr:pr/ocr-cap-notes
Open

bloosqr wants to merge 2 commits into
Drakonis96:mainfrom
bloosqr:pr/ocr-cap-notes

Conversation

@bloosqr

@bloosqr bloosqr commented Sep 30, 2026

Copy link
Copy Markdown
Contributor

A scanned book longer than the OCR page limit silently lost its tail, and the note said the missing pages had no text.

What broke

extractPdf OCRs at most ocrMaxPages of the blank and low-quality pages ([...blanks, ...lowQuality].slice(0, maxPages)) and reports every blank page it did not recover as "N página(s) sin texto omitidas". Pages beyond the cap were never OCR'd, but read as blank. Found on a real library: a fully scanned 1,342-page textbook with the cap at 1,000 carried "1000 página(s) recuperadas por OCR. 342 página(s) sin texto omitidas." — a quarter of the book absent from search, reported as empty pages. Nothing in the interface showed that a book was scanned at all; the notes were only stored.

Fix

  • Extraction notes (textExtractor.ts): pages left out by the cap are counted separately — "N página(s) no procesadas: superan el límite de OCR (M páginas por documento)." — and "sin texto omitidas" counts only pages OCR tried and could not read.
  • shared/textProvenance.ts: reads the notes back — OCR pages, capped pages and the cap, blank pages — and recognises notes written before this change (OCR stopping at a round figure ≥ 300 with pages left over is the cap). isScannedWork marks a work read mostly by OCR.
  • Work status window (WorkStatusModal.tsx, the "citable" row): "scanned: N pages read by OCR" and "N pages not processed because of the OCR page limit (M pages): raise it in Settings and re-extract the text". Eleven languages (src/i18n.textProvenance.ts).

Testing

  • scripts/test-text-provenance.mjs (new): current and old-format notes (including several attachments' notes in one), the cap-cut case, scanned vs a digital book with a few OCR pages.
  • test-i18n-coverage.mjs, test-i18n-no-duplicate-keys.mjs pass.
  • The same rules applied to a real library's notes: 177 of 204 PDFs digital, 25 digital with a few OCR pages, 2 mostly scanned — one of them the capped textbook above.
  • Full local CI on Node 22 (citation:check, npm ci, lint, build, build:server-web, test:ci, all four e2e suites) — all pass. (e2e-smoke timed out once at the study "improve" dialog, which this change does not touch, and passed on a rerun; e2e-argument-map needs an unlocked screen for native fullscreen — it fails on v5.7.2 too while the session is locked — and passes with the screen unlocked.)

🤖 Generated with Claude Code

bloosqr and others added 2 commits September 29, 2026 18:22
McMurry's Organic Chemistry 7e (1,342 scanned pages) lost its last 342 to the
1,000-page OCR cap, and the note reported them as blank pages. The note now says
how many pages were not processed because of the cap, and the cap itself.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0138YUCVRCxHiRyY1TFpmufj
The work status window's text row now says how many pages were read by OCR and
how many the OCR page cap left unprocessed (with the cap and what to do), read
back from the extraction notes by shared/textProvenance. Older notes, where OCR
stopped at a round cap and the rest was reported as blank, are recognised too
(McMurry 7e: 1,000 OCR pages, 342 not processed). 11 languages.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0138YUCVRCxHiRyY1TFpmufj

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant