fix(ocr): render scanned PDF pages before embedded images - #2348
Elioooon (Elioooon) wants to merge 2 commits into
Conversation
|
@microsoft-github-policy-service agree |
72a4f5d to
728a982
Compare
c9ce9a1 to
085ba17
Compare
|
I found a recovery regression at 085ba17 when full-page rendering fails. For an empty-text page, the new branch catches I compared the parent and this head with the same controlled case: Could full-page OCR remain the preferred path, but fall back to embedded-image extraction when rendering fails? A regression comparing that failure path would preserve recovery without returning to the decorative-image-first behavior this PR fixes. |
|
Addressed in Full-page OCR remains the preferred path for empty-text pages. If page rendering fails, the converter now tries OCR on usable embedded images from that same page before returning the render error, so the existing recovery path is retained without restoring image-first behavior. The regression test injects a Validation: |
|
Thanks, verified 559933d. The render-failure regression now passes and preserves the embedded-image OCR block. I also ran the full OCR plugin suite after installing its test dependency: 111 passed on Python 3.12 (16 in test_pdf_converter.py). No remote OCR calls were used. I will drop the duplicate local fix for this finding. |
|
Independently reproduced this on real-world data (a Turkish customs-style scanned |
559933d to
59058d5
Compare
Summary
Fixes #2343.
This keeps mixed PDFs intact: text pages still interleave extracted text and embedded-image OCR, while scanned pages no longer lose their body content to small decorative images.
Verification
python -m pytest packages/markitdown-ocr/tests/test_pdf_converter.py -q -k 'scanned_page_uses_full_page_ocr or convert_reuses_open_page_for_image_extraction'- 2 passedpython -m pytest packages/markitdown-ocr/tests -q- 37 passed, 1 pre-existing failure:test_pdf_multipageassumes pdfplumber rejects the fixture, but the current dependency version extracts it successfully; reproduced unchanged onmainblack==23.7.0 --check packages/markitdown-ocr/src/markitdown_ocr/_pdf_converter_with_ocr.py packages/markitdown-ocr/tests/test_pdf_converter.py- passedgit diff --check- passedRisk
Scoped to the OCR plugin's PDF path. Text-only conversion and non-PDF converters are unchanged. A page is treated as scanned only when
page.extract_text()is empty or whitespace.