Plan v4: PDF → Markdown extraction via docling-serve (as implemented)¶
Status: Superseded by plan-v5.md (scanned PDFs no longer split; timeout per page). Supersedes plan-v3.md. Date: 2026-10-08
Changes from v3¶
New requirements came up during implementation, and testing on real documents forced some design changes.
| Change | Reason |
|---|---|
E13 – one log per PDF (logs/<stem>.log, replaced each time that PDF is processed) instead of one cumulative logs/extract.log |
Requested: the log should cover the file being extracted |
E14 – failed PDFs move to errors/, and the run continues |
Requested. A server outage is the exception: the run stops and leaves the files in input/ |
| E15 – 6–7 GB RAM, CPU only | Requested for cost. The container now uses the docling-serve-cpu image with a 6.5 GB limit and 2 workers × 2 threads |
pdf_backend=pypdfium2 instead of docling's default docling_parse |
With docling_parse, a single page of the user's JBIG2-scanned PDF filled 6.35 GB in about 5 minutes and crashed the container. With pypdfium2, the same page took 19 s and peaked at 2.6 GB |
OCR is forced automatically on scanned PDFs (--ocr auto) |
The scanner's embedded text layer ran words together. docling's own OCR produced clean text and was faster (18 s compared with 35–47 s for 5 pages) |
--chunk-pages default lowered from 10 to 5 |
Smaller jobs need less memory at their peak |
--max-inflight default lowered from 3 to 2 |
Matches the 2 workers on the server |
--mem-high default lowered from 0.85 to 0.75 |
The idle container with models loaded uses about 2.6 GB, and each job adds about 1 GB |
| Status checks tolerate a busy server. A timeout means "keep waiting". Pages are resent only if the container's restart count changed | In v3, timeouts caused duplicate submissions, which doubled the load and led to an out-of-memory crash |
1. Expectations¶
| ID | Expectation | Where it's met |
|---|---|---|
| E1 | Minimal Python script | One file, stdlib + requests + pypdf |
| E2 | Read PDFs from input/ |
plan_docs() |
| E3 | Documents are taken one at a time, in order | The chunk queue is ordered by document |
| E4 | Parallel processing | ThreadPoolExecutor(--max-inflight) |
| E5 | Conversion uses the docling container | /v1/convert/file/async |
| E6 | The container keeps running; the script never starts or stops it | Only docker stats and docker inspect are called |
| E7 | Markdown stored in JSON | outputs/<stem>.json |
| E8 | Output written to outputs/ |
--output |
| E9 | Finished PDF moves to completed/ |
finish_doc(), after the JSON is written |
| E10 | Monitoring: work in progress, time taken, page count | SUBMIT, PROGRESS, CHUNK_DONE and DOC_DONE lines |
| E11 | Page-range splitting when it helps | --chunk-pages, benchmarked in A16 |
| E12 | Concurrency limited by container memory | MemoryMonitor plus the --mem-high throttle |
| E13 | Log per extracted file, not cumulative | doc_logger() → logs/<stem>.log (mode w) |
| E14 | Failed PDFs move to errors/; other files continue |
fail_doc() |
| E15 | Runs in 6–7 GB RAM, CPU only | docker-compose.yml: docling-serve-cpu, memory: 6500M |
2. Design summary¶
- Input queue:
input/*.pdf, sorted. Each PDF's page count and scanned-page share are read withpypdf. A page counts as scanned if it holds an image at least as wide as the page. - Chunks:
--chunk-pages(default 5) consecutive pages each, sent withpage_range=[start, end],pdf_backend=pypdfium2, andforce_ocr=truewhen at least 50% of the pages are scanned. - Dispatch: at most
--max-inflight(2) chunks run at once. While container memory is at or above--mem-high(75%), nothing new is submitted, unless nothing is running at all. - Per chunk: submit → poll every 3 s (a timeout means the server is busy, so keep waiting) → fetch the result.
- Container restart: detected through
docker inspectRestartCount or a 404 for a task the server no longer knows. The script waits up to 10 minutes for/health, then resends the chunk, up to--retries(2) times. - Success: join the chunks' Markdown in page order → write
outputs/<stem>.json(temp file, then rename) → move the PDF tocompleted/. - Document failure (unreadable PDF, docling error, retries used up): move the PDF to
errors/. The reason is inlogs/<stem>.log. - Server failure (no
/healthafter a restart):ABORTEDis logged, unfinished PDFs stay ininput/, and the script exits with code 2. - Exit codes:
0all succeeded ·1some PDFs moved toerrors/·2server unreachable or down.
3. Acceptance criteria¶
The checks are the same as in v3, plus A22–A24 below. Results are recorded in the README.
| ID | Covers | Check | Pass condition |
|---|---|---|---|
| A9 | E14 | Corrupt bad.pdf in input/ |
Moved to errors/; DOC_FAILED in logs/bad.log; other PDFs succeed; exit code 1 |
| A12 | E14 | After A9 | input/ is empty; bad.pdf is in errors/ |
| A14 | E13 | ls logs/ after a run |
One <stem>.log per PDF, each containing only that PDF's lines |
| A22 | E15 | docker inspect restart count before and after the full run |
Unchanged; memory stays under 6.5 GB |
| A23 | OCR auto | DOC_START line in each log |
force_ocr=True only for the scanned PDF |
| A24 | Restart | docker restart mid-run (A21) |
Chunks are resent; PDFs end up in completed/ |