Forked from microsoft/markitdown.
This fork extends the upstream project with better handling of images and PDFs, plus a vision-LLM-based PDF pipeline. Changes added on top of upstream:
-
Keep original images in the output Markdown. Images embedded in the source document are saved out and referenced from the generated Markdown instead of being dropped, via a shared image sink. Applies to both the EPUB and PDF converters.
-
EPUB conversion improvements. Embedded images are extracted with correct image types and emitted into the Markdown output.
-
PDF page ranges. Convert a subset of a PDF by passing start and end page.
-
Vector-image detection in PDFs. Detect vector graphics (charts, diagrams) on a page so they are captured rather than missed by text extraction.
-
Whole-page vision-LLM PDF converter (
PdfConverterLLMFullPage). Transcribes each PDF page as a full-page image through a vision LLM. The model emits Markdown with image placeholders carrying pixel bounding boxes; the converter crops those regions out of the rendered page and rewrites the placeholders to point at the saved crops. Designed for PDFs whose layout (multi-column, vector charts, mixed figures) defeats coordinate-ordered text extraction. Includes optimized page-image resolution and uses the preceding page's Markdown as context to keep formatting consistent across page breaks. -
glue_pages.pyscript. Stitches page-broken Markdown back into a continuous document: rejoins paragraphs and sentences split across page breaks, relocates figure/quote/sidebar blocks that bisected a paragraph, and uses the original PDF text (via pdfplumber) as a reference to correct obvious OCR errors. Processing is sequential and resumable via a sidecar JSON cache.