likhit is Jawafdehi's public MarkItDown plugin for Nepal-specific document support.
It extends MarkItDown with Nepal-specific PDF repair, layout-aware Markdown assembly, optional OCR fallback for image-dominant PDFs, and legacy .doc support. For PDFs, likhit now evaluates multiple extraction paths and returns the best result instead of relying on a single fixed pipeline.
Owned and maintained by Jawafdehi.
pip install likhit- Website: https://jawafdehi.org/
- GitHub: https://github.com/Jawafdehi/likhit/
- Contact: inquiry@jawafdehi.org
likhit is primarily used as a MarkItDown plugin.
Once installed, enable plugins when creating a MarkItDown instance:
from markitdown import MarkItDown
md = MarkItDown(enable_plugins=True)
result = md.convert("path/to/nepali-document.pdf")
print(result.text_content)You can also use likhit through the standard MarkItDown CLI:
markitdown --use-plugins path/to/nepali-document.pdfTo write the output to a file:
markitdown --use-plugins path/to/nepali-document.pdf -o output.mdTo verify the plugin is registered:
markitdown --list-pluginsYou should see likhit in the output.
This package also installs a small helper CLI that runs MarkItDown with the likhit plugin enabled and writes Markdown files for you:
likhit-save path/to/nepali-document.pdf --out output.mdConvert multiple files into a directory:
likhit-save samples/pressrelease.pdf samples/kanunpatrika.pdf --out-dir converted/Extract only one page or a page range from a PDF:
likhit-save path/to/nepali-document.pdf --pages 5 --out page-5.md
likhit-save path/to/nepali-document.pdf --pages 2-4 --out pages-2-4.mdlikhit adds behavior beyond MarkItDown in these places:
- PDF:
likhitintercepts PDF inputs, runs the default MarkItDown PDF converter first, and then decides whether to keep that result, retry with Nepal-specific extraction, or add an OCR candidate for image-dominant pages. It prefers directlikhitextraction immediately when known Nepali repair fonts are detected. - DOC: Legacy Microsoft Word
.docfiles are handled bylikhit's own extraction pipeline. - DOCX:
.docxfiles are still handled by MarkItDown's built-in Word converter, even when plugins are enabled.
- PDFs, including Nepal-specific born-digital PDFs and image-dominant PDFs that may need OCR
- Legacy
.docfiles .docxpassthrough via MarkItDown
For image-dominant or scanned PDFs, likhit can add an OCR extraction candidate through markitdown-ocr when OCR is configured.
Required model configuration:
export MARKITDOWN_OCR_MODEL="your-model-name"You can also provide the model through OPENAI_MODEL or GEMINI_MODEL.
Authentication options:
- OpenAI-compatible provider with a standard OpenAI key:
export OPENAI_API_KEY="your-api-key"- OpenAI-compatible provider with a custom base URL:
export OPENAI_API_KEY="your-api-key"
export OPENAI_BASE_URL="https://your-provider.example/v1/"
export MARKITDOWN_OCR_MODEL="your-model-name"- Gemini using the OpenAI compatibility endpoint:
export GEMINI_API_KEY="your-gemini-api-key"
export GEMINI_MODEL="gemini-2.5-flash"When GEMINI_API_KEY is set, likhit automatically uses Gemini's OpenAI-compatible base URL unless you explicitly override OPENAI_BASE_URL.
Optional variables:
export MARKITDOWN_OCR_PROMPT="Custom OCR instructions"The high-level PDF pipeline is:
- MarkItDown loads the plugin when
enable_plugins=Trueor--use-pluginsis used. - For PDF inputs,
likhitreads the file and optionally slices it to the requested page range. likhitscans embedded fonts. If it detects known Nepali repair fonts such as Kalimati broken-CMap fonts or legacy remap fonts, it tries the Nepal-specific extraction pipeline immediately.likhitalso runs the default MarkItDown PDF converter and keeps that result as a candidate.likhitanalyzes the PDF pages. If the file looks image-dominant with a suspicious text layer and OCR is configured, it adds an OCR candidate.- If the default Markdown output looks suspicious for Nepali text,
likhitretries extraction with its own PDF pipeline. - The Nepal-specific PDF pipeline can apply:
- Kalimati broken-CMap repair
- Devanagari reordering
- Devanagari spacing normalization
- Legacy-font remapping through
npttf2utf
- After extraction,
likhitchecks whether the document matches a whole-document semantic structure such as a single-column notice. - PDF layout ordering is assigned locally while assembling content blocks, so single-column, row-aligned, and two-column regions can coexist in one file.
- If multiple candidate outputs exist,
likhitscores them and returns the best one.
src/likhit/_plugin.py: MarkItDown plugin entry point and converter registrationsrc/likhit/converters/: plugin converters for PDF and legacy DOC inputssrc/likhit/nepali_pdf_repair.py: reusable Nepal-specific PDF repair layersrc/likhit/markdown_assembly.py: generic Markdown assembly for the default conversion pathsrc/likhit/extractors/: extraction strategies (PDF, DOC)font_based.py: PDF extraction with Nepali font repairdocx_based.py: legacy DOC text extraction
src/likhit/handlers/: structure-aware handlers and detection logicsrc/likhit/renderers/: Markdown renderingtests/: conversion, extraction, and plugin coveragetests/integration/: end-to-end integration teststests/integration/test_data/: committed test fixtures (PDF, DOCX, DOC samples)
likhit uses uv for dependency management, and the
rest of Astral's toolchain — ruff for linting and
formatting, ty for type checking.
Install the project together with its dev dependencies:
uv syncInstall the pre-commit hooks once per clone:
uv run pre-commit installRun all tests:
uv run pytestRun only the end-to-end integration suite:
uv run pytest tests/integrationCI runs the following. Run them yourself before opening a pull request:
uv run ruff check . # lint — gated
uv run ruff format --check . # formatting — gated
uv run pytest # tests — gated
uv run ty check # types — ADVISORY, see pyproject.tomlty is advisory rather than a gate: it is pre-1.0, and several of this package's
dependencies (PyMuPDF among them) ship no type information, so the rule families
it cannot yet see through are silenced in pyproject.toml. Read its output; do
not ignore it.
- MarkItDown: https://github.com/microsoft/markitdown
- MarkItDown sample plugin: https://github.com/microsoft/markitdown/tree/main/packages/markitdown-sample-plugin
Licensed under the Hippocratic License 3.0, an Ethical Source license. See LICENSING.md for details.
likhit is owned and maintained by Jawafdehi.
- Organization: Jawafdehi
- Website: https://jawafdehi.org/
- GitHub: https://github.com/Jawafdehi/likhit/
- Contact: inquiry@jawafdehi.org