Skip to content
View Xia-gx's full-sized avatar

Block or report Xia-gx

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.

Python 87,697 11,176 Updated Jul 22, 2026

EDSL code

Python 19 2 Updated Mar 19, 2022

pix2tex: Using a ViT to convert images of equations into LaTeX code.

Python 16,540 1,312 Updated Jan 18, 2025

PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.

Python 10,478 779 Updated Aug 13, 2026

Open-source code for RFCNLP paper.

Promela 56 10 Updated Nov 9, 2022

Implementation of the paper "MAZE: Data-Free Model Stealing Attack Using Zeroth-Order Gradient Estimation".

Python 31 6 Updated Dec 12, 2021

Xidian University TeX Suite 西安电子科技大学LaTeX套装

TeX 1,133 100 Updated May 4, 2025

A Unified Toolkit for Deep Learning Based Document Image Analysis

Python 5,770 534 Updated Aug 15, 2024

Convert a PDF via OCR to a TXT file in UTF-8 encoding

Python 161 33 Updated Oct 3, 2023

[python3.6] 运用tf实现自然场景文字检测,keras/pytorch实现ctpn+crnn+ctc实现不定长场景文字OCR识别

Python 2,954 941 Updated Aug 13, 2019

公式图片ocr,输入图片输出对应的latex表达式

HTML 293 74 Updated Apr 11, 2020

Python based Open Source ETL tools for file crawling, document processing (text extraction, OCR), content analysis (Entity Extraction & Named Entity Recognition) & data enrichment (annotation) pipe…

Python 284 70 Updated Oct 9, 2022

Python code to read text from a PDF file (OCR).

Python 69 21 Updated May 26, 2020

Detect text blocks and OCR poorly scanned PDFs in bulk. Python module available via pip.

Python 1,279 101 Updated Dec 1, 2020

Math formula recognition (Images to LaTeX strings)

Jupyter Notebook 308 64 Updated Oct 3, 2023

Call mathpix API to make Mathpix snipping tool.

Python 33 15 Updated Apr 30, 2021

Extract tables from scanned image PDFs using Optical Character Recognition.

Python 277 65 Updated Jun 9, 2020