Stars
Ground truth files from Biodiversity Heritage Library (BHL)
A pilot project with the Biodiversity Heritage Library to find handwritten text pages
OCR model that handles complex tables, forms, handwriting with full layout.
The International Chronostratigraphic Chart from International Commission on Stratigraphy, in computer readable formats
A collection of BGS's vocabularies, formulated using SKOS, serialised as RDF files.
Open source annotation tool for machine learning practitioners.
gpt-oss-120b and gpt-oss-20b are two open-weight language models by OpenAI
GPT Image 2 prompt gallery, image prompt library, agentic skill, and CLI for OpenAI image generation/editing
The best-benchmarked open-source AI memory system. And it's free.
TaxoNERD : recognizing taxonomic entities using deep models
JanusGraph: an open-source, distributed graph database
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
100+ AI Agents, Agent Skills and RAG Apps - Free and Open Source.
RAGFlow is a leading open-source Retrieval-Augmented Generation (RAG) engine that fuses cutting-edge RAG with Agent capabilities to create a superior context layer for LLMs
Hands-On Graph Neural Networks Using Python, published by Packt
A modular graph-based Retrieval-Augmented Generation (RAG) system
LLMs4OM: Matching Ontologies with Large Language Models
Run any open-source LLMs, such as DeepSeek and Llama, as OpenAI compatible API endpoint in the cloud.
Get up and running with Kimi-K2.6, GLM-5.2, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.
Plumb a PDF for detailed information about each char, rectangle, line, et cetera — and easily extract text and tables.
A Unified Toolkit for Deep Learning Based Document Image Analysis
A pure-python PDF library capable of splitting, merging, cropping, and transforming the pages of PDF files
Table structure recognition dataset of the paper: Complicated Table Structure Recognition
Table Transformer (TATR) is a deep learning model for extracting tables from unstructured documents (PDFs and images). This is also the official repository for the PubTables-1M dataset and GriTS ev…
Companion code to the paper "Extracting Scientific Figures with Distantly Supervised Neural Networks" 🤖
[NeurIPS 2025] SpatialLM: Training Large Language Models for Structured Indoor Modeling