- All languages
- Assembly
- AutoHotkey
- B4X
- C
- C#
- C++
- CMake
- CSS
- Cuda
- Dart
- Dockerfile
- Fancy
- Fluent
- Go
- HTML
- Haskell
- Java
- JavaScript
- Jinja
- Jupyter Notebook
- Kotlin
- Lua
- MATLAB
- Makefile
- Markdown
- Mustache
- Objective-C
- PHP
- Perl
- Python
- Rich Text Format
- Roff
- Ruby
- Rust
- SCSS
- Scala
- Shell
- Svelte
- Swift
- TeX
- TypeScript
- Visual Basic 6.0
- Vue
- XSLT
- Zig
Starred repositories
Unified Schema-Based Information Extraction
精选的中国开放文档格式(OFD)资源列表,包括标准规范、库、SDK、转换工具、阅读器和教程,为开发者和研究者提供全面参考。
0.3B OCR model which emits structured markup ---- math as latex, table as html, text as markdown.
Smart Document API turns PDFs, Office documents, and images into structured layouts, OCR text, tables, formulas, and Markdown for AI-powered document understanding.
Agent skill for beautiful, verifiable architecture, workflow, sequence, data-flow, and lifecycle diagrams—self-contained HTML with motion and crisp export.
VelOCR (飞舟OCR) — 基于 Tauri (Rust + React) 构建的跨平台智能 OCR 桌面应用。集成 PaddleOCR 本地推理引擎,支持图片/PDF 文字检测与识别,提供多语言翻译、Markdown 渲染、i18n 国际化等功能,注重隐私保护,无需云端 API 即可在本地完成 OCR 识别。
slime is an LLM post-training framework for RL Scaling.
Lightweight PaddleOCR-VL inference with ONNXRuntime layout and ROCm-backed OpenAI-compatible VLM serving.
[NeurlPS 2025] A Unified Benchmark & Codebase for All-Domain Fake Image Detection and Localization
A Rust implementation of the XYCut++ algorithm. https://arxiv.org/pdf/2504.10258
LLM speculative inference server for consumer hardware & heterogeneous computing
[ECCV 2026] StrucTab: A Structured Optimization Framework for Table Parsing
EasyNLP: A Comprehensive and Easy-to-use NLP Toolkit
An Open-Source Package for Neural Relation Extraction (NRE)
Repository for ECCV2026 P-MTP: Efficient Document Parsing via Multi-Token Prediction with Progressive Depth Scaling
High-performance, token-efficient JSON superset with Zen Grid tabular encoding
Dingo: A Comprehensive AI Data, Model and Application Quality Evaluation Tool
Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM
Fastest enterprise AI gateway (50x faster than LiteLLM) with adaptive load balancer, cluster mode, guardrails, 1000+ models support & <100 µs overhead at 5k RPS.
中国专利.skill:从项目文档到交底书编写(挖点·查新·脱敏成文)+ 专利通俗解读(叙事·图谱·Obsidian 私库)。
TokenSpeed is a speed-of-light LLM inference engine.
Llama Agents + Workflows are an event-driven, async-first, step-based way to control the execution flow of AI applications like agents.
An optimized quantization and inference library for running LLMs locally on modern consumer-class GPUs