Lists (29)
Sort Name ascending (A-Z)
chatgpt
csc
CWS
embedding
IE
instruct following
knowledge graph
LLM agent
ner
nlp工具包
nlp数据预处理
处理数据时,除了分句可能还要先清洗特殊的数据格式,如微博,HTML代码,URL,Email等,HarvestText包含一批常用的数据预处理和清洗操作。temporal relation event
公式识别
图像检测分割
大模型python解释器调用
大模型文档处理
工业化nlp pipeline
布局分析
形变图片增强
文本标注
文档合成
文档序列化
文档形变校正
文档阅读顺序
模型裁剪
模糊匹配
读取预处理图片文本文档
预训练数据清洗
预训练模型
Starred repositories
📚 《从零开始构建智能体》——从零开始的智能体原理与实践教程
🔥🔥超过1000本的计算机经典书籍、个人笔记资料以及本人在各平台发表文章中所涉及的资源等。书籍资源包括C/C++、Java、Python、Go语言、数据结构与算法、操作系统、后端架构、计算机系统知识、数据库、计算机网络、设计模式、前端、汇编以及校招社招各种面经~
Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
The agent that grows with you
An educational resource to help anyone learn deep reinforcement learning.
PyTorch version of Stable Baselines, reliable implementations of reinforcement learning algorithms.
A training framework for Stable Baselines3 reinforcement learning agents, with hyperparameter optimization and pre-trained agents included.
📑 PageIndex: Document Index for Vectorless, Reasoning-based RAG
Specification and documentation for Agent Skills
The official code for "OG-HFYOLO :Orientation Gradient Guidance and Heterogeneous Feature Fusion For Deformation Table Cell Instance Segmentation"
img2table is a table identification and extraction Python Library for PDF and images, based on OpenCV image processing
Convert PDF to markdown + JSON quickly with high accuracy
OCR, layout analysis, reading order, table recognition in 90+ languages
Get your documents ready for gen AI
Ready-to-use OCR with 80+ supported languages and all popular writing scripts including Latin, Chinese, Arabic, Devanagari, Cyrillic and etc.
Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
The official repository of the dots.vlm1 instruct models proposed by rednote-hilab.
Multilingual Document Layout Parsing in a Single Vision-Language Model
拼好RAG:手搓并融合了GraphRAG、LightRAG、Neo4j-llm-graph-builder进行知识图谱构建以及搜索;整合DeepSearch技术实现私域RAG的推理;自制针对GraphRAG的评估框架| Integrate GraphRAG, LightRAG, and Neo4j-llm-graph-builder for knowledge graph construct…
本项目旨在分享大模型相关技术原理以及实战经验(大模型工程化、大模型应用落地)
CypherBench: Towards Precise Retrieval over Full-scale Modern Knowledge Graphs in the LLM Era
Simple wrapper of tabula-java: extract table from PDF into pandas DataFrame
A pure-python PDF library capable of splitting, merging, cropping, and transforming the pages of PDF files
Tabula is a tool for liberating data tables trapped inside PDF files
Community maintained fork of pdfminer - we fathom PDF
Math OCR model that outputs LaTeX and markdown
Ingest, parse, and optimize any data format ➡️ from documents to multimedia ➡️ for enhanced compatibility with GenAI frameworks