Stars
Memori is agent-native memory infrastructure. A LLM-agnostic layer that turns agent execution and conversation into structured, persistent state for production systems. Built for enterprise, Memori…
Official Repo for ICML 2024 paper "Executable Code Actions Elicit Better LLM Agents" by Xingyao Wang, Yangyi Chen, Lifan Yuan, Yizhe Zhang, Yunzhu Li, Hao Peng, Heng Ji.
[ICCV 2025 Fingdings] IDMR: Towards Instance-Driven Precise Visual Correspondence in Multimodal Retrieval
[CVPR 2025] A Comprehensive Benchmark for Document Parsing and Evaluation
(CVPR 2025 highlight✨) Official repository of paper "LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models"
Leaderboard Comparing LLM Performance at Producing Hallucinations when Summarizing Short Documents
[COLM 2025] Open-Qwen2VL: Compute-Efficient Pre-Training of Fully-Open Multimodal LLMs on Academic Resources
The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible.
Understanding R1-Zero-Like Training: A Critical Perspective
A collection of tutorials on state-of-the-art computer vision models and techniques. Explore everything from foundational architectures like ResNet to cutting-edge models like RF-DETR, YOLO11, SAM …
Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable pipelines based on verified research papers.
OpenOCR: An Open-Source Toolkit for General-OCR Research and Applications, integrates a unified training and evaluation benchmark, commercial-grade OCR and Document Parsing systems, and faithful re…
Python tool for converting files and office documents to Markdown.
A curated list of awesome synthetic data for text location and recognition
Easy-to-Use RAG Framework; CCF AIOps International Challenge 2024 Top3 Solution; CCF AIOps 国际挑战赛 2024 季军方案
Enhances Tesseract OCR output using LLMs (local or API) for error correction, smart chunking, and markdown formatting of scanned PDFs
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
A Faster LayoutReader Model based on LayoutLMv3, Sort OCR bboxes to reading order.
Minimalistic large language model 3D-parallelism training
RAGFlow is a leading open-source Retrieval-Augmented Generation (RAG) engine that fuses cutting-edge RAG with Agent capabilities to create a superior context layer for LLMs
Silero VAD: pre-trained enterprise-grade Voice Activity Detector
PDF to markdown using vision LLMs — tables, layouts, and structure preserved
Alpaca Chinese Dataset -- 中文指令微调数据集