Stars
All-in-one intelligent assistant powered by LlamaIndex — RAG, GraphRAG, NL2SQL, Skills & Memory with multimodal support.
Dataflow-LoopAI is an intelligent system with self-optimization capabilities that automatically detects and evaluates generation deficiencies in LLMs within specific domains. Through dialog-based a…
DataFlow Knowledge Graph -- Knowledge graph data preparation with DataFlow style operators and pipelines
[ECCV 2026] A diffusion-based framework for document OCR that replaces autoregressive decoding with block-level parallel diffusion decoding.
A flexible framework for orchestrating deep learning models with Ray . It dynamically schedules and serves multiple models — from NLP (e.g., FastText) to CV (e.g., YOLO, SAM) — enabling scalable, d…
Automated system for LLM evaluation via agents. Doc as below:
Unified Codebase for Advanced World Models.
The First Unified Agent Data Synthesis Framework for Custom Agentic Task with all-in-one envrionment
A Python package for interacting with the MinerU Vision-Language Model.
Ray-powered accelerator for MinerU, turning PDF → Markdown into a scalable, cluster-ready data infrastructure. 基于 Ray 的 MinerU 加速层,将 PDF → Markdown 构建为可扩展、面向集群的数据基础设施。
Open-source implementation of AI-powered academic writing workspace inspired by OpenAI Prism, featuring LaTeX editing, PDF preview, and intelligent AI assistance
An NL2Pipeline Harness for building AI-ready data workflows.
Connect people, agents, and teams for the next era of human-AI collaboration.
Easy Data Preparation with latest LLMs-based Operators and Pipelines.
Dataflow-MM, multi-media operators for Dataflow. We aim to prepare data for Multimodal Large Language Models.
PKU-DAIR / Hetu
Forked from Hsword/HetuA high-performance distributed deep learning system targeting large-scale and automated distributed training.
Turn paper/text/topic into editable research figures, technical route diagrams, and presentation slides.
Self-evolving memory OS for LLM & AI Agents: ultra-persistent memory, hybrid-retrieval, and cross-task skill reuse, with 35.24% token savings
Data-centric LLM training with dynamic sample selection, domain mixture optimization, and example reweighting inside the LLaMA-Factory training loop.
AI Database for unified, scalable SQL + vector data management, search and analytics
PKU-DAIR / DataFlow
Forked from OpenDCAI/DataFlowEasy Data Preparation with latest LLMs-based Operators and Pipelines.
Official repository of RARE: Retrieval-Augmented Reasoning Modeling [KDD 2026 Research Track]
A comprehensive guide for beginners in the field of data management and artificial intelligence.
李宏毅2021/2022/2023春季机器学习课程课件及作业