Lists (32)
Sort Name ascending (A-Z)
Backbones
Book
Change detection
Classification
CUDA
Dataset
Diffusion Model
Face
GAN
HuaWei
Image restoration
Kaggle
MM
Object Detection
Object Tracking
OCR
ONNX
OpenCV
PaddlePaddle
Prune
Pytorch
Rotated Object Detection
Segmentation
Siamese
Study
Super-Resolution
TensorFlow
TensorRT
text2image
Tools
Transformer
YOLO
Starred repositories
Tongyi Deep Research, the Leading Open-source Deep Research Agent
Convert Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, and PDF to clean Markdown. Built in Rust, with Node.js and Python bindings.
End-to-End Vision-Language Pretraining Without Negatives
The end of web parsing. The beginning of scalable pixel-native search. link: https://pixelrag.ai/
一个基于多模态向量模型及视觉多模态模型构建的图片搜索引擎&管理系统,实现精准的以文搜文,文搜图、以图搜图多种智能检索方式。An image search engine management system built upon multimodal vector models and visual multimodal models, implementing multiple intelli…
Convert any file to LLM-ready Markdown - MCP Server + Claude Code Skill
World's first open-source, agentic video production system. 12 production pipelines, 100+ tools, 700+ agent skill and production-knowledge files. Turn your AI coding assistant into a full video pro…
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
📑 PageIndex: Document Index for Vectorless, Reasoning-based RAG
《深入理解 AI Agent:设计原理与工程实践》(李博杰 著)开源主仓库:全书正文、编译版 PDF 与按章配套代码
A new SOTA for RAG — an original retrieval architecture and an open-source knowledge base for humans and agents.
TreeRAG: Hierarchical RAG platform that indexes dense PDFs into logical trees for precise retrieval, 100% traceability, and LLM-guided deep traversal. Production-ready with enhanced security and st…
Self-Evolving Vision-Language Models for Image Quality Assessment via Voting and Ranking (ICLR 2026)
Self-supervised learning for spatial perception
A blazing-fast, Rust-based code analysis CLI powered by tree-sitter. Analyze code structure, track dependencies, and explore call graphs across multiple languages. Optimized for AI coding agents
Sub2API 一站式开源中转服务,让 Claude、Openai 、Gemini、Grok订阅统一接入,支持拼车共享,更高效分摊成本,原生工具无缝使用。
📚 《Deep Agents 实战》—— LangChain 官方大使出品,基于 LangChain / LangGraph 生态,从零构建生产级 AI Agent 的完整指南
High-Performance Face Recognition Framework (models trained on MS1MV2) | In PyTorch >> ONNX Runtime Inference
Pixel-in-Pixel Net: Towards Efficient Facial Landmark Detection in the Wild | PyTorch & ONNX Runtime Inference
Production-grade engineering skills for AI coding agents.
UniFace: A Unified Face Analysis Library for Python | Detection, alignment, landmarks, face-mesh, recognition, parsing, gaze, attributes and anti-spoofing under one API.
Efficient Universal Perception Encoder: a single on-device vision encoder with versatile representations that match or exceed specialized experts across multiple task domains.
[CVPR 2026] A Closer Look at Cross-Domain Few-Shot Object Detection: Fine-Tuning Matters and Parallel Decoder Helps