Stars
Build and share delightful machine learning apps, all in Python. 🌟 Star to support our work!
Pioneering Automated GUI Interaction with Native Agents
Build Agentic workflows, RAG pipelines, with rich AI model and tool support on one collaborative workspace. Deploy on cloud, VPC, or self-hosted, so teams move from prototype to production without …
No fortress, purely open ground. OpenManus is Coming.
Open-Source Chrome extension for AI-powered web automation. Run multi-agent workflows using your own LLM API key. Alternative to OpenAI Operator.
🌐 Make websites accessible for AI agents. Automate tasks online with ease.
Di♪♪Rhythm: Blazingly Fast and Embarrassingly Simple End-to-End Full-Length Song Generation with Latent Diffusion
Wan: Open and Advanced Large-Scale Video Generative Models
The python library for real-time communication
Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
The official repo of MiniMax-Text-01 and MiniMax-VL-01, large-language-model & vision-language-model based on Linear Attention
The most powerful and modular diffusion model GUI, api and backend with a graph/nodes interface.
TimesFM (Time Series Foundation Model) is a pretrained time-series foundation model developed by Google Research for time-series forecasting.
A framework for building realtime voice AI agents 🤖🎙️📹
Open-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio event detection.
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
An AI-Powered Speech Processing Toolkit and Open Source SOTA Pretrained Models, Supporting Speech Enhancement, Separation, and Target Speaker Extraction, etc.
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
Let your Claude able to think
text and image to video generation: CogVideoX (2024) and CogVideo (ICLR 2023)
Training-free Regional Prompting for Diffusion Transformers 🔥
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
The official repo for “DocScanner: Robust Document Image Rectification with Progressive Learning”, IJCV, 2025.