Lists (4)
Sort Name ascending (A-Z)
Stars
An open-source long-horizon SuperAgent harness that researches, codes, and creates. With the help of sandboxes, memories, tools, skill, subagents and message gateway, it handles different levels of…
An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
Library for conversion between Traditional and Simplified Chinese
Depth Pro: Sharp Monocular Metric Depth in Less Than a Second.
EMO-SUPERB: a reproducible speech emotion recognition benchmark with leakage-free splits for 6 datasets and 15 speech SSL models (IEEE SLT 2024)
Run macOS VM in a Docker! Run near native OSX-KVM in Docker! X11 Forwarding! CI/CD for OS X Security Research! Docker mac Containers.
End-to-end realtime stack for connecting humans and AI
WebRTC and ORTC implementation for Python using asyncio
Pretrain, finetune ANY AI model of ANY size on 1 or 10,000+ GPUs with zero code changes.
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
21 Lessons, Get Started Building with Generative AI
Qwen3 is the large language model series developed by Qwen team, Alibaba Cloud.
Use Microsoft Edge's online text-to-speech service from Python WITHOUT needing Microsoft Edge or Windows or an API key
Open-Sora: Democratizing Efficient Video Production for All
[CVPR 2024] Real-Time Open-Vocabulary Object Detection
X-AnyLabeling: A lightweight, efficient, and unified cross-platform desktop application for annotating text, image, video, and multimodal data, combining versatile built-in tools with state-of-the-…
Ultralytics YOLO26, YOLO11, YOLOv8 — object detection, instance segmentation, semantic segmentation, image classification, pose estimation, object tracking
EmotiVoice 😊: a Multi-Voice and Prompt-Controlled TTS Engine
Question and Answer based on Anything.
A python package to build AI-powered real-time audio applications
vits2 backbone with multilingual-bert
VITS2: Improving Quality and Efficiency of Single-Stage Text-to-Speech with Adversarial Learning and Architecture Design
VITS: Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech
PaddlePaddle Code Convert Toolkit. 『飞桨』深度学习代码转换工具