Stars
Autoprompt is a coding-agent skill that cuts failures by 45% on agentic coding tasks.
DeepSeek Harness: Everything is a Plugin.
An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Andr…
TradingAgents: Multi-Agents LLM Financial Trading Framework
An open-source AI coding agent that lives in your terminal.
ACE-Step: A Step Towards Music Generation Foundation Model
An implementation of Shazam's song recognition algorithm.
Toolkit to segment text into sentences or other semantic units in a robust, efficient and adaptable way.
Wunjo CE: Face Swap, Lip Sync, Control Remove Objects & Text & Background, Restyling, Audio Separator, Clone Voice, Video Generation. Open Source, Local & Free.
ComfyUI as a serverless API on Runpod
An AI-Powered Speech Processing Toolkit and Open Source SOTA Pretrained Models, Supporting Speech Enhancement, Separation, and Target Speaker Extraction, etc.
The first open-source harness builder for AI coding. Make AI coding deterministic and repeatable.
A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and textual embeddings.
This package is designed to bypass puppeteer's bot-detecting captchas such as Cloudflare. It acts like a real browser and can be managed with puppeteer.
Integration of FastAPI framework supported by Pydantic with SQLAlchemy ORM and PostgreSQL on asyncpg driver
🚀 Builder An API with Node + Express + Prisma + PostgreSQL
This workflow for ComfyUI will allow you to transfering subjects into new pictures while retaining their original features.
The open source AI engineering platform for agents, LLMs, and ML models. MLflow enables teams of all sizes to debug, evaluate, monitor, and optimize production-quality AI applications while control…
PyTorch Re-Implementation of "The Sparsely-Gated Mixture-of-Experts Layer" by Noam Shazeer et al. https://arxiv.org/abs/1701.06538
A react-based starter app for using the Live API over websockets with Gemini
Semantic segmentation models with 500+ pretrained convolutional and transformer-based backbones.
Codebase for the paper: "TIM: A Time Interval Machine for Audio-Visual Action Recognition"
A Python library for extracting color palettes from supplied images.