Skip to content
View knitvoger's full-sized avatar

Block or report knitvoger

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Docker configuration for running VLLM on dual DGX Sparks

Shell 2,087 358 Updated Aug 13, 2026

📚 《从零开始构建智能体》——从零开始的智能体原理与实践教程

Python 72,871 9,073 Updated Aug 12, 2026

"ViMax: Agentic Video Generation (Director, Screenwriter, Producer, and Video Generator All-in-One)"

Python 11,921 1,781 Updated Jul 29, 2026

Qwen-Image-Lightning: Speed up Qwen-Image model with distillation

Python 1,351 46 Updated Jan 1, 2026

VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning

Python 35,645 4,083 Updated Aug 12, 2026

利用 AI 大模型和自动化工作流,根据主题或关键词一键生成高清短视频。Generate HD short videos from a topic or keyword with an automated AI workflow.

Python 103,231 15,633 Updated Aug 13, 2026

AI-powered animated comic generator — transform scripts into fully animated videos with AI-driven character design, storyboarding, and video synthesis.

TypeScript 1,771 307 Updated Apr 27, 2026

Implement a ChatGPT-like LLM in PyTorch from scratch, step by step

Jupyter Notebook 102,623 15,727 Updated Aug 10, 2026

《Build a Large Language Model (From Scratch)》是一本深入探讨大语言模型原理与实现的电子书,适合希望深入了解 GPT 等大模型架构、训练过程及应用开发的学习者。为了让更多中文读者能够接触到这本极具价值的教材,我决定将其翻译成中文,并通过 GitHub 进行开源共享。

HTML 3,943 655 Updated Aug 10, 2026

Learning Long-Term Spatial-Temporal Graphs for Active Speaker Detection (ECCV 2022)

Python 67 9 Updated Oct 29, 2023

An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Python 22,849 2,774 Updated Aug 13, 2026

[KDD 2026] Voxlect: A Speech Foundation Model Benchmark for Modeling Dialects and Regional Languages Around the Globe

Python 45 4 Updated Aug 10, 2025

Code for the Active Speakers in Context Paper (CVPR2020)

Python 58 14 Updated May 19, 2021

Graph learning framework for long-term video understanding

Python 72 12 Updated Jul 13, 2026

The AVA dataset densely annotates 80 atomic visual actions in 351k movie clips with actions localized in space and time, resulting in 1.65M action labels with multiple labels per human occurring fr…

347 29 Updated Feb 9, 2022

ACM MM 2021: 'Is Someone Speaking? Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection'

Python 496 102 Updated Oct 23, 2023

📚 从零开始构建大模型

Jupyter Notebook 32,949 3,125 Updated Aug 8, 2026

基于AI的图片/视频硬字幕去除、文本水印去除,无损分辨率生成去字幕、去水印后的图片/视频文件。无需申请第三方API,本地实现。AI-based tool for removing hard-coded subtitles and text-like watermarks from videos or Pictures.

Python 12,326 1,581 Updated Jun 30, 2026

Netflix-level subtitle cutting, translation, alignment, and even dubbing - one-click fully automated AI video subtitle team | Netflix级字幕切割、翻译、对齐、甚至加上配音,一键全自动视频搬运AI字幕组

Python 18,148 1,999 Updated Jul 2, 2026

An AI-Powered Speech Processing Toolkit and Open Source SOTA Pretrained Models, Supporting Speech Enhancement, Separation, and Target Speaker Extraction, etc.

Python 4,405 359 Updated Aug 14, 2025

Code for the paper Hybrid Spectrogram and Waveform Source Separation

Python 3,022 281 Updated Jul 11, 2026

Code for the paper Hybrid Spectrogram and Waveform Source Separation

Python 10,356 1,563 Updated Apr 24, 2024

【C++面试&C++学习指南】 这里整理了C++后端研发工程师面试和工作必备的知识点 。

3,056 425 Updated Apr 14, 2025

📚 C/C++ 技术面试基础知识总结,包括语言、程序库、数据结构、算法、系统、网络、链接装载库等知识及面试经验、招聘、内推等信息。This repository is a summary of the basic knowledge of recruiting job seekers and beginners in the direction of C/C++ technology, in…

C++ 38,121 8,081 Updated Aug 24, 2025

An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and Chatbot Arena.

Python 39,518 4,788 Updated May 1, 2026

Simple script for downloading Youtube comments without using the Youtube API

Python 1,246 261 Updated Jul 30, 2026

Instant voice cloning by MIT and MyShell. Audio foundation model.

Python 37,143 4,145 Updated Apr 19, 2025

Complete API layer for private AI applications on local models: RAG, skills, tools, MCP, text-to-sql, and more. Works with any OpenAI-compatible inference server.

Python 57,435 7,610 Updated Aug 13, 2026

🌟 The Multi-Agent Framework: First AI Software Company, Towards Natural Language Programming

Python 69,812 8,880 Updated Jan 21, 2026

Whisper realtime streaming for long speech-to-text transcription and translation

Python 3,662 409 Updated Nov 12, 2025
Next