Stars
Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Andr…
High-Quality Voice Cloning TTS for 600+ Languages
Build your own Claude Code from scratch. 🔍 Claude Code 开源了 50 万行代码,读不动?用 ~5000 行 TypeScript / Python 从零复现核心架构,11 章分步教程带你理解 coding agent 精髓
Deep dive into Claude Code internals — architecture, agent loop, context engineering, and more. / 深入解析 Claude Code 源码:架构、Agent 循环、上下文工程、工具系统等
Academic Research Skills for Claude Code: research → write → review → revise → finalize
Qwen3 is the large language model series developed by Qwen team, Alibaba Cloud.
Stop renting your intelligence. Own it with AnythingLLM. Everything you need for a powerful local-first agent experience
Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
Train transformer language models with reinforcement learning.
VILA is a family of state-of-the-art vision language models (VLMs) for diverse multimodal AI tasks across the edge, data center, and cloud.
[NeurIPS2024] Cross-video Identity Correlating for Person Re-identification Pre-training
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
Ultralytics YOLO26, YOLO11, YOLOv8 — object detection, instance segmentation, semantic segmentation, image classification, pose estimation, object tracking
FaceChain is a deep-learning toolchain for generating your Digital-Twin.
This repository contains the official implementation of the research papers, "MobileCLIP" CVPR 2024 and "MobileCLIP2" TMLR August 2025
Official inference framework for 1-bit LLMs
The repository provides code for running inference with the Meta Segment Anything Model 2 (SAM 2), links for downloading the trained model checkpoints, and example notebooks that show how to use th…
The code for "TokenPacker: Efficient Visual Projector for Multimodal LLM", IJCV2025
MultimodalC4 is a multimodal extension of c4 that interleaves millions of images with text.
A generative speech model for daily dialogue.
llama3 implementation one matrix multiplication at a time
🔥🔥 LLaVA++: Extending LLaVA with Phi-3 and LLaMA-3 (LLaVA LLaMA-3, LLaVA Phi-3)
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
LLaVA-UHD v3: Progressive Visual Compression for Efficient Native-Resolution Encoding in MLLMs
[ECCV 2024 Oral] Code for paper: An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models