Skip to content
View JimmyHHua's full-sized avatar
🎯
Focusing
🎯
Focusing
  • 深圳, China

Block or report JimmyHHua

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Andr…

C++ 14,198 1,629 Updated Aug 13, 2026

High-Quality Voice Cloning TTS for 600+ Languages

Python 9,142 1,502 Updated Aug 10, 2026

Build your own Claude Code from scratch. 🔍 Claude Code 开源了 50 万行代码,读不动?用 ~5000 行 TypeScript / Python 从零复现核心架构,11 章分步教程带你理解 coding agent 精髓

Python 2,565 526 Updated Jul 9, 2026

Deep dive into Claude Code internals — architecture, agent loop, context engineering, and more. / 深入解析 Claude Code 源码:架构、Agent 循环、上下文工程、工具系统等

3,448 691 Updated Jul 12, 2026

Academic Research Skills for Claude Code: research → write → review → revise → finalize

Python 42,597 3,392 Updated Aug 15, 2026

Qwen3 is the large language model series developed by Qwen team, Alibaba Cloud.

Python 27,508 2,034 Updated Jan 9, 2026

Stop renting your intelligence. Own it with AnythingLLM. Everything you need for a powerful local-first agent experience

JavaScript 64,739 7,134 Updated Aug 13, 2026

Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.

Jupyter Notebook 19,790 1,830 Updated Jan 30, 2026

Train transformer language models with reinforcement learning.

Python 19,079 2,910 Updated Aug 15, 2026

VILA is a family of state-of-the-art vision language models (VLMs) for diverse multimodal AI tasks across the edge, data center, and cloud.

Python 3,854 330 Updated Mar 12, 2026

[NeurIPS2024] Cross-video Identity Correlating for Person Re-identification Pre-training

Python 107 5 Updated Jun 20, 2025

🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.

Python 164,124 34,249 Updated Aug 15, 2026

Ultralytics YOLO26, YOLO11, YOLOv8 — object detection, instance segmentation, semantic segmentation, image classification, pose estimation, object tracking

Python 60,646 11,588 Updated Aug 15, 2026

FaceChain is a deep-learning toolchain for generating your Digital-Twin.

Jupyter Notebook 9,506 875 Updated Jun 6, 2025

This repository contains the official implementation of the research papers, "MobileCLIP" CVPR 2024 and "MobileCLIP2" TMLR August 2025

Python 1,620 130 Updated Apr 15, 2026

Official inference framework for 1-bit LLMs

C++ 40,087 3,697 Updated Jul 27, 2026

The repository provides code for running inference with the Meta Segment Anything Model 2 (SAM 2), links for downloading the trained model checkpoints, and example notebooks that show how to use th…

Jupyter Notebook 19,703 2,528 Updated May 30, 2026

The code for "TokenPacker: Efficient Visual Projector for Multimodal LLM", IJCV2025

Python 279 9 Updated May 26, 2025

MultimodalC4 is a multimodal extension of c4 that interleaves millions of images with text.

Python 954 38 Updated Mar 19, 2025

LLM101n: Let's build a Storyteller

37,509 2,076 Updated Aug 1, 2024

A generative speech model for daily dialogue.

Python 39,767 4,258 Updated Apr 10, 2026

llama3 implementation one matrix multiplication at a time

Jupyter Notebook 15,223 1,282 Updated May 23, 2024
Python 4,715 470 Updated Jun 15, 2026

HPT - Open Multimodal LLMs from HyperGAI

Python 313 22 Updated Jun 6, 2024

🔥🔥 LLaVA++: Extending LLaVA with Phi-3 and LLaMA-3 (LLaVA LLaMA-3, LLaVA Phi-3)

Python 842 58 Updated Aug 5, 2025

The official Meta Llama 3 GitHub site

Python 29,257 3,526 Updated Jan 26, 2025

One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks

Python 4,364 639 Updated Aug 6, 2026

LLaVA-UHD v3: Progressive Visual Compression for Efficient Native-Resolution Encoding in MLLMs

Python 423 20 Updated Jul 6, 2026

[ECCV 2024 Oral] Code for paper: An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Python 593 31 Updated Jan 4, 2025
Next