-
Harbin Institute of Technology
- Shanghai China
- 1838635312@qq.com
Starred repositories
An awesome list of Agent Harness engineering resources, including GitHub projects, tools, benchmarks, and practical guides.
40+ tips for getting the most out of Claude Code, from basics to advanced - includes a custom status line script and Claude Code running itself in a container. Also includes the dx plugin: skills f…
12 Weeks, 24 Lessons, AI for All!
AIMET is a library that provides advanced quantization and compression techniques for trained neural network models.
A list of papers, docs, codes about model quantization. This repo is aimed to provide the info for model quantization research, we are continuously improving the project. Welcome to PR the works (p…
Bjontegaard metric calculation. Include BD-PSNR and BD-rate
Fast and differentiable MS-SSIM and SSIM for pytorch.
Accelerated Python package for computing the PSNR-HVS-M image metric
Comparison of IQA models in Perceptual Optimization
Code for Paper (Preserving Diversity in Supervised Fine-tuning of Large Language Models)
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
Kimi-Audio, an open-source audio foundation model excelling in audio understanding, generation, and conversation
Automatically Update Text-to-speech (TTS) Papers Daily using Github Actions (Update Every 12th hours)
LLaSA: Scaling Train-time and Inference-time Compute for LLaMA-based Speech Synthesis
This is an evolving repo for the paper "Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey".
torchange - A Unified Change Representation Learning Benchmark Library
[CVPR 2025] MatAnyone: Stable Video Matting with Consistent Memory Propagation
X-AnyLabeling: A lightweight, efficient, and unified cross-platform desktop application for annotating text, image, video, and multimodal data, combining versatile built-in tools with state-of-the-…
A ComfyUI custom node designed for advanced image background removal and object, face, clothes, and fashion segmentation, utilizing multiple models including RMBG-2.0, INSPYRENET, BEN, BEN2, BiRefN…
[ICCV 2025, Highlight] ZIM: Zero-Shot Image Matting for Anything
Official code for "F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching"
[CVPRW 2022] Attentions Help CNNs See Better: Attention-based Hybrid Image Quality Assessment Network
Speech Human Evaluation Estimation Toolkit (SHEET)