- Singapore
- https://x.com/FeitengLi
- @FeitengLi
Lists (1)
Sort Name ascending (A-Z)
Stars
A framework for building realtime voice AI agents 🤖🎙️📹
KVAE-Audio: a continuous full-band audio waveform autoencoder
A TTS that fits in your CPU (and pocket)
A lightweight, self-hostable tool for conducting subjective listening tests in a browser.
[ICML'26] Code and website for Self-Flow: Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis
Benchmark forced alignment models (Qwen3-FA, WhisperX, Seamless UnitY2) on FLEURS via implicit WER on cropped audio.
A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downs…
[ICLR 2026] The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution
Everything I know about running LLMs locally
Audio-Oscar is a multi-agent framework for generating long-form, controllable audio from complex audio scene descriptions.
MiniCPM5-1B: A SOTA 1B on-device LLM, small yet powerful.
A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.
Qwen3.6 is the large language model series developed by Qwen team, Alibaba Group.
DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms
A vector index built on TurboQuant, written in Rust with Python bindings
Target-level automatic benchmark for raw-input Chinese news TTS pronunciation accuracy
[CVPR 2026] Humanoid-GPT: Scaling Data and Structure for Zero-Shot Motion Tracking
Official implementation of paper "Vocoder is not all you need".
Code to pretrain, fine-tune, and evaluate DreamZero and run sim & real-world evals
Interactive World Model papers organized by core research challenges.
Dataflow-Oriented Reinforcement Learning for (Multi-)Agentic LLMs
Fine-tune Gemma 4 and 3n with audio, images and text on Apple Silicon, using PyTorch and Metal Performance Shaders.
Open-source LLM-based ASR model family for Chinese, dialect, accent, and multilingual speech, with FunASR, vLLM, streaming, and llama.cpp runtimes.
The most accurate natural language detection library for Python, suitable for short text and mixed-language text