- Brooklyn, New York
- https://zxxwxyyy.github.io/
- in/liqian-zhang-0572b91a7
Stars
gpt-oss-120b and gpt-oss-20b are two open-weight language models by OpenAI
Learn Low Level Design (LLD) and prepare for interviews using free resources.
Metrics for evaluating music and audio generative models – with a focus on long-form, full-band, and stereo generations.
Mora: More like Sora for Generalist Video Generation
A high-throughput and memory-efficient inference and serving engine for LLMs
[EMNLP 2023 Demo] Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
A tool to describe the content of videos and suggest similar scenes in other videos/films.
[CVPR2024 Highlight][VideoChatGPT] ChatGPT with video understanding! And many more supported LMs such as miniGPT4, StableLM, and MOSS.
OpenMMLab's Next Generation Video Understanding Toolbox and Benchmark
Repository for paper templates in ISMIR Proceedings
MU-LLaMA: Music Understanding Large Language Model
Carnatic singing voice separation trained with in-domain data with leakage
State-of-the-art audio codec with 90x compression factor. Supports 44.1kHz, 24kHz, and 16kHz mono/stereo audio.
A Streamilt web app for music source separation & karaoke
Generative models for conditional audio generation