Lists (17)
Sort Name ascending (A-Z)
Stars
zll961020 / MOSS-Audio
Forked from OpenMOSS/MOSS-AudioMOSS-Audio is an open-source foundation model for unified audio understanding, enabling speech, sound, music, captioning, QA, and reasoning in real-world scenarios.
zll961020 / Mega-ASR
Forked from xzf-thu/Mega-ASRFirst foundation ASR built for the real world - 7 atomic acoustic conditions, 54 compound scenarios, 2.6M samples, and up to ~30% gains over SOTA where every other model falls apart. **You'll come …
仅需Python基础,从0构建自己的具身智能机器人;从0逐步构建VLA/OpenVLA/SmolVLA/Pi0, 深入理解具身智能
A Large-scale Wu Dialect Speech Corpus with Multi-dimensional Annotations
zll961020 / MimicKit
Forked from xbpeng/MimicKitA lightweight suite of motion imitation methods for training controllers.
Murmur: An Efficient Inference System for Long-Form ASR
Official repository for the WenetSpeech-Chuan dataset.
zll961020 / multinerf
Forked from google-research/multinerfA Code Release for Mip-NeRF 360, Ref-NeRF, and RawNeRF
zll961020 / CLAP
Forked from LAION-AI/CLAPContrastive Language-Audio Pretraining
zll961020 / hello-agents
Forked from datawhalechina/hello-agents📚 《从零开始构建智能体》——从零开始的智能体原理与实践教程
zll961020 / VibeVoice
Forked from microsoft/VibeVoiceOpen-Source Frontier Voice AI
zll961020 / Qwen3-ASR
Forked from QwenLM/Qwen3-ASRQwen3-ASR is an open-source series of ASR models developed by the Qwen team at Alibaba Cloud, supporting stable multilingual speech/music/song recognition, language detection and timestamp prediction.
zll961020 / DiariZen
Forked from BUTSpeechFIT/DiariZenA toolkit for speaker diarization.
Claude Code v2.1.88 Source Code
zll961020 / claude-howto
Forked from luongnv89/claude-howtoA visual, example-driven guide to Claude Code — from basic concepts to advanced agents, with copy-paste templates that bring immediate value.
zll961020 / deepagents
Forked from langchain-ai/deepagentsAgent harness built with LangChain and LangGraph. Equipped with a planning tool, a filesystem backend, and the ability to spawn subagents - well-equipped to handle complex agentic tasks.
zll961020 / deer-flow
Forked from bytedance/deer-flowAn open-source long-horizon SuperAgent harness that researches, codes, and creates. With the help of sandboxes, memories, tools, skill, subagents and message gateway, it handles different levels of…
来自于文章Paraformer-v2: An improved non-autoregressive transformer for noise-robust speech recognition
Variational Bayes HMM over x-vectors diarization
zll961020 / ROLL
Forked from alibaba/ROLLAn Efficient and User-Friendly Scaling Library for Reinforcement Learning with Large Language Models
Open-source LLM-based ASR model family for Chinese, dialect, accent, and multilingual speech, with FunASR, vLLM, streaming, and llama.cpp runtimes.
zll961020 / r1-aqa
Forked from xiaomi-research/r1-aqa🤗 R1-AQA Model: mispeech/r1-aqa
zll961020 / SALMONN
Forked from bytedance/SALMONNSALMONN family: A suite of advanced multi-modal LLMs
zll961020 / Qwen3-Omni
Forked from QwenLM/Qwen3-OmniQwen3-omni is a natively end-to-end, omni-modal LLM developed by the Qwen team at Alibaba Cloud, capable of understanding text, audio, images, and video, as well as generating speech in real time.
zll961020 / HTGS
Forked from nerficg-project/HTGSOfficial code release for "Efficient Perspective-Correct 3D Gaussian Splatting Using Hybrid Transparency"