Skip to content
View azraelkuan's full-sized avatar
🎯
Focusing
🎯
Focusing
  • Tencent
  • ShenZhen, China
  • 11:23 (UTC +08:00)
  • LinkedIn in/azraelkuan

Highlights

  • Pro

Block or report azraelkuan

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

FinceptTerminal is a modern finance application offering advanced market analytics, investment research, and economic data tools, designed for interactive exploration and data-driven decision-makin…

C++ 31,948 4,524 Updated Sep 19, 2026

PersonaPlex code.

Python 10,570 1,472 Updated Mar 2, 2026

Fast audio super resolution from 16khz to 48khz.

Python 223 20 Updated Jan 3, 2026

VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning

Python 37,925 4,312 Updated Sep 2, 2026

Text-audio foundation model from Boson AI

Python 8,359 641 Updated Jun 5, 2026

PyTorch implementation of MeanFlow & iMF (one-step generative modeling).

Python 1,203 70 Updated Jul 1, 2026

所有小初高、大学PDF教材。

Roff 82,275 18,663 Updated Oct 18, 2025

Cosmos-Reason1 models understand the physical common sense and generate appropriate embodied decisions in natural language through long chain-of-thought reasoning processes.

Python 962 84 Updated Jun 7, 2026

Qwen2.5-Omni is an end-to-end multimodal model by Qwen team at Alibaba Cloud, capable of understanding text, audio, vision, video, and performing real-time speech generation.

Jupyter Notebook 4,076 329 Updated Jun 12, 2025

[ICLR 2025 Oral] Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models

Python 1,034 79 Updated Jul 10, 2025

Spark-TTS Inference Code

Python 11,000 1,151 Updated Apr 9, 2025

Muon is Scalable for LLM Training

1,547 103 Updated Aug 3, 2025

FlashMLA: Efficient Multi-head Latent Attention Kernels

C++ 12,952 1,160 Updated Sep 15, 2026

Official PyTorch implementation of "Paralinguistics-Aware Speech-Empowered LLMs for Natural Conversation" (NeurIPS 2024)

Python 95 4 Updated Dec 3, 2024

A family of state-of-the-art Transformer-based audio codecs for low-bitrate high-quality audio coding.

Python 445 30 Updated Jul 17, 2026

A fast multimodal LLM for real-time voice

Python 4,567 382 Updated Dec 12, 2025

Official implementation of "Lotus: Diffusion-based Visual Foundation Model for High-quality Dense Prediction"

Python 837 52 Updated Nov 28, 2025

A Survey of Spoken Dialogue Models (60 pages)

317 17 Updated Nov 28, 2024

Local realtime voice AI

Python 2,495 150 Updated Nov 26, 2025

Interface for OuteTTS models.

Python 1,440 117 Updated Mar 23, 2026

[ICCV 2025] SimVQ: Addressing Representation Collapse in Vector Quantized Models with One Linear Layer

Python 334 12 Updated Dec 29, 2024

A suite of image and video neural tokenizers

Jupyter Notebook 1,729 95 Updated Feb 11, 2025

Tencent Hunyuan3D-1.0: A Unified Framework for Text-to-3D and Image-to-3D Generation

Python 3,488 280 Updated Nov 19, 2025

open-source multimodal large language model that can hear, talk while thinking. Featuring real-time end-to-end speech input and streaming audio output conversational capabilities.

Python 3,581 310 Updated Nov 5, 2024

GLM-4 series: Open Multilingual Multimodal Chat LMs | 开源多语言多模态对话模型

Python 7,071 616 Updated Aug 5, 2026

GLM-4-Voice | 端到端中英语音对话模型

Python 3,237 289 Updated Dec 5, 2024

Official inference framework for 1-bit LLMs

C++ 40,339 3,743 Updated Jul 27, 2026

NanoGPT (124M) in 90 seconds

Python 5,820 892 Updated Sep 18, 2026
Next