Stars
FinceptTerminal is a modern finance application offering advanced market analytics, investment research, and economic data tools, designed for interactive exploration and data-driven decision-makin…
Fast audio super resolution from 16khz to 48khz.
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
Text-audio foundation model from Boson AI
PyTorch implementation of MeanFlow & iMF (one-step generative modeling).
Cosmos-Reason1 models understand the physical common sense and generate appropriate embodied decisions in natural language through long chain-of-thought reasoning processes.
Qwen2.5-Omni is an end-to-end multimodal model by Qwen team at Alibaba Cloud, capable of understanding text, audio, vision, video, and performing real-time speech generation.
[ICLR 2025 Oral] Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models
FlashMLA: Efficient Multi-head Latent Attention Kernels
Official PyTorch implementation of "Paralinguistics-Aware Speech-Empowered LLMs for Natural Conversation" (NeurIPS 2024)
A family of state-of-the-art Transformer-based audio codecs for low-bitrate high-quality audio coding.
Official implementation of "Lotus: Diffusion-based Visual Foundation Model for High-quality Dense Prediction"
[ICCV 2025] SimVQ: Addressing Representation Collapse in Vector Quantized Models with One Linear Layer
A suite of image and video neural tokenizers
Tencent Hunyuan3D-1.0: A Unified Framework for Text-to-3D and Image-to-3D Generation
open-source multimodal large language model that can hear, talk while thinking. Featuring real-time end-to-end speech input and streaming audio output conversational capabilities.
GLM-4 series: Open Multilingual Multimodal Chat LMs | 开源多语言多模态对话模型
Official inference framework for 1-bit LLMs