Skip to content
View fffeifang's full-sized avatar
🏠
Working from home
🏠
Working from home

Highlights

  • Pro

Block or report fffeifang

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Faster attention kernels for serving TML's Inkling model on vLLM. 2.7x over the shipping path on H100, and the only implementation that runs on A100.

Python 5 1 Updated Jul 26, 2026

A Hitchhiker's Guide to ML PhD Job Hunting

JavaScript 49 4 Updated Jul 26, 2026

SGLang Omni: High-Performance Multi-Stage Pipeline Framework for Omni Models

Python 698 287 Updated Jul 27, 2026

A Streaming-Native Serving Engine for TTS/STS Models

Python 74 9 Updated Jun 20, 2026

SOTA Open Source TTS

Python 31,388 2,694 Updated Jul 26, 2026

X-Talk is an open-source full-duplex cascaded spoken dialogue system framework enabling low-latency, interruptible, and human-like speech interaction with a lightweight, pure-Python, production-rea…

Python 233 33 Updated Jul 26, 2026

Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond

Python 1,117 129 Updated Jul 24, 2026

Official PyTorch+CUDA Full-functional Web Demo for MiniCPM-o 4.5

Python 300 58 Updated Jul 23, 2026

Open Source framework for voice and multimodal conversational AI

Python 13,747 2,376 Updated Jul 27, 2026

End-to-end realtime stack for connecting humans and AI

Go 19,988 2,191 Updated Jul 27, 2026

The Prometheus monitoring system and time series database.

Go 65,332 10,707 Updated Jul 27, 2026

The open and composable observability and data visualization platform. Visualize metrics, logs, and traces from multiple sources like Prometheus, Loki, Elasticsearch, InfluxDB, Postgres and many mo…

TypeScript 75,815 14,377 Updated Jul 27, 2026

Vendor-agnostic orchestration for training, inference and agentic workloads across NVIDIA, AMD, TPU, and Tenstorrent on clouds, Kubernetes, and bare metal.

Python 2,196 241 Updated Jul 27, 2026

Covo-Audio is a 7B-parameter end-to-end large audio language model that directly processes continuous audio inputs and generates audio outputs within a single unified architecture.

Python 173 17 Updated Mar 17, 2026

PersonaPlex code.

Python 10,268 1,429 Updated Mar 2, 2026

An asynchronous streaming data management module for efficient post-training.

Python 120 42 Updated Jul 12, 2026

Moshi is a speech-text foundation model and full-duplex spoken dialogue framework. It uses Mimi, a state-of-the-art streaming neural audio codec.

Python 10,740 997 Updated May 16, 2026

htop-like TUI for real-time RDMA network monitoring.

Rust 77 5 Updated Jul 23, 2026

Can LLMs Write Correct and Efficient GPU Communication Code?

Python 47 2 Updated Jul 7, 2026

The high-performance distributed tensor layer — load once, share everywhere.

C++ 30 Updated Jun 23, 2026
Python 35 4 Updated Jan 27, 2026

[MLsys2026]: RAG on Everything with LEANN. Enjoy 97% storage savings while running a fast, accurate, and 100% private RAG application on your personal device.

Python 12,735 1,148 Updated Jul 24, 2026

[ICLR 2023] ReAct: Synergizing Reasoning and Acting in Language Models

Jupyter Notebook 4,083 393 Updated Feb 6, 2024

Platform for stateful agents: AI with advanced memory that can learn and self-improve over time.

Python 23,985 2,553 Updated Jul 22, 2026

open-source code for paper: Retrieval Head Mechanistically Explains Long-Context Factuality

Python 241 27 Updated Aug 2, 2024

A Lightweight LLM Post-Training Library

Python 2,387 325 Updated Jul 27, 2026

LongBench v2 and LongBench (ACL 25'&24')

Python 1,215 136 Updated Jan 15, 2025

KV cache store for distributed LLM inference

C++ 424 43 Updated Nov 13, 2025
Python 21 5 Updated Jul 13, 2026

Modular and structured prompt caching for low-latency LLM inference

Python 115 14 Updated Nov 9, 2024
Next