-
UC Santa Cruz
- San Jose
-
09:13
(UTC -07:00) - https://fffeifang.github.io
- https://orcid.org/0009-0006-3709-6749
Highlights
- Pro
Lists (10)
Sort Name ascending (A-Z)
Stars
Faster attention kernels for serving TML's Inkling model on vLLM. 2.7x over the shipping path on H100, and the only implementation that runs on A100.
A Hitchhiker's Guide to ML PhD Job Hunting
SGLang Omni: High-Performance Multi-Stage Pipeline Framework for Omni Models
A Streaming-Native Serving Engine for TTS/STS Models
X-Talk is an open-source full-duplex cascaded spoken dialogue system framework enabling low-latency, interruptible, and human-like speech interaction with a lightweight, pure-Python, production-rea…
Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond
Official PyTorch+CUDA Full-functional Web Demo for MiniCPM-o 4.5
Open Source framework for voice and multimodal conversational AI
End-to-end realtime stack for connecting humans and AI
The Prometheus monitoring system and time series database.
The open and composable observability and data visualization platform. Visualize metrics, logs, and traces from multiple sources like Prometheus, Loki, Elasticsearch, InfluxDB, Postgres and many mo…
Vendor-agnostic orchestration for training, inference and agentic workloads across NVIDIA, AMD, TPU, and Tenstorrent on clouds, Kubernetes, and bare metal.
Covo-Audio is a 7B-parameter end-to-end large audio language model that directly processes continuous audio inputs and generates audio outputs within a single unified architecture.
An asynchronous streaming data management module for efficient post-training.
Moshi is a speech-text foundation model and full-duplex spoken dialogue framework. It uses Mimi, a state-of-the-art streaming neural audio codec.
htop-like TUI for real-time RDMA network monitoring.
Can LLMs Write Correct and Efficient GPU Communication Code?
The high-performance distributed tensor layer — load once, share everywhere.
[MLsys2026]: RAG on Everything with LEANN. Enjoy 97% storage savings while running a fast, accurate, and 100% private RAG application on your personal device.
[ICLR 2023] ReAct: Synergizing Reasoning and Acting in Language Models
Platform for stateful agents: AI with advanced memory that can learn and self-improve over time.
open-source code for paper: Retrieval Head Mechanistically Explains Long-Context Factuality
KV cache store for distributed LLM inference
Modular and structured prompt caching for low-latency LLM inference