Skip to content
View mingfeima's full-sized avatar
:octocat:
i do not stand by in the presence of evil
:octocat:
i do not stand by in the presence of evil
  • Intel Asia-Pacific R&D

Block or report mingfeima

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

RTX 6000 Pro Wiki — Running Large LLMs (Qwen3.5-397B, Kimi-K2.5, GLM-5) on PCIe GPUs without NVLink

Python 839 57 Updated Aug 15, 2026

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python 11,162 1,713 Updated Aug 16, 2026

SGLang-Omni empowers high-performance serving for TTS, ASR, speech and omni models.

Python 820 341 Updated Aug 16, 2026

Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.

Python 19,857 1,990 Updated Aug 14, 2026

OpenAI Triton backend for Intel® GPUs

MLIR 265 106 Updated Aug 15, 2026

Instant, Concurrent, Secure & Lightweight Sandbox for AI Agents.

Rust 11,134 1,030 Updated Aug 14, 2026

[HPCA 2026] AI Accelerator Benchmark focuses on evaluating AI Accelerators from a practical production perspective, including the ease of use and versatility of software and hardware.

Python 374 130 Updated Apr 22, 2026

Machine Learning Engineering Open Book

Python 18,626 1,199 Updated Aug 14, 2026

Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models

Python 4,591 351 Updated Jan 14, 2026

SGLang kernel library for Intel XPU

C++ 30 45 Updated Aug 15, 2026

verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework

Python 22,972 4,410 Updated Aug 15, 2026

Accelerating MoE with IO and Tile-aware Optimizations

Python 741 96 Updated Aug 15, 2026

High-performance inference framework for large language models, focusing on efficiency, flexibility, and availability.

Python 3,014 273 Updated Aug 13, 2026

Nano vLLM

Python 15,007 2,462 Updated Apr 26, 2026
Python 890 52 Updated Sep 15, 2025

gpt-oss-120b and gpt-oss-20b are two open-weight language models by OpenAI

Python 20,309 2,137 Updated Jul 24, 2026
C++ 382 43 Updated Jan 28, 2026
C++ 548 47 Updated Jul 14, 2026

Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.

Python 77,700 6,543 Updated Aug 14, 2026

Toward High-Accuracy Open-Source Biomolecular Structure Prediction.

Python 2,020 300 Updated Aug 1, 2026

FlashMLA: Efficient Multi-head Latent Attention Kernels

C++ 12,845 1,123 Updated Jul 28, 2026

Accelerate local LLM inference and finetuning (LLaMA, Mistral, ChatGLM, Qwen, DeepSeek, Mixtral, Gemma, Phi, MiniCPM, Qwen-VL, MiniCPM-V, etc.) on Intel XPU (e.g., local PC with iGPU and NPU, discr…

Python 8,862 1,429 Updated Jan 28, 2026

LMDeploy is a toolkit for compressing, deploying, and serving LLMs.

Python 8,008 725 Updated Aug 14, 2026

[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.

Cuda 3,645 479 Updated Jan 17, 2026

Official inference framework for 1-bit LLMs

C++ 40,087 3,697 Updated Jul 27, 2026

A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations

Python 19,240 1,524 Updated Aug 15, 2026

A throughput-oriented high-performance serving framework for LLMs

Jupyter Notebook 975 52 Updated Mar 29, 2026

Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.

C++ 6,283 1,085 Updated Aug 15, 2026
Next