Skip to content
View mingfeima's full-sized avatar
:octocat:
i do not stand by in the presence of evil
:octocat:
i do not stand by in the presence of evil
  • Intel Asia-Pacific R&D

Block or report mingfeima

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

RTX 6000 Pro Wiki — Running Large LLMs (Qwen3.5-397B, Kimi-K2.5, GLM-5) on PCIe GPUs without NVLink

Python 819 55 Updated Aug 12, 2026

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python 11,115 1,690 Updated Aug 12, 2026

SGLang Omni: High-Performance Multi-Stage Pipeline Framework for Omni Models

Python 780 324 Updated Aug 12, 2026

Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.

Python 19,794 1,984 Updated Aug 11, 2026

OpenAI Triton backend for Intel® GPUs

MLIR 265 105 Updated Aug 12, 2026

Instant, Concurrent, Secure & Lightweight Sandbox for AI Agents.

Rust 11,049 1,020 Updated Aug 12, 2026

[HPCA 2026] AI Accelerator Benchmark focuses on evaluating AI Accelerators from a practical production perspective, including the ease of use and versatility of software and hardware.

Python 372 130 Updated Apr 22, 2026

Machine Learning Engineering Open Book

Python 18,594 1,198 Updated Aug 12, 2026

Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models

Python 4,587 350 Updated Jan 14, 2026

SGLang kernel library for Intel XPU

C++ 29 44 Updated Aug 12, 2026

verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework

Python 22,921 4,386 Updated Aug 12, 2026

Accelerating MoE with IO and Tile-aware Optimizations

Python 739 95 Updated Jul 4, 2026

High-performance inference framework for large language models, focusing on efficiency, flexibility, and availability.

Python 3,073 272 Updated Aug 11, 2026

Nano vLLM

Python 14,962 2,448 Updated Apr 26, 2026
Python 890 52 Updated Sep 15, 2025

gpt-oss-120b and gpt-oss-20b are two open-weight language models by OpenAI

Python 20,306 2,134 Updated Jul 24, 2026
C++ 383 43 Updated Jan 28, 2026
C++ 548 47 Updated Jul 14, 2026

Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.

Python 77,404 6,521 Updated Aug 11, 2026

Toward High-Accuracy Open-Source Biomolecular Structure Prediction.

Python 2,019 298 Updated Aug 1, 2026

FlashMLA: Efficient Multi-head Latent Attention Kernels

C++ 12,835 1,120 Updated Jul 28, 2026

Accelerate local LLM inference and finetuning (LLaMA, Mistral, ChatGLM, Qwen, DeepSeek, Mixtral, Gemma, Phi, MiniCPM, Qwen-VL, MiniCPM-V, etc.) on Intel XPU (e.g., local PC with iGPU and NPU, discr…

Python 8,863 1,428 Updated Jan 28, 2026

LMDeploy is a toolkit for compressing, deploying, and serving LLMs.

Python 7,999 722 Updated Aug 12, 2026

[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.

Cuda 3,621 475 Updated Jan 17, 2026

Official inference framework for 1-bit LLMs

C++ 40,009 3,689 Updated Jul 27, 2026

A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations

Python 19,226 1,519 Updated Aug 8, 2026

A throughput-oriented high-performance serving framework for LLMs

Jupyter Notebook 974 52 Updated Mar 29, 2026

Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.

C++ 6,246 1,069 Updated Aug 12, 2026
Next