-
Intel Asia-Pacific R&D
Stars
RTX 6000 Pro Wiki — Running Large LLMs (Qwen3.5-397B, Kimi-K2.5, GLM-5) on PCIe GPUs without NVLink
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
SGLang Omni: High-Performance Multi-Stage Pipeline Framework for Omni Models
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
OpenAI Triton backend for Intel® GPUs
Instant, Concurrent, Secure & Lightweight Sandbox for AI Agents.
[HPCA 2026] AI Accelerator Benchmark focuses on evaluating AI Accelerators from a practical production perspective, including the ease of use and versatility of software and hardware.
Machine Learning Engineering Open Book
Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
Accelerating MoE with IO and Tile-aware Optimizations
High-performance inference framework for large language models, focusing on efficiency, flexibility, and availability.
gpt-oss-120b and gpt-oss-20b are two open-weight language models by OpenAI
Transforms complex documents like PDFs and Office docs into LLM-ready markdown/JSON for your Agentic workflows.
Toward High-Accuracy Open-Source Biomolecular Structure Prediction.
FlashMLA: Efficient Multi-head Latent Attention Kernels
Accelerate local LLM inference and finetuning (LLaMA, Mistral, ChatGLM, Qwen, DeepSeek, Mixtral, Gemma, Phi, MiniCPM, Qwen-VL, MiniCPM-V, etc.) on Intel XPU (e.g., local PC with iGPU and NPU, discr…
LMDeploy is a toolkit for compressing, deploying, and serving LLMs.
[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.
Official inference framework for 1-bit LLMs
A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations
A throughput-oriented high-performance serving framework for LLMs
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.