Popular repositories Loading
-
qwen38-k100ai-int8-optimization
qwen38-k100ai-int8-optimization PublicQwen3.8-27B INT8/W8A8 optimization for Hygon K100AI: SGLang, TP1/TP4, DFlash2 and long-context Agent workloads
-
qwen36-k100ai-w8a8-optimization
qwen36-k100ai-w8a8-optimization PublicReproducible single-GPU Qwen3.6-35B-A3B W8A8 inference tuning for Hygon K100AI (gfx928) on vLLM 0.18.1.
Python 2
-
minimax-h3-k100ai-optimization
minimax-h3-k100ai-optimization PublicMiniMax H3 INT8 ConvRot optimization and dual-K100AI QKV/Attention inference for Hygon K100AI (gfx928)
Python
-
qwen38-k100ai-w8a8-optimization
qwen38-k100ai-w8a8-optimization PublicReproducible Qwen3.8-27B SmoothQuant W8A8/INT8 optimization for Hygon K100AI (R054, 262K, Prefix Cache, adaptive MTP)
Python
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.