Skip to content
View NMsasa's full-sized avatar
  • Red Hat
  • Boston

Organizations

@neuralmagic

Block or report NMsasa

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Evaluate and Enhance Your LLM Deployments for Real-World Inference Needs

Python 1,528 213 Updated Aug 20, 2026

Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | 开源持续推理基准研究平台…

Python 1,414 261 Updated Aug 20, 2026

Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM

Python 3,706 625 Updated Aug 20, 2026

Synthetic Data Generation Toolkit for LLMs

Python 157 62 Updated Aug 19, 2026

Achieve state of the art inference performance with modern accelerators on Kubernetes

Shell 4,076 690 Updated Aug 20, 2026

A high-throughput and memory-efficient inference and serving engine for LLMs

Python 89,514 20,953 Updated Aug 20, 2026
PowerShell 2 Updated Apr 1, 2024
Python 210 27 Updated May 5, 2025

A high-throughput and memory-efficient inference and serving engine for LLMs

Python 267 10 Updated Dec 4, 2025

Refine high-quality datasets and visual AI models

TypeScript 11,019 813 Updated Aug 20, 2026

Open-source AI orchestration framework for building context-engineered, production-ready LLM applications. Design modular pipelines and agent workflows with explicit control over retrieval, routing…

Python 26,259 3,028 Updated Aug 20, 2026

Sparsity-aware deep learning inference runtime for CPUs

Python 3,157 193 Updated Jun 2, 2025

Top-level directory for documentation and general content

MDX 120 7 Updated Jun 2, 2025

ML model optimization product to accelerate inference.

Python 325 32 Updated Jun 2, 2025

Neural network model repository for highly sparse and sparse-quantized models with matching sparsification recipes

Python 388 28 Updated Jun 2, 2025

Libraries for applying sparsification recipes to neural networks with a few lines of code, enabling faster and smaller models

Python 2,145 156 Updated Jun 2, 2025