-
Neural Magic
- Boston
- https://neuralmagic.com/
Stars
A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM
General Information, model certifications, and benchmarks for nm-vllm enterprise distributions
Neural network model repository for highly sparse and sparse-quantized models with matching sparsification recipes
Libraries for applying sparsification recipes to neural networks with a few lines of code, enabling faster and smaller models
Sparsity-aware deep learning inference runtime for CPUs
ML model optimization product to accelerate inference.