-
RedHat
- Boston, MA, USA
- https://redhat.com
- in/alex-matveev-42bb8580
Stars
Achieve state of the art inference performance with modern accelerators on Kubernetes
Transformers-compatible library for applying various compression algorithms to LLMs for optimized deployment with vLLM
Top-level directory for documentation and general content
Neural network model repository for highly sparse and sparse-quantized models with matching sparsification recipes
Sparsity-aware deep learning inference runtime for CPUs
ML model optimization product to accelerate inference.
Libraries for applying sparsification recipes to neural networks with a few lines of code, enabling faster and smaller models