Neural Magic
Neural Magic (Acquired by Red Hat) empowers developers to optimize & deploy LLMs at scale. Our model compression & acceleration enable top performance with vLLM
Pinned Loading
Repositories
Showing 10 of 103 repositories
- model-validation-configs Public
- vllm Public Forked from vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
- nm-vllm-omni-ent Public Forked from vllm-project/vllm-omni
A framework for efficient model inference with omni-modality models
- tokenspeed Public Forked from lightseekorg/tokenspeed
TokenSpeed is a speed-of-light LLM inference engine.
- nyann-bench Public
- tpu-inference Public Forked from vllm-project/tpu-inference
TPU inference for vLLM, with unified JAX and PyTorch support.
- eval-hub-contrib Public Forked from eval-hub/eval-hub-contrib
Community-contributed evaluation framework adapters for eval-hub
Top languages
Loading…
Most used topics
Loading…