#
fp8
Here are 3 public repositories matching this topic...
An inference engine written in Rust for GGUF, AWQ, and FP8 (W8A8) checkpoints, with hand-written CUDA and Metal kernels — no PyTorch, no libtorch, no ggml
-
Updated
Sep 11, 2026 - Rust
Benchmark vLLM serving endpoints with a fast Rust client. Use parallel dataset generation, low memory overhead, and static binaries without Python dependencies.
kubernetes cli aws benchmark terraform grafana inference triton safety datasets robustness claude eks promotheus llm fp8 multimodal-llm
-
Updated
Sep 22, 2026 - Rust
Add this topic to your repo
To associate your repository with the fp8 topic, visit your repo's landing page and select "manage topics."