- Pacific North West
Highlights
- Pro
Pinned Loading
-
-
vllm-qwen36
vllm-qwen36 PublicServe Qwen3.6 NVFP4 on Blackwell GPUs with vLLM - full 262K context, fp8 KV cache, and MTP speculative decoding via Docker Compose
Shell
-
llama-qwen36
llama-qwen36 PublicDockerized llama.cpp Vulkan server setup for running Qwen3.6 27B GGUF on AMD GPUs.
-
nvfp4-vllm
nvfp4-vllm PublicQuantize HuggingFace models to NVFP4 and serve them with vLLM on NVIDIA Blackwell GPUs.
Python 1
-
llm-serve
llm-serve PublicSelf-hosted OpenAI-compatible LLM inference server for NVFP4 models on NVIDIA Blackwell GPUs, powered by vLLM and Docker.
Shell
-
slugvision
slugvision PublicTiny VLMs that turn an image + optional article title into a 3-5 word permalink slug — distillation pipeline, training, eval, and GGUF export
Python
If the problem persists, check the GitHub status page or contact support.