Skip to content
#

gguf

Here are 68 public repositories matching this topic...

Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.

  • Updated Sep 24, 2026
  • Go
Quartermaster

Local AI for your whole house: chat, images and voice on your own GPU. Computes each model's flags, fits them to your VRAM, and hot-swaps them behind one OpenAI- and Anthropic-compatible API.

  • Updated Sep 24, 2026
  • Go

基于 llama.cpp 的跨平台本地大模型客户端:一键调优 GGUF 模型,多模型共享一个 OpenAI 兼容端点,内置模型下载、本地聊天与实时监控。Windows / Linux / Android。| A friendly cross-platform local-LLM client powered by llama.cpp — one-click GGUF auto-tuning, all your models behind one OpenAI-compatible endpoint, with built-in downloads, local chat and live monitoring. Windows / Linux / Android.

  • Updated Sep 17, 2026
  • Go

Add this topic to your repo

To associate your repository with the gguf topic, visit your repo's landing page and select "manage topics."

Learn more