gguf
Here are 68 public repositories matching this topic...
An open source DevOps tool from the CNCF for packaging and versioning AI/ML models, datasets, code, and configuration into an OCI Artifact.
-
Updated
Sep 21, 2026 - Go
Go with your own intelligence - Write Go applications that directly integrate llama.cpp for local inference using hardware acceleration on Linux, macOS, Windows, & WebAssembly.
-
Updated
Sep 24, 2026 - Go
Go library for embedded vector search and semantic embeddings using llama.cpp
-
Updated
Sep 19, 2026 - Go
Review/Check GGUF files and estimate the memory usage and maximum tokens per second.
-
Updated
Sep 21, 2026 - Go
llama.cpp/ik_llama.cpp launcher: loads big MoE models across mismatched multi-GPU rigs by exact VRAM math.
-
Updated
Sep 24, 2026 - Go
Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.
-
Updated
Sep 24, 2026 - Go
Run your AI models at home in one binary: chat, persistent memory, web access, browser control, MCP tools, encryption at rest and end-to-end encrypted remote access. Linux, macOS, Windows. FR/EN docs.
-
Updated
Sep 24, 2026 - Go
A tiny neural-network framework in pure Go with AVX2 SIMD kernels (GOEXPERIMENT=simd)
-
Updated
Sep 25, 2026 - Go
Agentic Runtime
-
Updated
Sep 25, 2026 - Go
Universal local LLM runtime. Drop-in Ollama replacement with HuggingFace integration, pipelines, structured output. Free & open source.
-
Updated
Jun 17, 2026 - Go
Self-hosted semantic code search platform — Go server with web dashboard, CLI, and AI-agent skills. Search code by meaning, not text: hybrid BM25 + dense embeddings via llama.cpp.
-
Updated
Sep 21, 2026 - Go
Local AI for your whole house: chat, images and voice on your own GPU. Computes each model's flags, fits them to your VRAM, and hot-swaps them behind one OpenAI- and Anthropic-compatible API.
-
Updated
Sep 24, 2026 - Go
Local-first semantic code search for repository annotations via GGUF embeddings and sqlite-vec.
-
Updated
Jul 8, 2026 - Go
基于 llama.cpp 的跨平台本地大模型客户端:一键调优 GGUF 模型,多模型共享一个 OpenAI 兼容端点,内置模型下载、本地聊天与实时监控。Windows / Linux / Android。| A friendly cross-platform local-LLM client powered by llama.cpp — one-click GGUF auto-tuning, all your models behind one OpenAI-compatible endpoint, with built-in downloads, local chat and live monitoring. Windows / Linux / Android.
-
Updated
Sep 17, 2026 - Go
Add this topic to your repo
To associate your repository with the gguf topic, visit your repo's landing page and select "manage topics."