Find files on Windows by name, by content, or by what they mean. Fully local — no cloud, no account, one .exe.
-
Updated
Jul 30, 2026 - C++
Find files on Windows by name, by content, or by what they mean. Fully local — no cloud, no account, one .exe.
My GitHub profile README — Tokyo Night terminal banner, typing taglines, stats, contribution snake
A lightweight, end-to-end LLM fine-tuning pipeline for Llama-3.2-1B-Instruct. Optimized with Unsloth (QLoRA) and deployed via Hugging Face Spaces using GGUF for efficient CPU inference.
A Flutter toolkit for local and cloud LLM apps. Run GGUF models on Android and iOS, connect to OpenAI, Gemini, Claude and others through one unified API, and build fully local RAG pipelines. Supports streaming, multimodal vision models, and mobile-first AI workflows.
A Rust library and CLI for parsing GGUF model file headers — extract metadata, architecture, quantization, and tensor info without loading weights.
Windows-based, high-performance, full stack, local LLM chat application written in C# .NET 10 using the Lethe AI library.
Auto GGUF Converter for HuggingFace Hub Models with Multiple Quantizations (GGUF Format)
Offline local LLM terminal for desktop. GPU accelerated, no internet, no cloud. Runs any GGUF model.
Download GGUF Models in Clean CSV File with Updater
Local-first RAG over your files in pure Go; llama.cpp, a from-scratch HNSW index, and SSE streaming. No cloud, no CGO.
llama.cpp + Qwen3.6-27B (Q8_0 GGUF) OpenAI-compatible inference server on AMD Strix Halo (Ryzen AI Max+ 395, gfx1151). 256K context, ~7.5 t/s decode via TheRock ROCm Docker.
GenPark AI Agent Skill - Calculates tensor parallelism column and row weight matrix partition splits across distributed GPU clusters.
GenPark AI Agent Skill - Rejection sampling and logit verification engine for speculative decoding acceleration.
llama.cpp fleet manager with orchestration routing — cut AI coding tool costs by routing tool-loop turns to local GPU models and frontier APIs (Copilot, OpenRouter). GPU pinning, heterogeneous pools, browser dashboard, OpenAI-compatible API.
Serving LLMs on-prem with no GPU: benchmark (CPU + single GPU) across models, GGUF quants, engines, concurrency — and a decision guide
To associate your repository with the gguf topic, visit your repo's landing page and select "manage topics."