Stars
GPUDirect Storage (GDS) for the NVIDIA CMP 170HX on modern Linux kernels (6.18+). nvidia-fs NVMe registration hooks for the new blk_rq_dma_map_iter API that NVIDIA MOFED/DOCA does not cover. GPU-ag…
Hand-written NVFP4 W4A16 CUDA kernels and chain-MTP speculative serving — Qwen3.6-27B at up to 366 tok/s on four Tesla V100s, hardware with no FP4 support
DeepSeek-V4-Flash-0731 on 4x CMP 170HX (sm_80): 98 tok/s decode, ~5300 tok/s prefill. DSpark speculative decoding under pipeline parallelism, which vLLM does not support upstream.
Tuning and qualification harness for the NVIDIA CMP 170HX (GA100). Measures, gates and recovers clock/voltage settings, and treats a completed benchmark as evidence of nothing.
haohervchb / sglang-V100
Forked from sgl-project/sglangSGLang for cool dudes using V100
CMP 170hx 64gb LLM benchmarks. Qwen3.6 27b, Qwen35b a3b. vLLM
A tool to unlobotomize your NVIDIA card!
Container stack for reusing an Octominer/Octofan chassis as an AI inference server enclosure.
Orchestrate multiple coding agents from desktop and mobile
Monitoring for Proxmox, Docker, Kubernetes, TrueNAS, and vSphere that watches your infrastructure for you: smart alerts, AI patrols that catch silent failures, and verified fixes
zephyrq-z / ompcot
Forked from shixin-guo/picotLocal Codex-style desktop GUI for the Pi coding agent
shixin-guo / picot
Forked from deflating/tauLocal Codex-style desktop GUI for the Pi coding agent
An efficient AI coding agent extendable by neovim-like Lua plugins
AI agent framework, written from scratch (not based on openclaw), focused on stripping it down to the bare necessities, optimizing token count, reducing security risks. modular so you can enable on…
SmarterRouter: An intelligent LLM gateway and VRAM-aware router for Ollama, llama.cpp, and OpenAI. Features semantic caching, model profiling, and automatic failover for local AI labs.
A Python framework for self-hosted LLM tool-calling and multi-step agentic workflows
Lossless DFlash speculative decoding for MLX on Apple Silicon
A ~9M parameter LLM that talks like a small fish.
Local UI to run and train LLMs and diffusion models, including Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, DeepSeek-V4, FLUX and more.
Smallest transformer that can add two 10-digit numbers
A reverse engineering project for the TP-Link Tether Management Protocol (TMP), which allows authenticated configuration of TP-Link routers as well as other use cases.
Local AI assistant, dreaming explorable worlds.
Ultra-lightweight, open-source, self-hosted personal AI agent framework in Python with WebUI, tools, memory, MCP, multi-agent workflows, automation, and chat apps
Zero-installation web application that lets you explore, back up, and manage your ESP32… right from your browser
Linux service that provides the Native ESPHome API with a plugin architecture and Bluetooth proxy support natively