Highlights
Lists (1)
Sort Name ascending (A-Z)
Starred repositories
Proof-carrying cyber immunity for Linux containers: signed evidence, deterministic policy, local approval and exact TTL containment.
Point it at a local inference server: reports what it can actually do — measured, not claimed. Catches capacity a server silently fails to deliver.
DeepSeek-V4-Flash across two Strix Halo boxes over 100GbE RDMA: llama.cpp patches, launch config, and measured results. 273 t/s prefill / 21.5 t/s decode at 8k.
Reproducible local-LLM benchmark harness: llama.cpp on AMD Strix Halo (gfx1151, Ryzen AI Max+ 395) and NVIDIA DGX Spark — frozen corpora, quality gates with unit tests, sealed run bundles. Apache-2.0
Multi-slot LLM inference on AMD Strix Halo: recipes + honest benchmarks (236 tok/s @ 32 streams, llama.cpp Vulkan)
A list of Free Software network services and web applications which can be hosted on your own servers
Locally fine-tuned Russian RAG models: document splitter, query expansion, retrieval embedder. Teacher distillation → GGUF → llama.cpp on AMD Vulkan. Commercial-OK licenses.
Private LLM/RAG platform in one command for NVIDIA DGX Spark / GB10 (arm64). Validated on real hardware.
One-command local AI/RAG installer for macOS (Metal): Dify, Open WebUI, Ollama, Weaviate/Qdrant, Postgres, Redis. 230 tests.
WAKE.md for AI agents: compile project state so agents stop starting cold.
Self-hosted LLM/RAG stack in one command — AMD Strix Halo / x86_64 (ROCm/Vulkan, Docker Compose)