Skip to content
View drumih's full-sized avatar

Block or report drumih

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Pythonic binding to the Apple Neural Engine

Python 34 10 Updated Aug 6, 2026

A 3D mesh viewer for the Linux terminal written in C!

C 25 2 Updated Jun 3, 2023

LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar

Python 18,516 1,598 Updated Aug 7, 2026

Gemma 4 26B-A4B SSD-streaming inference for Windows/Linux — Rust port of TurboFieldfare

Rust 2 Updated Aug 2, 2026

TurboFieldfare 로컬 추론 서버를 위한 인증, TLS, 요청 제한 및 감사 로그 운영 게이트웨이

Go 4 1 Updated Jul 31, 2026

Minimal tensor computation framework in pure Go with SIMD assembly, inspired by tinygrad

Go 8 1 Updated Aug 7, 2026

Run Qwen3.6 35B on Apple M1-M5 with low RAM usage using SSD/NVM streaming.

Swift 7 Updated Aug 7, 2026

Fine-tune LLMs from one YAML. Layer streaming trains an 8B model on a 4 GB laptop GPU.

Python 398 47 Updated Aug 7, 2026

AirLLM 70B inference with single 4GB GPU

Jupyter Notebook 29,926 3,192 Updated Aug 6, 2026
Swift 446 19 Updated Aug 7, 2026

Swift + Metal MoE inference for Apple Silicon: Qwen 3.6 35B at 23.5–29.3 tok/s decode with 2.20× faster long-prompt prefill on a 24 GB M5; Gemma 4 26B in ~2 GB, DeepSeek-V4-Flash 284B, Inkling-Smal…

Swift 78 3 Updated Aug 6, 2026

Run the full 2.78-trillion-parameter Kimi K3 model beyond available RAM by streaming activated weights directly from NVMe. A dependency-free, embeddable C inference engine.

C 1,880 142 Updated Aug 7, 2026

Ouroboros — self-creating AI agent. Born Feb 16, 2026.

Python 1,000 577 Updated Aug 7, 2026

Ultra-minimalist macOS recording + transcription.

Swift 3,731 238 Updated Jul 30, 2026

Run large Mixture-of-Experts language models well on the memory you have.

Python 6 1 Updated Jul 19, 2026

Memory-aware GGUF inference runtime with OpenAI-compatible serving

Rust 26 2 Updated Aug 7, 2026

DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm

C 20,897 1,867 Updated Aug 5, 2026

A high-performance multi-language bindings generator for Rust, up to 1,000x faster than UniFFI. Ship Rust libraries that feels native to Python, Swift, Kotlin, and more

Rust 850 50 Updated Aug 5, 2026

Runs 405B LLMs on 8GB VRAM

Jupyter Notebook 3,055 237 Updated Apr 2, 2026

A cross-platform, safe, pure-Rust graphics API.

Rust 17,745 1,383 Updated Aug 8, 2026

MentraOS is the leading smart glasses OS. See live captions, stream your view, talk to AI, and capture photos hands-free on compatible glasses.

TypeScript 2,288 331 Updated Aug 8, 2026
Python 3,759 481 Updated Aug 7, 2026

A curated list of awesome platforms, tools, practices and resources that helps run LLMs locally

2,501 312 Updated Aug 4, 2026

Multi-harness control plane for Claude Code, Codex, Cursor, and OpenCode: quota-aware rotation across multiple Claude/Codex subscriptions, shared thread context, and cross-model review.

TypeScript 378 33 Updated Aug 7, 2026

Language model tokenization at GB/s

Rust 3,936 204 Updated Aug 6, 2026

A brief computer graphics / rendering course

C++ 24,070 2,288 Updated Jul 29, 2026

Kali Linux on the Youyeetoo X1S: NVMe rebuild, thermal testing, local AI benchmarks, and a loopback security workflow

Shell 4 Updated Jul 17, 2026

Seven Minutes Is All I Can Spare To Play With You.

Python 12 Updated Jul 23, 2026
Next