Skip to content
View imkow's full-sized avatar

Block or report imkow

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Starred repositories

Showing results

Local natural-language-to-shell command generator. A 941 MB fine-tuned Qwen2.5-Coder-1.5B running on CPU in ~1s.

Python 463 26 Updated Aug 15, 2026

Hand-written NVFP4 W4A16 CUDA kernels and chain-MTP speculative serving — Qwen3.6-27B at up to 366 tok/s on four Tesla V100s, hardware with no FP4 support

Python 32 3 Updated Aug 11, 2026

A tool to unlobotomize your NVIDIA card!

Shell 325 168 Updated Aug 13, 2026

Caching experts made model agnostic for llama. Drastically speeds up generation speed on MoE models that fully fit in the RAM but partially in Vram. WIP: Same speedup but for models that don't fit …

C++ 17 Updated Aug 6, 2026

Run the full 2.78-trillion-parameter Kimi K3 model beyond available RAM by streaming activated weights directly from NVMe. A dependency-free, embeddable C inference engine.

C 2,159 161 Updated Aug 13, 2026

Run MoE models bigger than your RAM. A 284B on a 12 GB phone, CPU only, lossless, on stock llama.cpp

C++ 365 31 Updated Aug 4, 2026

cyberneurova-DeepSeek-V4-Flash-abliterated-aligned local inference engine with M5 Metal support

C 33 5 Updated Aug 14, 2026

Vane is an AI-powered answering engine.

TypeScript 36,186 4,006 Updated Apr 11, 2026

~95% on SimpleQA (e.g. Qwen3.6-27B on a 3090). Supports all local and cloud LLMs (llama.cpp, Ollama, Google, ...). 10+ search engines - arXiv, PubMed, your private documents. Everything Local & En…

Python 8,929 788 Updated Aug 16, 2026

Run full Kimi K3 on a single device. And an OpenAI-compatible API server for local chat and coding agents.

Rust 757 90 Updated Aug 6, 2026

Python-based stock analysis tool that combines traditional technical analysis with AI prediction capabilities. Providing comprehensive stock analysis and forecasting using K-line charts, technical …

Python 340 95 Updated Aug 6, 2025

AirLLM 70B inference with single 4GB GPU

Jupyter Notebook 31,235 3,322 Updated Aug 15, 2026

DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm

C 21,442 1,962 Updated Aug 9, 2026

MLX (Apple Silicon) port of poolside/Laguna-S-2.1 (118B-A8B MoE). Quants: pipenetwork/Laguna-S-2.1-MLX-*

Python 4 1 Updated Jul 22, 2026

Community benchmarks and scripts for running Poolside Laguna S 2.1

Python 127 2 Updated Jul 22, 2026

Giant MoE models on a single consumer GPU by streaming experts from SSD. CUDA fork of antirez/ds4: runs GLM-5.2 (743B), Tencent Hy3 (295B), and DeepSeek 4 Flash, with io_uring expert streaming, LFU…

C 40 3 Updated Jul 12, 2026

Metal FP32 Vs BF16 Vs FP16 benchmark

Swift 10 1 Updated May 11, 2026

Run frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦

C 25,080 2,720 Updated Aug 15, 2026

Infinite Worlds with Versatile Interactions

Python 1,517 110 Updated Jul 14, 2026

A Julia implementation of Interactive Brokers API

Julia 10 4 Updated Mar 29, 2026

Differentiable Neural Computer in TensorFlow

Jupyter Notebook 27 11 Updated Mar 27, 2017

Differentiable Neural Computer (DNC) implementation in PyTorch.

Python 2 2 Updated Dec 12, 2019

Neural Turing Machine (NTM) & Differentiable Neural Computer (DNC) with pytorch & visdom

Python 279 51 Updated Feb 20, 2018

Differentiable Neural Computers, Sparse Access Memory and Sparse Differentiable Neural Computers, for Pytorch

Python 350 61 Updated Jul 28, 2026

A TensorFlow implementation of the Differentiable Neural Computer.

Python 2,531 446 Updated Jul 23, 2021

Compile programs directly into transformer weights. Includes a 2D convex-hull KV cache with O(log n) inference.

Python 215 46 Updated Jun 1, 2026

Large Language Models (Transformer deep learning architecture)

Jupyter Notebook 2 Updated Jun 6, 2025

Load nanoGPT-style transformers in Julia. Code ported from @karpathy's llama2.c

Julia 65 2 Updated Oct 8, 2023

The most atomic way to train and run inference for a GPT in 100 lines of pure, dependency-free Julia.

Julia 112 5 Updated May 23, 2026

Julia Implementation of Transformer models

Julia 571 83 Updated Jul 31, 2026
Next