Fine-tune LLMs from one YAML. Layer streaming trains an 8B model on a 4 GB laptop GPU.
-
Updated
Sep 21, 2026 - Python
Fine-tune LLMs from one YAML. Layer streaming trains an 8B model on a 4 GB laptop GPU.
One-click Windows installer for Z-Image Turbo AI image generation. Optimized for low-VRAM GPUs (4GB+). Features Gradio web UI, automatic setup, and GGUF model support.
WeeLLM runs large diffusion models with as little as 4 GB of VRAM, without any quantization. It dynamically determines how many layers can fit within the available VRAM and streams the text encoder and transformer layers to the GPU layer by layer, enabling inference on hardware with limited VRAM. It supports both safetensors and GGUF models.
A ComfyUI Workflow for low vram users
Hierarchical RAG architecture scaling to 693K chunks on consumer hardware (4GB VRAM). Features 3-address routing, hybrid vector+graph fusion, and SetFit classification.
Taiwanese Hokkien (Taigi) speech-to-text transcriber - MediaTek Breeze-ASR-26 with faster-whisper, tuned for RTX 3050 4GB low-VRAM GPUs. Gradio UI, CLI, Docker, SRT/VTT/TXT/JSON.
"Adaptive Hybrid Quantization Framework for deploying 7B+ LLMs on low-VRAM devices (e.g., GTX 1050). Features surgical block alignment and Numba-accelerated inference.
SCAIL-2 (Wan 2.1) Low-VRAM motion transfer — endless videos from a single reference image. 8+ GB GPUs.
llama.cpp fork tuned for running modern models (Gemma-4, Qwen3.x) at full context on 12 GB Turing GPUs (RTX 2060/2070/2080, T4). TurboQuant KV cache (KTQ+VTQ, 2.78 bpw f16-quality), SWA-aware KV, MTP+n-gram speculation.
Native Hunyuan3D 2.1 Full extension for Modly, optimized for low-VRAM NVIDIA GPUs with INT8, FP8 and FP16 support.
Modly extension for Hunyuan3D 2.1 patched for Windows AMD and modest PCs
Contains the notebooks and workflows configured to run inference from Wan 2.2 Animate with ComfyUI on Kaggle T4 GPUs smoothly
Simple FP16 image upscaler for all GPUs (low-mid end users)
Lightweight 6GB VRAM Gradio web app with auto-installer for running AuraFlow locally — no cloud, no clutter.
在 RTX 3060 6GB 上复现 SmolVLA × LIBERO,并开展低显存 LoRA 微调、多随机种子评测与遗忘控制实验。Low-VRAM SmolVLA × LIBERO reproduction with LoRA adaptation, fixed multi-seed evaluation, and forgetting controls.
Three MiniMax H3 ComfyUI workflows for 8GB laptop GPUs (20 / 8 / 4 steps). Includes the model list, launch flags, prompt-structure pitfalls, the resolution trap, and measured evidence for why you must restart ComfyUI before every run.
Unofficial AMD ROCm low-VRAM fork of Hunyuan3D-2.1 — 6-view PBR texture at ~10.5 GB peak on 20 GB AMD. See README_AMD_ROCM.md.
A workbench for running large Mixture-of-Experts LLMs locally on consumer hardware with a tight VRAM budget.
Perkunas AI Training Platform is a memory-aware model training and serving system for serious language model experimentation under tight hardware limits. It combines streaming training, rich telemetry, guarded recovery, checkpoint export, and OpenAI-compatible serving.
To associate your repository with the low-vram topic, visit your repo's landing page and select "manage topics."