Distribute and run LLMs with a single file.
-
Updated
Sep 22, 2026 - C++
Distribute and run LLMs with a single file.
High-speed Large Language Model Serving for Local Deployment
Mano-P: Open-source GUI-VLA agent for edge devices. #1 on OSWorld (specialized, 58.2%). Runs locally on Apple M4 Mac mini/MacBook — no data leaves your device.Mano-P 是一个开源 GUI-VLA 项目,支持在 Mac mini/MacBook 上或通过算力棒本地运行推理,实现纯视觉驱动的跨平台 GUI 自动化操作。数据完全本地处理,支持复杂多步骤任务规划与执行。
Tools, ComfyUI workflows and benchmark configs from a 4x RTX 3090 local-inference rig
[ICLR'25] Fast Inference of MoE Models with CPU-GPU Orchestration
Modern desktop application (Rust + Tauri v2 + Svelte 5 + Candle (HF)) for communicating with AI models that runs completely locally on your computer. No subscriptions, no data sent to the internet — just you and your personal AI assistant
InferrLM - On-device AI for iOS & Android
Offline desktop client for GLM-5.3-Flash — run private, local chat sessions without a server connection.
Real browser automation with local GLiNER2 inference and open weights. Work in progress. Contributions welcome.
Turbo Ultimate Field Fare is a MacOS app that lets users run models like Qwen, Gemma, and GPT-OSS models with expert-streaming, allowing for large models on devices without a lot of memory.
An open-source, model-agnostic agent harness for local LLMs. Define agents in YAML (tools, memory, deny-first permissions) and run them against any OpenAI-compatible endpoint: vLLM, Ollama, LM Studio, or llama.cpp.
Local LLM chat, diffusion, and on-device training on Xbox Series S|X — llama.cpp (GGUF) + ORT GenAI, LAN API, CPU and DirectML.
Notolog Markdown Editor
A fully browser-native RAG application for document Q&A, powered by Rust and WebAssembly with local vector search, embeddings, and in-browser LLM inference.
Local AI music generator with smart lyrics: Gradio web UI for HeartMuLa + Ollama/OpenAI, tags, history, and high-fidelity audio.
A lightweight CUDA-based local inference platform built around Z-Image Turbo by Tongyi
Run a 177B model (Qwen3.8-Flash-Next) on an 8 GB laptop GPU. Measured: 6.6 GiB VRAM, 47.8 GiB RAM, 34-35 tok/s. The n-gram table stays on disk; the experts run on CPU.
Provides an offline AI tutoring application built with Electron that runs locally on your computer using Ollama or Llama models, ensuring complete privacy without requiring internet connectivity.
Tool for test diferents large language models without code.
Local inference server for Apple Silicon — hot-swaps MLX models (LLM, vision, embeddings, TTS, STT) via OpenAI API
To associate your repository with the local-inference topic, visit your repo's landing page and select "manage topics."