Stars
Low-level unprivileged sandboxing tool used by Flatpak and similar projects
A vLLM patch + hand‑written SM120 SASS kernels: 2‑bit MoE experts + an FP4 "delta" cache that recovers precision — matching the official (NV)FP4 checkpoint's quality on consumer Blackwell cards
A runtime substrate that turns an agent's execution into a reversible, Git-like trace, so meta-agents can observe, fork, replay, and revert any run. Couples agent and environments in a copy-on-writ…
A ~9M parameter LLM that talks like a small fish.
Context Gateway is an agentic proxy that enhances any AI agent workflow with instant history compaction and context optimization tools
Beautiful, open source, WebGPU-based charting library
Run sandboxed code environments on Cloudflare's edge network
A fully customizable and self-hosted sandboxing solution for AI agent code execution and computer use. It features out-of-the-box support for backtracking, a simple REST API and Python SDK, automat…
Open-source, secure environment with real-world tools for enterprise-grade agents.
A lightweight LMM-based Document Parsing Model
Multilingual Document Layout Parsing in a Single Vision-Language Model
High-performance In-browser LLM Inference Engine
A new chunking strategy developed by ZeroEntropy for general semantic chunking using Llama-70B.
This is the repo for the LegalBench-RAG Paper: https://arxiv.org/abs/2408.10343.
Identity-aware VPN and tunneled reverse proxy for remote access based on WireGuard®.
High-performance MLX-based LLM inference engine for macOS with native Swift implementation
🚀 FlexLLama - Lightweight self-hosted tool for running multiple llama.cpp server instances with OpenAI v1 API compatibility and multi-GPU support
A simple clone of lovable.dev using BAML / FastMCP / Beam
A next-gen FOSS self-hosted unified zero trust secure access platform that can operate as a remote access VPN, a ZTNA platform, API/AI/MCP gateway, a PaaS, an ngrok-alternative and a homelab infras…
A lightweight tool for deploying and managing containerised applications across a network of Docker hosts. Bridging the gap between Docker and Kubernetes ✨
MiniMax-M1, the world's first open-weight, large-scale hybrid-attention reasoning model.
A reimplementation of Stable Diffusion 3.5 in pure PyTorch
A comprehensive Model Context Protocol (MCP) server implementing the latest specification.
🚀🤖 Crawl4AI: Open-source LLM Friendly Web Crawler & Scraper. Don't be shy, join here: https://discord.gg/jP8KfhDhyN