- Munich, Germany
Stars
Rust client for the huggingface hub aiming for minimal subset of features over `huggingface-hub` python package
DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms
AI inference, packed simply. A blazing-fast, zero-dependency WebGPU runtime to run GGUF models directly in the browser. Features a symmetric API for seamless local execution and cloud provider rout…
Tile primitives for speedy kernels
aidendle94 / vllm
Forked from vllm-project/vllmA high-throughput and memory-efficient inference and serving engine for LLMs
Bleeding edge vLLM Docker image for the NVIDIA DGX Spark (GB10 / sm_121a).
DGX Spark / GB10 vLLM Docker stack for large-model serving, presets, patches, and validation notes.
A vector index built on TurboQuant, written in Rust with Python bindings
local-inference-lab / vllm
Forked from vllm-project/vllmA high-throughput and memory-efficient inference and serving engine for LLMs
Build compute kernels and load them from the Hub.
ComfyUI-Manager is an extension designed to enhance the usability of ComfyUI. It offers management functions to install, remove, disable, and enable various custom nodes of ComfyUI. Furthermore, th…
Zerocopy makes zero-cost memory manipulation effortless. We write `unsafe` so you don’t have to.
A contact solver for physics-based simulations involving 👚 shells, 🪵 solids, 🪢 rods, 🧱 rigid bodies and ⏳ sand.
Ghidra is a software reverse engineering (SRE) framework
antirez/ds4-style hybrid quant DeepSeek V4 Flash on a single DGX Spark via vLLM
DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm
A harness optimized to smaller LLMs
FlashQLA TileLang GDN kernels ported to NVIDIA Blackwell consumer (GB10 / DGX Spark)
high-performance linear attention kernel library built on TileLang
Train transformer language models with reinforcement learning.
Your self-hosted, globally interconnected microblogging community
[WIP] Civ VI inspired game, written in Rust using the Bevy game engine
Model compression toolkit engineered for enhanced usability, comprehensiveness, and efficiency.
Parameter-Efficient Fine-Tuning For Edge-Cloud Collabrative Computing
A curated collection of official and community-built Claude Skills – extend Anthropic's Claude with powerful, modular capabilities for productivity, creativity, coding, and more.
Qwen3-TTS is an open-source series of TTS models developed by the Qwen team at Alibaba Cloud, supporting stable, expressive, and streaming speech generation, free-form voice design, and vivid voice…