Highlights
- Pro
Lists (1)
Sort Name ascending (A-Z)
Starred repositories
SkillOpt is a text-space optimizer that trains reusable natural-language skills for frozen LLM agents through trajectory-driven edits, validation-gated updates, and deployable best_skill.md artifacts.
Tools for merging pretrained large language models.
Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
A vector indexing library to bring fast, fresh and filtered search to your database
🤗 PEFT: State-of-the-art Parameter-Efficient Fine-Tuning.
alesha-pro / llama.cpp
Forked from ggml-org/llama.cppLLM inference in C/C++
RSR-core: A High-Performance Engine for Low-Bit Matrix-Vector Multiplication
Model compression toolkit engineered for enhanced usability, comprehensiveness, and efficiency.
AirLLM 70B inference with single 4GB GPU
QuixiAI / QuixiCore-CUDA
Forked from HazyResearch/ThunderKittensTile primitives for speedy kernels
A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downs…
Browser automation CLI for AI agents
Control panel for VLLM, Sglang, llama.cpp, exllamav3
BitPolar: near-optimal vector quantization — 3-8 bit compression with zero training. 58 integrations across every major AI framework.
Build and run agents you can see, understand and trust.
Build distributed, production-grade, long-running agents.
Coreutils for Windows: Installer & Packaging
A vector index built on TurboQuant, written in Rust with Python bindings
NVIDIA Linux open GPU with P2P support
[CVPR 2026] UniCorrn: Unified Correspondence Transformer Across 2D and 3D
LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.
Hermes Agent setup, migration, LightRAG, Telegram, and skill creation guide
Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currently shipping Qwen3.6-27B Qwen3.6 35B Gemma 4 26B Gemma 4 31B c…
replacement for unraid-plg-geminicli