Skip to content
View Neiko2002's full-sized avatar

Sponsoring

@ENTERPILOT

Highlights

  • Pro

Organizations

@ESCRIBA @bytedeco @WhenPerformanceMatters @Visual-Computing

Block or report Neiko2002

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Starred repositories

Showing results

SkillOpt is a text-space optimizer that trains reusable natural-language skills for frozen LLM agents through trajectory-driven edits, validation-gated updates, and deployable best_skill.md artifacts.

Python 14,962 1,377 Updated Jul 24, 2026

Tools for merging pretrained large language models.

Python 7,261 768 Updated Jun 17, 2026

Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.

Python 62,310 4,705 Updated Jul 25, 2026

Optimized VLLM for Ampere

Python 11 1 Updated Jul 23, 2026

A vector indexing library to bring fast, fresh and filtered search to your database

Rust 1,882 436 Updated Jul 25, 2026

🤗 PEFT: State-of-the-art Parameter-Efficient Fine-Tuning.

Python 21,449 2,401 Updated Jul 24, 2026

LLM inference in C/C++

C++ 47 7 Updated Jul 23, 2026

RSR-core: A High-Performance Engine for Low-Bit Matrix-Vector Multiplication

Python 7 2 Updated Apr 1, 2026

Model compression toolkit engineered for enhanced usability, comprehensiveness, and efficiency.

Python 1,483 168 Updated Jul 23, 2026

AirLLM 70B inference with single 4GB GPU

Jupyter Notebook 24,012 2,707 Updated Jul 23, 2026

Tile primitives for speedy kernels

Cuda 20 Updated Jul 24, 2026

A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downs…

Python 3,306 511 Updated Jul 25, 2026

Browser automation CLI for AI agents

Rust 39,170 2,560 Updated Jul 24, 2026

Control panel for VLLM, Sglang, llama.cpp, exllamav3

TypeScript 1,497 122 Updated Jul 24, 2026
Python 14 Updated Jul 21, 2026

Upskill your model: Flash price. Pro performance.

Shell 86 13 Updated Jun 20, 2026

BitPolar: near-optimal vector quantization — 3-8 bit compression with zero training. 58 integrations across every major AI framework.

Python 20 1 Updated Jul 14, 2026

Build and run agents you can see, understand and trust.

Python 28,253 3,250 Updated Jul 23, 2026

Build distributed, production-grade, long-running agents.

Java 4,697 1,043 Updated Jul 24, 2026
347 33 Updated Jun 9, 2026

Coreutils for Windows: Installer & Packaging

Rust 4,890 94 Updated Jul 22, 2026
Python 44 2 Updated Jun 18, 2026

A vector index built on TurboQuant, written in Rust with Python bindings

Python 14,151 1,264 Updated Jul 24, 2026

NVIDIA Linux open GPU with P2P support

C 356 47 Updated Jul 13, 2026

[CVPR 2026] UniCorrn: Unified Correspondence Transformer Across 2D and 3D

Python 214 14 Updated Jul 12, 2026

LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.

Python 1,214 195 Updated Jul 25, 2026

Hermes Agent setup, migration, LightRAG, Telegram, and skill creation guide

Shell 547 46 Updated Jul 17, 2026

Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currently shipping Qwen3.6-27B Qwen3.6 35B Gemma 4 26B Gemma 4 31B c…

Python 1,780 107 Updated Jul 24, 2026

replacement for unraid-plg-geminicli

PHP 10 2 Updated Jul 19, 2026
Next