Skip to content
View Neiko2002's full-sized avatar

Sponsoring

@ENTERPILOT

Highlights

  • Pro

Organizations

@ESCRIBA @bytedeco @WhenPerformanceMatters @Visual-Computing

Block or report Neiko2002

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Starred repositories

Showing results

SkillOpt is a text-space optimizer that trains reusable natural-language skills for frozen LLM agents through trajectory-driven edits, validation-gated updates, and deployable best_skill.md artifacts.

Python 15,504 1,430 Updated Aug 2, 2026

Tools for merging pretrained large language models.

Python 7,275 774 Updated Jun 17, 2026

Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.

Python 64,074 4,875 Updated Aug 2, 2026

Optimized VLLM for Ampere

Python 10 1 Updated Jul 29, 2026

A vector indexing library to bring fast, fresh and filtered search to your database

Rust 1,889 440 Updated Aug 1, 2026

🤗 PEFT: State-of-the-art Parameter-Efficient Fine-Tuning.

Python 21,481 2,414 Updated Aug 1, 2026

LLM inference in C/C++

C++ 48 8 Updated Jul 28, 2026

RSR-core: A High-Performance Engine for Low-Bit Matrix-Vector Multiplication

Python 7 2 Updated Apr 1, 2026

Model compression toolkit engineered for enhanced usability, comprehensiveness, and efficiency.

Python 1,504 171 Updated Aug 1, 2026

AirLLM 70B inference with single 4GB GPU

Jupyter Notebook 25,654 2,883 Updated Jul 29, 2026

Tile primitives for speedy kernels

Cuda 20 1 Updated Aug 1, 2026

A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downs…

Python 3,370 523 Updated Aug 2, 2026

Browser automation CLI for AI agents

Rust 39,775 2,605 Updated Aug 2, 2026

Control panel for VLLM, Sglang, llama.cpp, exllamav3

TypeScript 1,525 129 Updated Aug 2, 2026
Python 17 Updated Aug 1, 2026

Upskill your model: Flash price. Pro performance.

Shell 89 14 Updated Jul 26, 2026

BitPolar: near-optimal vector quantization — 3-8 bit compression with zero training. 58 integrations across every major AI framework.

Python 20 1 Updated Jul 14, 2026

Build and run agents you can see, understand and trust.

Python 28,495 3,284 Updated Aug 1, 2026

Build distributed, production-grade, long-running agents.

Java 4,846 1,099 Updated Aug 2, 2026
350 34 Updated Jun 9, 2026

Coreutils for Windows: Installer & Packaging

Rust 4,941 95 Updated Jul 28, 2026
Python 44 2 Updated Jun 18, 2026

A vector index built on TurboQuant, written in Rust with Python bindings

Rust 14,579 1,293 Updated Aug 2, 2026

NVIDIA Linux open GPU with P2P support

C 375 47 Updated Jul 13, 2026

[CVPR 2026] UniCorrn: Unified Correspondence Transformer Across 2D and 3D

Python 217 15 Updated Jul 12, 2026

LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.

Python 1,221 197 Updated Aug 2, 2026

Hermes Agent setup, migration, LightRAG, Telegram, and skill creation guide

Shell 563 45 Updated Aug 2, 2026

Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currently shipping Qwen3.6-27B Qwen3.6 35B Gemma 4 26B Gemma 4 31B c…

Python 1,860 112 Updated Aug 2, 2026

replacement for unraid-plg-geminicli

PHP 10 2 Updated Jul 30, 2026
Next