Skip to content
View Neiko2002's full-sized avatar

Sponsoring

@ENTERPILOT

Highlights

  • Pro

Organizations

@ESCRIBA @bytedeco @WhenPerformanceMatters @Visual-Computing

Block or report Neiko2002

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Starred repositories

Showing results

A self-improving RLM agent for coding workflows and long-running autonomous tasks.

TypeScript 12,900 1,305 Updated Aug 10, 2026

SkillOpt is a text-space optimizer that trains reusable natural-language skills for frozen LLM agents through trajectory-driven edits, validation-gated updates, and deployable best_skill.md artifacts.

Python 15,855 1,473 Updated Aug 9, 2026

Tools for merging pretrained large language models.

Python 7,287 777 Updated Jun 17, 2026

Compress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.

Python 65,801 5,023 Updated Aug 10, 2026

Optimized VLLM for Ampere

Python 13 1 Updated Jul 29, 2026

A vector indexing library to bring fast, fresh and filtered search to your database

Rust 1,898 442 Updated Aug 10, 2026

🤗 PEFT: State-of-the-art Parameter-Efficient Fine-Tuning.

Python 21,522 2,428 Updated Aug 6, 2026

LLM inference in C/C++

C++ 53 9 Updated Aug 8, 2026

RSR-core: A High-Performance Engine for Low-Bit Matrix-Vector Multiplication

Python 7 2 Updated Aug 10, 2026

Model compression toolkit engineered for enhanced usability, comprehensiveness, and efficiency.

Python 1,516 173 Updated Aug 7, 2026

AirLLM 70B inference with single 4GB GPU

Jupyter Notebook 30,557 3,256 Updated Aug 10, 2026

Tile primitives for speedy kernels

Cuda 20 1 Updated Aug 10, 2026

A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downs…

Python 3,418 539 Updated Aug 10, 2026

Browser automation CLI for AI agents

Rust 40,349 2,657 Updated Aug 10, 2026

Control panel for VLLM, Sglang, llama.cpp, exllamav3

TypeScript 1,639 137 Updated Aug 10, 2026
Python 18 Updated Aug 10, 2026

Upskill your model: Flash price. Pro performance.

Shell 90 16 Updated Jul 26, 2026

BitPolar: near-optimal vector quantization — 3-8 bit compression with zero training. 58 integrations across every major AI framework.

Python 20 1 Updated Jul 14, 2026

Build and run agents you can see, understand and trust.

Python 28,775 3,328 Updated Aug 10, 2026

Build distributed, production-grade, long-running agents.

Java 4,985 1,152 Updated Aug 10, 2026
356 35 Updated Jun 9, 2026

Coreutils for Windows: Installer & Packaging

Rust 4,993 99 Updated Jul 28, 2026
Python 44 2 Updated Jun 18, 2026

A vector index built on TurboQuant, written in Rust with Python bindings

Rust 14,710 1,315 Updated Aug 10, 2026

NVIDIA Linux open GPU with P2P support

C 392 48 Updated Jul 13, 2026

[CVPR 2026] UniCorrn: Unified Correspondence Transformer Across 2D and 3D

Python 219 15 Updated Jul 12, 2026

LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.

Python 1,224 199 Updated Aug 8, 2026

Hermes Agent setup, migration, LightRAG, Telegram, and skill creation guide

Shell 579 47 Updated Aug 2, 2026

Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currently shipping Qwen3.6-27B Qwen3.6 35B Gemma 4 26B Gemma 4 31B c…

Python 1,922 117 Updated Aug 10, 2026
Next