Skip to content
View sar's full-sized avatar

Block or report sar

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Starred repositories

Showing results

Fast and accurate DRAM power and energy estimation tool

C++ 210 61 Updated Jul 22, 2026

Train GPT-2 large language model (150M to 774M) series on Google Colab using K80 GPUs

Jupyter Notebook 4 Updated Jan 4, 2020

AMD ROCm AI Inference & Training Solutions Library for RNDA2/GFX1030/Radeon Pro v620 includes: compiling vLLM Source, Llama.cpp, HIPFire, and Fine-Tuning paths.

Dockerfile 1 Updated Jul 14, 2026

Slurm on Kubernetes Architecture Solution for Fine-tuning LLMs, Inference, and Eval across distributed NVIDIA L40s X8 GPU Cluster on Nebius AI Cloud

HCL 1 Updated Dec 17, 2025

Run Slurm in Kubernetes

Go 403 60 Updated Jul 23, 2026

HIP: C++ Heterogeneous-Compute Interface for Portability

C++ 4,380 588 Updated Jul 8, 2026

super repo for rocm libraries

Assembly 390 349 Updated Jul 24, 2026

A high-throughput and memory-efficient inference and serving engine for LLMs

Python 87,031 19,788 Updated Jul 24, 2026

RDNA-native LLM inference engine in Rust.

Rust 489 50 Updated Jul 24, 2026
C++ 51 1 Updated Jul 1, 2026

The main repository for building Pascal-compatible versions of ML applications and libraries.

Shell 214 33 Updated Aug 23, 2025

A fast high-compression read-only file system for Linux, FreeBSD, macOS and Windows

C++ 2,585 88 Updated Jul 23, 2026

Achieve state of the art inference performance with modern accelerators on Kubernetes

Shell 3,870 633 Updated Jul 24, 2026

Kimi Code CLI is your next CLI agent.

Python 10,729 1,251 Updated Jul 16, 2026

Jobs scraper library for LinkedIn, Indeed, Glassdoor, Google, ZipRecruiter & more

Python 3,939 781 Updated Feb 18, 2026

Qwen3-Coder is the code version of Qwen3, the large language model series developed by Qwen team.

Python 16,742 1,230 Updated Mar 24, 2026

An open-source AI coding agent that lives in your terminal.

TypeScript 26,275 2,713 Updated Jul 24, 2026

Flexible I/O Tester

C 6,302 1,415 Updated Jul 23, 2026

A fast JSON parser/generator for C++ with both SAX/DOM style API

C++ 15,108 3,648 Updated Feb 5, 2025

NVIDIA Inference Xfer Library (NIXL)

C++ 1,148 376 Updated Jul 24, 2026

LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

Python 10,853 1,611 Updated Jul 24, 2026

Distributed reliable key-value store for the most critical data of a distributed system

Go 52,023 10,427 Updated Jul 23, 2026

The etcd-cpp-apiv3 is a C++ library for etcd's v3 client APIs, i.e., ETCDCTL_API=3.

C++ 389 152 Updated Mar 28, 2025

IOR and mdtest

C 482 199 Updated Apr 3, 2026

Magnum IO community repo

C++ 117 23 Updated Jun 22, 2026

NVIDIA GPUDirect Storage Driver

C 367 65 Updated Jun 1, 2026

llama.cpp fork with additional SOTA quants and improved performance

C++ 2,959 386 Updated Jul 23, 2026

Diffusion model(SD,Flux,Wan,Qwen Image,Z-Image,...) inference in pure C/C++

C++ 6,583 705 Updated Jul 23, 2026

Minimal CLI coding agent by Mistral

Python 4,730 608 Updated Jul 23, 2026

Home for "How To Scale Your Model", a short blog-style textbook about scaling LLMs on TPUs

HTML 1,292 183 Updated Jul 13, 2026
Next