Highlights
- All languages
- ASL
- Assembly
- BitBake
- C
- C#
- C++
- CSS
- Clojure
- Crystal
- Cuda
- Dockerfile
- Elm
- F#
- Go
- Go Template
- Groovy
- HCL
- HTML
- Haskell
- Java
- JavaScript
- Jupyter Notebook
- Lua
- Makefile
- Markdown
- Nix
- Objective-C
- PHP
- Perl
- PowerShell
- Python
- Ruby
- Rust
- SCSS
- Shell
- Svelte
- TypeScript
- Vala
- Verilog
- Vim Script
- Vue
- Wikitext
- Zig
Starred repositories
Fast and accurate DRAM power and energy estimation tool
Train GPT-2 large language model (150M to 774M) series on Google Colab using K80 GPUs
AMD ROCm AI Inference & Training Solutions Library for RNDA2/GFX1030/Radeon Pro v620 includes: compiling vLLM Source, Llama.cpp, HIPFire, and Fine-Tuning paths.
Slurm on Kubernetes Architecture Solution for Fine-tuning LLMs, Inference, and Eval across distributed NVIDIA L40s X8 GPU Cluster on Nebius AI Cloud
HIP: C++ Heterogeneous-Compute Interface for Portability
A high-throughput and memory-efficient inference and serving engine for LLMs
The main repository for building Pascal-compatible versions of ML applications and libraries.
A fast high-compression read-only file system for Linux, FreeBSD, macOS and Windows
Achieve state of the art inference performance with modern accelerators on Kubernetes
Kimi Code CLI is your next CLI agent.
Jobs scraper library for LinkedIn, Indeed, Glassdoor, Google, ZipRecruiter & more
Qwen3-Coder is the code version of Qwen3, the large language model series developed by Qwen team.
An open-source AI coding agent that lives in your terminal.
A fast JSON parser/generator for C++ with both SAX/DOM style API
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
Distributed reliable key-value store for the most critical data of a distributed system
The etcd-cpp-apiv3 is a C++ library for etcd's v3 client APIs, i.e., ETCDCTL_API=3.
llama.cpp fork with additional SOTA quants and improved performance
Diffusion model(SD,Flux,Wan,Qwen Image,Z-Image,...) inference in pure C/C++
Home for "How To Scale Your Model", a short blog-style textbook about scaling LLMs on TPUs