- All languages
- Assembly
- C
- C#
- C++
- CSS
- Clojure
- Cuda
- Dart
- GDScript
- Go
- HLSL
- HTML
- Handlebars
- Haskell
- Java
- JavaScript
- Jupyter Notebook
- Just
- KiCad Layout
- Kotlin
- LLVM
- Lua
- MDX
- MLIR
- Makefile
- Markdown
- Mojo
- Nix
- Objective-C++
- PHP
- Prolog
- Python
- QML
- Ruby
- Rust
- SCSS
- Scala
- Shell
- Smali
- Smarty
- Solidity
- Starlark
- Svelte
- Swift
- TeX
- TypeScript
- Verilog
- Vim Script
- Vue
Starred repositories
A no-fluff and highly practical masterclass that reignites engineering curiosity and helps SDE-2, SDE-3, and above become great at designing, implementing, and shipping scalable, fault-tolerant, an…
NVIDIA Linux open GPU kernel module source
Multi-agent deep research for Claude Code. Zero API keys. Paste one line, type /research.
Professional Claude Code skills marketplace featuring production-ready skills for enhanced development workflows.
General-purpose deep research skill for AI agents — subagent-driven, source-backed, cited reports inline. skills.sh-compliant.
A curated list of awesome skills for Cursor
Deploy autonomous AI agents as your digital twins across 10 social platforms
Up to 3× faster LLM decoding on Apple Silicon, lossless. Native MLX port of DeepSeek's DSpark & z-lab's DFlash speculative decoding — Gemma-4, Qwen3, Ornith-1.0, ternary Bonsai-27B.
Build a compiler to solve Anthropic's interview challenge.
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
Inferno aims to be a super lightweight, highly efficient Rust inference engine for running open weights models on Apple Silicon with Metal, targeting machines such as a MacBook Pro with 64 GB of un…
Run frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦
LiteRT, successor to TensorFlow Lite. is Google's On-device framework for high-performance ML & GenAI deployment on edge platforms, via efficient conversion, runtime, and optimization
High-efficiency floating-point neural network inference operators for mobile, server, and Web
A close-to-metal Python API for programming AMD Ryzen™ AI NPUs (AI Engines), built on an open-source MLIR-based compiler toolchain.
ypapadop-amd / ggml
Forked from ggml-org/ggmlTensor library for machine learning
Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails. Self-hostable. Apache 2.0.
A converter for transferring gguf Q4_0, Q4_1 to FLM Q4NX
Run LLMs on AMD Ryzen™ AI NPUs in minutes. Just like Ollama - but purpose-built and deeply optimized for the AMD NPUs.
Awesome list and survey website for agents in the era of experience
A feed-forward 3D foundation model for reconstructing scenes from streaming data
DeepSpec: a full-stack codebase for training and evaluating speculative decoding algorithms