Starred repositories
Step-by-step optimization of CUDA SGEMM
My learning notes for ML SYS.
[EMNLP'23, ACL'24] To speed up LLMs' inference and enhance LLM's perceive of key information, compress the prompt and KV-Cache, which achieves up to 20x compression with minimal performance loss.
[ICML 2024] LLMCompiler: An LLM Compiler for Parallel Function Calling
A guidance language for controlling large language models.
This repository contains tutorials and examples for Triton Inference Server
You like pytorch? You like micrograd? You love tinygrad! ❤️
A retargetable MLIR-based machine learning compiler and runtime toolkit.
Learn how to design large-scale systems. Prep for the system design interview. Includes Anki flashcards.
An open-source efficient deep learning framework/compiler, written in python.
⭐ A Curated List of Awesome WebAssembly Applications
A Software Framework for Neuromorphic Computing
Learn about the Neumorphic engineering process of creating large-scale integration (VLSI) systems containing electronic analog circuits to mimic neuro-biological architectures.
A list of awesome compiler projects and papers for tensor computation and deep learning.
Reinforcement learning environments for compiler and program optimization tasks
Infrastructure for Machine Learning Guided Optimization (MLGO) in LLVM.
Code and documentation to train Stanford's Alpaca models, and generate the data.
An Extensible Toolkit for Finetuning and Inference of Large Foundation Models. Large Models for All.
Development repository for the Triton language and compiler
A listing of compiler, language and runtime teams for people looking for jobs in this area
All Coursework from my CS61c (Great Ideas in Computer Architecture / Machine Structures) Course at UC Berkeley