Stars
Modern CUDA Learn Notes with PyTorch for Beginners, 200+ CUDA Kernels, Tensor Cores, HGEMM, FA-2 MMA.
A Easy-to-understand TensorOp Matmul Tutorial
Samples for CUDA Developers which demonstrates features in CUDA Toolkit
LLMPerf is a library for validating and benchmarking LLMs
A high-throughput and memory-efficient inference and serving engine for LLMs
DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
校招、秋招、春招、实习好项目!带你从零实现一个高性能的深度学习推理库,支持大模型 llama2 、Unet、Yolov5、Resnet等模型的推理。Implement a high-performance deep learning inference library step by step
A list of awesome compiler projects and papers for tensor computation and deep learning.
Tensors and Dynamic neural networks in Python with strong GPU acceleration
Ongoing research training transformer models at scale
Development repository for the Triton language and compiler
🔥🔥超过1000本的计算机经典书籍、个人笔记资料以及本人在各平台发表文章中所涉及的资源等。书籍资源包括C/C++、Java、Python、Go语言、数据结构与算法、操作系统、后端架构、计算机系统知识、数据库、计算机网络、设计模式、前端、汇编以及校招社招各种面经~
Implementation of our NeurIPS 2021 paper "A Bi-Level Framework for Learning to Solve Combinatorial Optimization on Graphs".
GraphMAE: Self-Supervised Masked Graph Autoencoders in KDD'22
Awesome machine learning for combinatorial optimization papers.
Official code for the ICML2022 paper -- GNNRank: Learning Global Rankings from Pairwise Comparisons via Directed Graph Neural Networks
Attention based model for learning to solve different routing problems
Rosetta: A Realistic High-level Synthesis Benchmark Suite for Software Programmable FPGAs (FPGA'18)
⏰ Agenticly track worldwide conference deadlines (Website, Python Cli, Wechat Applet)