Skip to content
View dogansagbili's full-sized avatar

Highlights

  • Pro

Organizations

@ParCoreLab

Block or report dogansagbili

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

LLM training in simple, raw C/CUDA

Cuda 30,724 3,719 Updated Jun 26, 2025

LLM inference in C/C++

C++ 122,649 21,307 Updated Aug 4, 2026

Optimized primitives for collective multi-GPU communication

C++ 4,927 1,360 Updated Aug 3, 2026

Distributed Communication-Optimal Matrix-Matrix Multiplication Algorithm

C++ 217 33 Updated Jul 29, 2026

AI prompts for accelerating the research workflow.

192 28 Updated Jun 5, 2026
C++ 3 2 Updated Jul 10, 2026

Triton-based Symmetric Memory operators and examples

Python 109 14 Updated May 15, 2026

Uniconn is a unified, portable high-level C++ communication library that supports both point-to-point and collective operations across GPU clusters. Uniconn enables seamless switching between backe…

Cuda 3 Updated Dec 17, 2025

Modern C++ Programming Course (C++03/11/14/17/20/23/26)

HTML 15,972 1,120 Updated Apr 19, 2026

Professionally written C++ function traits library (single header-only) for retrieving info about any function (arg types, arg count, return type, etc.)

C++ 50 6 Updated Sep 3, 2025

AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming

Python 195 44 Updated Aug 3, 2026

Perplexity open source garden for inference technology

Rust 611 66 Updated May 27, 2026

A suite of microbenchmarks developed for systems with multi-GPU per node.

Cuda 11 3 Updated Jan 22, 2026

MSCCL++: A GPU-driven communication stack for scalable AI applications

C++ 545 102 Updated Aug 3, 2026

Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels

Python 7,118 673 Updated Aug 4, 2026

Modular RDMA Interface

C++ 165 72 Updated Aug 4, 2026

torchcomms: a modern PyTorch communications API

C++ 384 172 Updated Aug 4, 2026

COCCL: Compression and precision co-aware collective communication library

C++ 38 5 Updated Jul 20, 2026

MiniAMR Adaptive Mesh Refinement (AMR) Mini-App

C 39 28 Updated Nov 12, 2024

Distributed MoE in a Single Kernel [NeurIPS '25]

Cuda 281 40 Updated May 5, 2026

A comprehensive hands-on project for learning GPU programming with CUDA and HIP, covering fundamental concepts through advanced optimization techniques.

C++ 38 4 Updated Nov 20, 2025

NVIDIA NVSHMEM is a parallel programming interface for NVIDIA GPUs based on OpenSHMEM. NVSHMEM can significantly reduce multi-process communication and coordination overheads by allowing programmer…

C++ 567 98 Updated Jul 29, 2026

VIP cheatsheet for Stanford's CME 295 Transformers and Large Language Models

4,622 663 Updated May 25, 2026

DeepEP: an efficient expert-parallel communication library

Cuda 9,943 1,359 Updated Aug 4, 2026

Perplexity GPU Kernels

C++ 597 100 Updated Nov 7, 2025

[DEPRECATED] Moved to ROCm/rocm-systems repo

C++ 145 44 Updated Jul 20, 2026

Examples demonstrating available options to program multiple GPUs in a single node or a cluster

Cuda 909 153 Updated Sep 26, 2025

RAJA Performance Portability Layer (C++)

C++ 590 114 Updated Aug 3, 2026

Lightweight C++ command line option parser

C++ 4,797 653 Updated Jul 13, 2026
Next