Skip to content
View dogansagbili's full-sized avatar

Highlights

  • Pro

Organizations

@ParCoreLab

Block or report dogansagbili

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

Optimized primitives for collective multi-GPU communication

C++ 4,905 1,355 Updated Jul 23, 2026

Distributed Communication-Optimal Matrix-Matrix Multiplication Algorithm

C++ 215 33 Updated Apr 18, 2026

AI prompts for accelerating the research workflow.

192 28 Updated Jun 5, 2026
C++ 3 2 Updated Jul 10, 2026

Triton-based Symmetric Memory operators and examples

Python 106 14 Updated May 15, 2026

Uniconn is a unified, portable high-level C++ communication library that supports both point-to-point and collective operations across GPU clusters. Uniconn enables seamless switching between backe…

Cuda 3 Updated Dec 17, 2025

Modern C++ Programming Course (C++03/11/14/17/20/23/26)

HTML 15,946 1,116 Updated Apr 19, 2026

Professionally written C++ function traits library (single header-only) for retrieving info about any function (arg types, arg count, return type, etc.)

C++ 50 6 Updated Sep 3, 2025

AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming

Python 193 44 Updated Jul 23, 2026

Perplexity open source garden for inference technology

Rust 604 64 Updated May 27, 2026

A suite of microbenchmarks developed for systems with multi-GPU per node.

Cuda 11 3 Updated Jan 22, 2026

MSCCL++: A GPU-driven communication stack for scalable AI applications

C++ 542 102 Updated Jul 25, 2026

Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels

Python 6,902 652 Updated Jul 25, 2026

Modular RDMA Interface

C++ 157 68 Updated Jul 25, 2026

torchcomms: a modern PyTorch communications API

C++ 380 165 Updated Jul 25, 2026

COCCL: Compression and precision co-aware collective communication library

C++ 38 4 Updated Jul 20, 2026

MiniAMR Adaptive Mesh Refinement (AMR) Mini-App

C 39 28 Updated Nov 12, 2024

Distributed MoE in a Single Kernel [NeurIPS '25]

Cuda 273 40 Updated May 5, 2026

A comprehensive hands-on project for learning GPU programming with CUDA and HIP, covering fundamental concepts through advanced optimization techniques.

C++ 38 4 Updated Nov 20, 2025

NVIDIA NVSHMEM is a parallel programming interface for NVIDIA GPUs based on OpenSHMEM. NVSHMEM can significantly reduce multi-process communication and coordination overheads by allowing programmer…

C++ 564 97 Updated Jul 20, 2026

VIP cheatsheet for Stanford's CME 295 Transformers and Large Language Models

4,579 662 Updated May 25, 2026

DeepEP: an efficient expert-parallel communication library

Cuda 9,888 1,338 Updated Jul 14, 2026

Perplexity GPU Kernels

C++ 594 100 Updated Nov 7, 2025

[DEPRECATED] Moved to ROCm/rocm-systems repo

C++ 146 44 Updated Jul 20, 2026

Examples demonstrating available options to program multiple GPUs in a single node or a cluster

Cuda 909 152 Updated Sep 26, 2025

RAJA Performance Portability Layer (C++)

C++ 590 113 Updated Jul 24, 2026

Lightweight C++ command line option parser

C++ 4,795 650 Updated Jul 13, 2026

Collaborative Collection of C++ Best Practices. This online resource is part of Jason Turner's collection of C++ Best Practices resources. See README.md for more information.

8,790 903 Updated Aug 6, 2024

CMake for C++ Best Practices

CMake 1,787 203 Updated Jul 25, 2026
Next