Skip to content
View dogansagbili's full-sized avatar

Highlights

  • Pro

Organizations

@ParCoreLab

Block or report dogansagbili

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

LLM training in simple, raw C/CUDA

Cuda 30,772 3,722 Updated Jun 26, 2025

LLM inference in C/C++

C++ 123,348 21,527 Updated Aug 10, 2026

Optimized primitives for collective multi-GPU communication

C++ 4,977 1,371 Updated Aug 10, 2026

Distributed Communication-Optimal Matrix-Matrix Multiplication Algorithm

C++ 217 33 Updated Jul 29, 2026

AI prompts for accelerating the research workflow.

192 28 Updated Jun 5, 2026
C++ 3 2 Updated Jul 10, 2026

Triton-based Symmetric Memory operators and examples

Python 110 14 Updated May 15, 2026

Uniconn is a unified, portable high-level C++ communication library that supports both point-to-point and collective operations across GPU clusters. Uniconn enables seamless switching between backe…

Cuda 3 Updated Dec 17, 2025

Modern C++ Programming Course (C++03/11/14/17/20/23/26)

HTML 15,995 1,121 Updated Apr 19, 2026

Professionally written C++ function traits library (single header-only) for retrieving info about any function (arg types, arg count, return type, etc.)

C++ 50 6 Updated Sep 3, 2025

AMD RAD's multi-GPU Triton-based framework for seamless multi-GPU programming

Python 195 46 Updated Aug 8, 2026

Perplexity open source garden for inference technology

Rust 612 66 Updated May 27, 2026

A suite of microbenchmarks developed for systems with multi-GPU per node.

Cuda 11 3 Updated Jan 22, 2026

MSCCL++: A GPU-driven communication stack for scalable AI applications

C++ 546 103 Updated Aug 10, 2026

Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels

Python 7,181 686 Updated Aug 10, 2026

Modular RDMA Interface

C++ 166 76 Updated Aug 10, 2026

torchcomms: a modern PyTorch communications API

C++ 385 175 Updated Aug 10, 2026

COCCL: Compression and precision co-aware collective communication library

C++ 38 5 Updated Aug 10, 2026

MiniAMR Adaptive Mesh Refinement (AMR) Mini-App

C 39 28 Updated Nov 12, 2024

Distributed MoE in a Single Kernel [NeurIPS '25]

Cuda 281 40 Updated May 5, 2026

A comprehensive hands-on project for learning GPU programming with CUDA and HIP, covering fundamental concepts through advanced optimization techniques.

C++ 38 4 Updated Nov 20, 2025

NVIDIA NVSHMEM is a parallel programming interface for NVIDIA GPUs based on OpenSHMEM. NVSHMEM can significantly reduce multi-process communication and coordination overheads by allowing programmer…

C++ 571 99 Updated Aug 10, 2026

VIP cheatsheet for Stanford's CME 295 Transformers and Large Language Models

4,633 665 Updated May 25, 2026

DeepEP: an efficient expert-parallel communication library

Cuda 9,972 1,370 Updated Aug 5, 2026

Perplexity GPU Kernels

C++ 601 99 Updated Nov 7, 2025

[DEPRECATED] Moved to ROCm/rocm-systems repo

C++ 145 44 Updated Aug 10, 2026

Examples demonstrating available options to program multiple GPUs in a single node or a cluster

Cuda 910 153 Updated Sep 26, 2025

RAJA Performance Portability Layer (C++)

C++ 590 114 Updated Aug 10, 2026

Lightweight C++ command line option parser

C++ 4,797 652 Updated Jul 13, 2026
Next