Open-source book with Modern CUDA Learn Notes for Beginners, includes FP16/BF16, FP8, HGEMM, FlashAttention, CuTe, etc.
-
Updated
Sep 23, 2026 - Cuda
Open-source book with Modern CUDA Learn Notes for Beginners, includes FP16/BF16, FP8, HGEMM, FlashAttention, CuTe, etc.
Pre-compiled custom CUDA extension for Block Sparse Attention (Python 3.11 / PyTorch 2.6.0+cu124).
Generate narrated CUDA course videos with animated slides and AI avatars using Remotion, Gemini, and ElevenLabs TTS for automated production.
OpenAI Triton compiler fork for IBM POWER9/POWER10 (ppc64le) with CUDA 12.4. Tesla V100 sm70 GPU kernels for PyTorch and SGLang.
Evaluation, latency benchmarks, and voice cloning lab for OmniVoice zero-shot TTS on local desktop GPUs.
PyTorch 2.12 fork for IBM POWER9/POWER10 (ppc64le) with CUDA 12.4 and Triton. Tesla V100 sm70, GPU training and LLM inference.
To associate your repository with the cuda-12 topic, visit your repo's landing page and select "manage topics."