-
Sandia National Laboratories
- https://vlkale.github.io
- @vivek_lkale
Lists (5)
Sort Name ascending (A-Z)
Starred repositories
Caliper is an instrumentation and performance profiling library
Experimental Explicit Communications API for Kokkos
Userspace eBPF runtime for Observability, Network, GPU & General Extensions Framework
Reverse Engineering: Decompiling Binary Code with Large Language Models
HIP: C++ Heterogeneous-Compute Interface for Portability
Development repository for the Triton language and compiler
Anthropic's original performance take-home, now open for you to try!
A Python-embedded DSL that makes it easy to write fast, scalable ML kernels with minimal boilerplate.
The NVIDIA NeMo Agent toolkit is an open-source library for efficiently connecting and optimizing teams of AI agents.
Code, labs, and resources for O'Reilly AI Systems Performance Engineering: GPU optimization, distributed training, inference scaling, and full-stack tuning.
A multi-level dataflow tracer for capturing I/O calls from workflows.
Explain Compiler Explorer output using AI
Run compilers interactively from your web browser and interact with the assembly
NVIDIA NVSHMEM is a parallel programming interface for NVIDIA GPUs based on OpenSHMEM. NVSHMEM can significantly reduce multi-process communication and coordination overheads by allowing programmer…
OSV-SCALIBR: A library for Software Composition Analysis
Claude Code for CUDA. Free AI assistant that actually understands GPU architecture
AITemplate is a Python framework which renders neural network into high performance CUDA/HIP C++ code. Specialized for FP16 TensorCore (NVIDIA GPU) and MatrixCore (AMD GPU) inference.
The Torch-MLIR project aims to provide first class support from the PyTorch ecosystem to the MLIR ecosystem.
FB (Facebook) + GEMM (General Matrix-Matrix Multiplication) - https://code.fb.com/ml-applications/fbgemm/
Must read research papers and links to tools and datasets that are related to using machine learning for compilers and systems optimisation
abouteiller / rocSHMEM
Forked from ROCm/rocSHMEMrocSHMEM intra-kernel networking runtime for AMD dGPUs on the ROCm platform.
Python implementation of algorithms from Russell And Norvig's "Artificial Intelligence - A Modern Approach"