Reverse engineering NVIDIA SASS instruction dictionary, kernel audits and pattern recognition across GPU architectures.
-
Updated
May 18, 2026 - Cuda
Reverse engineering NVIDIA SASS instruction dictionary, kernel audits and pattern recognition across GPU architectures.
Pre-built wheels for llama-cpp-python across platforms and CUDA versions
A cleaner, modernized reboot of Real-ESRGAN focused on best practices, up-to-date hardware support, and pragmatic developer ergonomics. Supports RTX 50-series (Blackwell) GPUs with automatic device selection, CPU fallback, and improved inference pipelines for both images and videos.
Building PaddlePaddle Inference C++ from source for NVIDIA Blackwell (RTX 50-series, sm_120) with CUDA 13 and TensorRT 10
Prebuilt spconv v2.3.8 wheels for CUDA 12.8 / 13.0 with native Blackwell (RTX 50-series, sm_120) kernels — for default PyPI torch (cu130) or torch +cu128
A lightweight wrapper for the Chatterbox SoTA TTS providing an OpenAI-compatible API schema.
LLM Inference Optimization for RTX 5070 Ti (Blackwell) - Benchmarks, quantization, serving configs for vLLM/SGLang/TensorRT-LLM/llama.cpp
To associate your repository with the blackwell-architecture topic, visit your repo's landing page and select "manage topics."