HIP/ROCm fork optimized for AMD RDNA2 (gfx1030) with PrismML Q1_0_G128 1-bit quant support, RotorQuant, TurboQuant, EAGLE3 and P-EAGLE speculative decoding, and full Wave32 kernel optimizations.
-
Updated
Jun 23, 2026 - C++
HIP/ROCm fork optimized for AMD RDNA2 (gfx1030) with PrismML Q1_0_G128 1-bit quant support, RotorQuant, TurboQuant, EAGLE3 and P-EAGLE speculative decoding, and full Wave32 kernel optimizations.
AMD ROCm (gfx1030) inference fork with RotorQuant/TurboQuant KV compression, PHANTOM-X zero-copy draft speculation, EAGLE3 speculative decoding, 12 RDNA2 crash fixes, and PrismML Bonsai Q1_0_G128 1-bit GGUF support.
Experimental open-source compatibility research for running DLSS Neural Rendering workloads on AMD RDNA2/gfx1030. Includes ISA translation, host emulation, proof-oriented validation, and bounded hardware testing. Early development: no end-to-end game frame yet.
Exact FlashAttention for AMD RDNA2 (gfx1030 / RX 6800 XT), written in HIP. Drop-in replacement for flash-attn: 19.8 TFLOP/s forward, decode at 99% of HBM bandwidth.
Custom SPIR-V kernel factory for PHANTOM speculative decoding — LLVM IR to GPU (SPIR-V/HIP) and CPU (native x86) cross-target compilation, RDNA2/gfx1030 optimized pre-compiled kernels, dynamic kernel swapping, zero-JIT inference pipeline
ComfyUI + ZLUDA on Windows for AMD RDNA2 desktop GPUs (RX 6950/6900/6800 XT, RX 6700/6750 XT) — fixes HIP/ZLUDA version mismatches, torch DLL loading, the mem-efficient attention backend that resets the display driver, and installs rocBLAS kernels on gfx1031.
To associate your repository with the gfx1030 topic, visit your repo's landing page and select "manage topics."