Agenium Scale vectorization library for CPUs and GPUs
-
Updated
Oct 21, 2021 - C
Agenium Scale vectorization library for CPUs and GPUs
A high-performance RCCL / NCCL (ROCm Communication Collectives Library) plugin for Thunderbolt 5 that enables GPU-to-GPU communication across Thunderbolt connections with RDMA support.
A modular program analysis tool framework for accelerators (NVIDIA, AMD, and DL workloads).
GPU-accelerated ARC4 40-bit key recovery for DMR Enhanced Privacy. CUDA (NVIDIA) and ROCm/HIP (AMD). Windows GUI + Linux CLI.
Tensor-parallel DeepSeek V4 Flash inference on dual AMD Strix Halo over OdinLink USB4/TB5 RDMA or Mellanox RoCE v2.
DeepSeek V4 Flash 284B on AMD Strix Halo (gfx1151) — up to 32 tok/s decode & ~250 tok/s prefill via ROCmFPX, DSpark & ROCm 7.2
A generic thermal observability framework for CPU, GPU, board, and platform telemetry across vendor APIs, kernel interfaces, and runtime correlation layers.
Gen 1 C host substrate for openOODA. 26 capability tokens, Landlock sandbox, PQC (ML-KEM-768 + ML-DSA-65), GPU HIP/ROCm.
Lightweight, dependency-free terminal UI for monitoring AI training runs. Pure C, no ncurses, no Python. Pluggable profiles (BasicSR, Lightning, HuggingFace, custom) and multi-GPU backend (AMD sysfs, NVIDIA nvidia-smi).
To associate your repository with the rocm topic, visit your repo's landing page and select "manage topics."