vLLM + Qwen3.6-27B (BF16) OpenAI-compatible inference server on AMD Strix Halo (Ryzen AI Max+ 395, gfx1151). Vision input, 256K context, /v1/responses with separated reasoning, via TheRock ROCm.
-
Updated
Apr 26, 2026 - Python
vLLM + Qwen3.6-27B (BF16) OpenAI-compatible inference server on AMD Strix Halo (Ryzen AI Max+ 395, gfx1151). Vision input, 256K context, /v1/responses with separated reasoning, via TheRock ROCm.
Open-source bring-up + verified matmul on the first-gen AMD XDNA1 (Phoenix/Hawk Point) NPU on Linux — the gen FastFlowLM/Lemonade skip. RyzenAI-npu1, mlir-aie/IRON, XRT.
Xian-VL monorepo - Multilingual Assistant for Gaming Environments 🧙♂️ and other AI translation/research tools. Powered by Lemonade. 🍋
AMD ISP4 camera driver module for Ryzen AI laptops
llama.cpp + Qwen3.6-27B (Q8_0 GGUF) OpenAI-compatible inference server on AMD Strix Halo (Ryzen AI Max+ 395, gfx1151). 256K context, ~7.5 t/s decode via TheRock ROCm Docker.
Fix internal speakers on the ASUS ProArt PX13 (HN7306, Strix Halo) under Linux: TAS2783 firmware + ALSA UCM Speaker profile + ACP/SoundWire reset for boot and suspend/resume.
ComfyUI on AMD Strix Halo (RDNA 3.5 / gfx1151) via Docker. Ubuntu 26.04 LTS + uv-managed Python 3.12 + pinned TheRock ROCm 7.13 wheels. Fixes the silent CPU fallback Debian / Python 3.13 images hit on gfx1151.
Stable Diffusion image generation on AMD Ryzen AI NPUs for Linux
Talos-O (Omni): A sovereign, embodied agentic organism forged on AMD Strix Halo. Integrating the Chimera Kernel (Linux 7.0), Zero-Copy Introspection, and the Phronesis Engine. Built from First Principles.
ROCmFPX llama.cpp fork for Windows 🏆 — native build, headless OpenAI-compatible server & benchmarks. Tested on AMD Strix Halo (gfx1151), runs on other GPUs too.
Docker infrastructure for AMD Strix Halo (RDNA 3.5 / gfx1151): PyTorch + ROCm base container and a separate Ollama LLM service. Two folders, two Compose files, one Strix Halo box.
A lightweight TUI monitor for AMD Ryzen AI NPUs
ComfyUI for AMD Strix Halo (gfx1151) on Ubuntu 26.04 via distrobox + rootless podman
Text embeddings on AMD XDNA NPU (Strix Halo) — OpenAI-compatible /v1/embeddings endpoint powered by all-MiniLM-L6-v2 running on the NPU tile array
Experiments, notes, and benchmarks for AMD Ryzen AI (Krackan Point NPU + Radeon iGPU) on Linux
Verified, reproducible recipe + tools to run real compute (i32/bf16 matmul) on a first-gen AMD Ryzen AI NPU (XDNA1 / Phoenix, e.g. 7840U) on Linux via iree-amd-aie — gotchas, benchmarks, 5-language docs
A fully offline, air-gapped desktop AI workspace for legal, medical, and financial professionals. Runs entirely on the AMD Ryzen NPU with zero cloud calls and zero telemetry. All AI processing happens locally so no data leaves the device, ensuring complete privacy and confidentiality.
Add a description, image, and links to the ryzen-ai topic page so that developers can more easily learn about it.
To associate your repository with the ryzen-ai topic, visit your repo's landing page and select "manage topics."