Skip to content

ARM64 builds: CUDA on aarch64, the small service images, and device validation #115

Description

@itigges22

Why

Every current Dockerfile targets x86_64 only. As ARM64 hardware proliferates the install matrix needs to follow... NVIDIA DGX Spark (GB10 Grace + Blackwell, ARM CPU + Blackwell GPU, shipping Q1-Q2 2026), Snapdragon X Elite laptops (Adreno GPU + Hexagon NPU), Apple Silicon containers, Jetson Orin/Thor, Raspberry Pi 5, AWS Graviton... none of these can even pull our images today.

Pairs with PC-114 (#114). Vulkan covers the GPU diversity axis, ARM64 covers the CPU arch axis... together they unlock basically every modern compute target outside the macOS-native install path (#32).

What changes

  1. .github/workflows/build-images.yml gets docker buildx + QEMU setup, builds for linux/amd64,linux/arm64, pushes multi-arch GHCR manifests. Users keep pulling the same tag and get the right arch automatically.

  2. Per-Dockerfile arm64 audit:

    • Dockerfile.v31 (CUDA): NVIDIA ships nvidia/cuda:*-ubuntu22.04 arm64 variants for CUDA 12.x (GB10 is on this line). Need to verify the exact tag we pin has arm64 + that the patched llama.cpp build works on aarch64 CUDA.
    • Dockerfile.rocm: AMD doesn't ship arm64 ROCm at all currently. Skip arm64 for this image until/unless that changes.
    • Dockerfile.vulkan (feat: Vulkan universal backend (one image, covers NVIDIA + AMD + Intel + Apple + CPU) #114): ubuntu:22.04 + mesa-vulkan-drivers, both arm64-native. Should Just Work but needs validation.
    • proxy/Dockerfile: Go binary, just GOOS=linux GOARCH=arm64 in the buildx target.
    • geometric-lens/Dockerfile, v3-service/Dockerfile, sandbox/Dockerfile: Python + apt, all arm64-native on Ubuntu.
  3. tier.py: detect_gpu() is mostly arch-agnostic but the nvidia-smi/rocm-smi output parsers might trip on the GB10's Grace+Blackwell topology... verify on actual hardware when available.

  4. atlas doctor: docker-pull tests should add --platform when a specific service has only amd64 (so users on arm64 hosts get a clear 'no arm64 build for X' message instead of pulling the wrong arch and crashing).

  5. Docs: SETUP.md gets an arm64 matrix table (which Dockerfile supports which arch per release).

What this unlocks

Hardware Today After this
NVIDIA DGX Spark (GB10, arm64) can't pull images CUDA install
Snapdragon X Elite laptop can't pull images Vulkan install via Adreno
Apple Silicon (Docker route) qemu CPU emulation (very slow) native arm64 containers (still MoltenVK for GPU but no CPU emu)
Jetson Orin / Thor can't pull images CUDA install
Pi 5 (8GB) can't pull images Vulkan via V3DV (slow but boots)
Ampere/Graviton cloud can't pull images CPU-only Vulkan lavapipe

Hardware testing matrix

Each combo needs a tester...

  • DGX Spark + CUDA: waiting on hardware (NVIDIA Q1-Q2 2026)
  • Apple Silicon + Vulkan-in-Docker: need a Mac dev
  • Snapdragon X Elite + Vulkan: need a Snapdragon laptop owner
  • Jetson Orin + CUDA: any Jetson user willing
  • Pi 5 + Vulkan: cheap, can probably pick one up

Out of scope

  • Per-arch performance tuning (one ticket per arch if perf is bad)
  • Apple Silicon NATIVE install path (feat: Apple Silicon (Metal) deployment guide #32) is the no-Docker fast path, tracked separately
  • Windows ARM64 is a separate Docker-on-Windows story

Dependencies

Status update (2026-09-28)

Done: llama-vulkan builds for linux/arm64 on release tags (#116), and
3.1.4 shipped it. Remaining:

  • a CUDA llama image for aarch64 (sbsa or l4t base);
  • arm64 for the four small service images (proxy, v3, lens, sandbox),
    which is a matrix change once a device validates end to end;
  • an architecture table in SETUP.md;
  • one end-to-end run on a real arm64 device. SUPPORT_MATRIX.md lists
    Linux/arm64 as Preview until then.

Activity

  1. itigges22 commented on May 22, 2026

    @itigges22
    CollaboratorAuthor

    arm64 detection validated on real hardware today — lead's 32GB M2 Pro mbp showed up cleanly through the whole pipeline:

    ✓ arch    aarch64 (Apple Silicon) — Metal hybrid path supported (#32)
    ✓ gpu     [apple] Apple M2 Pro (32.0 GB unified) — Metal hybrid path supported (#32)
    

    tier.arch_detect() correctly returned aarch64, atlas init correctly recommended the Metal hybrid path instead of trying to write ATLAS_BACKEND=rocm (which would have collided with the AMD-on-arm64 carve-out from this ticket). So at least one of the five target devices (Apple Silicon) is now end-to-end validated.

    still gated for the other four:

    • DGX Spark (Grace Blackwell sbsa) — would validate the BUILDER_IMAGE override path in Dockerfile.v31
    • Snapdragon X Elite — would validate the Adreno vulkan driver path
    • Jetson Orin — would validate the l4t-jetpack base image swap
    • Pi 5 — would validate the lavapipe CPU fallback under the arm64 buildx (ci: add linux/arm64 to buildx matrix for Dockerfile.vulkan #116 is the prereq there)

    prebuilt arm64 ghcr images still not published... that's #116 (arm64 buildx matrix). caching for those is #117.

    keeping this open until at least one of the four arm64-Linux devices reports in. if anyone has hardware in any of those categories please drop output here.

  2. itigges22 commented on Jul 2, 2026

    @itigges22
    CollaboratorAuthor

    Status (2026-07-02): partially delivered via #116 — build-images.yml builds llama-vulkan for linux/arm64 on tag pushes (QEMU; amd64-only on dev pushes to keep iteration fast). Remaining: CUDA/aarch64 base-image variant (sbsa/l4t), arm64 for the four small service images (one-line matrix change once an arm64 device validates end-to-end), and a SETUP arch matrix. SUPPORT_MATRIX.md currently lists Linux/arm64 as Preview for exactly this reason — no end-to-end device validation yet. Contributors: don't redo the vulkan half.

  3. changed the title [-]feat: ARM64 multi-arch builds (DGX Spark, Snapdragon X Elite, Apple Silicon, Jetson, Pi 5)[/-] [+]ARM64 builds: CUDA on aarch64, the small service images, and device validation[/+] on Sep 28, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions