Skip to content

feat: Apple Silicon (Metal) deployment guide #32

Description

@itigges22

Description

llama.cpp supports Metal for Apple Silicon acceleration. We need a deployment guide and potentially a Dockerfile for macOS users.

Requirements

  • Test Qwen3.5-9B-Q6_K on Apple Silicon (M1/M2/M3/M4 with 16GB+ unified memory)
  • Document bare-metal setup for macOS (Homebrew llama.cpp + Python services)
  • Verify self-embeddings work via Metal backend
  • Benchmark generation speed vs CUDA
  • Add macOS section to docs/SETUP.md
  • Consider Docker Desktop for Mac limitations (no GPU passthrough) — may need bare-metal-only guide

Context

Many developers run macOS. Metal support in llama.cpp is mature. The main question is whether Docker or bare-metal is the right approach (Docker Desktop doesn't support GPU passthrough on macOS).

Activity

  1. itigges22 commented on May 12, 2026

    @itigges22
    CollaboratorAuthor

    Status — on the V3.1.1 roadmap. Apple Metal is one of the three V3.1.1 deliverables (alongside macOS/Windows installers via #7-style accelerator work, and the formal 9B benchmark pass). Hard-prereq is the macOS installer landing first so the rest of the stack can run; ATLAS is a Docker-Compose-first install today and macOS users currently have to manually run bare-metal mode.

  2. itigges22 commented on May 19, 2026

    @itigges22
    CollaboratorAuthor

    partial coverage — Vulkan-via-MoltenVK landed in #114 (f88a717) so mac users have a docker path now... but it's slow. MoltenVK translates vulkan calls to metal at runtime, so you're paying double overhead (vulkan abstraction + the translation layer) on top of docker desktop's hypervisor cost. fine for "i wanna try atlas on my macbook to see if it works", not the path you'd actually want to run a real workload on.

    the right end state for mac is still a native install that bypasses docker entirely:

    • brew install for system deps (cmake, ninja, the apple developer command-line tools)
    • uv for the python side so atlas-cli + the geometric-lens trainer install cleanly into a venv
    • llama.cpp built locally with -DGGML_METAL=ON against the system metal-cpp headers, no container in the path
    • a setup-macos.sh that wraps all of the above plus an atlas init --backend metal flag
    • atlas doctor would need a _check_metal_native() path that just verifies the metal device shows up via system_profiler SPDisplaysDataType

    that gets us actual M-series perf (the unified memory architecture on M3/M4 is genuinely good for 9B-scale inference) and skips the docker desktop tax entirely.

    tied to #115 on the container side too, since the vulkan/MoltenVK docker fallback needs arm64 multi-arch builds to actually work on Apple Silicon hosts. for now if you're on a mac and want to try atlas, the vulkan compose file works on intel macs only... arm64 mac support waits on #115 + this issue's native path.

  3. itigges22 commented on May 22, 2026

    @itigges22
    CollaboratorAuthor

    shipped + validated! end-to-end install just ran clean on a 32GB M2 Pro mbp.

    what's in:

    • scripts/atlas-setup-macos.sh — installs Xcode CLT + brew deps (cmake, git, python@3.12, pipx, go), pulls llama.cpp at the pinned SHA, applies the PC-202 patch + spec-decode sed, builds with -DGGML_METAL=ON -DGGML_METAL_USE_BF16=ON, installs the binary to ~/.atlas/macos/bin/llama-server-metal. Also handles atlas-cli install via pipx (dodges Homebrew Python's PEP 668 enforcement which blocks plain pip install) + builds atlas-tui via go.
    • scripts/atlas-llama-macos.sh — foreground launcher, same flags as entrypoint-v3.1-9b.sh so the runtime is identical to the linux+cuda/rocm path.
    • docker-compose.macos.yml — overlay that swaps llama-server for a tiny alpine/socat forwarding llama-server:8080 -> host.docker.internal:8080. zero changes to the base compose file or the 4 service containers.
    • atlas init --backend metal (auto-recommended on darwin+aarch64 hosts), atlas doctor check_metal_native, fresh docs/SETUP_MACOS.md walkthrough, new hardware-nav table at the top of SETUP.md.

    what i hit during the actual install (all fixed in dev):

    • atlas init on Mac before docker compose up -> auto-launched local proxy collided with the docker stack on :8090 (fixed in 6d253f9 by detecting the docker stack + waiting)
    • TUI launch via execv orphaned the local proxy because atexit didn't fire (fixed in 6d253f9 by calling _stop_local_proxy() explicitly before execv)
    • alpine/socat doesn't ship netcat so the healthcheck failed -> swapped to grep socat /proc/1/cmdline (14aa0b2)
    • compose overlay inherited the base file's 8080:8080 port publish + collided with native llama-server on :8080 -> added ports: !reset [] (14aa0b2)
    • atlas doctor's check_metal_native gated on --help returning 0 but llama-server returns 1 by convention -> now checks for usage markers in output (593522d)
    • Lens + ASA artifacts weren't auto-downloading because the registry knew which files SHOULD exist but had no URL for them -> wired lens_artifact_url_base + asa_artifact_url_base pointing at itigges22/ATLAS dataset, atlas model install now fetches them after the gguf (40987b9)

    doctor output on the validated install: 22 passed, 1 info-warn (cpu_cores < 16 on M2 Pro for xlarge tier... cosmetic, inference is GPU-bound on Apple Silicon), e2e_smoke green.

    closing! follow-ups already filed:

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    help wantedExtra attention is needed

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions