Everything in this project except DepthMap is a stage upstream AliceVision runs on the CPU for
every vendor. The GPU matcher, GPU SIFT, the depth-map filter, the meshing votes and min-cut and
visibilities, the texturing - NVIDIA users do not get any of it from upstream either. The source
was written in CUDA dialect and compiled as HIP, so it should build as CUDA with no porting at all.
This is the work to find out, and to answer one open question properly: does the memory bridge earn its place on a vendor whose driver already migrates pages?
Not a preference - a ceiling. CUDA 13's release notes: "Removed support for Maxwell, Pascal, and Volta GPUs, corresponding to compute capabilities earlier than 7.5." Both test cards are compute 6.1, so the build has to stay on 12.x.
This section said 12.9.2 until 2026-09-21, which was wrong: the toolkit every measurement in this
document was taken with reports V12.9.86, and that is 12.9.1's compiler (12.9.0 ships
12.9.41). Nothing under 12.9.2/ on NVIDIA's download host resolves, while 12.9.1's installers do.
The number in the prose was never the number on disk.
Verified rather than assumed - nvcc -arch=sm_61 compiles on the installed toolkit:
Cuda compilation tools, release 12.9, V12.9.86
sm_61: OK
| test cards | GTX 1080 Ti (11 GB, house-pc, Linux), GTX 1050 Ti (4 GB, bench-pc, Windows) |
| both | Pascal, compute 6.1, so one sm_61 target covers them |
| drivers | 580.173.02 and 581.57, both far above 12.9's floor |
No driver rollback is needed. CUDA's minor-version compatibility means a newer driver runs an older runtime, and both boxes report a maximum of CUDA 13.0, so a 12.9 build runs as-is.
Pascal sm_61 |
our host compiler | |
|---|---|---|
| CUDA 13.x | removed | MSVC 14.50 / GCC 13 fine |
| CUDA 12.9.1 | supported | Windows: _MSC_VER < 1950, i.e. MSVC 14.1x-14.4x |
WSL Ubuntu 24.04 ships GCC 13.3 and 12.9 accepts GCC 6.x-14.x, so the Linux side has no friction at all. Windows needs VS 2022 Build Tools installed alongside VS 2026. That is why Linux went first.
The exact bound, measured on 2026-09-21 rather than taken from the release notes -
crt/host_config.h reads #if _MSC_VER < 1910 || _MSC_VER >= 1950, so 12.9 accepts 1910 to
1949, which is MSVC 14.1x through 14.4x. This document said "MSVC 193x only", which was too
strict by a whole toolset generation:
| toolset | cl |
_MSC_VER |
nvcc 12.9 |
|---|---|---|---|
| VS 2022 Build Tools 14.44 | 19.44.35229 | 1944 | compiles sm_61 |
| VS 2026 Build Tools 14.50 | 19.50.35729 | 1950 | rejected: "Only the versions between 2017 and 2022 (inclusive) are supported" |
So the machine's own toolset misses by exactly one version, and VS 2022 Build Tools with the
14.44 toolset is both necessary and sufficient. -allow-unsupported-compiler would override the
check, but that is a way to get a build rather than a working one.
A risk to smoke-test early on the Windows side rather than discover at final link: the prebuilt
vcpkg dependencies were built with MSVC 14.50, and linking them against .cu objects compiled
through an MSVC 193x host leans on Microsoft's binary-compatibility guarantee holding across that
boundary.
The reference runs used Meshroom 2023.3's own prebuilt binaries - CUDA 11.3 on Linux, 11.6 on
Windows. Matching those would mean installing GCC 10 and a VS 2019-era MSVC, and it still would not
give bit-identity, because Meshroom ships AliceVision ~3.0 and this tree is 3.4-dev plus the patches
in patches/. Different source.
It is also not the comparison worth having. README.md already records that CUDA is not
bit-identical to CUDA across versions: the same 1080 Ti under 11.6/Windows and 11.3/Linux differs on
4 of 41 views, one by 7.2 % of pixels.
The clean measurement is this tree built as CUDA against this tree built as HIP - same source, same patches, same machine, only the backend differing. Nobody else can make that comparison. The Meshroom reference stays for the user-facing "faster than what you have today" number, with the version drift stated.
The GPU sources ported in v0.2.5 through v0.2.16 are CUDA dialect throughout. The HIP build reaches them by force-including the compat header, and that is applied by generator expression:
add_compile_options("$<$<COMPILE_LANGUAGE:HIP>:-include.../cheshire/cuda_to_hip.h>")COMPILE_LANGUAGE:HIP - so a CUDA build never receives it and compiles the original dialect
directly. The references to cuda_to_hip.h inside src/aliceVision/fuseCut/gpu/* are comments
explaining this, not includes.
hip/compat/include/cheshire/bridge.h says of itself: "Header-only, HIP-only." It includes
<hip/hip_runtime.h> and calls hipMalloc, hipMemGetInfo, hipMallocPitch, hipHostFree
directly. Porting it is a thin shim over about ten entry points - a day, not a project.
Whether it is needed is the interesting part, and the answer is not obvious in our favour.
docs/02 finding #2 is the bridge's whole reason for existing:
managed=0:hipMallocManagedsucceeds but behaves like mapped host memory (26 GB/s, no page migration into VRAM). Do not build the bridge on managed memory.
That is an AMD fact. CUDA's unified memory does migrate pages, with hardware fault handling, at page granularity - finer than the bridge's whole-buffer placement by class. And a GTX 1050 Ti already ran the 41-view set in 4 GB through upstream CUDA with no out-of-memory.
So the order is: build without the bridge, find where memory actually binds on 11 GB and on 4 GB, and port the bridge only if something binds. Then the comparison is buffer-class placement against page migration, which is a real question with a publishable answer either way - page migration can thrash a volume that SGM sweeps repeatedly, and whole-buffer placement cannot.
Building it first and measuring afterwards would answer the wrong question.
The build needed no porting, exactly as predicted: 674/674 steps, zero failures, all 14 .cu
files compiled unmodified. On house-pc's GTX 1080 Ti, 41 views, --downscale 2:
depth maps vs ref-cuda113 |
41 / 41 byte-identical |
sim maps vs ref-cuda113 |
41 / 41 byte-identical |
| wall time | 211 s vs the reference's 379 s (1.80x) |
Byte-identical, not "within tolerance". This tree built with CUDA 12.9 for sm_61 reproduces
Meshroom 2023.3 (AliceVision 3.1, CUDA 11.3) exactly, which means three things that were open
questions until now:
- every patch this project carries is numerically neutral on CUDA - provably, not plausibly;
- AliceVision 3.1 vs 3.4-dev contributes zero to depth-map numerics;
- CUDA 11.3 vs 12.9 contributes zero on Pascal.
The 1.80x is pure scheduling - one process with 16 simultaneous tiles against Meshroom's four sequential chunks - not a numerical shortcut. The output is the same bytes.
With the CUDA side pinned to the reference bit-for-bit, HIP-vs-CUDA is now unambiguous: same source, same patches, same machine, only the backend differing.
| our CUDA vs our HIP, matching parameters | |
|---|---|
| mask agreement | 41 / 41 views >= 0.95 |
| median relative depth error | 0.0000 |
| median view within 1 % | 0.9806 |
| worst view | 0.9233 |
| max p95 | 6.2 % |
docs/04 guessed this residual was AliceVision version drift plus FMA/texture-filter
differences. The version half is now disproven: it is entirely AMD hardware and compiler float
behaviour.
The first comparison scored 0.823 median-view within 1 %, worse than the HIP build managed against the same reference - backwards, and it looked like a real defect in the CUDA port. Four experiments chased it, each rebuilding and rerunning the 41 views:
| experiment | result |
|---|---|
| rerun the same build | 41 / 41 bit-identical - deterministic, not a race |
CHESHIRE_SGM_LEGACY (upstream's three-kernel SGM) |
41 / 41 bit-identical to the fused kernel |
buffer-copy mip levels instead of surf2Dwrite |
41 / 41 bit-identical |
| submodule + patch audit across v0.2.5..HEAD | depth-map estimation unchanged |
None of them moved a single byte. The actual cause was sgmDepthListPerTile: Meshroom's
node default is 1, the bare CLI default is 0, and the reference plus the HIP control were both
produced through Meshroom's full parameter set while the CUDA run was invoked with four
arguments. Comparing a defaults run against a Meshroom-parameters reference was never
apples-to-apples, and the flag changes how SGM samples depth.
The blunt lesson: scripts/linux/run-depthmap.sh already encodes the node's standard
preset, --sgmDepthListPerTile True included, and it is what produced every validated HIP
run in docs/04. The CUDA test hand-rolled its own four-argument invocation instead of using
it. Use the runner.
The recovery lesson, for when a comparison does look wrong: diff the two runs' own parameter
dumps before diffing their output. AliceVision prints every parameter with a (default)
marker; normalising that marker away and diffing the two logs found in one step what four
rebuilds did not. The rebuilds were not wasted - three toggles of our own patches producing
byte-identical output is a stronger correctness statement for the port than the comparison
being attempted - but they answered a question that did not need asking.
For the record, since it is the sort of thing that gets assumed: surf2Dwrite into 16-bit
float arrays works correctly on CUDA 12.9 / Pascal. The buffer-copy mip path is an AMD
workaround only, and carrying it costs nothing either way.
Ported, and it answers the question above - though not in the shape the question assumed.
The bridge's logic has nothing backend-specific in it: ten entry points, all with direct CUDA
equivalents. Rather than rewrite bridge.h in CUDA dialect and risk changing behaviour that is
validated on five AMD cards, cheshire/hip_to_cuda.h maps those ten names onto CUDA - the exact
mirror of cuda_to_hip.h, and safe for the same reason: neither runtime's headers declare the
other's names.
Reaching AliceVision's allocations is the part that does not mirror. On HIP, cuda_to_hip.h
shadows cudaMalloc / cudaMallocPitch / cudaMalloc3D / cudaFree with inline functions,
which works only because the CUDA names are absent there. Two approaches that look like they
should work in a CUDA build, and do not:
- a function-like macro -
cudaMallocPitch<Type>(&buf, ...)does not expand, because the token after the macro name is<and not(. AliceVision's main device allocation would have silently bypassed the bridge. - overloads inside
namespace aliceVision::depthMap- the idea being that unqualified lookup searches enclosing namespaces before the global one. It does, but ADL adds the global namespace straight back in: every argument type (float2,__half,cudaPitchedPtr,cudaExtent) lives in::, so all five call sites become ambiguous with the real CUDA declarations. The compiler said so twenty times.
So the five call sites in memory.hpp name the allocator through CHESHIRE_DEV_MALLOC and
friends (patch step 1j), which resolve to cheshire::devmem on CUDA, to cheshire::bridge on
HIP, and to plain CUDA if neither is compiled in. No CMake change is needed: CHESHIRE_HIP is
defined by cuda_to_hip.h itself, so memory.hpp can tell which backend it is in straight after
including <cuda_runtime.h>.
cheshire/managed_cuda.h adds the third arm, CHESHIRE_CUDA_MANAGED=1, routing the same five
sites to cudaMallocManaged. It emulates the pitched forms by hand, since unified memory has no
pitched allocator.
41 views, downscale 2, GTX 1080 Ti. Every row byte-identical to run-cuda-41-dlpt, which is
itself byte-identical to the CUDA 11.3 reference.
| allocator | time | tiles | depth maps |
|---|---|---|---|
plain cudaMalloc |
225 s | 16 | 41 / 41 identical |
| bridge | 226 s | 14 | 41 / 41 identical |
| unified memory | 219 s | 22 | 41 / 41 identical |
Every natural cap spilled zero bytes. The planner cut tile parallelism to fit instead:
| VRAM cap | depth maps x tiles | volume peak | time | depth maps |
|---|---|---|---|---|
| none | 3 x 14 | 6118 MB | 225 s | 41 / 41 identical |
| 4000 MB | 1 x 5 | 2185 MB | 236 s | 41 / 41 identical |
| 2000 MB | 1 x 2 | 874 MB | 236 s | 41 / 41 identical |
| 1000 MB | 1 x 1 | 437 MB | 238 s | 41 / 41 identical |
42 concurrent tiles down to 1 costs 6 %. That is docs/02's "tile parallelism costs nothing
to give up", now confirmed on NVIDIA, and it is the bridge's cheap lever: the same job in a
quarter of the VRAM for 6 %.
Tile count pinned at 1 by a 1000 MB cap, so the delta is placement and nothing else:
| forced to host | time | vs VRAM | AMD equivalent (docs/02) |
|---|---|---|---|
| nothing (control) | 238 s | 1.00x | - |
| volume | 972 s | 4.08x | 5.6x |
| map | 912 s | 3.83x | 1.9x |
| volume + map | 1595 s | 6.70x | - |
All byte-identical. Volumes are cheaper to spill on NVIDIA than on the RX 9070, maps dearer - though the map row forces 1304 allocations totalling 58 MB, because naming a class bypasses the 4 MB minimum-spill threshold. It is a stress test, not a policy the defaults would produce.
On HIP-Windows the camera mipmaps are emulated with buffers (mipmap_emu.h), which is why the
bridge has an Image class at all. On CUDA they are real cudaMipmappedArrays and never pass
through cudaMalloc, so the bridge cannot see or account for them and the image reserve is
inert. The bridge manages volumes and maps - by docs/02's measurements the cheap classes to
spill, not the expensive one.
This has one sharp edge, found by measurement. CHESHIRE_BRIDGE_VRAM_MB is read by both the
bridge's cap and the tile planner's budget. Set above the card's physical memory, the planner
commits more tiles than the card holds and the bridge fills VRAM to 10966 MB of 11165 trying to
honour the cap. Its own allocations spill correctly (117 spills, 0 spill failures) - and then the
camera arrays, which it does not own, hit a hard OOM with ~200 MB left. bridge.h now clamps a
cap that exceeds the device's total memory down to the default fraction of free VRAM. A cap
below total is a deliberate budget and is left exactly as asked, so the bridge-v2 matrices that
go down to 500 MB are untouched.
Verified on the configuration that failed:
[cheshire] bridge: requested vram cap 20000 MB exceeds this device (11165 MB total); clamping to 9911 MB
exit=0 225 s 41 depth maps 41/41 byte-identical (0 spills, 0 spill failures)
Same command that previously stopped after 43 s with CUDA Error: out of memory.
The comparison cannot be run, and the reason is the answer. CHESHIRE_BRIDGE_VRAM_MB is both the
cap and the planner's budget, so there is no way to over-commit the tile count for the bridge
while leaving its cap honest. The two are not competing strategies for one situation:
- the bridge prevents pressure - it plans the job to fit the budget, and spills only when a single tile plus its images will not;
- unified memory absorbs pressure - it commits whatever is asked and migrates pages on fault.
Two attempts to force genuine oversubscription on an 11 GB card both failed, which is worth
recording so nobody repeats them. At downscale 2, 24 tiles x 437 MB came to ~10.5 GB, just under
the card, so unified memory never migrated a page. Downscale 1 does not help either:
tileBufferWidth/Height fix the tile size, so a finer downscale produces more tiles rather than
bigger ones, and per-tile cost stays ~437 MB. Forcing the issue needs a working set that genuinely
exceeds the card - many more views at downscale 1, or a raised sgmMaxDepths to inflate the
volume per tile.
So on this hardware and these datasets, memory never binds, exactly as this document predicted before the port. What the bridge buys on CUDA is not survival, it is choice: the same output in a quarter of the VRAM for 6 %, leaving the card free for something else. Unified memory is a genuine alternative here in a way it never was on AMD - free when nothing binds, and it tolerates an over-committed planner where the bridge needs its budget told the truth.
DepthMap is not why an NVIDIA user would want this build - upstream already gives them that.
The value is the stages upstream CUDA does not have: the GPU descriptor matcher, the
depth-map filter, meshing (votes, min-cut, visibility kNN) and texturing. All of those had only
ever run as HIP. They compiled as CUDA, which is the cheapest evidence there is, so they were run
from the bundle with each port's CHESHIRE_*_CHECK toggle on - those answer the same queries on
the host and count disagreements, which is the difference between "the stage executed" and "the
stage is right".
| stage | result on the GTX 1080 Ti |
|---|---|
FeatureExtraction (sift / PopSIFT) |
2 s vs 22 s for CPU dspsift - 11x, 22939 keypoints |
| FeatureMatching (GPU matcher) | GPU brute-force L2 2-NN on NVIDIA GeForce GTX 1080 Ti |
| DepthMap | byte-identical to the reference (above); bridge planner live |
| DepthMapFilter (GPU filter) | GPU votes, ~1.17 M pixels per view |
| Meshing (votes / min-cut / kNN) | 153 s, 392 788 vertices, 781 463 faces |
| Texturing (GPU) | 7 atlases; the port profiles its own rasterisation |
The meshing self-checks are the strongest statement available, and they are exact:
filterByPixSize check: identical to single-threaded upstream on all 1472035 slots
visibilities on the GPU: 152201929 queries, 36931967 votes, 0 answered on the host
GPU knn check: identical to nanoflann on all 152201929 queries
neighbour table check: 21386816 entries, lists differing from upstream's construction: 0
facet weight check: 21386816 facets, differing from the sequential computation: 0
Every libpopsift.so on these machines was a HIP build (they declare libamdhip64), so the CUDA
AliceVision had ALICEVISION_USE_POPSIFT=OFF. PopSIFT is a CUDA project to begin with -
third_party/popsift is a clone of alicevision/popsift v0.10.0, and apply_popsift_patch.py is
what adapts it to HIP - so the CUDA side just wants the pristine source:
scripts/linux/build-popsift-cuda.sh. build-alicevision-cuda.sh now finds it automatically and
says so loudly when it cannot, instead of silently producing a bundle without GPU SIFT.
Switching PopSIFT on broke the bundle target outright: cannot resolve item 'libpopsift.so.0.10.0', then READ_ELF given FILE ... that does not exist. The cause is one line
of upstream's top-level CMakeLists.txt, with two defects in it:
-DBUNDLE_LIBS_PATHS=${BUNDLE_LIBS_PATHS} # unquoted, and no VERBATIM- Unquoted, the CMake list expands into separate command-line arguments, so
MakeBundle.cmakereceives only the first path and silently drops the rest. The CUDA toolkit'slib64has never actually been on that search path. - Quoting alone is not enough: ninja runs the command through
sh, where the semicolons separate commands, soshtries to execute the paths -/bin/sh: 1: /opt/AliceVision_deps/lib64: Permission denied.
The COMMAND form with VERBATIM has CMake escape each argument for the native shell, which is
also correct for the Windows bundles. Applied as patch step 3c; an upstream bug rather than ours,
and a candidate for a PR alongside #2179 / #2181. It stayed invisible for as long as every
dependency happened to be resolvable another way; PopSIFT, installed outside the deps prefix, is
the first one that was not.
The FeatureExtraction failure was not PopSIFT. It was first reported here as "the CUDA bundle
has no GPU SIFT, so --describerTypes sift dies". The real error, once PopSIFT was present and
the exception could surface, is that the enginebay cache's cameraInit.sfm holds Windows D:
paths and those photos are not on the Linux box at all. The original bare std::runtime_error
with no message was that same exception swallowed by terminate inside an OpenMP region. The
PopSIFT work was still necessary and the 11x above is real, but the crash had a different cause,
and it was attributed on a plausible story rather than on evidence.
Texturing first "passed" while generating nothing. It wrote texturedMesh.obj and .mtl,
exited 0, and produced no textures at all, because --colorMappingFileType defaults to NONE and
the run passed the SfM sfm.abc where Meshroom passes Meshing's densePointCloud.abc (which is
what carries visibility). Counting output files is not verifying a stage - the exact blind spot
verify_bundle_stages.py exists to prevent, still present in the checker itself. Each stage there
now has to emit a specific log line as well, and Texturing was added, wired the way Meshroom
wires it.
Builds, links, installs, packages and runs. 396 / 396 targets, 0 unresolved symbols, and the result on bench-pc's GTX 1050 Ti is bit-identical pixel data to the Linux CUDA build - which is itself byte-identical to the Meshroom CUDA 11.3 reference:
| our Windows CUDA (MSVC, GTX 1050 Ti, 4 GB) | bit-identical pixels to |
| our Linux CUDA (GCC 13.3, GTX 1080 Ti, 11 GB) | byte-identical to |
| Meshroom 2023.3 (AliceVision 3.1, CUDA 11.3, GTX 1080 Ti) |
Different operating system, different host compiler, different GPU, and the depth values are the
same bits. The files themselves are not byte-identical - that difference is EXR container
metadata, confirmed by comparing the decoded pixel arrays directly rather than trusting
compare_depthmaps.py's four-decimal output, where "0.0000" covers anything under 5e-5.
The six-view sample was not representative, which is the reason for running the rest rather than extrapolating from it: 37 of 41 views are pixel bit-identical, and four are not.
| Windows vs Linux (ours) | CUDA 11.6 Win vs CUDA 11.3 Linux (upstream's own) | |
|---|---|---|
| bit-identical views | 37 / 41 | 37 / 41 |
| mask agreement >= 0.95 | 41 / 41 | 41 / 41 |
| median view within 1 % | 100 % | 100 % |
| worst view | 92.85 % | 92.9 % |
So our cross-platform divergence is indistinguishable from the divergence upstream's own binaries show between the same two platforms - and ours carries an extra variable theirs does not, a different GPU (1050 Ti vs 1080 Ti).
Two of the four differ trivially (max relative 6e-4). The two that differ substantially are
1227295871 and 1430763847 - precisely the views docs/04 already names as the least stable
in every cross-build comparison (mask agreement 0.971, and the worst view at 0.917 respectively).
They are the ones that flip when anything about the floating-point path changes.
compare_depthmaps.py reports FAIL, because its strict criterion is >= 98 % of pixels within 1 %
on every view and the worst is 92.85 %. README already records that bar as sitting below the
algorithm's own noise floor.
41 views on the GTX 1050 Ti: 640.5 s against 211 s on the 11 GB 1080 Ti. More usefully, the planner reported
Cheshire bridge planner 1: VRAM budget 2417.33 MB, 0 full R cameras + 4 tiles
Zero full R cameras. On the 1080 Ti it allocated two or three full camera sets plus tiles, and every artificial cap was absorbed by cutting tile parallelism without spilling a byte. Here it could not fit even one, and fell back to tiles alone. That is the regime the bridge exists for, reached honestly rather than by inflating a budget - the oversubscription test that could not be constructed on 11 GB (see above) simply happens on 4 GB.
nvcc drives cl.exe, so this build is MSVC rather than the clang-cl the HIP build needs for
ROCm. scripts/windows/build-alicevision-cuda.cmd.
The ordering trap, which cost an hour: scripts/env.cmd calls vcvarsall for VS 2026 whenever
VCToolsInstallDir is unset, and vcvars refuses to re-initialise a shell it has already
configured - it returns 0 and changes nothing. Calling env.cmd before vcvars64 -vcvars_ver=14.44 therefore leaves 14.50 on PATH, nvcc picks it up, and the configure fails
with "Only the versions between 2017 and 2022 (inclusive) are supported" - which reads exactly
like VS 2022 not being installed. 2022 is now initialised first, CMAKE_CUDA_HOST_COMPILER is
pinned rather than left to PATH order, and the script aborts if cl is not a 19.4x.
-
The vcpkg STL gap
docs/18predicted as the ABI risk:__std_adjacent_find_4and__std_unique_4, unresolved in CoinUtils when linkingaliceVision_lInftyComputerVision. Two symbols across 1204 targets.hip/compat/stlcompat/std_vector_algorithms.cpp. -
OpenMesh's unguarded
/bigobj. vcpkg'sOpenMeshConfig.cmakesetsINTERFACE_COMPILE_OPTIONS "/bigobj"with no language genex - pybind11 in the same tree writes$<$<COMPILE_LANGUAGE:CXX>:/bigobj>, which is what it should be. It reaches every source of any target linking OpenMesh, and nvcc on Windows reads a leading/as a file path:nvcc fatal : A single input file is required for a non-link phasemeshis the only CUDA-bearing target that links OpenMesh, and the GPU texturing port is the first CUDA source ever to live there, so this could not have bitten anyone before. -
libomp140.x86_64.dllmissing from the package. Every binary died with0xC0000135(STATUS_DLL_NOT_FOUND) before writing a log line, while the same package ran on the machine that built it. CMake'sFindOpenMPselects-openmp:llvm(OpenMP 3.0+, which AliceVision needs; the redistributablevcomp140is 2.0 only), and that runtime ships with Visual Studio rather than the VC++ redistributable. The build machine had it insystem32; bench-pc did not. The shipped HIP Windows packages already bundle it, so the CUDA package matches them.Worth flagging rather than deciding quietly:
libomp140.x86_64.dlllives under adebug_nonredistpath in the VS tree. That applies equally to the HIP packages already published, so it is a pre-existing question about both, not one this build introduces.
The first comparison used out-cuda-mini6 as the yardstick without checking what it was: a run
of the monstree-mini6 dataset, while this run took monstree-full --rangeSize 6, i.e. the
first 6 of 41 views. Not one view id overlapped. The same failure as the sgmDepthListPerTile
episode - check what the reference actually is before comparing against it.
And a Test-Path "...\$d" sent through ssh quoting never expanded $d, so it tested the
directory, returned True six times, and was reported as "the runtime DLLs are in the package".
They were not. Remote PowerShell goes in a script file, not inside nested quotes.
bridge.h and memory.hpp are shipping AMD code: the allocation macros now call
cheshire::bridge::* directly instead of going through cuda_to_hip.h's cudaMalloc shim, and
bridge.h gained the VRAM-cap clamp. The argument was that this is the same function and
therefore neutral, and the compiler agreed it builds - but neither of those is a measurement, and
until 2026-09-21 no AMD card had executed a build from these commits.
Rebuilt for gfx1201 and run on the RX 9070: 476/476 targets, 0 errors, and the six-view
monstree-mini6 output is byte-identical to a run of the same views on the same card from
before any of these changes. Against the CUDA reference it lands where docs/04 says the
validated HIP build lands - median error 0, 98-99 % within 1 %, simMAD 0.12-0.25.
Still open: that verification ran from build/av-gfx1201-install, not from a bundle produced by
the patched bundle target. The VERBATIM fix changes packaging for every platform including the
AMD packages already published, and a rebuilt binary is not a rebuilt bundle.