Minor optimization by exploiting demag kernel symmetry - #411
Open
JonathanMaes wants to merge 5 commits into
Open
Minor optimization by exploiting demag kernel symmetry#411JonathanMaes wants to merge 5 commits into
JonathanMaes wants to merge 5 commits into
Conversation
Squashed snapshot of the CUDA-Graph-accelerated mumax3 fork (based on mumax/3 commit 3fe3d41, "Update gpus.txt"). Adds a transparent split-graph CUDA Graph capture/replay path for Run()/Steps() across Heun, RK23, RK45DP, RK56 and BackwardEuler, with safety guards for Temp, time-dependent B_ext/material parameters, NoDemagSpins and custom field terms, plus the conv_kernmul launch-range and lazy MaxTorque/MemCpyDtoH micro-optimizations. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Replaces the upstream mumax3 README with a fork-focused README covering: - Performance table (4K–1M cell speedup measurements on RTX 5060 Ti) - How the CUDA Graph capture/replay path works - Safety guard list - Pre-built Windows binary download link (GitHub Releases) - Build-from-source instructions (Windows CUDA 13.2 + Linux) - Summary table of all source changes vs upstream 3fe3d41 - Upstream citation and license notice Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR captures part of the changes introduced in mumax³-CO, as presented in C.-Y. You, J. Magn. 31, 204-213 (2026).
In the demag kernel multiplication (during convolution), the symmetry of the demag kernel is not fully utilized to optimize performance. The CUDA launch grid can be halved (Ny/2+1 rows), with each thread then using this symmetry to compute both a row and its mirror row. This results in a 0.75% speedup on a 2048×2048×1 grid, and a negligible difference on small/medium grids.