A drop-in replacement for the standard residual connection x = x + layer(x).
WHC replaces the single residual path with n parallel lanes and learnable mixing matrices, delivering:
- 6–15% faster training than standard residual
- Lower final loss on benchmark tasks
- Proven stability — spectral norm ≤ 1
- 82 parameters for the mixing matrix (vs mHC's 32,768)
- 2× faster than mHC in wall-clock time
from whc import WormholeHyperconnection
whc = WormholeHyperconnection(
dim=512,
expansion_rate=4,
num_layers=12,
)
x = whc.expand(x) # (B, T, dim) -> (B, T, n, dim)
for i, layer in enumerate(transformer_blocks):
x = whc(x, layer, layer_idx=i) # replaces `x = x + layer(x)`
x = whc.reduce(x) # (B, T, n, dim) -> (B, T, dim)pip install whcgit clone https://github.com/fardinsabid/wHC.git
cd wHC
pip install -e .- Python >= 3.8
- PyTorch >= 2.0.0
| Feature | Description |
|---|---|
| Drop-in replacement | Replace x = x + layer(x) with x = whc(x, layer=layer) |
| n parallel lanes | Instead of 1 fixed path |
| Learnable mixing | Network learns how lanes interact |
| Proven stability | Spectral norm ≤ 1 guarantees no explosion |
| No iterations | Closed-form kernel, no Sinkhorn-Knopp |
| Minimal overhead | 82 parameters for the mixing matrix |
| GPU ready | Runs on CUDA, CPU, MPS |
| Method | Mean Final Gain | Max Final Gain |
|---|---|---|
| Unconstrained HC | 9.54×10⁵ | 4.77×10⁶ |
| mHC (DeepSeek) | 0.42 | 0.68 |
| WHC | 0.48 | 0.63 |
WHC and mHC stay bounded; unconstrained HC explodes.
| Method | Final Loss | Time (s) |
|---|---|---|
| Standard Residual | 0.3765 | 0.91 |
| mHC-static | 0.6966 | 11.01 |
| WHC | 0.3590 | 5.54 |
WHC achieves lower loss and is 2× faster than mHC.
| Component | Parameters |
|---|---|
| WormholeKernel (n=8) | 82 |
| mHC H_res generator (dim=512, n=8) | 32,768 |
WHC uses ~400× fewer parameters for the mixing matrix.
from whc import WormholeHyperconnection
# Create WHC wrapper
whc = WormholeHyperconnection(dim=512, expansion_rate=4, num_layers=12)
# Forward pass
x = whc.expand(x)
for i, layer in enumerate(transformer_blocks):
x = whc(x, layer, layer_idx=i)
x = whc.reduce(x)whc = WormholeHyperconnection(dim=512, expansion_rate=4, num_layers=1)
x = whc.expand(x)
x = whc(x, layer, layer_idx=0)
x = whc.reduce(x)layer_out = layer(x)
x = whc(x, layer_out=layer_out)whc = WormholeHyperconnection(
dim=768,
expansion_rate=8, # Number of parallel lanes
num_layers=24, # Number of layers in your stack
manifold_dim=3, # Dimension of the wormhole manifold
shared_kernel=False, # Share kernel across layers?
)| Example | Description | Run |
|---|---|---|
simple_mlp.py |
MLP with WHC replacing residuals | python examples/simple_mlp.py |
transformer_whc.py |
Transformer with WHC replacing residuals | python examples/transformer_whc.py |
resnet_whc.py |
ResNet with WHC replacing residuals | python examples/resnet_whc.py |
pytest tests/ -v| Test File | What It Tests |
|---|---|
test_kernel.py |
Kernel stability, spectral norm ≤ 1 |
test_whc.py |
Shape, forward pass, parameters |
test_gradients.py |
Gradient flow, no NaNs, no explosion |
All 21 tests passed ✅
The full mathematical derivation, stability proof, and experimental results are in:
Key contributions:
- Spectral normalization (closed-form, no iteration)
- Wormhole kernel (physics-inspired, interpretable)
- 82 parameters for H_res (vs mHC's 32,768)
- 2× faster than mHC
- Better loss than standard residual
whc/
├── whc.py # Core implementation
├── README.md # This file
├── LICENSE # MIT License
├── setup.py # Package installer
├── pyproject.toml # Build config
├── requirements.txt # Dependencies
├── examples/
│ ├── simple_mlp.py # MLP with WHC
│ ├── transformer_whc.py # Transformer with WHC
│ └── resnet_whc.py # ResNet with WHC
├── papers/
│ └── whc.pdf # Research paper
├── assets/
│ ├── whc_stability.png # Stability comparison
│ └── whc_training.png # Training loss curves
└── tests/
├── test_kernel.py # Kernel unit tests
├── test_whc.py # WHC unit tests
└── test_gradients.py # Gradient stability tests
If you use WHC in your research, please cite:
@misc{sabid2026whc,
author = {Fardin Sabid},
title = {Wormhole Hyperconnections: A Physics-Inspired Framework for Stable Deep Residual Learning},
year = {2026},
publisher = {GitHub},
howpublished = {\url{https://github.com/fardinsabid/wHC}}
}MIT License — see LICENSE for details.
Fardin Sabid
- GitHub: @fardinsabid
- Research: Deep Learning Optimization, Physics-Inspired Architectures
WHC builds on the foundation of:
- He et al. (2016) — Residual connections
- Zhu et al. (2025) — Hyper-Connections (HC)
- Xie et al. (2025/2026) — Manifold-Constrained Hyper-Connections (mHC)
- Kipf & Welling (2017) — Graph Convolutional Networks
If you find WHC useful, please ⭐ star the repository!
The standard residual was 2016. WHC is 2026.