Reproducing the paper "PADAM: Closing The Generalization Gap of Adaptive Gradient Methods In Training Deep Neural Networks" for the ICLR 2019 Reproducibility Challenge
tensorflow optimization keras wide-residual-networks adam-optimizer tensorflow-eager amsgrad sgd-momentum padam
-
Updated
Apr 13, 2019 - Python