The network trained with the AdamW optimizer has a faster reduction of loss and better convergence and is always in a decreasing trend, while the network trained with the Adagrad optimizer has a faster loss reduction in the early stage and an unstable and flattening reduction in the later stage under the influence of the dynamic adjustment of the learning rate.
← all excerpts
Fourier Ptychographic Microscopic Reconstruction Method Based on Residual Hybrid Attention Network.
1
—
—