In-depth paper analyses with visualizations of optimization algorithms and deep learning concepts
Detailed analysis of the Adam optimization algorithm: first moment (momentum) and second moment (adaptive scale) with visualizations.
Deep Residual Learning for Image Recognition. Residual learning, identity shortcuts, and the degradation problem.
Why does deep and cheap learning work so well? Physics-inspired analysis of symmetry, locality, and compositionality.
Full analysis of the Transformer architecture: scaled dot-product attention, multi-head attention, and positional encoding.
Step-by-step visualization of Adam optimization path and gradient dynamics.
Interactive parameter exploration for the Adam optimizer.
Mathematical animation of the Adam optimization algorithm.
Visualization of residual learning and skip connections.
Animation of deep learning concepts from Lin, Tegmark & Rolnick.
Animation of the Transformer architecture and attention mechanism.