🧠 Adam Optimization · 3D Core Math
Loss surface · Parameter trajectory ·
1st moment
(momentum) ·
2nd moment
(adaptive scale)
⚙️ Hyperparameters
Learning Rate (η)
0.12
Step size.
Too high → divergence. Too low → slow convergence.
Beta 1 (β₁)
0.900
Momentum decay.
Controls the exponential moving average of the gradient. Default 0.9.
Beta 2 (β₂)
0.999
Adaptive scale decay.
Controls the moving average of squared gradients. Default 0.999.
Animation Speed
200 ms
Step interval.
Higher = slower, easier to observe.
⟳ Apply & Restart
Trajectory
Momentum (1st moment m)
Adaptive scale (√v)
Gradient ∇f
Iteration
0 / 60
Current η
0.12