🌊 GELU Visualizer

Every input gets multiplied by its own probability β€” that's the whole idea of GELU.

πŸ“‹ How To Use This Lab
  1. Probability Gate: drag the input point β€” watch the bell curve's shaded area (its probability) multiply into the output live.
  2. The Dip: drag to a small negative x β€” a small negative Γ— small positive probability makes the curve dip just below zero.
  3. Magnifying Glass: compare ReLU's sharp corner and GELU's smooth glide at the origin, side by side.
  4. Gradient Highway: watch GELU's derivative flow smoothly through zero while ReLU's flow hits a wall.
x = 2.00
Ξ¦(x) = 0.977
GELU(x) = 1.95
1 Β· Probability Gate β€” the Bell Curve Decides How Much Passes GELU(x) = x Γ— Ξ¦(x) 2 Β· Magnifying Glass β€” Zoomed at the Origin ReLU β€” sharp, jagged corner GELU β€” smooth, gliding curve 3 Β· Gradient Highway β€” Derivative Flow
🌊 GELLYDrag the input β€” I'll multiply it by its own bell-curve probability, live.
πŸ§‘β€πŸ« PROF. TORCHGPT and other Transformers use GELU specifically because this smoothness means no sudden jolts in the gradient during training.