π GELU Visualizer
Every input gets multiplied by its own probability β that's the whole idea of GELU.
π How To Use This Lab
Probability Gate:
drag the input point β watch the bell curve's shaded area (its probability) multiply into the output live.
The Dip:
drag to a small negative x β a small negative Γ small positive probability makes the curve dip just below zero.
Magnifying Glass:
compare ReLU's sharp corner and GELU's smooth glide at the origin, side by side.
Gradient Highway:
watch GELU's derivative flow smoothly through zero while ReLU's flow hits a wall.
x = 2.00
Ξ¦(x) = 0.977
GELU(x) = 1.95
1 Β· Probability Gate β the Bell Curve Decides How Much Passes
GELU(x) = x Γ Ξ¦(x)
2 Β· Magnifying Glass β Zoomed at the Origin
ReLU β sharp, jagged corner
GELU β smooth, gliding curve
3 Β· Gradient Highway β Derivative Flow
β Previous
Next βΆ
Auto Run βΆβΆ
βΊ Reset
π GELLY
Drag the input β I'll multiply it by its own bell-curve probability, live.
π§βπ« PROF. TORCH
GPT and other Transformers use GELU specifically because this smoothness means no sudden jolts in the gradient during training.