// writing
Notes on AI, infrastructure and what I’m learning.
Why scaling attention scores by √d_k keeps softmax from saturating and its gradients from vanishing.
// writing
Why scaling attention scores by √d_k keeps softmax from saturating and its gradients from vanishing.