programdemos/backprop/backprop.xtl
One step of training, every array shown: a small network (2 inputs, 4 tanh units, softmax over 3 classes) on six points of a spiral. The forward pass, the loss, each layer's gradient as one expression, the step; then every gradient checked against a finite difference.
The batch and the starting weights, read from data/
Forward
Backward: each layer's gradient, one expression each
ᵘo̲nes : Float -> Float
ᵘo̲nes ← { x → x c̲at₂ ((t̲ally x) c̲at 1) r̲eshape 1.0 } ⍝ a column of 1s: the bias input
D1 : Float
D1 ← (D2 '+ '× i̲nner o̲\ -1 d̲rop W2) × 1.0 − H × H ⍝ back through W2 and tanh
Used in: G1
The step
ᵘl̲oss : Float -> Float -> Float
ᵘl̲oss ← { w1 w2 → Y ⁿⁿc̲rossEntropy ⁿⁿs̲oftmax (ⁿⁿt̲anh X ⁿⁿd̲ense w1) ⁿⁿd̲ense w2 }
e : Float
Each analytic gradient against a central finite difference, (L(w + e) - L(w - e)) / 2e for every weight: the largest difference of all 27.
e ← 0.00001
ᵘb̲ump : a -> Int -> Float
ᵘb̲ump ← { w i → e × f̲loat (s̲hape w) r̲eshape i = r̲ange t̲ally r̲avel w }
F1 : Float
F1 ← (s̲hape W1) r̲eshape '{ i → (((W1 + W1 ᵘb̲ump i) ᵘl̲oss W2) − (W1 − W1 ᵘb̲ump i) ᵘl̲oss W2) ÷ 2.0 × e } e̲ach r̲ange 12
Used in: gap