sourcedemos/train-live/train-live.xtl

1⍝!/usr/bin/env xetal 2⍝# Training in X_eTaL: a network (2 inputs, 16 tanh units, softmax over 3⍝# 3) learns which of three spiral arms a point is on. The points are 4⍝# made here; the backward pass is the backprop microscope's; Adam, 5⍝# written out, moves the weights; p_ower repeats the step. 6 7ⁿⁿ⁼u̲se< "NN" 8 9⍝## The spiral: 100 points on each of three arms 10n ← 300 11i ← o̲ffsets n 12arm ← i d̲iv 100 13t ← (0.5 + f̲loat i m̲od 100) ÷ 100.0 14a ← (2.0944 × f̲loat arm) + (5.5 × t) + 0.1 × (f̲loat (r̲oll! n r̲eshape 1001) − 501) ÷ 500.0 15X ← o̲\ (2 c̲at n) r̲eshape ((0.1 + 0.9 × t) × c̲os a) c̲at (0.1 + 0.9 × t) × s̲in a 16Y ← 3 ⁿⁿo̲neHot 1 + arm 17 18⍝## The network: W1 (3 x 16), W2 (17 x 3), each bias its last row 19ᵘo̲nes ← { x → x c̲at₂ ((t̲ally x) c̲at 1) r̲eshape 1.0 } 20ᵘf̲orward ← { W1 W2 → ⁿⁿs̲oftmax (ⁿⁿt̲anh X ⁿⁿd̲ense W1) ⁿⁿd̲ense W2 } 21⍝# The gradients of the loss by W1 and by W2, a pair (the backprop 22⍝# microscope's four lines). 23ᵘg̲rad ← { W1 W2 → 24 H ← ⁿⁿt̲anh X ⁿⁿd̲ense W1 25 D2 ← ((ⁿⁿs̲oftmax H ⁿⁿd̲ense W2) − Y) ÷ f̲loat n 26 D1 ← (D2 '+ '× i̲nner o̲\ -1 d̲rop W2) × 1.0 − H × H 27 ((o̲\ ᵘo̲nes X) '+ '× i̲nner D1, (o̲\ ᵘo̲nes H) '+ '× i̲nner D2) 28} 29⍝## Adam 30⍝ The state is a tuple, (W1, W2, M1, M2, V1, V2, k): the weights, the 31⍝ running averages of gradients and of their squares, the step count; 32⍝ p_ower iterates it as one value. 33⍝# The learning rate. 34lr ← 0.02 35⍝# How far Adam moves a weight array, from its averages (m, v) at step k. 36ᵘm̲ove ← { (m, v) k → (m ÷ 1.0 − 0.9 ^ k) ÷ 0.00000001 + (v ÷ 1.0 − 0.999 ^ k) ^ 0.5 } 37⍝# One step of Adam. 38ᵘa̲dam ← { (W1, W2, M1, M2, V1, V2, k) → 39 k ← 1.0 + k 40 (G1, G2) ← W1 ᵘg̲rad W2 41 M1 ← (0.9 × M1) + 0.1 × G1 42 M2 ← (0.9 × M2) + 0.1 × G2 43 V1 ← (0.999 × V1) + 0.001 × G1 × G1 44 V2 ← (0.999 × V2) + 0.001 × G2 × G2 45 (W1 − lr × (M1, V1) ᵘm̲ove k, W2 − lr × (M2, V2) ᵘm̲ove k, M1, M2, V1, V2, k) 46} 47⍝# The starting state: small random weights, zero averages, step 0. 48w0 ← (f̲loat (r̲oll! 99 r̲eshape 2001) − 1001) ÷ 1000.0 49W1 ← 3 16 r̲eshape 48 t̲ake w0 50W2 ← 17 3 r̲eshape 48 d̲rop w0 51s0 ← (W1, W2, 0.0 × W1, 0.0 × W2, 0.0 × W1, 0.0 × W2, 0.0) 52⍝ -- end of the core ------------------------------------------ 53 54⍝# The loss, and the share of points right, of a state. 55ᵘl̲oss ← { (W1, W2, _, _, _, _, _) → Y ⁿⁿc̲rossEntropy W1 ᵘf̲orward W2 } 56ᵘr̲ight ← { (W1, W2, _, _, _, _, _) → (1 + arm) ⁿⁿa̲ccuracy W1 ᵘf̲orward W2 } 57⍝# The loss and the share of points right: before, after 100 steps, 58⍝# after 400. 59s1 ← 100 'ᵘa̲dam p̲ower s0 60s2 ← 300 'ᵘa̲dam p̲ower s1 61ᵘr̲ound ← { a → (f̲loat f̲loor 0.5 + 1000.0 × a) ÷ 1000.0 } 62ᵘr̲ound (ᵘl̲oss s0) c̲at (ᵘl̲oss s1) c̲at ᵘl̲oss s2 63ᵘr̲ound (ᵘr̲ight s0) c̲at (ᵘr̲ight s1) c̲at ᵘr̲ight s2