sourcedemos/train-live/train-live.xtl
1⍝!/usr/bin/env xetal
2⍝# Training in X_eTaL: a network (2 inputs, 16 tanh units, softmax over
3⍝# 3) learns which of three spiral arms a point is on. The points are
4⍝# made here; the backward pass is the backprop microscope's; Adam,
5⍝# written out, moves the weights; p_ower repeats the step.
6
7ⁿⁿ⁼u̲se< "NN"
8
9⍝## The spiral: 100 points on each of three arms
10n ← 300
11i ← o̲ffsets n
12arm ← i d̲iv 100
13t ← (0.5 + f̲loat i m̲od 100) ÷ 100.0
14a ← (2.0944 × f̲loat arm) + (5.5 × t) + 0.1 × (f̲loat (r̲oll! n r̲eshape 1001) − 501) ÷ 500.0
15X ← o̲\ (2 c̲at n) r̲eshape ((0.1 + 0.9 × t) × c̲os a) c̲at (0.1 + 0.9 × t) × s̲in a
16Y ← 3 ⁿⁿo̲neHot 1 + arm
17
18⍝## The network: W1 (3 x 16), W2 (17 x 3), each bias its last row
19ᵘo̲nes ← { x → x c̲at₂ ((t̲ally x) c̲at 1) r̲eshape 1.0 }
20ᵘf̲orward ← { W1 W2 → ⁿⁿs̲oftmax (ⁿⁿt̲anh X ⁿⁿd̲ense W1) ⁿⁿd̲ense W2 }
21⍝# The gradients of the loss by W1 and by W2, a pair (the backprop
22⍝# microscope's four lines).
23ᵘg̲rad ← { W1 W2 →
24 H ← ⁿⁿt̲anh X ⁿⁿd̲ense W1
25 D2 ← ((ⁿⁿs̲oftmax H ⁿⁿd̲ense W2) − Y) ÷ f̲loat n
26 D1 ← (D2 '+ '× i̲nner o̲\ -1 d̲rop W2) × 1.0 − H × H
27 ((o̲\ ᵘo̲nes X) '+ '× i̲nner D1, (o̲\ ᵘo̲nes H) '+ '× i̲nner D2)
28}
29⍝## Adam
30⍝ The state is a tuple, (W1, W2, M1, M2, V1, V2, k): the weights, the
31⍝ running averages of gradients and of their squares, the step count;
32⍝ p_ower iterates it as one value.
33⍝# The learning rate.
34lr ← 0.02
35⍝# How far Adam moves a weight array, from its averages (m, v) at step k.
36ᵘm̲ove ← { (m, v) k → (m ÷ 1.0 − 0.9 ^ k) ÷ 0.00000001 + (v ÷ 1.0 − 0.999 ^ k) ^ 0.5 }
37⍝# One step of Adam.
38ᵘa̲dam ← { (W1, W2, M1, M2, V1, V2, k) →
39 k ← 1.0 + k
40 (G1, G2) ← W1 ᵘg̲rad W2
41 M1 ← (0.9 × M1) + 0.1 × G1
42 M2 ← (0.9 × M2) + 0.1 × G2
43 V1 ← (0.999 × V1) + 0.001 × G1 × G1
44 V2 ← (0.999 × V2) + 0.001 × G2 × G2
45 (W1 − lr × (M1, V1) ᵘm̲ove k, W2 − lr × (M2, V2) ᵘm̲ove k, M1, M2, V1, V2, k)
46}
47⍝# The starting state: small random weights, zero averages, step 0.
48w0 ← (f̲loat (r̲oll! 99 r̲eshape 2001) − 1001) ÷ 1000.0
49W1 ← 3 16 r̲eshape 48 t̲ake w0
50W2 ← 17 3 r̲eshape 48 d̲rop w0
51s0 ← (W1, W2, 0.0 × W1, 0.0 × W2, 0.0 × W1, 0.0 × W2, 0.0)
52⍝ -- end of the core ------------------------------------------
53
54⍝# The loss, and the share of points right, of a state.
55ᵘl̲oss ← { (W1, W2, _, _, _, _, _) → Y ⁿⁿc̲rossEntropy W1 ᵘf̲orward W2 }
56ᵘr̲ight ← { (W1, W2, _, _, _, _, _) → (1 + arm) ⁿⁿa̲ccuracy W1 ᵘf̲orward W2 }
57⍝# The loss and the share of points right: before, after 100 steps,
58⍝# after 400.
59s1 ← 100 'ᵘa̲dam p̲ower s0
60s2 ← 300 'ᵘa̲dam p̲ower s1
61ᵘr̲ound ← { a → (f̲loat f̲loor 0.5 + 1000.0 × a) ÷ 1000.0 }
62ᵘr̲ound (ᵘl̲oss s0) c̲at (ᵘl̲oss s1) c̲at ᵘl̲oss s2
63ᵘr̲ound (ᵘr̲ight s0) c̲at (ᵘr̲ight s1) c̲at ᵘr̲ight s2