sourcedemos/net-macro/train-it.xtl

1⍝!/usr/bin/env xetal 2⍝# Training a network from its one-line spec: the Net macro library 3⍝# writes, from "2 16 relu 16 relu 3 softmax", the network's backward 4⍝# pass and an Adam step (net:t_rain<) and the starting state 5⍝# (net:s_tate<); p_ower repeats the step. The task is net-macro's: which 6⍝# of three spiral arms is a point on? Here the 300 training points are 7⍝# made in X_eTaL, and the weights start small and random. 8 9ⁿⁿ⁼u̲se< "NN" 10ⁿᵉᵗ⁼u̲se< "Net" 11 12⍝## The spiral: 100 points on each of three arms 13n ← 300 14i ← o̲ffsets n 15arm ← i d̲iv 100 16t ← (0.5 + f̲loat i m̲od 100) ÷ 100.0 17a ← (2.0944 × f̲loat arm) + (5.5 × t) + 0.1 × (f̲loat (r̲oll! n r̲eshape 1001) − 501) ÷ 500.0 18X ← o̲\ (2 c̲at n) r̲eshape ((0.1 + 0.9 × t) × c̲os a) c̲at (0.1 + 0.9 × t) × s̲in a 19Y ← 3 ⁿⁿo̲neHot 1 + arm 20lr ← 0.02 21 22⍝## The network and its training, from one spec 23"u:s_tep X Y lr" ⁿᵉᵗt̲rain< "2 16 relu 16 relu 3 softmax"
ⁿᵉᵗt̲rain< expands to
ᵘs̲tep ← { (W1, W2, W3, M1, M2, M3, V1, V2, V3, k) →
  A0 ← X
  A1 ← ⁿⁿr̲elu A0 ⁿⁿd̲ense W1
  A2 ← ⁿⁿr̲elu A1 ⁿⁿd̲ense W2
  A3 ← ⁿⁿs̲oftmax A2 ⁿⁿd̲ense W3
  D3 ← (A3 − Y) ÷ f̲loat t̲ally X
  D2 ← (D3 '+ '× i̲nner o̲\ -1 d̲rop W3) × f̲loat A2 > 0.0
  D1 ← (D2 '+ '× i̲nner o̲\ -1 d̲rop W2) × f̲loat A1 > 0.0
  G1 ← (o̲\ A0 c̲at₂ ((t̲ally A0) c̲at 1) r̲eshape 1.0) '+ '× i̲nner D1
  G2 ← (o̲\ A1 c̲at₂ ((t̲ally A1) c̲at 1) r̲eshape 1.0) '+ '× i̲nner D2
  G3 ← (o̲\ A2 c̲at₂ ((t̲ally A2) c̲at 1) r̲eshape 1.0) '+ '× i̲nner D3
  k ← 1.0 + k
  M1 ← (0.9 × M1) + 0.1 × G1
  V1 ← (0.999 × V1) + 0.001 × G1 × G1
  W1 ← W1 − lr × (M1 ÷ 1.0 − 0.9 ^ k) ÷ 0.00000001 + (V1 ÷ 1.0 − 0.999 ^ k) ^ 0.5
  M2 ← (0.9 × M2) + 0.1 × G2
  V2 ← (0.999 × V2) + 0.001 × G2 × G2
  W2 ← W2 − lr × (M2 ÷ 1.0 − 0.9 ^ k) ÷ 0.00000001 + (V2 ÷ 1.0 − 0.999 ^ k) ^ 0.5
  M3 ← (0.9 × M3) + 0.1 × G3
  V3 ← (0.999 × V3) + 0.001 × G3 × G3
  W3 ← W3 − lr × (M3 ÷ 1.0 − 0.9 ^ k) ÷ 0.00000001 + (V3 ÷ 1.0 − 0.999 ^ k) ^ 0.5
  (W1, W2, W3, M1, M2, M3, V1, V2, V3, k)
}
24⍝# Small random weights of a shape (inputs + 1 by outputs), then the state. 25ᵘr̲andom ← { sh → sh r̲eshape (f̲loat (r̲oll! ('× r̲/ sh) r̲eshape 2001) − 1001) ÷ 2000.0 } 26w1 ← ᵘr̲andom 3 16 27w2 ← ᵘr̲andom 17 16 28w3 ← ᵘr̲andom 17 3 29s0 ← @ ⁿᵉᵗs̲tate< "w1 w2 w3"
ⁿᵉᵗs̲tate< expands to
((w1, w2, w3, 0.0 × w1, 0.0 × w2, 0.0 × w3, 0.0 × w1, 0.0 × w2, 0.0 × w3, 0.0))
30ᵘs̲tart ← "2 16 relu 16 relu 3 softmax" ⁿᵉᵗn̲etwork< "w1 w2 w3"
ⁿᵉᵗn̲etwork< expands to
({ x → ⁿⁿs̲oftmax (ⁿⁿr̲elu (ⁿⁿr̲elu x ⁿⁿd̲ense w1) ⁿⁿd̲ense w2) ⁿⁿd̲ense w3 })
31⍝ -- end of the core ------------------------------------------ 32 33⍝# Four hundred steps; then the trained weights, out of the state, and 34⍝# the network on them. 35(w1, w2, w3, _, _, _, _, _, _, _) ← 400 'ᵘs̲tep p̲ower s0 36ᵘt̲rained ← "2 16 relu 16 relu 3 softmax" ⁿᵉᵗn̲etwork< "w1 w2 w3"
ⁿᵉᵗn̲etwork< expands to
({ x → ⁿⁿs̲oftmax (ⁿⁿr̲elu (ⁿⁿr̲elu x ⁿⁿd̲ense w1) ⁿⁿd̲ense w2) ⁿⁿd̲ense w3 })
37⍝# The loss and the share of the 300 points right, before and after. 38ᵘr̲ound ← { a → (f̲loat f̲loor 0.5 + 10000.0 × a) ÷ 10000.0 } 39ᵘr̲ound (Y ⁿⁿc̲rossEntropy ᵘs̲tart X) c̲at Y ⁿⁿc̲rossEntropy ᵘt̲rained X 40ᵘr̲ound ((1 + arm) ⁿⁿa̲ccuracy ᵘs̲tart X) c̲at (1 + arm) ⁿⁿa̲ccuracy ᵘt̲rained X