macro librarylibs/Net/src/Net.xtlm

Net: a network written as one line -- its layer sizes and activations become ordinary NN calls

Import it with an alias of your choice: "net:" u_se< "Net". Put libs/Net/src on XETAL_PATH ("just path"); the reference is libs/Net/docs. Names with m: are macros; the others are private to this file.

A macro is a function from the source text written left and right of its call to the source that replaces the call, before the program is compiled (X_eTaL MC10); xetal expand FILE shows the result. A spec is a list of words: sizes and activations, the input's size first: "784 128 relu 10 softmax". A size after the first is a dense layer (nn:d_ense, with its weights: one array per layer, the bias its last row); an activation is a function of NN (relu, sigmoid, tanh, softmax, logSoftmax). The expansion calls NN under the alias nn: (an expansion cannot ask which alias the caller chose), so import NN as "nn:" u_se< "NN".

source

Reading a spec (private)

ʰw̲ords : Char -> Box Char

function (private) · line 22

The words of t: runs of characters other than white space, boxed (after X_eTaL's Combinators.xtlm).

ʰw̲ords ← { t → (n̲ot t m̲ember? " \n\t") p̲artition t }

ʰj̲oin : Box Char -> Char

function (private) · line 24

The boxed texts b joined into one text ("" for none).

ʰj̲oin ← { b →
  0 = t̲ally b ? ""
  d̲isclose '{ x y → e̲nclose (d̲isclose x) c̲at d̲isclose y } r̲/ b
}

ʰs̲ize? : Truthy a => Char -> a

function (private) · line 29

Whether the text w is a whole number (a layer's size).

ʰs̲ize? ← { w → '∧ r̲/ w m̲ember? "0123456789" }

ʰa̲ct? : Truthy a => Char -> a

function (private) · line 31

Whether the text w names an activation.

ʰa̲ct? ← { w → '∨ r̲/ '{ a → w m̲atch d̲isclose a } e̲ach ʰw̲ords "relu sigmoid tanh softmax logSoftmax" }
Used in: ʰw̲rong

ʰa̲ct : Char -> Char

function (private) · line 33

An activation's function in NN: relu -> nn:r_elu.

ʰa̲ct ← { w → "nn:" c̲at (1 t̲ake w) c̲at "_" c̲at 1 d̲rop w }

ʰs̲izes : Box Char -> Int

function (private) · line 35

The spec's sizes, as Ints.

ʰs̲izes ← { ws → f̲loor n̲umbers ʰj̲oin '{ w → (ʰs̲ize? d̲isclose w) ? (d̲isclose w) c̲at " "◆ "" } m̲ap ws }

specShape : Char

value (private) · line 37

What a spec looks like, for the messages.

specShape ← "a spec is sizes and activations, the input's size first: \"784 128 relu 10 softmax\""
Used in: ʰw̲rong

ʰw̲rong : Box Char -> Char

function (private) · line 39

What is wrong with the spec's words ws, or "".

ʰw̲rong ← { ws →
  2 > t̲ally ws ? specShape
  n̲ot ʰs̲ize? d̲isclose f̲irst ws ? specShape
  bad ← '{ w → n̲ot (ʰs̲ize? d̲isclose w) ∨ ʰa̲ct? d̲isclose w } e̲ach ws
  0 < '+ r̲/ 0 + bad ? "not a size or an activation (relu, sigmoid, tanh, softmax, logSoftmax): " c̲at d̲isclose f̲irst (w̲here bad) s̲elect ws
  ""
}

Writing the forward pass (private)

ʰf̲unction : Box Char -> Box Char -> Char

function (private) · line 51

The forward function for the spec's words ws and the weight names ns: a lambda of nn: calls, the input x through each layer in turn.

ʰf̲unction ← { ws ns →
  size ← '{ w → ʰs̲ize? d̲isclose w } e̲ach ws
  k ← '+ s̲\ 0 + size
  pre ← '{ i →
    n̲ot i s̲elect size ? (ʰa̲ct d̲isclose i s̲elect ws) c̲at " "
    2 = i s̲elect k ? ""
    "("
  } m̲ap 1 d̲rop r̲ange t̲ally ws
  post ← '{ i →
    n̲ot i s̲elect size ? ""
    d ← " nn:d_ense " c̲at d̲isclose ((i s̲elect k) − 1) s̲elect ns
    2 = i s̲elect k ? d
    ")" c̲at d
  } m̲ap 1 d̲rop r̲ange t̲ally ws
  "{ x -> " c̲at (ʰj̲oin r̲ev pre) c̲at "x" c̲at (ʰj̲oin post) c̲at " }"
}

ʰc̲ount : Box Char -> a -> Char

function (private) · line 68

What is wrong with the weight names ns for the spec's words ws, or "".

ʰc̲ount ← { ws ns →
  dense ← -1 + '+ r̲/ '{ w → 0 + ʰs̲ize? d̲isclose w } e̲ach ws
  (t̲ally ns) = dense ? ""
  "name one weight array per dense layer: " c̲at (f̲ormat dense) c̲at " for this spec, " c̲at (f̲ormat t̲ally ns) c̲at " given"
}

Networks

ᵐn̲etwork< : Char -> Char -> Char

macro · line 80

A forward function for the network the spec on the left describes, using the weight arrays named on the right, one per dense layer in order: an ordinary lambda of nn: calls, a batch (one example per row) in, the last layer's output out.

ᵐn̲etwork< ← { spec weights →
  ws ← ʰw̲ords spec
  ns ← ʰw̲ords weights
  0 < t̲ally ʰw̲rong ws ? "bad-macro-argument left" ⎕R̲EJECT ʰw̲rong ws
  0 < t̲ally ws ʰc̲ount ns ? "bad-macro-argument right" ⎕R̲EJECT ws ʰc̲ount ns
  ws ʰf̲unction ns
}

ᵐm̲odel< : Char -> Char -> Char

macro · line 96

A whole model from one spec: on the left the function's name and a prefix ("u:d_eep c"), on the right the spec. It writes, as statements, each dense layer's weights read from data/<prefix>K.txt and reshaped as the spec says (inputs + 1 by outputs; the program stops with an error naming the layer when a file does not hold that many numbers, since r_eshape alone would silently repeat or cut the data), then the function, as n_etwork< writes it. The spec is the one place the network is described.

ᵐm̲odel< ← { what spec →
  ws ← ʰw̲ords spec
  0 < t̲ally ʰw̲rong ws ? "bad-macro-argument right" ⎕R̲EJECT ʰw̲rong ws
  nm ← ʰw̲ords what
  2 ≠ t̲ally nm ? "bad-macro-argument left" ⎕R̲EJECT "write the function's name and the weight files' prefix: \"u:d_eep c\" (for data/c1.txt, data/c2.txt, ...)"
  name ← d̲isclose f̲irst nm
  prefix ← d̲isclose 2 s̲elect nm
  s ← ʰs̲izes ws
  ns ← '{ i → prefix c̲at f̲ormat i } m̲ap r̲ange -1 + t̲ally s
  load ← '{ i →
    w ← d̲isclose i s̲elect ns
    need ← (1 + i s̲elect s) × (i + 1) s̲elect s
    w c̲at " := { v -> " c̲at (f̲ormat need) c̲at " = t_ally v ? " c̲at (f̲ormat 1 + i s̲elect s) c̲at " " c̲at (f̲ormat (i + 1) s̲elect s) c̲at " r_eshape v; []P_ANIC \"data/" c̲at w c̲at ".txt holds \" c_at (f_ormat t_ally v) c_at \" numbers; layer " c̲at (f̲ormat i) c̲at " of \\\"" c̲at spec c̲at "\\\" (" c̲at (f̲ormat i s̲elect s) c̲at " inputs and a bias, " c̲at (f̲ormat (i + 1) s̲elect s) c̲at " outputs) needs " c̲at (f̲ormat need) c̲at "\" } n_umbers []N_GET \"data/" c̲at w c̲at ".txt\"\n"
  } m̲ap r̲ange -1 + t̲ally s
  (ʰj̲oin load) c̲at name c̲at " := " c̲at ws ʰf̲unction ns
}
Used in: a1, b1, c1

Writing backpropagation (private)

ʰa̲ctFn : Char -> Char

function (private) · line 116

The function of NN for an activation word (relu -> nn:r_elu); "" for none.

ʰa̲ctFn ← { w → 0 = t̲ally w ? ""◆ (ʰa̲ct w) c̲at " " }
Used in: ʰb̲ackprop

ʰs̲lope : Char -> Char -> Char

function (private) · line 119

Its slope at its output A (the name of the array), times which the error goes back through it; "" when there is none.

ʰs̲lope ← { w a →
  w m̲atch "relu" ? " * f_loat " c̲at a c̲at " > 0.0"
  w m̲atch "tanh" ? " * 1.0 - " c̲at a c̲at " * " c̲at a
  w m̲atch "sigmoid" ? " * " c̲at a c̲at " * 1.0 - " c̲at a
  ""
}
Used in: ʰb̲ackprop

ʰa̲cts : Box Char -> Box Char

function (private) · line 127

The activation after each dense layer of the spec's words ws (a box per layer, "" for none).

ʰa̲cts ← { ws →
  size ← '{ w → ʰs̲ize? d̲isclose w } e̲ach ws
  at ← 1 d̲rop w̲here size
  '{ p → (p < t̲ally ws) ∧ (n̲ot (p + 1) s̲elect size) ? d̲isclose (p + 1) s̲elect ws◆ "" } m̲ap at
}

ʰt̲xt : a -> Char

function (private) · line 133

N as text.

ʰt̲xt ← { n → f̲ormat n }

ʰt̲uple : Char -> Int -> Char

function (private) · line 137

The names P1 .. PL as X_eTaL's tuple text, "(P1, P2)"; one alone is not a tuple, "P1" (X_eTaL has no one-element tuple).

ʰt̲uple ← { p L →
  L = 1 ? p c̲at "1"
  "(" c̲at (ʰj̲oin '{ l → (l = 1) ? p c̲at ʰt̲xt l◆ ", " c̲at p c̲at ʰt̲xt l } m̲ap r̲ange L) c̲at ")"
}

ʰb̲ackprop : Box Char -> Box Char -> Char

function (private) · line 146

The lines that run the spec's network forward from the inputs (the weights W1 .. WL bound by the caller) and back from the targets, and bind each layer's gradient G1 .. GL. ws: the spec's words; xy: the inputs' and targets' names, boxed.

ʰb̲ackprop ← { ws xy →
  x ← d̲isclose 1 s̲elect xy
  y ← d̲isclose 2 s̲elect xy
  acts ← ʰa̲cts ws
  L ← t̲ally acts
  r ← r̲ange L
  fwd ← ʰj̲oin '{ l → "  A" c̲at (ʰt̲xt l) c̲at " := " c̲at (ʰa̲ctFn d̲isclose l s̲elect acts) c̲at "A" c̲at (ʰt̲xt l − 1) c̲at " nn:d_ense W" c̲at (ʰt̲xt l) c̲at "\n" } m̲ap r
  dL ← "  D" c̲at (ʰt̲xt L) c̲at " := (A" c̲at (ʰt̲xt L) c̲at " - " c̲at y c̲at ") / f_loat t_ally " c̲at x c̲at "\n"
  back ← ʰj̲oin '{ l → "  D" c̲at (ʰt̲xt l) c̲at " := (D" c̲at (ʰt̲xt l + 1) c̲at " '+ '* i_nner o_\\ -1 d_rop W" c̲at (ʰt̲xt l + 1) c̲at ")" c̲at ((d̲isclose l s̲elect acts) ʰs̲lope "A" c̲at ʰt̲xt l) c̲at "\n" } m̲ap r̲ev -1 d̲rop r
  grads ← ʰj̲oin '{ l → "  G" c̲at (ʰt̲xt l) c̲at " := (o_\\ A" c̲at (ʰt̲xt l − 1) c̲at " c_at_2 ((t_ally A" c̲at (ʰt̲xt l − 1) c̲at ") c_at 1) r_eshape 1.0) '+ '* i_nner D" c̲at (ʰt̲xt l) c̲at "\n" } m̲ap r
  "  A0 := " c̲at x c̲at "\n" c̲at fwd c̲at dL c̲at back c̲at grads
}

ʰt̲rainable : Box Char -> Char

function (private) · line 159

What is wrong with a trainable spec's words ws, or "".

ʰt̲rainable ← { ws →
  0 < t̲ally ʰw̲rong ws ? ʰw̲rong ws
  n̲ot (d̲isclose f̲irst r̲ev ws) m̲atch "softmax" ? "a network trained with cross-entropy ends in softmax"
  ""
}

Training

ᵐg̲radient< : Char -> Char -> Char

macro · line 173

The gradient of the cross-entropy loss by each layer's weights, for the network the spec on the right describes: on the left the function's name, the inputs' and the targets' (one-hot rows): "u:g_rad X Y". The function takes the weights as a tuple (W1, .., WL) and gives the gradients as a tuple (G1, .., GL), each the shape of its layer's weights: backpropagation, written out.

ᵐg̲radient< ← { what spec →
  ws ← ʰw̲ords spec
  0 < t̲ally ʰt̲rainable ws ? "bad-macro-argument right" ⎕R̲EJECT ʰt̲rainable ws
  nm ← ʰw̲ords what
  3 ≠ t̲ally nm ? "bad-macro-argument left" ⎕R̲EJECT "write the function's name, then the inputs' and the targets': \"u:g_rad X Y\""
  L ← t̲ally ʰa̲cts ws
  (d̲isclose 1 s̲elect nm) c̲at " := { " c̲at ("W" ʰt̲uple L) c̲at " ->\n" c̲at (ws ʰb̲ackprop 1 d̲rop nm) c̲at "  " c̲at ("G" ʰt̲uple L) c̲at "\n}"
}

ᵐt̲rain< : Char -> Char -> Char

macro · line 190

A training step for the network the spec on the right describes: on the left the function's name, the inputs', the targets' (one-hot rows) and the learning rate's ("u:s_tep X Y lr"). The step takes the state, a tuple of 3L + 1 arrays (each layer's weights, Adam's running averages of gradients and of their squares, the step count; @ net:s_tate< writes the first), and gives the next: the forward pass, each layer's gradient by backpropagation (as g_radient< writes it), and Adam. The spec must end in softmax.

ᵐt̲rain< ← { what spec →
  ws ← ʰw̲ords spec
  0 < t̲ally ʰt̲rainable ws ? "bad-macro-argument right" ⎕R̲EJECT ʰt̲rainable ws
  nm ← ʰw̲ords what
  4 ≠ t̲ally nm ? "bad-macro-argument left" ⎕R̲EJECT "write the step's name, then the inputs', the targets' and the learning rate's: \"u:s_tep X Y lr\""
  lr ← d̲isclose 4 s̲elect nm
  L ← t̲ally ʰa̲cts ws
  r ← r̲ange L
  adam ← ʰj̲oin '{ l →
    i ← ʰt̲xt l
    "  M" c̲at i c̲at " := (0.9 * M" c̲at i c̲at ") + 0.1 * G" c̲at i c̲at "\n  V" c̲at i c̲at " := (0.999 * V" c̲at i c̲at ") + 0.001 * G" c̲at i c̲at " * G" c̲at i c̲at "\n  W" c̲at i c̲at " := W" c̲at i c̲at " - " c̲at lr c̲at " * (M" c̲at i c̲at " / 1.0 - 0.9 ^ k) / 0.00000001 + (V" c̲at i c̲at " / 1.0 - 0.999 ^ k) ^ 0.5\n"
  } m̲ap r
  names ← ʰj̲oin '{ l → "W" c̲at (ʰt̲xt l) c̲at ", " } m̲ap r
  names ← names c̲at (ʰj̲oin '{ l → "M" c̲at (ʰt̲xt l) c̲at ", " } m̲ap r) c̲at ʰj̲oin '{ l → "V" c̲at (ʰt̲xt l) c̲at ", " } m̲ap r
  state ← "(" c̲at names c̲at "k)"
  (d̲isclose 1 s̲elect nm) c̲at " := { " c̲at state c̲at " ->\n" c̲at (ws ʰb̲ackprop 1 d̲rop 3 t̲ake nm) c̲at "  k := 1.0 + k\n" c̲at adam c̲at "  " c̲at state c̲at "\n}"
}

ᵐs̲tate< : Unit -> Char -> Char

macro · line 210

The starting state for a training step (t_rain<) from the weight arrays named on the right: the weights, zero averages, step 0.

ᵐs̲tate< ← { @ weights →
  ns ← ʰw̲ords weights
  0 = t̲ally ns ? "bad-macro-argument right" ⎕R̲EJECT "name the starting weight arrays, one per dense layer"
  ws ← ʰj̲oin '{ w → (d̲isclose w) c̲at ", " } m̲ap ns
  zs ← ʰj̲oin '{ w → "0.0 * " c̲at (d̲isclose w) c̲at ", " } m̲ap ns
  "(" c̲at ws c̲at zs c̲at zs c̲at "0.0)"
}
Used in: %2, s0

Counting and checking

ᵐp̲arams< : Char -> Unit -> Char

macro · line 226

How many numbers the network of the spec on the left has to learn (weights and biases), worked out when the program is compiled: the call becomes that number. Nothing is written on the right: @.

      ⁿᵉᵗ⁼u̲se< "Net"
      "2 16 relu 16 relu 3 softmax" ⁿᵉᵗp̲arams< @
371
ᵐp̲arams< ← { spec @ →
  ws ← ʰw̲ords spec
  0 < t̲ally ʰw̲rong ws ? "bad-macro-argument left" ⎕R̲EJECT ʰw̲rong ws
  s ← ʰs̲izes ws
  f̲ormat '+ r̲/ (1 + -1 d̲rop s) × 1 d̲rop s
}

ᵐs̲hapes< : Char -> Char -> Char

macro · line 241

Whether the weight arrays named on the right have the shapes the spec on the left says (a layer from m inputs to n outputs is m + 1 by n: the bias is the last row): 1 or 0.

      ⁿᵉᵗ⁼u̲se< "Net"
      w1 ← 3 4 r̲eshape 0.0
      w2 ← 5 3 r̲eshape 0.0
      "2 4 tanh 3 softmax" ⁿᵉᵗs̲hapes< "w1 w2"
1
ᵐs̲hapes< ← { spec weights →
  ws ← ʰw̲ords spec
  ns ← ʰw̲ords weights
  0 < t̲ally ʰw̲rong ws ? "bad-macro-argument left" ⎕R̲EJECT ʰw̲rong ws
  s ← ʰs̲izes ws
  0 < t̲ally ws ʰc̲ount ns ? "bad-macro-argument right" ⎕R̲EJECT ws ʰc̲ount ns
  one ← '{ i → "((s_hape " c̲at (d̲isclose i s̲elect ns) c̲at ") m_atch " c̲at (f̲ormat 1 + i s̲elect s) c̲at " " c̲at (f̲ormat (i + 1) s̲elect s) c̲at ")" } m̲ap r̲ange t̲ally ns
  ʰj̲oin '{ i → (i = 1) ? d̲isclose i s̲elect one◆ " & " c̲at d̲isclose i s̲elect one } m̲ap r̲ange t̲ally ns
}