macro librarylibs/Net/src/Net.xtlm
Net: a network written as one line -- its layer sizes and activations become ordinary NN calls
Import it with an alias of your choice: "net:" u_se< "Net". Put libs/Net/src on XETAL_PATH ("just path"); the reference is libs/Net/docs. Names with m: are macros; the others are private to this file.
A macro is a function from the source text written left and right of its call to the source that replaces the call, before the program is compiled (X_eTaL MC10); xetal expand FILE shows the result. A spec is a list of words: sizes and activations, the input's size first: "784 128 relu 10 softmax". A size after the first is a dense layer (nn:d_ense, with its weights: one array per layer, the bias its last row); an activation is a function of NN (relu, sigmoid, tanh, softmax, logSoftmax). The expansion calls NN under the alias nn: (an expansion cannot ask which alias the caller chose), so import NN as "nn:" u_se< "NN".
Reading a spec (private)
ʰw̲ords : Char -> Box Char
The words of t: runs of characters other than white space, boxed (after X_eTaL's Combinators.xtlm).
ʰw̲ords ← { t → (n̲ot t m̲ember? " \n\t") p̲artition t }
ʰj̲oin : Box Char -> Char
The boxed texts b joined into one text ("" for none).
ʰj̲oin ← { b → 0 = t̲ally b ? "" d̲isclose '{ x y → e̲nclose (d̲isclose x) c̲at d̲isclose y } r̲/ b }
ʰs̲ize? : Truthy a => Char -> a
Whether the text w is a whole number (a layer's size).
ʰs̲ize? ← { w → '∧ r̲/ w m̲ember? "0123456789" }
ʰa̲ct? : Truthy a => Char -> a
Whether the text w names an activation.
ʰa̲ct? ← { w → '∨ r̲/ '{ a → w m̲atch d̲isclose a } e̲ach ʰw̲ords "relu sigmoid tanh softmax logSoftmax" }
ʰa̲ct : Char -> Char
An activation's function in NN: relu -> nn:r_elu.
ʰa̲ct ← { w → "nn:" c̲at (1 t̲ake w) c̲at "_" c̲at 1 d̲rop w }
ʰs̲izes : Box Char -> Int
The spec's sizes, as Ints.
ʰs̲izes ← { ws → f̲loor n̲umbers ʰj̲oin '{ w → (ʰs̲ize? d̲isclose w) ? (d̲isclose w) c̲at " "◆ "" } m̲ap ws }
specShape : Char
What a spec looks like, for the messages.
specShape ← "a spec is sizes and activations, the input's size first: \"784 128 relu 10 softmax\""
ʰw̲rong : Box Char -> Char
What is wrong with the spec's words ws, or "".
ʰw̲rong ← { ws → 2 > t̲ally ws ? specShape n̲ot ʰs̲ize? d̲isclose f̲irst ws ? specShape bad ← '{ w → n̲ot (ʰs̲ize? d̲isclose w) ∨ ʰa̲ct? d̲isclose w } e̲ach ws 0 < '+ r̲/ 0 + bad ? "not a size or an activation (relu, sigmoid, tanh, softmax, logSoftmax): " c̲at d̲isclose f̲irst (w̲here bad) s̲elect ws "" }
Writing the forward pass (private)
ʰf̲unction : Box Char -> Box Char -> Char
The forward function for the spec's words ws and the weight names ns: a lambda of nn: calls, the input x through each layer in turn.
ʰf̲unction ← { ws ns → size ← '{ w → ʰs̲ize? d̲isclose w } e̲ach ws k ← '+ s̲\ 0 + size pre ← '{ i → n̲ot i s̲elect size ? (ʰa̲ct d̲isclose i s̲elect ws) c̲at " " 2 = i s̲elect k ? "" "(" } m̲ap 1 d̲rop r̲ange t̲ally ws post ← '{ i → n̲ot i s̲elect size ? "" d ← " nn:d_ense " c̲at d̲isclose ((i s̲elect k) − 1) s̲elect ns 2 = i s̲elect k ? d ")" c̲at d } m̲ap 1 d̲rop r̲ange t̲ally ws "{ x -> " c̲at (ʰj̲oin r̲ev pre) c̲at "x" c̲at (ʰj̲oin post) c̲at " }" }
ʰc̲ount : Box Char -> a -> Char
What is wrong with the weight names ns for the spec's words ws, or "".
ʰc̲ount ← { ws ns → dense ← -1 + '+ r̲/ '{ w → 0 + ʰs̲ize? d̲isclose w } e̲ach ws (t̲ally ns) = dense ? "" "name one weight array per dense layer: " c̲at (f̲ormat dense) c̲at " for this spec, " c̲at (f̲ormat t̲ally ns) c̲at " given" }
Networks
ᵐn̲etwork< : Char -> Char -> Char
A forward function for the network the spec on the left describes, using the weight arrays named on the right, one per dense layer in order: an ordinary lambda of nn: calls, a batch (one example per row) in, the last layer's output out.
ᵐn̲etwork< ← { spec weights → ws ← ʰw̲ords spec ns ← ʰw̲ords weights 0 < t̲ally ʰw̲rong ws ? "bad-macro-argument left" ⎕R̲EJECT ʰw̲rong ws 0 < t̲ally ws ʰc̲ount ns ? "bad-macro-argument right" ⎕R̲EJECT ws ʰc̲ount ns ws ʰf̲unction ns }
ᵐm̲odel< : Char -> Char -> Char
A whole model from one spec: on the left the function's name and a prefix ("u:d_eep c"), on the right the spec. It writes, as statements, each dense layer's weights read from data/<prefix>K.txt and reshaped as the spec says (inputs + 1 by outputs; the program stops with an error naming the layer when a file does not hold that many numbers, since r_eshape alone would silently repeat or cut the data), then the function, as n_etwork< writes it. The spec is the one place the network is described.
ᵐm̲odel< ← { what spec → ws ← ʰw̲ords spec 0 < t̲ally ʰw̲rong ws ? "bad-macro-argument right" ⎕R̲EJECT ʰw̲rong ws nm ← ʰw̲ords what 2 ≠ t̲ally nm ? "bad-macro-argument left" ⎕R̲EJECT "write the function's name and the weight files' prefix: \"u:d_eep c\" (for data/c1.txt, data/c2.txt, ...)" name ← d̲isclose f̲irst nm prefix ← d̲isclose 2 s̲elect nm s ← ʰs̲izes ws ns ← '{ i → prefix c̲at f̲ormat i } m̲ap r̲ange -1 + t̲ally s load ← '{ i → w ← d̲isclose i s̲elect ns need ← (1 + i s̲elect s) × (i + 1) s̲elect s w c̲at " := { v -> " c̲at (f̲ormat need) c̲at " = t_ally v ? " c̲at (f̲ormat 1 + i s̲elect s) c̲at " " c̲at (f̲ormat (i + 1) s̲elect s) c̲at " r_eshape v; []P_ANIC \"data/" c̲at w c̲at ".txt holds \" c_at (f_ormat t_ally v) c_at \" numbers; layer " c̲at (f̲ormat i) c̲at " of \\\"" c̲at spec c̲at "\\\" (" c̲at (f̲ormat i s̲elect s) c̲at " inputs and a bias, " c̲at (f̲ormat (i + 1) s̲elect s) c̲at " outputs) needs " c̲at (f̲ormat need) c̲at "\" } n_umbers []N_GET \"data/" c̲at w c̲at ".txt\"\n" } m̲ap r̲ange -1 + t̲ally s (ʰj̲oin load) c̲at name c̲at " := " c̲at ws ʰf̲unction ns }
Writing backpropagation (private)
ʰa̲ctFn : Char -> Char
The function of NN for an activation word (relu -> nn:r_elu); "" for none.
ʰa̲ctFn ← { w → 0 = t̲ally w ? ""◆ (ʰa̲ct w) c̲at " " }
ʰs̲lope : Char -> Char -> Char
Its slope at its output A (the name of the array), times which the error goes back through it; "" when there is none.
ʰs̲lope ← { w a → w m̲atch "relu" ? " * f_loat " c̲at a c̲at " > 0.0" w m̲atch "tanh" ? " * 1.0 - " c̲at a c̲at " * " c̲at a w m̲atch "sigmoid" ? " * " c̲at a c̲at " * 1.0 - " c̲at a "" }
ʰa̲cts : Box Char -> Box Char
The activation after each dense layer of the spec's words ws (a box per layer, "" for none).
ʰa̲cts ← { ws → size ← '{ w → ʰs̲ize? d̲isclose w } e̲ach ws at ← 1 d̲rop w̲here size '{ p → (p < t̲ally ws) ∧ (n̲ot (p + 1) s̲elect size) ? d̲isclose (p + 1) s̲elect ws◆ "" } m̲ap at }
ʰt̲uple : Char -> Int -> Char
The names P1 .. PL as X_eTaL's tuple text, "(P1, P2)"; one alone is not a tuple, "P1" (X_eTaL has no one-element tuple).
ʰt̲uple ← { p L → L = 1 ? p c̲at "1" "(" c̲at (ʰj̲oin '{ l → (l = 1) ? p c̲at ʰt̲xt l◆ ", " c̲at p c̲at ʰt̲xt l } m̲ap r̲ange L) c̲at ")" }
ʰb̲ackprop : Box Char -> Box Char -> Char
The lines that run the spec's network forward from the inputs (the weights W1 .. WL bound by the caller) and back from the targets, and bind each layer's gradient G1 .. GL. ws: the spec's words; xy: the inputs' and targets' names, boxed.
ʰb̲ackprop ← { ws xy → x ← d̲isclose 1 s̲elect xy y ← d̲isclose 2 s̲elect xy acts ← ʰa̲cts ws L ← t̲ally acts r ← r̲ange L fwd ← ʰj̲oin '{ l → " A" c̲at (ʰt̲xt l) c̲at " := " c̲at (ʰa̲ctFn d̲isclose l s̲elect acts) c̲at "A" c̲at (ʰt̲xt l − 1) c̲at " nn:d_ense W" c̲at (ʰt̲xt l) c̲at "\n" } m̲ap r dL ← " D" c̲at (ʰt̲xt L) c̲at " := (A" c̲at (ʰt̲xt L) c̲at " - " c̲at y c̲at ") / f_loat t_ally " c̲at x c̲at "\n" back ← ʰj̲oin '{ l → " D" c̲at (ʰt̲xt l) c̲at " := (D" c̲at (ʰt̲xt l + 1) c̲at " '+ '* i_nner o_\\ -1 d_rop W" c̲at (ʰt̲xt l + 1) c̲at ")" c̲at ((d̲isclose l s̲elect acts) ʰs̲lope "A" c̲at ʰt̲xt l) c̲at "\n" } m̲ap r̲ev -1 d̲rop r grads ← ʰj̲oin '{ l → " G" c̲at (ʰt̲xt l) c̲at " := (o_\\ A" c̲at (ʰt̲xt l − 1) c̲at " c_at_2 ((t_ally A" c̲at (ʰt̲xt l − 1) c̲at ") c_at 1) r_eshape 1.0) '+ '* i_nner D" c̲at (ʰt̲xt l) c̲at "\n" } m̲ap r " A0 := " c̲at x c̲at "\n" c̲at fwd c̲at dL c̲at back c̲at grads }
ʰt̲rainable : Box Char -> Char
What is wrong with a trainable spec's words ws, or "".
ʰt̲rainable ← { ws → 0 < t̲ally ʰw̲rong ws ? ʰw̲rong ws n̲ot (d̲isclose f̲irst r̲ev ws) m̲atch "softmax" ? "a network trained with cross-entropy ends in softmax" "" }
Training
ᵐg̲radient< : Char -> Char -> Char
The gradient of the cross-entropy loss by each layer's weights, for the network the spec on the right describes: on the left the function's name, the inputs' and the targets' (one-hot rows): "u:g_rad X Y". The function takes the weights as a tuple (W1, .., WL) and gives the gradients as a tuple (G1, .., GL), each the shape of its layer's weights: backpropagation, written out.
ᵐg̲radient< ← { what spec → ws ← ʰw̲ords spec 0 < t̲ally ʰt̲rainable ws ? "bad-macro-argument right" ⎕R̲EJECT ʰt̲rainable ws nm ← ʰw̲ords what 3 ≠ t̲ally nm ? "bad-macro-argument left" ⎕R̲EJECT "write the function's name, then the inputs' and the targets': \"u:g_rad X Y\"" L ← t̲ally ʰa̲cts ws (d̲isclose 1 s̲elect nm) c̲at " := { " c̲at ("W" ʰt̲uple L) c̲at " ->\n" c̲at (ws ʰb̲ackprop 1 d̲rop nm) c̲at " " c̲at ("G" ʰt̲uple L) c̲at "\n}" }
ᵐt̲rain< : Char -> Char -> Char
A training step for the network the spec on the right describes: on the left the function's name, the inputs', the targets' (one-hot rows) and the learning rate's ("u:s_tep X Y lr"). The step takes the state, a tuple of 3L + 1 arrays (each layer's weights, Adam's running averages of gradients and of their squares, the step count; @ net:s_tate< writes the first), and gives the next: the forward pass, each layer's gradient by backpropagation (as g_radient< writes it), and Adam. The spec must end in softmax.
ᵐt̲rain< ← { what spec → ws ← ʰw̲ords spec 0 < t̲ally ʰt̲rainable ws ? "bad-macro-argument right" ⎕R̲EJECT ʰt̲rainable ws nm ← ʰw̲ords what 4 ≠ t̲ally nm ? "bad-macro-argument left" ⎕R̲EJECT "write the step's name, then the inputs', the targets' and the learning rate's: \"u:s_tep X Y lr\"" lr ← d̲isclose 4 s̲elect nm L ← t̲ally ʰa̲cts ws r ← r̲ange L adam ← ʰj̲oin '{ l → i ← ʰt̲xt l " M" c̲at i c̲at " := (0.9 * M" c̲at i c̲at ") + 0.1 * G" c̲at i c̲at "\n V" c̲at i c̲at " := (0.999 * V" c̲at i c̲at ") + 0.001 * G" c̲at i c̲at " * G" c̲at i c̲at "\n W" c̲at i c̲at " := W" c̲at i c̲at " - " c̲at lr c̲at " * (M" c̲at i c̲at " / 1.0 - 0.9 ^ k) / 0.00000001 + (V" c̲at i c̲at " / 1.0 - 0.999 ^ k) ^ 0.5\n" } m̲ap r names ← ʰj̲oin '{ l → "W" c̲at (ʰt̲xt l) c̲at ", " } m̲ap r names ← names c̲at (ʰj̲oin '{ l → "M" c̲at (ʰt̲xt l) c̲at ", " } m̲ap r) c̲at ʰj̲oin '{ l → "V" c̲at (ʰt̲xt l) c̲at ", " } m̲ap r state ← "(" c̲at names c̲at "k)" (d̲isclose 1 s̲elect nm) c̲at " := { " c̲at state c̲at " ->\n" c̲at (ws ʰb̲ackprop 1 d̲rop 3 t̲ake nm) c̲at " k := 1.0 + k\n" c̲at adam c̲at " " c̲at state c̲at "\n}" }
ᵐs̲tate< : Unit -> Char -> Char
The starting state for a training step (t_rain<) from the weight arrays named on the right: the weights, zero averages, step 0.
ᵐs̲tate< ← { @ weights → ns ← ʰw̲ords weights 0 = t̲ally ns ? "bad-macro-argument right" ⎕R̲EJECT "name the starting weight arrays, one per dense layer" ws ← ʰj̲oin '{ w → (d̲isclose w) c̲at ", " } m̲ap ns zs ← ʰj̲oin '{ w → "0.0 * " c̲at (d̲isclose w) c̲at ", " } m̲ap ns "(" c̲at ws c̲at zs c̲at zs c̲at "0.0)" }
Counting and checking
ᵐp̲arams< : Char -> Unit -> Char
How many numbers the network of the spec on the left has to learn (weights and biases), worked out when the program is compiled: the call becomes that number. Nothing is written on the right: @.
ⁿᵉᵗ⁼u̲se< "Net"
"2 16 relu 16 relu 3 softmax" ⁿᵉᵗp̲arams< @ 371
ᵐp̲arams< ← { spec @ → ws ← ʰw̲ords spec 0 < t̲ally ʰw̲rong ws ? "bad-macro-argument left" ⎕R̲EJECT ʰw̲rong ws s ← ʰs̲izes ws f̲ormat '+ r̲/ (1 + -1 d̲rop s) × 1 d̲rop s }
ᵐs̲hapes< : Char -> Char -> Char
Whether the weight arrays named on the right have the shapes the spec on the left says (a layer from m inputs to n outputs is m + 1 by n: the bias is the last row): 1 or 0.
ⁿᵉᵗ⁼u̲se< "Net"
w1 ← 3 4 r̲eshape 0.0
w2 ← 5 3 r̲eshape 0.0
"2 4 tanh 3 softmax" ⁿᵉᵗs̲hapes< "w1 w2" 1
ᵐs̲hapes< ← { spec weights → ws ← ʰw̲ords spec ns ← ʰw̲ords weights 0 < t̲ally ʰw̲rong ws ? "bad-macro-argument left" ⎕R̲EJECT ʰw̲rong ws s ← ʰs̲izes ws 0 < t̲ally ws ʰc̲ount ns ? "bad-macro-argument right" ⎕R̲EJECT ws ʰc̲ount ns one ← '{ i → "((s_hape " c̲at (d̲isclose i s̲elect ns) c̲at ") m_atch " c̲at (f̲ormat 1 + i s̲elect s) c̲at " " c̲at (f̲ormat (i + 1) s̲elect s) c̲at ")" } m̲ap r̲ange t̲ally ns ʰj̲oin '{ i → (i = 1) ? d̲isclose i s̲elect one◆ " & " c̲at d̲isclose i s̲elect one } m̲ap r̲ange t̲ally ns }