librarylibs/NN/src/NN.xtl
NN: neural-network building blocks -- activations, softmax by row, dense layers, one-hot, loss, accuracy
Import it with an alias of your choice: "nn:" u_se< "NN". Put libs/NN/src on XETAL_PATH ("just path"); the reference is libs/NN/docs. Names with l: are exported; the others are private to this file.
A batch is a matrix, one example per row; a layer's weights are a
matrix with one row per input and one column per output, and its
bias is one more row at the bottom (so a layer is one array, and
x ⁿⁿd̲ense wb takes two arguments). Activations work item by item
on any shape. Softmax, log-softmax and argmax work along the last
axis of any array: a vector is one example, a matrix one per row, a
rank-3 array a batch of sequences.
Helpers (private)
ʰs̲pread : a -> b -> a
The row values r (one per row of m) spread across m's columns.
ʰs̲pread ← { r m → o̲\ ((-1 t̲ake s̲hape m) c̲at t̲ally r) r̲eshape r }
ʰr̲ows : a -> a
m as a matrix of its last-axis rows (a vector becomes one row).
ʰr̲ows ← { m → (('× r̲/ -1 d̲rop s̲hape m) c̲at -1 t̲ake s̲hape m) r̲eshape m }
Activations
ˡl̲eaky : Float -> Float -> Float
a l̲eaky x: leaky ReLU, x where positive and a times x elsewhere.
ⁿⁿ⁼u̲se< "NN"
0.1 ⁿⁿl̲eaky -2.0 3.0 -0.2 3.0
ˡl̲eaky ← { a x → (x m̲ax 0.0) + a × x m̲in 0.0 }
Softmax and argmax: along the last axis
ˡs̲oftmax : Num a => a -> Float
s̲oftmax m: each row into probabilities; the row's largest is taken
off first (the same result, no overflow), the row's sum divides.
ⁿⁿ⁼u̲se< "NN"
ⁿⁿs̲oftmax 1.0 2.0 3.0 0.09003057317038046 0.24472847105479764 0.6652409557748218
ˡs̲oftmax ← { a → m ← ʰr̲ows a e ← e̲xp m − ('m̲ax r̲/₂ m) ʰs̲pread m (s̲hape a) r̲eshape e ÷ ('+ r̲/₂ e) ʰs̲pread e }
ˡl̲ogSoftmax : Float -> Float
l̲ogSoftmax m: the logarithm of each row's softmax, computed stably.
ˡl̲ogSoftmax ← { a → m ← ʰr̲ows a z ← m − ('m̲ax r̲/₂ m) ʰs̲pread m (s̲hape a) r̲eshape z − (l̲og '+ r̲/₂ e̲xp z) ʰs̲pread z }
Layers
ˡd̲ense : Num a => a -> a -> a
x d̲ense wb: the layer x W + b for a batch x (one row each) and wb,
the weights W with the bias b as one more row at the bottom.
ⁿⁿ⁼u̲se< "NN"
(2 2 r̲eshape 1.0 2.0 3.0 4.0) ⁿⁿd̲ense 3 2 r̲eshape 1.0 0.0 0.0 2.0 10.0 20.0 11.0 24.0 13.0 28.0
ˡd̲ense ← { x wb → (x '+ '× i̲nner -1 d̲rop wb) + (o̲ffsets t̲ally x) 'r̲ight t̲able f̲irst -1 t̲ake wb }
Labels
ˡa̲rgmax : Num a => a -> Int
a̲rgmax a: for each row (along the last axis), the position (from
1) of its largest item (the first, on a tie); a vector gives one Int.
ⁿⁿ⁼u̲se< "NN"
ⁿⁿa̲rgmax 2 3 r̲eshape 1.0 2.0 3.0 -1.0 0.0 5.0 3 3
ˡa̲rgmax ← { a → m ← ʰr̲ows a n ← f̲irst -1 t̲ake s̲hape m hit ← f̲loat m = ('m̲ax r̲/₂ m) ʰs̲pread m (-1 d̲rop s̲hape a) r̲eshape f̲loor 'm̲in r̲/₂ (hit × (o̲ffsets t̲ally m) 'r̲ight t̲able f̲loat r̲ange n) + (1.0 − hit) × f̲loat n + 1 }
ˡo̲neHot : Int -> Int -> Float
k o̲neHot y: the labels y (1 to k) as rows of k Floats, 1.0 at the label.
ⁿⁿ⁼u̲se< "NN"
3 ⁿⁿo̲neHot 2 1 3 0.0 1.0 0.0 1.0 0.0 0.0 0.0 0.0 1.0
ˡo̲neHot ← { k y → f̲loat y '= t̲able r̲ange k }
Loss and accuracy
ˡc̲rossEntropy : Float -> Float -> Float
y c̲rossEntropy p: the mean over rows of -sum(y log p), for one-hot
(or probability) rows y and predicted probabilities p; p is clamped
at 1e-12 first, so a probability of 0 costs 27.6, not a log error.
ⁿⁿ⁼u̲se< "NN"
(2 ⁿⁿo̲neHot 1 2) ⁿⁿc̲rossEntropy 2 2 r̲eshape 0.9 0.1 0.2 0.8 0.164252033486018
ˡc̲rossEntropy ← { y p → ('+ r̲/ '+ r̲/₂ n̲eg y × l̲og p m̲ax 1e-12) ÷ f̲loat t̲ally p }
ˡm̲se : Float -> Float -> Float
ˡm̲se ← { y p → ('+ r̲/ r̲avel (y − p) ^ 2) ÷ f̲loat t̲ally r̲avel p }