librarylibs/NN/src/NN.xtl

NN: neural-network building blocks -- activations, softmax by row, dense layers, one-hot, loss, accuracy

Import it with an alias of your choice: "nn:" u_se< "NN". Put libs/NN/src on XETAL_PATH ("just path"); the reference is libs/NN/docs. Names with l: are exported; the others are private to this file.

A batch is a matrix, one example per row; a layer's weights are a matrix with one row per input and one column per output, and its bias is one more row at the bottom (so a layer is one array, and x ⁿⁿd̲ense wb takes two arguments). Activations work item by item on any shape. Softmax, log-softmax and argmax work along the last axis of any array: a vector is one example, a matrix one per row, a rank-3 array a batch of sequences.

source

Helpers (private)

ʰs̲pread : a -> b -> a

function (private) · line 18

The row values r (one per row of m) spread across m's columns.

ʰs̲pread ← { r m → o̲\ ((-1 t̲ake s̲hape m) c̲at t̲ally r) r̲eshape r }

ʰr̲ows : a -> a

function (private) · line 21

m as a matrix of its last-axis rows (a vector becomes one row).

ʰr̲ows ← { m → (('× r̲/ -1 d̲rop s̲hape m) c̲at -1 t̲ake s̲hape m) r̲eshape m }

Activations

ˡr̲elu : Float -> Float

function · line 29

r̲elu x: max(x, 0), item by item.

      ⁿⁿ⁼u̲se< "NN"
      ⁿⁿr̲elu -1.5 0.0 2.0
0.0 0.0 2.0
ˡr̲elu ← { x → x m̲ax 0.0 }

ˡl̲eaky : Float -> Float -> Float

function · line 35

a l̲eaky x: leaky ReLU, x where positive and a times x elsewhere.

      ⁿⁿ⁼u̲se< "NN"
      0.1 ⁿⁿl̲eaky -2.0 3.0
-0.2 3.0
ˡl̲eaky ← { a x → (x m̲ax 0.0) + a × x m̲in 0.0 }

ˡs̲igmoid : Num a => a -> Float

function · line 41

s̲igmoid x: 1 / (1 + e^-x), item by item.

      ⁿⁿ⁼u̲se< "NN"
      ⁿⁿs̲igmoid 0.0 2.0
0.5 0.8807970779778823
ˡs̲igmoid ← { x → 1.0 ÷ 1.0 + e̲xp n̲eg x }
Used in: ˡt̲anh

ˡt̲anh : Float -> Float

function · line 47

t̲anh x: the hyperbolic tangent, as 2 s_igmoid(2x) - 1.

      ⁿⁿ⁼u̲se< "NN"
      ⁿⁿt̲anh 0.0
0.0
ˡt̲anh ← { x → (2.0 × ˡs̲igmoid 2.0 × x) − 1.0 }

Softmax and argmax: along the last axis

ˡs̲oftmax : Num a => a -> Float

function · line 56

s̲oftmax m: each row into probabilities; the row's largest is taken off first (the same result, no overflow), the row's sum divides.

      ⁿⁿ⁼u̲se< "NN"
      ⁿⁿs̲oftmax 1.0 2.0 3.0
0.09003057317038046 0.24472847105479764 0.6652409557748218
ˡs̲oftmax ← { a →
  m ← ʰr̲ows a
  e ← e̲xp m − ('m̲ax r̲/₂ m) ʰs̲pread m
  (s̲hape a) r̲eshape e ÷ ('+ r̲/₂ e) ʰs̲pread e
}

ˡl̲ogSoftmax : Float -> Float

function · line 63

l̲ogSoftmax m: the logarithm of each row's softmax, computed stably.

ˡl̲ogSoftmax ← { a →
  m ← ʰr̲ows a
  z ← m − ('m̲ax r̲/₂ m) ʰs̲pread m
  (s̲hape a) r̲eshape z − (l̲og '+ r̲/₂ e̲xp z) ʰs̲pread z
}

Layers

ˡd̲ense : Num a => a -> a -> a

function · line 77

x d̲ense wb: the layer x W + b for a batch x (one row each) and wb, the weights W with the bias b as one more row at the bottom.

      ⁿⁿ⁼u̲se< "NN"
      (2 2 r̲eshape 1.0 2.0 3.0 4.0) ⁿⁿd̲ense 3 2 r̲eshape 1.0 0.0 0.0 2.0 10.0 20.0
11.0 24.0
13.0 28.0
ˡd̲ense ← { x wb →
  (x '+ '× i̲nner -1 d̲rop wb) + (o̲ffsets t̲ally x) 'r̲ight t̲able f̲irst -1 t̲ake wb
}

Labels

ˡa̲rgmax : Num a => a -> Int

function · line 88

a̲rgmax a: for each row (along the last axis), the position (from 1) of its largest item (the first, on a tie); a vector gives one Int.

      ⁿⁿ⁼u̲se< "NN"
      ⁿⁿa̲rgmax 2 3 r̲eshape 1.0 2.0 3.0 -1.0 0.0 5.0
3 3
ˡa̲rgmax ← { a →
  m ← ʰr̲ows a
  n ← f̲irst -1 t̲ake s̲hape m
  hit ← f̲loat m = ('m̲ax r̲/₂ m) ʰs̲pread m
  (-1 d̲rop s̲hape a) r̲eshape f̲loor 'm̲in r̲/₂ (hit × (o̲ffsets t̲ally m) 'r̲ight t̲able f̲loat r̲ange n) + (1.0 − hit) × f̲loat n + 1
}

ˡo̲neHot : Int -> Int -> Float

function · line 101

k o̲neHot y: the labels y (1 to k) as rows of k Floats, 1.0 at the label.

      ⁿⁿ⁼u̲se< "NN"
      3 ⁿⁿo̲neHot 2 1 3
0.0 1.0 0.0
1.0 0.0 0.0
0.0 0.0 1.0
ˡo̲neHot ← { k y → f̲loat y '= t̲able r̲ange k }

Loss and accuracy

ˡc̲rossEntropy : Float -> Float -> Float

function · line 111

y c̲rossEntropy p: the mean over rows of -sum(y log p), for one-hot (or probability) rows y and predicted probabilities p; p is clamped at 1e-12 first, so a probability of 0 costs 27.6, not a log error.

      ⁿⁿ⁼u̲se< "NN"
      (2 ⁿⁿo̲neHot 1 2) ⁿⁿc̲rossEntropy 2 2 r̲eshape 0.9 0.1 0.2 0.8
0.164252033486018
ˡc̲rossEntropy ← { y p → ('+ r̲/ '+ r̲/₂ n̲eg y × l̲og p m̲ax 1e-12) ÷ f̲loat t̲ally p }

ˡm̲se : Float -> Float -> Float

function · line 117

y m̲se p: the mean squared error over every item.

      ⁿⁿ⁼u̲se< "NN"
      1.0 2.0 ⁿⁿm̲se 1.0 4.0
2.0
ˡm̲se ← { y p → ('+ r̲/ r̲avel (y − p) ^ 2) ÷ f̲loat t̲ally r̲avel p }

ˡa̲ccuracy : Num a => Int -> a -> Float

function · line 123

y a̲ccuracy p: the fraction of rows of p whose argmax is the label y.

      ⁿⁿ⁼u̲se< "NN"
      3 2 ⁿⁿa̲ccuracy 2 3 r̲eshape 1.0 2.0 3.0 -1.0 0.0 5.0
0.5
ˡa̲ccuracy ← { y p → ('+ r̲/ f̲loat y = ˡa̲rgmax p) ÷ f̲loat t̲ally y }