librarylibs/Strings/src/Strings.xtl

Strings: text functions -- case, trimming, words, split and join, search, replace, padding. Import it with an alias of your choice: "t:" u_se< "Strings". Put libs/Strings/src on XETAL_PATH ("just path"); the reference is libs/Strings/docs. Names with l: are exported; those under h: are private to this file.

A string is a Char vector; a list of strings is a nested vector (Box Char), as "ab" "cde" is. Patterns, separators and widths go on the left, the text on the right: "," t:s_plit "a,b".

source

blanks : Char

value (private) · line 11
blanks ← " \t\n"                            ⍝ what trimming removes

Case and trimming

ˡu̲pper : Char -> Char

function · line 17

Case by character code ([]U_CS): a letter's two cases are 32 apart. ASCII letters only: []U_CHAR takes codes 0 to 127 (ask X16).

ˡu̲pper ← { t → c ← ⎕U̲CS t◆ ⎕U̲CHAR c − 32 × (c ≥ 97) ∧ c ≤ 122 }              ⍝ "Hello" to "HELLO"

ˡl̲ower : Char -> Char

function · line 18
ˡl̲ower ← { t → c ← ⎕U̲CS t◆ ⎕U̲CHAR c + 32 × (c ≥ 65) ∧ c ≤ 90 }               ⍝ "Hello" to "hello"
Used in: clean

ˡt̲rimStart : Char -> Char

function · line 21

Trimming: blanks (space, tab, newline) off the start, the end, or both.

ˡt̲rimStart ← { t → (n̲ot '∧ s̲\ t m̲ember? blanks) r̲eplicate t }

ˡt̲rimEnd : Char -> Char

function · line 22
ˡt̲rimEnd ← { t → r̲ev ˡt̲rimStart r̲ev t }
Used in: ˡt̲rim

ˡt̲rim : Char -> Char

function · line 23
ˡt̲rim ← { t → ˡt̲rimEnd ˡt̲rimStart t }

Words and joining

ˡw̲ords : Char -> Box Char

function · line 28

w_ords t: the words of t, split at runs of blanks (dfns words).

ˡw̲ords ← { t → (n̲ot t m̲ember? blanks) p̲artition t }

ˡj̲oin : Char -> Box Char -> Char

function · line 31

sep j_oin list: the strings of list with sep between them.

ˡj̲oin ← { sep b →
  0 = t̲ally b ? ""
  d̲isclose '{ e̲nclose (d̲isclose ⍺) c̲at sep c̲at d̲isclose ⍵ } r̲/ b
}

ˡs̲queeze : Char -> Char

function · line 37

s_queeze t: the words of t one space apart (J's deb).

ˡs̲queeze ← { t → " " ˡj̲oin ˡw̲ords t }

Finding and splitting

ˡf̲ind : Eq a => a -> a -> Int

function · line 43

p f_ind t: where the pattern p starts in t, overlaps included (APL's find, dfns ss's search).

ˡf̲ind ← { p t →
  m ← t̲ally p
  n ← t̲ally t
  (m = 0) ∨ m > n ? 0 r̲eshape 0
  cells ← ((r̲ange 1 + n − m) '+ t̲able o̲ffsets m) s̲elect t
  w̲here '∧ r̲/₂ cells = (s̲hape cells) r̲eshape p
}

ʰa̲part : Num a => a -> a -> a

function (private) · line 53

Starts ps of matches of length m, keeping only those that do not overlap an earlier one (left to right).

ʰa̲part ← { m ps →
  0 = t̲ally ps ? ps
  f ← f̲irst ps
  f c̲at m ʰa̲part (ps ≥ f + m) r̲eplicate ps
}

ˡo̲ccurrences : Eq a => a -> a -> Int

function · line 60

p o_ccurrences t: how many times p occurs in t, without overlaps.

ˡo̲ccurrences ← { p t → t̲ally (t̲ally p) ʰa̲part p ˡf̲ind t }

ˡs̲plit : Eq a => a -> a -> Box a

function · line 64

sep s_plit t: the pieces of t between the separators (any length); empty pieces are kept, so "a,,b" has three ("a" "" "b").

ˡs̲plit ← { sep t →
  m ← t̲ally sep
  m = 0 ? 1 r̲eshape e̲nclose t
  ps ← m ʰa̲part sep ˡf̲ind t
  starts ← 1 c̲at ps + m
  ends ← ps c̲at 1 + t̲ally t
  '{ i → ((i s̲elect ends) − i s̲elect starts) t̲ake ((i s̲elect starts) − 1) d̲rop t } m̲ap r̲ange t̲ally starts
}

ˡl̲ines : Char -> Box Char

function · line 74

l_ines t: the lines of t (a final newline does not make an empty line).

ˡl̲ines ← { t →
  e ← (0 < t̲ally t) ∧ ("\n" m̲atch -1 t̲ake t)
  "\n" ˡs̲plit (n̲eg e) d̲rop t
}
Used in: ˡr̲ows

ˡr̲eplace : Box Char -> Char -> Char

function · line 81

Old and new as a pair: "cat" "dog" t:r_eplace t, without overlaps (dfns ss).

ˡr̲eplace ← { pair t →
  (d̲isclose 2 s̲elect pair) ˡj̲oin (d̲isclose 1 s̲elect pair) ˡs̲plit t
}

Prefixes

ˡp̲refix? : (Match a, Truthy b) => a -> a -> b

function · line 89

p p_refix? t, p s_uffix? t, p i_nfix? t: whether p begins, ends or occurs in t.

ˡp̲refix? ← { p t → (t̲ally p) > t̲ally t ? 0 = 1◆ p m̲atch (t̲ally p) t̲ake t }

ˡs̲uffix? : (Match a, Truthy b) => a -> a -> b

function · line 90
ˡs̲uffix? ← { p t → (t̲ally p) > t̲ally t ? 0 = 1◆ p m̲atch (n̲eg t̲ally p) t̲ake t }

ˡi̲nfix? : (Eq a, Truthy b) => a -> a -> b

function · line 91
ˡi̲nfix? ← { p t → 0 < t̲ally p ˡf̲ind t }

Padding

ˡp̲adLeft : Int -> a -> a

function · line 98

n p_adLeft t, n p_adRight t, n c_enter t: t in a field n wide, filled with spaces (right-aligned, left-aligned, centered); a longer t is kept whole.

ˡp̲adLeft ← { n t → (n̲eg n m̲ax t̲ally t) t̲ake t }

ˡp̲adRight : Int -> a -> a

function · line 99
ˡp̲adRight ← { n t → (n m̲ax t̲ally t) t̲ake t }

ˡc̲enter : Int -> a -> a

function · line 100
ˡc̲enter ← { n t →
  w ← n m̲ax t̲ally t
  w t̲ake (n̲eg (t̲ally t) + (w − t̲ally t) d̲iv 2) t̲ake t
}

ˡr̲epeat : Int -> a -> a

function · line 106

n r_epeat t: t, n times over.

ˡr̲epeat ← { n t → 0 = t̲ally t ? t◆ (n × t̲ally t) r̲eshape t }

ˡm̲ix : Box a -> a

function · line 111

m_ix list: the texts of a list as a character matrix, one per row, padded with spaces to the longest (APL2's disclose of a list of strings, "mix").

ˡm̲ix ← { b →
  w ← 'm̲ax r̲/ '{ t̲ally d̲isclose ⍵ } e̲ach b
  all ← d̲isclose '{ e̲nclose (d̲isclose ⍺) c̲at d̲isclose ⍵ } r̲/ '{ s → w t̲ake d̲isclose s } m̲ap b
  ((t̲ally b) c̲at w) r̲eshape all
}