librarylibs/Strings/src/Strings.xtl
Strings: text functions -- case, trimming, words, split and join, search, replace, padding. Import it with an alias of your choice: "t:" u_se< "Strings". Put libs/Strings/src on XETAL_PATH ("just path"); the reference is libs/Strings/docs. Names with l: are exported; those under h: are private to this file.
A string is a Char vector; a list of strings is a nested vector (Box Char), as "ab" "cde" is. Patterns, separators and widths go on the left, the text on the right: "," t:s_plit "a,b".
Case and trimming
ˡu̲pper : Char -> Char
Case by character code ([]U_CS): a letter's two cases are 32 apart. ASCII letters only: []U_CHAR takes codes 0 to 127 (ask X16).
ˡu̲pper ← { t → c ← ⎕U̲CS t◆ ⎕U̲CHAR c − 32 × (c ≥ 97) ∧ c ≤ 122 } ⍝ "Hello" to "HELLO"
ˡl̲ower : Char -> Char
ˡl̲ower ← { t → c ← ⎕U̲CS t◆ ⎕U̲CHAR c + 32 × (c ≥ 65) ∧ c ≤ 90 } ⍝ "Hello" to "hello"
ˡt̲rimStart : Char -> Char
Trimming: blanks (space, tab, newline) off the start, the end, or both.
ˡt̲rimStart ← { t → (n̲ot '∧ s̲\ t m̲ember? blanks) r̲eplicate t }
ˡt̲rim : Char -> Char
ˡt̲rim ← { t → ˡt̲rimEnd ˡt̲rimStart t }
Words and joining
ˡw̲ords : Char -> Box Char
w_ords t: the words of t, split at runs of blanks (dfns words).
ˡw̲ords ← { t → (n̲ot t m̲ember? blanks) p̲artition t }
ˡj̲oin : Char -> Box Char -> Char
sep j_oin list: the strings of list with sep between them.
ˡj̲oin ← { sep b → 0 = t̲ally b ? "" d̲isclose '{ e̲nclose (d̲isclose ⍺) c̲at sep c̲at d̲isclose ⍵ } r̲/ b }
ˡs̲queeze : Char -> Char
s_queeze t: the words of t one space apart (J's deb).
ˡs̲queeze ← { t → " " ˡj̲oin ˡw̲ords t }
Finding and splitting
ˡf̲ind : Eq a => a -> a -> Int
p f_ind t: where the pattern p starts in t, overlaps included (APL's find, dfns ss's search).
ˡf̲ind ← { p t → m ← t̲ally p n ← t̲ally t (m = 0) ∨ m > n ? 0 r̲eshape 0 cells ← ((r̲ange 1 + n − m) '+ t̲able o̲ffsets m) s̲elect t w̲here '∧ r̲/₂ cells = (s̲hape cells) r̲eshape p }
ʰa̲part : Num a => a -> a -> a
Starts ps of matches of length m, keeping only those that do not overlap an earlier one (left to right).
ʰa̲part ← { m ps → 0 = t̲ally ps ? ps f ← f̲irst ps f c̲at m ʰa̲part (ps ≥ f + m) r̲eplicate ps }
ˡo̲ccurrences : Eq a => a -> a -> Int
p o_ccurrences t: how many times p occurs in t, without overlaps.
ˡo̲ccurrences ← { p t → t̲ally (t̲ally p) ʰa̲part p ˡf̲ind t }
ˡs̲plit : Eq a => a -> a -> Box a
sep s_plit t: the pieces of t between the separators (any length); empty pieces are kept, so "a,,b" has three ("a" "" "b").
ˡs̲plit ← { sep t → m ← t̲ally sep m = 0 ? 1 r̲eshape e̲nclose t ps ← m ʰa̲part sep ˡf̲ind t starts ← 1 c̲at ps + m ends ← ps c̲at 1 + t̲ally t '{ i → ((i s̲elect ends) − i s̲elect starts) t̲ake ((i s̲elect starts) − 1) d̲rop t } m̲ap r̲ange t̲ally starts }
ˡl̲ines : Char -> Box Char
l_ines t: the lines of t (a final newline does not make an empty line).
ˡl̲ines ← { t → e ← (0 < t̲ally t) ∧ ("\n" m̲atch -1 t̲ake t) "\n" ˡs̲plit (n̲eg e) d̲rop t }
ˡr̲eplace : Box Char -> Char -> Char
Old and new as a pair: "cat" "dog" t:r_eplace t, without overlaps (dfns ss).
ˡr̲eplace ← { pair t → (d̲isclose 2 s̲elect pair) ˡj̲oin (d̲isclose 1 s̲elect pair) ˡs̲plit t }
Prefixes
ˡp̲refix? : (Match a, Truthy b) => a -> a -> b
p p_refix? t, p s_uffix? t, p i_nfix? t: whether p begins, ends or occurs in t.
ˡp̲refix? ← { p t → (t̲ally p) > t̲ally t ? 0 = 1◆ p m̲atch (t̲ally p) t̲ake t }
ˡs̲uffix? : (Match a, Truthy b) => a -> a -> b
ˡs̲uffix? ← { p t → (t̲ally p) > t̲ally t ? 0 = 1◆ p m̲atch (n̲eg t̲ally p) t̲ake t }
Padding
ˡp̲adLeft : Int -> a -> a
n p_adLeft t, n p_adRight t, n c_enter t: t in a field n wide, filled with spaces (right-aligned, left-aligned, centered); a longer t is kept whole.
ˡp̲adLeft ← { n t → (n̲eg n m̲ax t̲ally t) t̲ake t }
ˡp̲adRight : Int -> a -> a
ˡp̲adRight ← { n t → (n m̲ax t̲ally t) t̲ake t }
ˡc̲enter : Int -> a -> a
ˡc̲enter ← { n t → w ← n m̲ax t̲ally t w t̲ake (n̲eg (t̲ally t) + (w − t̲ally t) d̲iv 2) t̲ake t }
ˡr̲epeat : Int -> a -> a
n r_epeat t: t, n times over.
ˡr̲epeat ← { n t → 0 = t̲ally t ? t◆ (n × t̲ally t) r̲eshape t }
ˡm̲ix : Box a -> a
m_ix list: the texts of a list as a character matrix, one per row, padded with spaces to the longest (APL2's disclose of a list of strings, "mix").
ˡm̲ix ← { b → w ← 'm̲ax r̲/ '{ t̲ally d̲isclose ⍵ } e̲ach b all ← d̲isclose '{ e̲nclose (d̲isclose ⍺) c̲at d̲isclose ⍵ } r̲/ '{ s → w t̲ake d̲isclose s } m̲ap b ((t̲ally b) c̲at w) r̲eshape all }