Skip to content

String Manipulation ​

All the methods are type-bound procedures of string. Unless stated otherwise they return a new string and they do not change the one they are called on.

Case conversion ​

f90
program case_conversion
!< Upper, lower and the like.
use stringifor
implicit none
type(string) :: s

s = 'Hello World'
print '(A)', s%upper()//''
print '(A)', s%lower()//''
print '(A)', s%swapcase()//''
print '(A)', s%capitalize()//''
print '(2(L1,1X))', s%is_upper(), s%is_lower()
endprogram case_conversion
$ letter_case
HELLO WORLD
hello world
hELLO wORLD
Hello world
F F

The word-level styles split the string in words, at the blanks or at sep:

f90
program word_case
!< Word-level case styles.
use stringifor
implicit none
type(string) :: s

s = 'the quick brown fox'
print '(A)', s%startcase()//''
print '(A)', s%camelcase()//''
print '(A)', s%snakecase()//''
s = 'the-quick-brown-fox'
print '(A)', s%camelcase(sep='-')//''
endprogram word_case
$ wordcase
The Quick Brown Fox
TheQuickBrownFox
the_quick_brown_fox
TheQuickBrownFox
MethodResult
upper(), lower(), swapcase()every letter changed
capitalize()first character upper case, the rest lower case
startcase([sep])every word capitalized
camelcase([sep])every word capitalized, separators removed
snakecase([sep])every word lower case, joined by _
is_upper(), is_lower()true if every character is an upper (lower) case letter

Character classes ​

f90
program char_classes
!< The class of the characters of some strings.
use stringifor
implicit none
type(string) :: words(5)
integer      :: w

words(1) = 'Fortran'
words(2) = 'F2023'
words(3) = 'c0ffee'
words(4) = '?!,'
words(5) = ' '
print '(A)', 'word      alpha alnum xdigit punct space'
do w = 1, size(words)
  print '(A,5L6)', words(w)%ljust(10)//'', words(w)%is_alpha(), words(w)%is_alnum(), words(w)%is_xdigit(), &
                   words(w)%is_punct(), words(w)%is_space()
enddo
endprogram char_classes
$ char_classes
word      alpha alnum xdigit punct space
Fortran        T     T     F     F     F
F2023          F     T     T     F     F
c0ffee         F     T     T     F     F
?!,            F     F     F     T     F
               F     F     F     F     T
MethodTrue if all the characters are
is_alpha()letters
is_alnum()letters or digits
is_digit()digits
is_xdigit()hexadecimal digits, 0-9, a-f, A-F
is_punct()punctuation: printable, not letters, digits or space (as C ispunct)
is_space()whitespace: space, tab, new line, vertical tab, form feed, carriage return (as C isspace)
is_lower(), is_upper()not uppercase, not lowercase letters

The classes are the ASCII ones. A null string is in no class, as in Python.

Cleaning and replacing ​

f90
program strip_blanks
!< Remove what surrounds a string.
use stringifor
implicit none
type(string) :: s

s = '   hello world   '
print '(A)', '['//s%strip()//']'              ! both ends
print '(A)', '['//s%trim()//']'               ! the end only
print '(A)', '['//s%adjustl()//']'            ! moved to the left, same length
s = '--=hello world=--'
print '(A)', '['//s%strip(remove='-=')//']'   ! any character of a set
print '(A)', '['//s%lstrip(remove='-=')//']'  ! the beginning only
print '(A)', '['//s%rstrip(remove='-=')//']'  ! the end only
s = achar(9)//' hello world'//new_line('a')
print '(A)', '['//s%strip(whitespace=.true.)//']'  ! tabs and new lines too
endprogram strip_blanks
$ strip
[hello world]
[   hello world]
[hello world      ]
[hello world]
[hello world=--]
[--=hello world]
[hello world]

strip([remove_nulls][, remove][, whitespace]) removes the blanks at both ends; with remove, a set of characters, it removes every leading and trailing character belonging to the set; whitespace=.true. removes tabs, new lines, vertical tabs, form feeds and carriage returns too, as Python str.strip(). lstrip([remove][, whitespace]) and rstrip([remove][, whitespace]) do the same at the beginning or at the end only. remove_nulls=.true. cuts the string at its first null character.

f90
program replace_text
!< Replace, collapse, insert.
use stringifor
implicit none
type(string) :: s

s = 'a-b-c-d'
print '(A)', s%replace(old='-', new=' + ')//''
print '(A)', s%replace(old='-', new='', count=2)//''   ! the first two only
s = 'too    many     blanks'
print '(A)', s%unique(' ')//''
s = 'Helo'
print '(A)', s%insert(substring='l', pos=3)//''
endprogram replace_text
$ replace
a + b + c + d
abc-d
too many blanks
Hello
MethodResult
replace(old, new[, count])every occurrence of old replaced by new, or the first count ones, left to right; the replaced text is not searched again
unique([substring])every run of substring (default a blank) collapsed to one occurrence
insert(substring, pos)substring inserted at position pos
escape(to_escape[, esc])the character to_escape preceded by a backslash, or by esc
unescape(to_unescape[, unesc])the backslash before to_unescape removed, or replaced by unesc
f90
program escape_characters
!< Escape a character, and back.
use stringifor
implicit none
type(string) :: s, escaped

s = 'C:\temp\new'
escaped = s%escape(to_escape='\')
print '(A)', escaped//''
print '(A)', escaped%unescape(to_unescape='\')//''
endprogram escape_characters
$ escape
C:\\temp\\new
C:\temp\new
f90
program tidy
!< Whitespace, repeated characters, tabs, character sets and quotes.
use stringifor
implicit none
type(string) :: s

s = '  too   many'//achar(9)//'blanks  '
print '(A)', '['//s%compact()//']'
print '(A)', '['//s%compact(sep='_')//']'
s = 'Mississippi --- yes!!!'
print '(A)', s%squeeze()//''
print '(A)', s%squeeze(set='-!')//''
s = 'id'//achar(9)//'name'//achar(9)//'value'
print '(A)', s%expand_tabs(8)//''
s = 'hello world'
print '(A)', s%transliterate('lo', 'LO')//''
print '(A)', s%transliterate('aeiou', '')//''
s = 'say "hi"'
s = s%quote()
print '(A)', s//''
print '(A)', s%unquote()//''
endprogram tidy
$ tidy
[too many blanks]
[too_many_blanks]
Misisipi - yes!
Mississippi - yes!
id      name    value
heLLO wOrLd
hll wrld
"say ""hi"""
say "hi"
MethodResult
compact([sep])the words separated by one sep (default a blank): whitespace runs collapsed, the ends removed, as Python sep.join(s.split())
squeeze([set])every run of a repeated character reduced to one, only for the characters of set if passed, as tr -s
expand_tabs([tab_size])the tabs replaced by blanks up to the next tab stop, every tab_size columns (default 8), as Python str.expandtabs
transliterate(old_set, new_set)each character of old_set replaced by the one at the same position in new_set; a shorter new_set repeats its last character, a null one deletes, as tr
quote([quote_char])the string between quotes (default "), the inner ones doubled, as Fortran list-directed output and CSV
unquote()the inverse of quote, for a string starting and ending with the same ' or "; any other string unchanged

Splitting and joining ​

f90
program split_words
!< Split a string into tokens.
use stringifor
implicit none
type(string)              :: s
type(string), allocatable :: tokens(:)
integer                   :: t

s = 'the quick  brown   fox'
call s%split(tokens=tokens)                         ! on blanks, repeated ones count as one
print '(I0,A)', size(tokens), ' words, the last one is "'//tokens(size(tokens))//'"'

s = '2001-07-14'
call s%split(tokens=tokens, sep='-')
print '(*(A,:,"/"))', (tokens(t)//'', t=size(tokens), 1, -1)

s = 'key = value = with = equals'
call s%split(tokens=tokens, sep=' = ', max_tokens=1)  ! split once
print '(A)', '['//tokens(1)//'] ['//tokens(2)//']'
endprogram split_words
$ split_words
4 words, the last one is "fox"
14/07/2001
[key] [value = with = equals]

split(tokens[, sep][, max_tokens][, keep_empty]) is a subroutine: it allocates tokens. Repeated separators count as one and the separators at the ends are ignored; max_tokens is the number of splits, the last token keeps the rest. split_chunked(tokens, chunks[, sep]) gives the same tokens splitting in chunks of chunks tokens.

With keep_empty=.true. nothing is collapsed and the empty fields are tokens, as Python str.split(sep): the fields of a CSV record keep their position.

f90
program csv_fields
!< The fields of a CSV record, the empty ones included.
use stringifor
implicit none
type(string)              :: record
type(string), allocatable :: fields(:)
integer                   :: f

record = 'Rossi,,42,'
call record%split(tokens=fields, sep=',', keep_empty=.true.)
do f = 1, size(fields)
  print '(I0,A)', f, ': ['//fields(f)//']'
enddo
call record%split(tokens=fields, sep=',')
print '(I0,A)', size(fields), ' tokens without keep_empty'
endprogram csv_fields
$ csv_fields
1: [Rossi]
2: []
3: [42]
4: []
2 tokens without keep_empty
f90
program partition_once
!< Split at the first separator, keeping the three parts.
use stringifor
implicit none
type(string) :: s, parts(3)

s = 'name = John = Smith'
parts = s%partition(sep=' = ')
print '(A)', '['//parts(1)//'] ['//parts(2)//'] ['//parts(3)//']'
endprogram partition_once
$ partition
[name] [ = ] [John = Smith]

partition([sep]) splits once, at the first separator, into three strings.

f90
program join_strings
!< Join an array into one string.
use stringifor
implicit none
type(string) :: glue, words(3), joined
character(5) :: chars(3)

words(1) = 'one' ; words(2) = 'two' ; words(3) = 'three'
glue = ', '
joined = glue%join(array=words)                 ! the string is the separator
print '(A)', joined//''
print '(A)', glue%join(array=words, sep='-')//'' ! or pass one

chars = ['alpha', 'beta ', 'gamma']
joined = strjoin(array=chars, sep='+')          ! no separator string needed; characters are trimmed
print '(A)', joined//''
endprogram join_strings
$ join_strings
one, two, three
one-two-three
alpha+beta+gamma

join(array[, sep]) joins an array of strings or of characters; the separator is sep, or the string itself. Elements not allocated, or empty characters, are skipped. strjoin(array[, sep][, is_trim][, is_col]) is the same as a function of the module, with no separator string; is_trim=.false. keeps the trailing blanks of the characters.

strjoin also joins a 2D array, by columns or by rows:

f90
program join_a_table
!< Join the columns, or the rows, of a 2D array.
use stringifor
implicit none
type(string) :: table(3,2), columns(2), rows(3)
integer      :: j

table(1,1) = 'a' ; table(2,1) = 'b' ; table(3,1) = 'c'
table(1,2) = 'd' ; table(2,2) = 'e' ; table(3,2) = 'f'
columns = strjoin(array=table, sep=',')                 ! each column joined
print '(*(A,:,1X))', (columns(j)//'', j=1, size(columns))
rows = strjoin(array=table, sep=',', is_col=.false.)    ! each row joined
print '(*(A,:,1X))', (rows(j)//'', j=1, size(rows))
endprogram join_a_table
$ strjoin2d
a,b,c d,e,f
a,d b,e c,f

Slicing and reversing ​

f90
program slice_a_string
!< Substrings, with a stride too.
use stringifor
implicit none
type(string) :: s

s = 'Hello World'
print '(A)', s%slice(first=1, last=5)
print '(A)', s%slice(first=7)                   ! to the end
print '(A)', s%slice(last=5)                    ! from the beginning
print '(A)', s%slice(stride=2)                  ! every other character
print '(A)', s%slice(stride=-1)                 ! backwards
print '(A)', '['//s%slice(first=20, last=30)//']' ! out of the string: clamped, never out of bounds
endprogram slice_a_string
$ slice
Hello
World
Hello
HloWrd
dlroW olleH
[]

slice([first][, last][, stride]) returns the section first:last:stride as a character. stride defaults to 1, first and last to the bounds of the string (1 and len, or len and 1 when the stride is negative). The bounds are clamped into the string: a slice never goes out of bounds, and it is empty when the section is.

f90
program reverse_things
!< Reverse the characters or the words.
use stringifor
implicit none
type(string) :: s

s = 'the sky is blue'
print '(A)', s%reverse()//''
print '(A)', s%reverse_words()//''
s = 'usr/local/lib'
print '(A)', s%reverse_words(sep='/')//''
endprogram reverse_things
$ reverse
eulb si yks eht
blue is sky the
lib/local/usr

Searching ​

f90
program search_text
!< Look for something in a string.
use stringifor
implicit none
type(string) :: s

s = 'report_2001_final.tar.gz'
print '(2(L1,1X))', s%start_with('report'), s%end_with('.gz')
print '(I0)', s%count('_')
print '(I0)', s%index('2001')
print '(I0)', s%index('.', back=.true.)
print '(I0)', s%scan('0123456789')              ! the first digit
print '(I0)', s%verify('abcdefghijklmnopqrstuvwxyz')  ! the first character that is not a letter
s = 'aaaa'
print '(I0,1X,I0)', s%count('aa'), s%count('aa', overlapping=.true.)
s = 'a.b.c.d'
print '(I0,1X,I0)', s%index('.', occurrence=2), s%index('.', back=.true., occurrence=2)
endprogram search_text
$ search
T T
2
8
22
8
7
2 3
4 4
MethodResult
start_with(prefix[, start][, end]), end_with(suffix[, start][, end][, ignore_null_eof])true if the string, or its part start:end, starts (ends) with it
count(substring[, ignore_isolated][, overlapping])number of occurrences, not overlapping (as Python str.count) unless overlapping=.true.
index(substring[, back][, occurrence])like the intrinsic; occurrence=k gives the k-th occurrence, not overlapping, from the end if back
scan(set[, back]), verify(set[, back])like the intrinsics

search returns the first text enclosed by two tags, tags included:

f90
program tagged_text
!< The text between two tags.
use stringifor
implicit none
type(string) :: s, found
integer      :: first, last

s = '<b>bold</b> plain <b>bold again</b>'
found = s%search(tag_start='<b>', tag_end='</b>', istart=first, iend=last)
print '(A,2(1X,I0))', found//'', first, last
endprogram tagged_text
$ tags
<b>bold</b> 1 11

search(tag_start, tag_end[, in_string][, in_character][, istart][, iend]) searches the string itself, or the one passed as in_string or in_character; istart and iend return where the text starts and ends.

Padding and layout ​

f90
program pad_a_string
!< Pad to a width.
use stringifor
implicit none
type(string) :: s

s = '42'
print '(A)', s%fill(width=6)//''                             ! zeros on the left
print '(A)', s%fill(width=6, right=.true.)//''
print '(A)', s%fill(width=6, filling_char='.')//''
s = 'name'
print '(A)', '['//s%fill(width=10, right=.true., filling_char=' ')//']'
endprogram pad_a_string
$ pad
000042
420000
....42
[name      ]

fill(width[, right][, filling_char]) pads to width characters, on the left unless right=.true., with zeros unless filling_char is passed. A string already as wide as width, or wider, is returned unchanged.

f90
program align
!< Strings centered or justified in a given width.
use stringifor
implicit none
type(string) :: s

s = 'title'
print '(A)', '['//s%center(11)//']'
print '(A)', '['//s%center(11, '=')//']'
print '(A)', '['//s%ljust(11, '.')//']'
print '(A)', '['//s%rjust(11)//']'
print '(A)', '['//s%center(3)//']'
endprogram align
$ align
[   title   ]
[===title===]
[title......]
[      title]
[title]

center(width[, fill_char]), ljust(width[, fill_char]) and rjust(width[, fill_char]) center, left justify and right justify the string in width characters, padding with blanks or with fill_char, as the Python methods of the same name: a string already as wide as width is returned unchanged, and an odd padding of center puts the extra character on the left if width is odd, on the right otherwise.

f90
program justify_text
!< Wrap a paragraph in justified lines.
use stringifor
implicit none
type(string)              :: text
type(string), allocatable :: lines(:)
integer                   :: l

text = 'Fortran is a general-purpose, compiled imperative programming language '// &
       'that is especially suited to numeric computation and scientific computing.'
call text%justify(lines=lines, width=36)
do l = 1, size(lines)
  print '(A)', '|'//lines(l)//'|'
enddo
print '(A,I0)', 'the last word has length ', text%len_last_word()
endprogram justify_text
$ justify
|Fortran    is   a   general-purpose,|
|compiled    imperative   programming|
|language  that  is especially suited|
|to     numeric    computation    and|
|scientific computing.               |
the last word has length 10

justify(lines, width) is a subroutine: it allocates lines. The words are packed greedily and the blanks spread between them, the leftmost gaps taking the extra ones; the last line and the lines of one word are left-justified and padded. A word longer than width is not broken. len_last_word([sep]) is the length of the last word, trailing separators ignored.

Comparing ​

f90
program common_prefix
!< What several strings start with.
use stringifor
implicit none
type(string) :: files(3), prefix

files(1) = 'src/lib/stringifor.F90'
files(2) = 'src/lib/stringifor_string_t.F90'
files(3) = 'src/tests/stringifor/test.f90'
prefix = files(1)%common_prefix(files(2))         ! of two strings
print '(A)', prefix//''
prefix = files(1)%common_prefix(array=files)      ! of an array
print '(A)', prefix//''
endprogram common_prefix
$ prefix
src/lib/stringifor
src/

common_prefix(other) takes a string or a character; common_prefix(array) an array of strings, and returns what the string and all its elements start with.

f90
program compare_versions
!< Which version is newer?
use stringifor
implicit none
type(string) :: version

version = '1.10.0'
print '(I0)', version%compare_version('1.9.3')    !  1: newer
print '(I0)', version%compare_version('1.10')     !  0: the same, missing fields are zero
print '(I0)', version%compare_version('2.0')      ! -1: older
version = '2024-03'
print '(I0)', version%compare_version('2024-11', sep='-')
endprogram compare_versions
$ version
1
0
-1
-1

compare_version(other[, sep]) returns -1, 0 or 1. The versions are compared field by field (the fields separated by ., or by sep): two fields made of digits only are compared as integers of any size, the others as text; missing fields count as zero.

WARNING

The comparison is not semantic-versioning aware: 1.0.0-rc1 compares greater than 1.0.0.