Utf8

Parsers and helpers specialized to UTF-8 byte lists and Roc Str values.

Use these with the combinators in Parser whenever the input is text. Parsers work on bytes (Bytes), so codeunit and digit match single bytes, and parse_str converts a Str to bytes and back for you.

parse_str : Parser(Bytes, a), Str -> Try(a, [ParseError({ message : Str, offset : U64 })])

Parse a whole Str using a Parser.

Fails with ParseError({ message, offset }), where offset is the byte offset of the furthest failure. Input left over after the parser succeeds is a failure too (see Parser.parse).

color : Parser(Utf8.Bytes, [Red, Green, Blue])
color =
    Parser.one_of([
        Parser.const(Red).skip(Utf8.string("red")),
        Parser.const(Green).skip(Utf8.string("green")),
        Parser.const(Blue).skip(Utf8.string("blue")),
    ])

expect Utf8.parse_str(color, "green") == Ok(Green)
parse_str_partial : Parser(Bytes, a), Str -> Try({ value : a, rest : Str }, [ParseError({ message : Str, offset : U64 })])

Runs a parser against the start of a string, allowing the parser to consume it only partially.

  • If the parser succeeds, returns the resulting value and the rest of the string. A rest that starts inside a multi-byte character is rendered with U+FFFD replacement characters rather than crashing.
  • If the parser fails, returns Err(ParseError({ message, offset })).
at_sign : Parser(Utf8.Bytes, [AtSign])
at_sign = Parser.const(AtSign).skip(Utf8.codeunit('@'))

expect Utf8.parse_str_partial(at_sign, "@").map_ok(|r| r.value) == Ok(AtSign)
expect Utf8.parse_str_partial(at_sign, "$").is_err()
parse_bytes : Parser(Bytes, a), Bytes -> Try(a, [ParseError({ message : Str, offset : U64 })])

Runs a parser against UTF-8 bytes, requiring the parser to consume them fully.

Fails with ParseError({ message, offset }) like Utf8.parse_str.

parse_bytes_partial : Parser(Bytes, a), Bytes -> Try({ value : a, rest : Bytes }, [ParseError({ message : Str, offset : U64 })])

Runs a parser against the start of UTF-8 bytes, allowing the parser to consume them only partially.

Returns the parsed value and the rest of the bytes, or Err(ParseError({ message, offset })).

codeunit_satisfies : (U8 -> Bool) -> Parser(Bytes, U8)

Match one UTF-8 code unit when it satisfies the given predicate.

Fails on empty input or when the predicate returns False.

is_digit : U8 -> Bool
is_digit = |b| b >= '0' and b <= '9'

expect Utf8.parse_str(Utf8.codeunit_satisfies(is_digit), "0") == Ok('0')
expect Utf8.parse_str(Utf8.codeunit_satisfies(is_digit), "*").is_err()
codeunit : U8 -> Parser(Bytes, U8)

Match one exact UTF-8 code unit.

For a character outside ASCII, which is several code units, use string.

at_sign : Parser(Utf8.Bytes, [AtSign])
at_sign = Parser.const(AtSign).skip(Utf8.codeunit('@'))

expect Utf8.parse_str(at_sign, "@") == Ok(AtSign)
expect Utf8.parse_str_partial(at_sign, "$").is_err()
utf8 : List(U8) -> Parser(Bytes, List(U8))

Match an exact sequence of UTF-8 bytes and return them.

string : Str -> Parser(Bytes, Str)

Match the given Str exactly (case-sensitive) and return it.

expect Utf8.parse_str(Utf8.string("Foo"), "Foo") == Ok("Foo")
expect Utf8.parse_str(Utf8.string("Foo"), "Bar").is_err()
any_codeunit : Parser(Bytes, U8)

Match any single U8 code unit; fails only on empty input.

expect Utf8.parse_str(Utf8.any_codeunit, "a") == Ok('a')
expect Utf8.parse_str(Utf8.any_codeunit, "$") == Ok('$')
rest : Parser(Bytes, Bytes)

Consume the rest of the input and return it as bytes; never fails.

expect {
    bytes = "consumes all the input".to_utf8()
    Utf8.rest.parse(bytes) == Ok(bytes)
}
rest_str : Parser(Bytes, Str)

Consume the rest of the input as a Str, failing if the bytes are not valid UTF-8.

digit : Parser(Bytes, U64)

Parse one ASCII decimal digit into a U64 from 0 through 9.

expect Utf8.parse_str(Utf8.digit, "0") == Ok(0)
expect Utf8.parse_str(Utf8.digit, "not a digit").is_err()
digits : Parser(Bytes, U64)

Parse one or more ASCII decimal digits into a U64, accepting leading zeroes.

Fails when the input does not start with a digit or the value does not fit in a U64. Signs and decimal points are not accepted.

expect Utf8.parse_str(Utf8.digits, "0123") == Ok(123)
expect Utf8.parse_str(Utf8.digits, "not a digit").is_err()
find_any : Bytes, U64, ByteClass -> U64

The index of the first byte at or after pos that is in class, or the length of bytes if there is none. Scans 16 bytes at a time.

markup = Utf8.ByteClass.from_bytes(['<', '&'])
expect Utf8.find_any("a&b<c".to_utf8(), 0, markup) == 1
expect Utf8.find_any("abc".to_utf8(), 0, markup) == 3
skip_class : Bytes, U64, ByteClass -> U64

The index of the first byte at or after pos that is not in class, or the length of bytes if every remaining byte is. Use it to skip a run of name characters, digits or whitespace.

digits = Utf8.ByteClass.from_predicate(|b| b >= '0' and b <= '9')
expect Utf8.skip_class("123abc".to_utf8(), 0, digits) == 3
find_line_end : Bytes, U64 -> U64

The index of the next \n or \r at or after pos, or the length of bytes if there is none. Scans 16 bytes at a time.

expect Utf8.find_line_end("ab\r\ncd".to_utf8(), 0) == 2
span_class : ByteClass -> Parser(Bytes, Bytes)

Consume the longest run of bytes in class, possibly empty, and return it as a slice of the input. It scans 16 bytes at a time, so prefer it to Parser.chomp_while for runs longer than a few bytes.

name_char = Utf8.ByteClass.from_predicate(|b| (b >= 'a' and b <= 'z') or b == '-')
expect Utf8.parse_str_partial(Utf8.span_class(name_char), "foo-bar=1").map_ok(|r| r.rest) == Ok("=1")
Bytes : List(U8)

UTF-8 input represented as a list of bytes.

ByteClass

Utf8.ByteClass :: # (opaque)

A set of byte values, for finding or skipping runs of bytes 16 at a time.

Build a class once, outside any loop, and reuse it: construction derives the lookup tables the vector scan uses. Any set of bytes works. A set whose high-nibble rows have at most eight distinct shapes (every ASCII class, and most practical ones) is scanned with two table lookups per 16 bytes; other sets fall back to a byte loop.

spaces = Utf8.ByteClass.from_bytes([' ', '\t'])
expect spaces.contains('\t')
expect !spaces.contains('x')
from_predicate : (U8 -> Bool) -> ByteClass

The class holding every byte b for which check(b) is true.

complement : ByteClass -> ByteClass

The class holding every byte that is not in this one.