Parser

Parser(input, a) :: # (opaque)

Generic parser combinators for transforming input into structured values.

A Parser(input, a) is a value that describes how to read an a from the front of an input. Combine small parsers with keep, skip, map, one_of, many and sep_by to build larger ones, then run the result with Utf8.parse_str (for Str) or Parser.parse (for any input type).

This parser turns "Game 1: 3 blue, 4 red; 1 red, 2 green, 6 blue; 2 green" into { id: 1, requirements: [[Blue(3), Red(4)], [Red(1), Green(2), Blue(6)], [Green(2)]] } (the same code is a test in Utf8.roc):

Requirement : [Green(U64), Red(U64), Blue(U64)]
RequirementSet : List(Requirement)
Game : { id : U64, requirements : List(RequirementSet) }

parse_game : Str -> Try(Game, [ParsingError])
parse_game = |s| {
    green = Parser.const(|x| Green(x)).keep(Utf8.digits).skip(Utf8.string(" green"))
    red = Parser.const(|x| Red(x)).keep(Utf8.digits).skip(Utf8.string(" red"))
    blue = Parser.const(|x| Blue(x)).keep(Utf8.digits).skip(Utf8.string(" blue"))

    requirement_set : Parser(_, RequirementSet)
    requirement_set = Parser.one_of([green, red, blue]).sep_by(Utf8.string(", "))

    requirements : Parser(_, List(RequirementSet))
    requirements = requirement_set.sep_by(Utf8.string("; "))

    game : Parser(_, Game)
    game =
        Parser.const(|id| |r| { id, requirements: r })
            .skip(Utf8.string("Game "))
            .keep(Utf8.digits)
            .skip(Utf8.string(": "))
            .keep(requirements)

    match Utf8.parse_str(game, s) {
        Ok(g) => Ok(g)
        Err(ParseError(_)) => Err(ParsingError)
    }
}

Alternatives backtrack: when one alternative of alt or one_of fails, the next one is tried on the original input, however much the failed one read.

Failures report the furthest position any parser reached. When a repetition such as many stops at a bad element, or an alternative fails after reading further than the one that succeeded, that failure is kept and reported if the whole parse later fails at or before it.

The representation is internal and may change to improve efficiency or error messages.

custom : (input -> ParseResult(input, a)) -> Parser(input, a) where [input.len : input -> U64]

Write a custom parser without using the provided combinators.

The function receives the remaining input and returns either the parsed value with the rest it did not consume, or Err(ParseError({ message, offset })), where offset is how far into the given input the failure is (usually 0).

run : Parser(input, a), input -> ParseResult(input, a) where [input.len : input -> U64]

Run a parser on the start of an input, returning the parsed value and the input it did not consume.

Most parsers consume part of input when they succeed. This allows you to string parsers together that run one after the other. The part of the input that the first parser did not consume, is used by the next parser.

On failure, the error is the furthest failure seen, with its offset counted from the start of input. This is mostly useful when creating your own parsing building blocks.

parse : Parser(input, a), input -> Try(a, [ParseError({ message : Str, offset : U64 })]) where [input.len : input -> U64]

Run a parser on the given input, expecting it to consume all of it.

The input type only needs a len method, which is used both to detect leftover input and to turn the furthest failure into an offset counted from the start of input. Leftover input is a failure too: if some parser failed at or beyond the leftover, its message is reported; otherwise the message is unexpected input. For UTF-8 text use Utf8.parse_str or Utf8.parse_bytes.

fail : Str -> Parser(input, a) where [input.len : input -> U64]

Parser that can never succeed, regardless of the given input. It will always fail with the given error message.

This is mostly useful as a 'base case' if all other parsers in a one_of or alt have failed, to provide some more descriptive error message.

const : a -> Parser(_, a)

Parser that will always produce the given a, without looking at the actual input.

This is the usual start of a pipeline: const supplies a (curried) constructor function and each keep feeds it one parsed value.

parse_u32 : Parser(Utf8.Bytes, U32)
parse_u32 = Parser.const(U64.to_u32_wrap).keep(Utf8.digits)

expect Utf8.parse_str(parse_u32, "123") == Ok(123.U32)
alt : Parser(input, a), Parser(input, a) -> Parser(input, a)

Try the first parser and (only) if it fails, try the second parser as fallback.

The second parser starts from the same input as first (backtracking). If both fail, the failure that reached further is reported; when both stopped at the same place, the second one is.

one_of : List(Parser(input, a)) -> Parser(input, a) where [input.len : input -> U64]

Try a list of parsers in turn, until one of them succeeds.

Each parser starts from the same input. An empty list always fails. For UTF-8 input, Parser.one_of behaves the same way.

color : Parser(Utf8.Bytes, [Red, Green, Blue])
color =
    Parser.one_of([
        Parser.const(Red).skip(Utf8.string("red")),
        Parser.const(Green).skip(Utf8.string("green")),
        Parser.const(Blue).skip(Utf8.string("blue")),
    ])

expect Utf8.parse_str(color, "green") == Ok(Green)
map : Parser(input, a), (a -> b) -> Parser(input, b)

Transforms the result of parsing into something else, using the given transformation function.

map2 : Parser(input, a), Parser(input, b), (a, b -> c) -> Parser(input, c)

Transforms the result of parsing into something else, using the given two-parameter transformation function.

map3 : Parser(input, a), Parser(input, b), Parser(input, c), (a, b, c -> d) -> Parser(input, d)

Transforms the result of parsing into something else, using the given three-parameter transformation function.

If you need transformations with more inputs, take a look at keep.

flatten : Parser(input, Try(a, Str)) -> Parser(input, a) where [input.len : input -> U64]

Removes a layer of Try from running the parser.

Use this to map functions that return a Try over the parser: an Err(msg) value becomes a failure with that message, reported where the inner parser started.

even : Parser(Utf8.Bytes, U64)
even =
    Utf8.digits
        .map(|n| if n % 2 == 0 { Ok(n) } else { Err("odd number") })
        .flatten()

expect Utf8.parse_str(even, "42") == Ok(42)
expect Utf8.parse_str(even, "7") == Err(ParseError({ message: "odd number", offset: 0 }))
and_then : Parser(input, a), (a -> Parser(input, b)) -> Parser(input, b)

Run first, then build the next parser from its value and run that.

Use this when what comes next depends on what was read, such as a length prefix. When the next parser is fixed, prefer keep, skip and map, which are simpler and let the parser be built once.

sized : Parser(Utf8.Bytes, List(U8))
sized = Utf8.digit.and_then(|n| Utf8.any_codeunit.many().map(|bytes| bytes.take_first(n)))
lazy : ({} -> Parser(input, a)) -> Parser(input, a)

Runs a parser lazily.

This is (only) useful when dealing with a recursive structure. For instance, consider a type Comment : { message : Str, responses : List(Comment) }. Without lazy, you would ask the compiler to build an infinitely deep parser. (Resulting in a compiler error.)

Mutually recursive top-level parser values currently require a lower-level workaround because of roc-lang/roc#10098.

maybe : Parser(input, a) -> Parser(input, Try(a, [Missing]))

Make a parser optional.

Returns Ok(value) when the given parser succeeds, and Err(Missing) without consuming input when it fails, so the result never fails.

many : Parser(input, a) -> Parser(input, List(a)) where [input.len : input -> U64]

A parser which runs the element parser zero or more times on the input, returning a list containing all the parsed elements.

Repetition stops at the first element that fails, or that succeeds without consuming any input; that element's value is not included. (Without this rule, a parser such as chomp_while or maybe(p) that can succeed on empty input would repeat forever.) The failure of the element that stopped the repetition is kept, so if the parse later fails there, that failure is what gets reported.

one_or_more : Parser(input, a) -> Parser(input, List(a)) where [input.len : input -> U64]

A parser which runs the element parser one or more times on the input, returning a list containing all the parsed elements.

Fails when the first element fails. Also see Parser.many.

between : Parser(input, a), Parser(input, open), Parser(input, close) -> Parser(input, a)

Runs a parser for an 'opening' delimiter, then your main parser, then the 'closing' delimiter, and only returns the result of your main parser.

Useful to recognize structures surrounded by delimiters (like braces, parentheses, quotes, etc.)

between_brackets = |parser| parser.between(Utf8.codeunit('['), Utf8.codeunit(']'))
sep_by_one_or_more : Parser(input, a), Parser(input, sep) -> Parser(input, List(a)) where [input.len : input -> U64]

Parse one or more values separated by separator. The separators are consumed and omitted from the result.

A trailing separator that is not followed by a value is left unconsumed.

sep_by : Parser(input, a), Parser(input, sep) -> Parser(input, List(a)) where [input.len : input -> U64]

Parse zero or more values separated by separator. The separators are consumed and omitted from the result.

parse_numbers : Parser(Utf8.Bytes, List(U64))
parse_numbers = Utf8.digits.sep_by(Utf8.codeunit(','))

expect Utf8.parse_str(parse_numbers, "1,2,3") == Ok([1, 2, 3])
ignore : Parser(input, a) -> Parser(input, {})

Discard a parser's value while preserving how much input it consumes.

Useful with many to skip a repeated token without collecting values.

keep : Parser(input, a -> b), Parser(input, a) -> Parser(input, b)

Run a parser producing a function, then a parser producing its argument, and return the function result.

Start with const(|a| |b| ...) and add one keep per argument, as in the module example. The function must be curried (|a| |b| ..., not |a, b| ...), because its arguments are applied one at a time.

skip : Parser(input, a), Parser(input, _) -> Parser(input, a)

Run two parsers in sequence, discarding the second parser's value.

Both parsers must succeed, and the input read by the second is consumed.

at_sign : Parser(Utf8.Bytes, [AtSign])
at_sign = Parser.const(AtSign).skip(Utf8.codeunit('@'))

expect Utf8.parse_str(at_sign, "@") == Ok(AtSign)
chomp_until : a -> Parser(List(a), List(a)) where [a.is_eq : a, a -> Bool]

Match zero or more codeunits until it reaches the given codeunit. The given codeunit is not included in the match and is not consumed. Fails if the codeunit never appears in the remaining input.

This can be used with Parser.skip to ignore text.

ignore_text : Parser(Utf8.Bytes, U64)
ignore_text =
    Parser.const(|d| d)
        .skip(Parser.chomp_until(':'))
        .skip(Utf8.codeunit(':'))
        .keep(Utf8.digits)

expect Utf8.parse_str(ignore_text, "ignore preceding text:123") == Ok(123)

This can be used with Parser.keep to capture a list of U8 codeunits.

capture_text : Parser(Utf8.Bytes, List(U8))
capture_text =
    Parser.const(|codeunits| codeunits)
        .keep(Parser.chomp_until(':'))
        .skip(Utf8.codeunit(':'))

expect Utf8.parse_str(capture_text, "Roc:") == Ok(['R', 'o', 'c'])

Use Str.from_utf8_lossy to turn the results into a Str.

Also see Parser.chomp_while.

chomp_while : (a -> Bool) -> Parser(List(a), List(a))

Match zero or more codeunits until the check returns false. The codeunit that returned false is not included in the match. Note: a chomp_while parser always succeeds, possibly consuming nothing.

This can be used with Parser.skip to ignore text. This is useful for chomping whitespace or variable names.

ignore_numbers : Parser(Utf8.Bytes, Str)
ignore_numbers =
    Parser.const(|str| str)
        .skip(Parser.chomp_while(|b| b >= '0' and b <= '9'))
        .keep(Utf8.string("TEXT"))

expect Utf8.parse_str(ignore_numbers, "0123456789876543210TEXT") == Ok("TEXT")

This can be used with Parser.keep to capture a list of U8 codeunits.

capture_numbers : Parser(Utf8.Bytes, List(U8))
capture_numbers =
    Parser.const(|codeunits| codeunits)
        .keep(Parser.chomp_while(|b| b >= '0' and b <= '9'))
        .skip(Utf8.string("TEXT"))

expect Utf8.parse_str(capture_numbers, "123TEXT") == Ok(['1', '2', '3'])

Use Str.from_utf8_lossy to turn the results into a Str.

Also see Parser.chomp_until.

span : Parser(List(item), a) -> Parser(List(item), List(item))

Run a parser and return the input it consumed, instead of its value.

The result is a slice of the input, not a copy, so recognising a token with small parsers and keeping its text costs no allocation.

identifier = Parser.span(Utf8.codeunit('$').skip(Parser.chomp_while(|b| b >= 'a' and b <= 'z')))
expect Utf8.parse_str(identifier, "$abc") == Ok("$abc".to_utf8())
ParseResult : Try({ value : a, rest : input }, [ParseError({ message : Str, offset : U64 })])

The result of running a parser on part of an input: the parsed value and the rest of the input, or a ParseError whose offset counts the input elements (bytes, for UTF-8 input) read before the failure.