Generic parser combinators
for transforming input into structured values.
A Parser(input, a) is a value that describes how to read an a from the
front of an input. Combine small parsers with keep, skip, map,
one_of, many and sep_by to build larger ones, then run the result with
Utf8.parse_str (for Str) or Parser.parse (for any input type).
This parser turns "Game 1: 3 blue, 4 red; 1 red, 2 green, 6 blue; 2 green"
into { id: 1, requirements: [[Blue(3), Red(4)], [Red(1), Green(2), Blue(6)], [Green(2)]] }
(the same code is a test in Utf8.roc):
Alternatives backtrack: when one alternative of alt or one_of fails, the
next one is tried on the original input, however much the failed one read.
Failures report the furthest position any parser reached. When a
repetition such as many stops at a bad element, or an alternative fails
after reading further than the one that succeeded, that failure is kept
and reported if the whole parse later fails at or before it.
The representation is internal and may change to improve efficiency or error messages.
custom : (input -> ParseResult(input, a)) -> Parser(input, a) where [input.len : input -> U64]
Write a custom parser without using the provided combinators.
The function receives the remaining input and returns either the parsed
value with the rest it did not consume, or
Err(ParseError({ message, offset })), where offset is how far into
the given input the failure is (usually 0).
run : Parser(input, a), input -> ParseResult(input, a) where [input.len : input -> U64]
Run a parser on the start of an input, returning the parsed value and
the input it did not consume.
Most parsers consume part of input when they succeed. This allows you to string parsers
together that run one after the other. The part of the input that the first
parser did not consume, is used by the next parser.
On failure, the error is the furthest failure seen, with its offset
counted from the start of input. This is mostly useful when creating
your own parsing building blocks.
Run a parser on the given input, expecting it to consume all of it.
The input type only needs a len method, which is used both to detect
leftover input and to turn the furthest failure into an offset counted
from the start of input. Leftover input is a failure too: if some
parser failed at or beyond the leftover, its message is reported;
otherwise the message is unexpected input. For UTF-8 text use
Utf8.parse_str or Utf8.parse_bytes.
fail : Str -> Parser(input, a) where [input.len : input -> U64]
Parser that can never succeed, regardless of the given input.
It will always fail with the given error message.
This is mostly useful as a 'base case' if all other parsers
in a one_of or alt have failed, to provide some more descriptive error message.
alt : Parser(input, a), Parser(input, a) -> Parser(input, a)
Try the first parser and (only) if it fails, try the second parser as fallback.
The second parser starts from the same input as first (backtracking).
If both fail, the failure that reached further is reported; when both
stopped at the same place, the second one is.
one_of : List(Parser(input, a)) -> Parser(input, a) where [input.len : input -> U64]
Try a list of parsers in turn, until one of them succeeds.
Each parser starts from the same input. An empty list always fails.
For UTF-8 input, Parser.one_of behaves the same way.
color : Parser(Utf8.Bytes, [Red, Green, Blue])
color =
Parser.one_of([
Parser.const(Red).skip(Utf8.string("red")),
Parser.const(Green).skip(Utf8.string("green")),
Parser.const(Blue).skip(Utf8.string("blue")),
])
expect Utf8.parse_str(color, "green") == Ok(Green)
map : Parser(input, a), (a -> b) -> Parser(input, b)
Transforms the result of parsing into something else,
using the given transformation function.
map2 : Parser(input, a), Parser(input, b), (a, b -> c) -> Parser(input, c)
Transforms the result of parsing into something else,
using the given two-parameter transformation function.
Transforms the result of parsing into something else,
using the given three-parameter transformation function.
If you need transformations with more inputs,
take a look at keep.
flatten : Parser(input, Try(a, Str)) -> Parser(input, a) where [input.len : input -> U64]
Removes a layer of Try from running the parser.
Use this to map functions that return a Try over the parser:
an Err(msg) value becomes a failure with that message, reported where
the inner parser started.
even : Parser(Utf8.Bytes, U64)
even =
Utf8.digits
.map(|n| if n % 2 == 0 { Ok(n) } else { Err("odd number") })
.flatten()
expect Utf8.parse_str(even, "42") == Ok(42)
expect Utf8.parse_str(even, "7") == Err(ParseError({ message: "odd number", offset: 0 }))
and_then : Parser(input, a), (a -> Parser(input, b)) -> Parser(input, b)
Run first, then build the next parser from its value and run that.
Use this when what comes next depends on what was read, such as a
length prefix. When the next parser is fixed, prefer keep, skip
and map, which are simpler and let the parser be built once.
lazy : ({} -> Parser(input, a)) -> Parser(input, a)
Runs a parser lazily.
This is (only) useful when dealing with a recursive structure.
For instance, consider a type Comment : { message : Str, responses : List(Comment) }.
Without lazy, you would ask the compiler to build an infinitely deep parser.
(Resulting in a compiler error.)
Mutually recursive top-level parser values currently require a lower-level
workaround because of roc-lang/roc#10098.
maybe : Parser(input, a) -> Parser(input, Try(a, [Missing]))
Make a parser optional.
Returns Ok(value) when the given parser succeeds, and Err(Missing)
without consuming input when it fails, so the result never fails.
many : Parser(input, a) -> Parser(input, List(a)) where [input.len : input -> U64]
A parser which runs the element parser zero or more times on the input,
returning a list containing all the parsed elements.
Repetition stops at the first element that fails, or that succeeds
without consuming any input; that element's value is not included.
(Without this rule, a parser such as chomp_while or maybe(p) that can
succeed on empty input would repeat forever.) The failure of the element
that stopped the repetition is kept, so if the parse later fails there,
that failure is what gets reported.
one_or_more : Parser(input, a) -> Parser(input, List(a)) where [input.len : input -> U64]
A parser which runs the element parser one or more times on the input,
returning a list containing all the parsed elements.
Fails when the first element fails. Also see Parser.many.
between : Parser(input, a), Parser(input, open), Parser(input, close) -> Parser(input, a)
Runs a parser for an 'opening' delimiter, then your main parser, then the 'closing' delimiter,
and only returns the result of your main parser.
Useful to recognize structures surrounded by delimiters (like braces, parentheses, quotes, etc.)
Discard a parser's value while preserving how much input it consumes.
Useful with many to skip a repeated token without collecting values.
keep : Parser(input, a -> b), Parser(input, a) -> Parser(input, b)
Run a parser producing a function, then a parser producing its argument,
and return the function result.
Start with const(|a| |b| ...) and add one keep per argument, as in
the module example. The function must be curried (|a| |b| ..., not
|a, b| ...), because its arguments are applied one at a time.
skip : Parser(input, a), Parser(input, _) -> Parser(input, a)
Run two parsers in sequence, discarding the second parser's value.
Both parsers must succeed, and the input read by the second is consumed.
Match zero or more codeunits until it reaches the given codeunit.
The given codeunit is not included in the match and is not consumed.
Fails if the codeunit never appears in the remaining input.
Match zero or more codeunits until the check returns false.
The codeunit that returned false is not included in the match.
Note: a chomp_while parser always succeeds, possibly consuming nothing.
This can be used with Parser.skip to ignore text.
This is useful for chomping whitespace or variable names.
ignore_numbers : Parser(Utf8.Bytes, Str)
ignore_numbers =
Parser.const(|str| str)
.skip(Parser.chomp_while(|b| b >= '0' and b <= '9'))
.keep(Utf8.string("TEXT"))
expect Utf8.parse_str(ignore_numbers, "0123456789876543210TEXT") == Ok("TEXT")
This can be used with Parser.keep to capture a list of U8 codeunits.
capture_numbers : Parser(Utf8.Bytes, List(U8))
capture_numbers =
Parser.const(|codeunits| codeunits)
.keep(Parser.chomp_while(|b| b >= '0' and b <= '9'))
.skip(Utf8.string("TEXT"))
expect Utf8.parse_str(capture_numbers, "123TEXT") == Ok(['1', '2', '3'])
Run a parser and return the input it consumed, instead of its value.
The result is a slice of the input, not a copy, so recognising a token
with small parsers and keeping its text costs no allocation.
identifier = Parser.span(Utf8.codeunit('$').skip(Parser.chomp_while(|b| b >= 'a' and b <= 'z')))
expect Utf8.parse_str(identifier, "$abc") == Ok("$abc".to_utf8())
ParseResult : Try({ value : a, rest : input }, [ParseError({ message : Str, offset : U64 })])
The result of running a parser on part of an input: the parsed value
and the rest of the input, or a ParseError whose offset counts the
input elements (bytes, for UTF-8 input) read before the failure.