CSV

RFC 4180-style CSV parsing and typed record decoding.

The dialect is RFC 4180 with the common relaxations of Python's csv module (strict mode) and Rust's csv crate:

  • Fields are separated by , and records by CRLF, LF or a lone CR. The line break after the last record is optional, and empty input has no records.
  • A field is any run of bytes other than ,, CR and LF, so fields may hold arbitrary UTF-8, tabs, spaces (kept verbatim) and control characters.
  • A field that starts with " is quoted: it may contain ,, CR and LF, a doubled "" stands for one ", and the closing quote must be followed by ,, a line break or the end of input. An unterminated quoted field is an error.
  • A " anywhere else in an unquoted field is an ordinary character.
  • A blank line is a record with one empty field, as RFC 4180's grammar says (Python's csv.reader returns an empty list for it instead).
  • Rows may have different numbers of fields. A byte order mark is not removed; it is part of the first field.

There are three ways to read a file:

  • CSV.parse decodes rows into records whose field names match the header row, with the row type chosen by type inference;
  • CSV.parse_with runs a hand-built record parser (record, field and Parser.keep) over every row, matching columns by position;
  • CSV.parse_records returns the raw fields.
Person : { name : Str, age : U64 }

people : Try(List(Person), [InvalidCsv(CSV.Error), MissingRequiredField(Str)])
people = CSV.parse("name,age\nAda,36\nAlan,41\n")

expect people == Ok([{ name: "Ada", age: 36 }, { name: "Alan", age: 41 }])
parse_records : Str -> Try(List(Record), [InvalidCsv(Error)])

Parse CSV text into raw records without decoding them.

expect CSV.parse_records("a,\"b,c\"\n") == Ok([["a".to_utf8(), "b,c".to_utf8()]])
parser : Parser(Bytes, List(Record))

A parser for a whole CSV file, for use inside larger parsers.

It reads records until its input ends, so it consumes all of it or fails with the offset of the malformed field. It accepts any bytes.

split_header : List(Record) -> { header : List(Str), rows : List(Record) }

Split off the first record as column names.

Names are decoded lossily and kept verbatim. Empty input gives no names and no rows.

expect {
    records = CSV.parse_records("name,age\nAda,36\n")?
    CSV.split_header(records) == { header: ["name", "age"], rows: [["Ada".to_utf8(), "36".to_utf8()]] }
}
parse_with : Parser(Record, a), Str -> Try(List(a), [InvalidCsv(Error)])

Parse CSV text and decode every record with a hand-built record parser.

Fails with InvalidCsv(error) at the first problem: text that is not valid CSV, a field that does not decode, a record with too few fields, or a record with fields the parser did not read.

user : Parser(CSV.Record, { name : Str, age : U64 })
user = CSV.record(|name| |age| { name, age }).keep(CSV.field(CSV.string)).keep(CSV.field(CSV.u64))

expect CSV.parse_with(user, "Ada,36\nAlan,41\n") == Ok([{ name: "Ada", age: 36 }, { name: "Alan", age: 41 }])
decode : Parser(Record, a), List(Record) -> Try(List(a), [InvalidCsv(Error)])

Decode already parsed records with a hand-built record parser.

Errors are as for CSV.parse_with, numbered from the first record given, with line and column set to 0.

record : a -> Parser(Record, a)

Start a record parser with a curried constructor for the desired value.

Add one .keep(CSV.field(...)) per column, in column order:

Name : { first : Str, last : Str }

name : Parser(CSV.Record, Name)
name = CSV.record(|first| |last| { first, last }).keep(CSV.field(CSV.string)).keep(CSV.field(CSV.string))

expect CSV.parse_with(name, "Ada,Lovelace") == Ok([{ first: "Ada", last: "Lovelace" }])
field : Parser(Bytes, a) -> Parser(Record, a)

Consume the next field of a CSV.Record using a parser for its bytes.

The field parser must consume the whole field. Fails when the record has no fields left.

string : Parser(Bytes, Str)

Parse one field as a valid UTF-8 string, kept verbatim (no trimming).

u64 : Parser(Bytes, U64)

Parse one field as an unsigned 64-bit integer.

The field must be ASCII decimal digits with an optional leading +, and fit in a U64 (the grammar of Rust's u64::from_str). Whitespace, underscores, signs other than + and 0x-style prefixes are rejected.

f64 : Parser(Bytes, F64)

Parse one field as a 64-bit floating-point number.

The field is an optional sign followed by inf, infinity or nan in any case, or by a decimal number with an optional fraction and exponent (12, 1.5, .5, 5., 1e-3, 2.5E+10); this is the grammar of Rust's f64::from_str. Magnitudes too large for an F64 become infinity and too small ones become zero. Whitespace, underscores and hexadecimal forms are rejected.

parse : Str -> Try(List(row), [InvalidCsv(Error), ..errs]) where [row.Parseable([InvalidCsv(Error), ..errs])]

Decode CSV text whose first record is a header row into rows of an inferred type, usually a record.

Each record field reads the column whose header is exactly its name; columns no field names are ignored, and when two columns share a name the last one wins. A field of type Try(a, [Missing]) is Err(Missing) when its column is absent, and a field of type Try(a, [Null]) is Err(Null) when its cell is empty. A required field with no column fails with MissingRequiredField(name), which the compiler adds to the error type.

Cells decode as Str (verbatim), Bool (true or false in any letter case), integers (ASCII digits with an optional sign, - only for signed types), Dec (an optional sign, digits and an optional fraction), F32 and F64 (as CSV.f64), and tags without payloads (the cell is the tag name). A field that is a record, a tuple or a tag with a payload fails when it is decoded; a field that is a list or a dictionary does not compile.

Blank lines are skipped. Every other record must have exactly as many fields as the header; a shorter or longer record is an error naming it.

Item : { sku : Str, count : U64, note : Try(Str, [Missing]), price : Try(Dec, [Null]) }

items : Try(List(Item), [InvalidCsv(CSV.Error), MissingRequiredField(Str)])
items = CSV.parse("price,sku,count\n1.50,A1,3\n,B2,0\n")

expect items == Ok([
    { sku: "A1", count: 3, note: Err(Missing), price: Ok(1.50) },
    { sku: "B2", count: 0, note: Err(Missing), price: Err(Null) },
])
parse_normalized : Str -> Try(List(row), [InvalidCsv(Error), ..errs]) where [row.Parseable([InvalidCsv(Error), ..errs])]

Like CSV.parse, but header names are normalized before they are matched: surrounding spaces are removed, ASCII letters are lowercased, and runs of spaces, - and _ become one _. A column headed First Name, first-name or FIRST_NAME fills a field first_name.

names : Try(List({ first_name : Str }), [InvalidCsv(CSV.Error), MissingRequiredField(Str)])
names = CSV.parse_normalized(" First Name \nAda\n")

expect names == Ok([{ first_name: "Ada" }])
parse_headerless : Str -> Try(List(row), [InvalidCsv(Error), ..errs]) where [row.Parseable([InvalidCsv(Error), ..errs])]

Decode CSV text without a header row into tuples, one element per column.

Every record (blank lines excepted) must have exactly as many fields as the tuple has elements. Cells decode as in CSV.parse.

pairs : Try(List((Str, U64)), [InvalidCsv(CSV.Error)])
pairs = CSV.parse_headerless("Ada,36\nAlan,41\n")

expect pairs == Ok([("Ada", 36), ("Alan", 41)])
Record : List(List(U8))

One CSV record: its fields as raw bytes, in column order.

Fields parsed from a Str are valid UTF-8; records built by hand may hold any bytes.

Error : { record : U64, field : U64, line : U64, column : U64, message : Str }

Where and why CSV text could not be read or decoded.

record and field are one-based and count every record of the input, including a header row and blank lines. line and column are the one-based physical line and byte column of the field (or of the record, when the field does not exist); they are 0 when the source text is not known, as in CSV.decode. message is bounded in length and quotes field text lossily, so it is safe to show for any input.

Parseable : row
    where [
        row.parser_for : Format -> (State -> Try({ value : row, rest : State }, errs)),
    ]

Names the requirement that a row type can be decoded from CSV, so a signature can say "CSV-parseable" without naming the internal format and state types.

Format

CSV.Format :: # (opaque)

The CSV format for type-directed decoding. Used through CSV.parse; you do not need to name it.

parse_str : Format, State -> Try({ value : Str, rest : State }, [InvalidCsv(Error)])

Read the cell as UTF-8 text, verbatim.

parse_bool : Format, State -> Try({ value : Bool, rest : State }, [InvalidCsv(Error)])

Read true or false in any letter case.

parse_u8 : Format, State -> Try({ value : U8, rest : State }, [InvalidCsv(Error)])

Read ASCII digits with an optional + as a U8.

parse_i8 : Format, State -> Try({ value : I8, rest : State }, [InvalidCsv(Error)])

Read ASCII digits with an optional sign as a I8.

parse_u16 : Format, State -> Try({ value : U16, rest : State }, [InvalidCsv(Error)])

Read ASCII digits with an optional + as a U16.

parse_i16 : Format, State -> Try({ value : I16, rest : State }, [InvalidCsv(Error)])

Read ASCII digits with an optional sign as a I16.

parse_u32 : Format, State -> Try({ value : U32, rest : State }, [InvalidCsv(Error)])

Read ASCII digits with an optional + as a U32.

parse_i32 : Format, State -> Try({ value : I32, rest : State }, [InvalidCsv(Error)])

Read ASCII digits with an optional sign as a I32.

parse_u64 : Format, State -> Try({ value : U64, rest : State }, [InvalidCsv(Error)])

Read ASCII digits with an optional + as a U64.

parse_i64 : Format, State -> Try({ value : I64, rest : State }, [InvalidCsv(Error)])

Read ASCII digits with an optional sign as a I64.

parse_u128 : Format, State -> Try({ value : U128, rest : State }, [InvalidCsv(Error)])

Read ASCII digits with an optional + as a U128.

parse_i128 : Format, State -> Try({ value : I128, rest : State }, [InvalidCsv(Error)])

Read ASCII digits with an optional sign as a I128.

parse_dec : Format, State -> Try({ value : Dec, rest : State }, [InvalidCsv(Error)])

Read an optional sign, digits and an optional fraction.

parse_f32 : Format, State -> Try({ value : F32, rest : State }, [InvalidCsv(Error)])

Read a float with the grammar of CSV.f64.

parse_f64 : Format, State -> Try({ value : F64, rest : State }, [InvalidCsv(Error)])

Read a float with the grammar of CSV.f64.

parse_null : Format, State -> Try(State, [InvalidCsv(Error)])

An empty cell is null, so a Try(a, [Null]) field is Err(Null).

parse_record_start : Format, State -> Try([Counted({ len : U64, rest : State }), Uncounted(State)], [InvalidCsv(Error)])

A row is one record; a record inside a row is rejected.

parse_record_field : Format, FieldNames(_shape), State -> Try(
    [
        Field({ field : FieldName(_shape), rest : State }),
        TryField({ name : Str, rest : State }),
        TryFieldCaseless({ name : Str, rest : State }),
        Continue(State),
        Done(State),
    ],
    [InvalidCsv(Error)],
)

Offer the next header name as the field to fill.

parse_record_after_field : Format, State -> Try([Continue(State), Done(State)], [InvalidCsv(Error)])

The record ends after the last header column.

skip_record_field : Format, State -> Try(State, [InvalidCsv(Error)])

The cell to skip is chosen by the column, so nothing is consumed.

parse_tuple_start : Format, State, U64 -> Try(State, [InvalidCsv(Error)])

A headerless row is a tuple with one element per field.

invalid_value : Format, State -> [InvalidCsv(Error)]

The error for a cell that holds no valid value.

State

CSV.State :: # (opaque)

The cursor for type-directed decoding: one record, the header names, and indices into them. Used through CSV.parse; you do not need to name it.