Xml

:= {
    declaration : Try(Declaration, [Missing]),
    root : Node,
}

XML document tree and parser based on the XML 1.0 (Fifth Edition) specification.

The parser checks well-formedness for documents without a document type declaration. It supports an optional XML declaration, nested elements, attributes, character data, CDATA sections, the five predefined entity references (< > & ' "), and decimal and hexadecimal character references. Comments and processing instructions are checked and then dropped from the tree.

The tree is normalised the way an XML processor reports it: line endings become \n (XML 1.0 section 2.11), attribute values are normalised as CDATA attributes (section 3.3.3), references are replaced by their text, and adjacent character data (including CDATA sections and text on both sides of a comment) is merged into one Text node.

Document type declarations (<!DOCTYPE ...>) are rejected, so entities other than the predefined ones are never declared and referring to one is a well-formedness error. Namespaces are not processed: a name such as svg:path is kept as written. Input is already decoded text, so a declared encoding is reported but not used to decode.

Read a document with Xml.parse_str, then walk it with pattern matching or the Xml.Node helpers name, attribute, children_named and text:

expect {
    xml = Xml.parse_str("<feed><title>News</title><link href='/a'/></feed>")?
    titles = xml.root.children_named("title").map(|title| title.text())
    links = xml.root.children_named("link").map(|link| link.attribute("href"))
    titles == ["News"] and links == [Ok("/a")]
}

Originally written by Johannes Maas.

is_eq : _

Compare two XML documents structurally (declaration and tree).

to_hash : _

Hash an XML document, so documents can be Dict keys and Set members.

parse_str : Str -> Try(Xml, [InvalidXml(Error)])

Parse one complete XML document.

Leading and trailing whitespace, comments and processing instructions around the root element are allowed; anything else after it is an error. A leading byte order mark is accepted.

expect Xml.parse_str("<a>x &amp; y</a>").map_ok(|xml| xml.root) == Ok(Element({ name: "a", attributes: [], children: [Text("x & y")] }))

expect Xml.parse_str("<a>&nbsp;</a>") == Err(InvalidXml({ line: 1, column: 4, message: "undeclared entity &nbsp;" }))
parser : Parser(Bytes, Xml)

Parse one XML document, including an optional declaration and any trailing comments, processing instructions and whitespace, and leave the input after that for the next parser.

A failure is ParseError({ message, offset }) with the same message Xml.parse_str reports and the byte offset of the problem. Use this to embed an XML document in a larger parser; otherwise prefer Xml.parse_str, which reports a line and column.

expect Utf8.parse_str(Xml.parser, "<a></b>") == Err(ParseError({ message: "end tag </b> does not match start tag <a>", offset: 3 }))

expect Utf8.parse_str(Xml.parser, "<a/> <b/>") == Err(ParseError({ message: "unexpected input", offset: 5 }))
Attribute : { name : Str, value : Str }

An XML attribute name and decoded value.

The name is kept as written, including any namespace prefix. The value has references replaced and whitespace normalised. Attributes keep their source order; duplicate names are a well-formedness error.

TextEncoding : [
    Utf8Encoding,
    OtherEncoding(Str),
]

Text encoding declared by an XML declaration. Encoding names are compared case-insensitively, so UTF-8 is Utf8Encoding.

Declaration : {
    version : Version,
    encoding : Try(TextEncoding, [Missing]),
}

Version and optional encoding from an XML declaration. An absent encoding is Err(Missing).

Version : { major : U64, minor : U64 }

An XML version from the declaration, such as { major: 1, minor: 0 } for version="1.0". Only 1.x versions are accepted, so major is always 1.

Node

:= [
    Element({ name : Str, attributes : List(Attribute), children : List(Node) }),
    Text(Str),
]

An XML element or text node.

An Element holds its name as written, its attributes in source order, and its children in document order. Text(str) is decoded character data; adjacent text, CDATA and references are merged into one Text, and whitespace-only text between elements is kept.

is_eq : _

Compare two XML nodes structurally.

to_hash : _

Hash an XML node, so nodes can be Dict keys and Set members.

name : Node -> Try(Str, [NotAnElement])

The name of an element, or Err(NotAnElement) for a text node.

expect Xml.parse_str("<svg:rect/>").map_ok(|xml| xml.root.name()) == Ok(Ok("svg:rect"))
attribute : Node, Str -> Try(Str, [Missing])

The value of the element's attribute with this exact name, or Err(Missing) when there is none or the node is text.

expect {
    root = Xml.parse_str("<a href='/x'/>")?.root
    root.attribute("href") == Ok("/x") and root.attribute("title") == Err(Missing)
}
children_named : Node, Str -> List(Node)

The element's direct child elements with this exact name, in document order. Empty for a text node.

expect {
    root = Xml.parse_str("<list><item>1</item><other/><item>2</item></list>")?.root
    root.children_named("item").map(|item| item.text()) == ["1", "2"]
}
text : Node -> Str

All character data in the node and its descendants, concatenated in document order (like the DOM's textContent).

expect Xml.parse_str("<p>Hello, <b>world</b>!</p>").map_ok(|xml| xml.root.text()) == Ok("Hello, world!")
Error : { line : U64, column : U64, message : Str }

Location and explanation of malformed XML. Lines and columns are one-based; columns count UTF-8 bytes, and \r\n, \r, and \n each end a line.