Writing parsers is hard

Arne Christian Beer

This is a story of how we got from something like this:

Failed to deserialize BUILDINFO file:
Parser failed with the following error:
unknown
       ^
invalid makepkg build environment option
expected `buildflags`, `ccache`, `check`, `color`, `distcc`, `sign`, `makeflags`

To this:

A parser error message with

When the ALPM project started, the goal was fairly simple: Better tooling to interact with the low-level plumbing of Arch Linux Package Management (ALPM). This meant binaries to read/write individual file types, better error messages, extensive specifications and of course exports to common dataformats like JSON.

The thing is, although some of us already had experience with low-level dataframe parsing, none of us had experience on how to parse whole files. So as can be expected, we had to learn some things the hard way.

Although this is a very technical topic, I'll try to keep it short for you to have an enjoyable read! Let's get started.

Serde

In the Rust ecosystem, there's this very convenient crate called serde. It's a generic de-/serialization library, which allows effortless transformation between "simple" dataformats like JSON or YAML.

The first data formats we wrote parsers for were PKGINFO and BUILDINFO files, which contain information about a package and its build environment, respectively. Those are fairly straight forward and basically simplified INI-style data formats:

format = 2
pkgname = mypkg
pkgbase = mupkg
pkgver = 1.2-1
pkgarch = x86_64
installed = acl-2.3.2-1-x86_64
installed = archlinux-keyring-20241203-1-any
buildenv = unknown

Our first approach was to simply parse these values with the winnow parser library to a key-value map and feed it into serde's generic data types. serde then did the heavy lifting and mapped this generic data onto our alpm-types via their FromString implementations. While this worked, the error messages were somewhat lacking:

Failed to deserialize BUILDINFO file:
Parser failed with the following error:
unknown
       ^
invalid makepkg build environment option
expected `buildflags`, `ccache`, `check`, `color`, `distcc`, `sign`, `makeflags`

This is an inherent design flaw of the approach chosen by us: We deserialize the data into a generic serde-compatible format, during which all information about the surrounding context (line/column number) is lost. Any errors that happen when mapping serde's data to our types are completely detached from the actual context of the file.

So while this worked, we learned that serde is just not designed to be used in a context-preserving way, which would allow for more helpful error messages.

Winnow

Our next attempt was to write the parser with the winnow library from start to finish.

We left the old parsers be and started work on a parser for the SRCINFO format. This format is quite a bit more complex! Although it's also an INI style format, it uses the special keywords pkgbase = mypkgname and pkgname = mypkgname to start sections which span until either another pkgname keyword or the EOF is hit. It's not straight forward.

So to summarize: This parser needs to be aware of the current section it's in, by keeping track of special keywords.

All that said, writing the parser with winnow worked surprisingly well and we managed to get parser errors with surrounding context:

File parsing error:
parse error at line 4, column 11
  |
4 |  pkgrel = 1-nope
  |           ^
expected end of package release value

That was a big step in the right direction! The parser became quite a bit more complex, as we no longer had serde to do all of the type mapping for us, but on the upside we had full control over everything.

We continued to write parsers this way until the end of last year, when we increasingly noticed issues with the quality of error messages. You can already see an artifact of this problem in the error message above: the caret should be on the - and not the 1, as that's the actual position in which the parser failed.

Definitely-not-forward-parsing

I'm sure there's a well-known terminology for the type of parsing behavior I'll present to you in a bit, but I'll just call it "forward-parsing".

So what we did until now was "definitely-not-forward-parsing". Take a look at this code example, which can parse the line pkgname = my_pkg_name:

"pkgname = ".parse_next(input)?;
till_line_end
    .and_then(Name::parser)
    .context("The error message would go here")
    .parse_next(input)?;

The first line simply consumes the slice pkgname = .

till_line_end then takes all of the remaining content from the current cursor position up until the next newline or the end of file (EOF) (i.e. my_pkg_name). .and_then(Name::parser) then calls the nested Name::parser, on the literal slice of content my_pkg_name.

Intuitively, I expected this to work out just fine, as Name::parser is also a winnow parser. However, we run into a similar issue as we did with the serde parsers: When using and_then, the nested Name::parser only operates on the my_pkg_name slice without any context of its surroundings.

Now, when an error occurs, the parent parser has no idea where exactly the error occurred and simply points to the start of the slice that we just tried to parse:

File parsing error:
parse error at line 120, column 11
    |
120 | pkgname = qemu-common$$$
    |           ^
invalid character in package name
expected ASCII alphanumeric character, `_`, `@`, `+`, `-`, `.`, the name of a package

This was a fundamental issue in most of our parsers. The parsing approach for all our types was guided by the line-based nature of our file parsers:

  1. Parse until the end of a line
  2. Ensure that content can be mapped to the type while being fully consumed

Turns out, what we had to do instead was:

  1. Consume all valid tokens for the given type and ensure they are valid.
  2. Check if there's the expected newline/delimiter/EOF, if not throw an error about unexpected trailing content.

Forward-parsing

To fix this issue, we had to refactor everything. All of our parsers would need to be restructured to follow this new paradigm.

The ideation for this happened in November 2025, with the first draft MR for the groundwork being opened shortly after. Several months of development and many headaches later, the final MR was merged (that MR contains pretty much all rationale, examples and up-/downsides, if you're interested in further details).

All parsers now attempted to consume the tokens their respective type expected, and, if valid, stopped afterwards. This allowed much finer control when handling parsers and made filetype specific error handling of the consumer libraries much better.

In practice, the call-sites now looked like this:

"name = ".parse_next(input)?;
let name = Name::parser.parse_next(input)?;
// Expect either the newline or the EOF
(eof, newline)
  .context("The error message would go here")
  .parse_next(input)?;

// We also introduced some helper functions, which allowed us
// to make the newline handling a bit more ergonomic:
let name = Name::parser_until_line_ending_inclusive(input)?;

Although this doesn't look like much of a difference, the error messages were finally correct:

File parsing error:
parse error at line 120, column 22
    |
120 | pkgname = qemu-common$$$
    |                      ^
invalid character in package name
expected ASCII alphanumeric character, `_`, `@`, `+`, `-`, `.`, the name of a package

Another great benefit of this refactoring is that all of our parsers now behave the exact same way. Previously, there were small deviations in behavior, but since we restructured all parsers around proper traits (interfaces), we know exactly how each parser will behave. They're just really really nice to use now.

But wait, there's more

Now that we got proper error positioning, I finally decided to tackle the issue of missing context in error messages.

While the errors are already pretty good now, they could be better!
There could be an error span, which highlights the specific sequence of characters that caused an issue. When parsing nested types or data formats, it would be awesome to see the different layers...
There could be colors.

Winnow's default error type is more of a debug type, and they explicitly state so in their documentation and tutorial. It became clear that if we wanted to have custom-tailored contextual errors, we would have to write our own parser error library.

After another two months of working on a new parser error library and refactoring almost all parsers again, we merged the last MR and we finally have really fancy errors:

A parser error message for an invalid package name. Multiple colors are used to highlight specific sections and the different parsing layers are part of the error message.

With support for layered context and error spans:

Another parser error message, which features an error span beneath the sequence of characters that caused the error in question.

While the error messages may not yet be 100% perfect, we now have all the tooling to make them perfect. It's just a matter of adjusting wordings and adding/removing layers of context. Small stuff.

And not just that, the new error type also allows us to translate our parser error messages 🎉 , which simply wasn't an option before.

Next steps

While most of the heavy lifting is now done, the two original serde-based parsers for BUILDINFO and PKGINFO still need to be migrated. However, at this point it should now be fairly straight-forward. It's even a good first issue if that's your kind of thing :D.

Either way, the parsers in alpm-srcinfo, [alpm-db] and [alpm-repo-db] are released and ready. And it would make a lot of sense to actually start using them.

For example, it would be super helpful to run these parsers by default in the [aurweb] application, so that new AUR package maintainers get early feedback if they made any mistakes.

We also regularly observe outdated or invalid SRCINFO files in the official package source repositories during our test runs. As such it might even make sense to call alpm-srcinfo or even alpm-lint somewhere in the devtools to catch such oversights early on.

Anyhow, we're excited to already have removed a huge chunk of technical debt in a six month effort and we're looking forward to writing new stuff again! Soonâ„¢