Skip to content

Should the parser stay ANTLR-generated? (intent of #565, needs-design) #898

Description

@Rafael-SOWNet

Splitting the intent of #565 out of the branch, which is being closed: it proposed replacing the ANTLR-generated parser with one built on Yoakke, and the question it raises is live even though the branch is not.

What the branch did

Deleted the whole of Core/Antlr — the grammar, the four generated files, the .interp/.tokens
artefacts and the bundled antlr-4.8-complete.jar — and replaced it with a hand-assembled parser:
+740 / −7185 across 26 files.

Why ANTLR is worth replacing

Not on parser-theory grounds. On the day-to-day cost of it:

  • The generated files are committed, so a grammar change is a two-step ritual — edit
    AngouriMath.g, run antlr_rerun.bat, then run a post-processing pass whose only job is to rewrite
    public to internal on the generated classes. Docs/Contributing/ImproveParser.md documents it,
    and the documentation has to warn you to regenerate the unmodified grammar first and check the diff
    is empty, so that a toolchain version difference is not mistaken for your change.
  • It needs a JDK to change the grammar at all, in a repository that otherwise needs only the .NET
    SDK.
  • A 1 MB jar is in the source tree.
  • The error messages are ANTLR's, which is why MissingOperatorParseException and friends exist to
    translate them, and why Docs/Usage/Syntax.md had to be written by hand as a separate statement of
    what the parser accepts.

Why it is not obviously worth doing

  • The parser works, and its contract is now tested. StringizeRoundTripTest holds printing to being
    parsing's inverse across every node type, and 2.0 fixed several node shapes that did not round-trip.
    A rewrite starts that guarantee from zero.
  • The grammar is the specification. Syntax.md describes it, but AngouriMath.g is the thing that
    decides, and it is 480-odd lines of readable declarative rules. A hand-written recursive-descent
    parser is more code and less obviously equivalent to a reader.
  • Yoakke is itself a dependency, and a small one — it would be trading a build-time dependency on a
    jar for a run-time dependency on a young library, which for a package with 313k downloads is a
    different kind of risk rather than less risk.
  • Nothing about the parser is currently a bug. The open parse issues are about what the grammar
    says — precedence, notation — not about how it is produced.

What would make this decidable

  1. Is the round trip preserved? StringizeRoundTripTest over every node type is the acceptance, and
    it did not exist when Yoakke parser rewrite #565 was written.
  2. Are the error messages better? The reason to hand-write a parser is control over failure; if the
    replacement's messages are no better than the translated ANTLR ones, the main benefit is only the
    build simplification.
  3. What does it cost at run time? Parsing is on the hot path for FromString, which caches by
    string precisely because it is not free.
  4. Does Syntax.md stay true, and can it be generated from the new parser rather than maintained
    beside it?

I have no recommendation. The four costs above are real and so are the four objections, and this is a
maintainer's call about what the library wants to own. Filed so that the reasoning survives the branch
rather than being rediscovered in another four years.

Activity

  1. Happypig375 commented on Aug 11, 2026

    @Happypig375
    Member

    Yoakke still seems to be inefficient. Unless there is proof of it becoming efficient, this rewrite is still out of the question.

  2. Rafael-SOWNet commented on Aug 16, 2026

    @Rafael-SOWNet
    MemberAuthor

    Four things this week's work turned up that bear on the question, all measured on master at a45a7256. None settles it; together they change what the trade-off is about.

    1. The generated parser is checked in, and nothing regenerates it. Sources/AngouriMath/Core/Antlr/AngouriMathParser.cs carries <auto-generated> and the project references Antlr4.Runtime.Standard only — there is no Antlr4BuildTasks, so the grammar in AngouriMath.g is not compiled at build time. Changing the grammar means running the ANTLR tool by hand and committing the output. That is a real maintenance cost and it is invisible from the csproj.

    2. It cost a real change this week. #873 (a quotient of two integer literals should parse as a Rational) would naturally be a grammar change. It is instead a post-parse tree rewrite in Parser.cs, precisely because regenerating was the higher-risk option. The fix works, but the shape of the fix was chosen by the toolchain rather than by the problem.

    3. Three open issues are blocked on the same missing grammar feature. #225 (quantifiers), #248 (Σ, Π) and #495's syntax half all need the parser to bind a variable — exists x in S, sum(i, …), a => … are one construct in different clothes. ConditionalSet already binds one on the entity side, so the concept exists in the tree and not in the grammar. Whichever is built first pays for that machinery.

    4. And a fourth wants the opposite of a grammar tweak. #286 and #495 both need implicit multiplication disabled for an identifier in call position, so f(x) is an application rather than f * x. Measured: f(x) * g(x) currently parses as f * x * g * x and differentiates to 2 * f * g * x — a silent wrong answer. Also measured: f(x, y) does not parse at all today, so the multi-argument form is free to claim.

    What I think this adds up to. The question is usually posed as ANTLR versus Yoakke, i.e. a library choice. The measurements say the live cost is not which generator but that the generated output is committed and regeneration is a manual, risky step — which makes every grammar change expensive regardless of the tool. Four issues are queued behind grammar changes right now.

    So the cheaper experiment before any rewrite: wire the existing grammar into the build (Antlr4BuildTasks), so a grammar edit is an edit rather than an operation. If that works, the four blocked issues get much cheaper without changing parser technology at all, and the Yoakke question can be decided later on its merits rather than under the pressure of an unusable toolchain. If it does not work, that failure is itself the strongest argument for replacing it.

    I have not attempted the build-time integration — it is a change I would want to verify on all three TFMs, and it deserves its own issue rather than being smuggled into this one.

  3. Rafael-SOWNet commented on Aug 16, 2026

    @Rafael-SOWNet
    MemberAuthor

    Correcting my comment above, and acting on the narrower thing it turned out to be.

    I wrote that regeneration was "a manual, risky step". Measured on 4b2c104b: it is manual and it is not risky. A full regeneration — the vendored antlr-4.13.1-complete.jar, then AntlrPostProcessorReplacePublicWithInternal — reproduces every committed file byte for byte. The jar is pinned in the repo, both commands are in antlr_rerun.bat, and Java is the only external requirement.

    So the committed parser has not drifted from its grammar, and my case for replacing the toolchain was weaker than I made it sound.

    What is real is forgetting to run it. #965 adds a CI job that regenerates and fails if the tree moves — verified both ways: it passes on a clean tree, and planting a new lexer token without regenerating makes it fail with 7 files and 443 lines of difference.

    I deliberately did not do the build-time integration I suggested earlier. It would put a Java runtime between a contributor and dotnet build, and it solves a problem the measurement says does not exist.

    None of this answers the question this issue actually asks — ANTLR or something else. It removes one of the reasons grammar changes feel expensive, which matters mainly because four issues are queued behind grammar work (#225, #248, #495's syntax half, #286). If the parser is eventually replaced, the check goes with the generated files it guards.

  4. Rafael-SOWNet commented on Sep 5, 2026

    @Rafael-SOWNet
    MemberAuthor

    Recommendation: keep ANTLR, and take the performance complaint seriously as a separate, much
    smaller change.

    What the objection actually was

    #565 was not about build friction. It was measured performance, and against the rewrite:
    "ParseEasy 6,815,103 ns vs 38,910 ns … (huge performance regression)", with the standing objection
    "Unless there is proof of it becoming efficient, this rewrite is still out of the question."

    That proof never arrived. Yoakke's performance PR closed unmerged in March 2025, its last commit is
    September 2025, and NuGet carries only nightlies, the most recent from 2022. Replacing a working
    4,952-line generated parser with a dependency in that state is not a trade worth making.

    The grammar is ambiguous, and that is where the cost is

    AngouriMath.g is 557 lines — 22 parser rules, 102 function-call alternatives, 155 token types.
    With LL_EXACT_AMBIG_DETECTION over the test corpus, ANTLR reports exact ambiguities in
    power_list (2,029), unary_expression (671) and negate_expression (243)
    . Each forces
    full-context prediction, which the DFA cannot cache.

    Those counts are structural and do not depend on the machine.

    The number I could not verify

    A two-stage strategy — SLL first, falling back to LL only on a bail — is the standard answer to
    exactly that, and is about ten lines in Parser.cs with no regeneration. There is a reported 43×
    improvement on ParseHard from it.

    I could not confirm that, and I am not going to quote it as if I had. My attempt to A/B it ran
    while a five-release benchmark was saturating the machine, and the baseline alone swung between
    9.7 µs and 61.8 µs per parse on identical code. Under that noise the comparison says nothing. What
    was stable across every run is the magnitude of the uncached hard parse: 1.5–1.9 ms, against
    about 2 µs when the parse cache answers, so a caller parsing many distinct expressions pays it.

    The measurement to make, on a quiet machine, both arms in one process:

    • MathS.FromString(hard, useCache: false) before and after two-stage;
    • the same for the easy input, because that is the common case and the one a regression would hurt;
    • StringizeRoundTripTest's 168 cases and the full suite, since a mispredicting SLL pass reports a
      syntax error rather than parsing slowly.

    What I would not do on this issue

    The other three items — round trip, error messages, a generated Syntax.md — are real, but none of
    them is an argument for a hand-written parser on its own, and the error-message one is where a
    replacement would actually earn its keep. If that is ever prioritised, decide it then, on a prototype
    that passes StringizeRoundTripTest and beats the two-stage numbers above rather than the numbers in
    #565.

  5. added this to the 2.6.0 milestone on Sep 18, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions