Repository navigation
Should the parser stay ANTLR-generated? (intent of #565, needs-design) #898
Description
Activity
Yoakke still seems to be inefficient. Unless there is proof of it becoming efficient, this rewrite is still out of the question.
Four things this week's work turned up that bear on the question, all measured on
masterata45a7256. None settles it; together they change what the trade-off is about.1. The generated parser is checked in, and nothing regenerates it.
Sources/AngouriMath/Core/Antlr/AngouriMathParser.cscarries<auto-generated>and the project referencesAntlr4.Runtime.Standardonly — there is noAntlr4BuildTasks, so the grammar inAngouriMath.gis not compiled at build time. Changing the grammar means running the ANTLR tool by hand and committing the output. That is a real maintenance cost and it is invisible from the csproj.2. It cost a real change this week. #873 (a quotient of two integer literals should parse as a
Rational) would naturally be a grammar change. It is instead a post-parse tree rewrite inParser.cs, precisely because regenerating was the higher-risk option. The fix works, but the shape of the fix was chosen by the toolchain rather than by the problem.3. Three open issues are blocked on the same missing grammar feature. #225 (quantifiers), #248 (Σ, Π) and #495's syntax half all need the parser to bind a variable —
exists x in S,sum(i, …),a => …are one construct in different clothes.ConditionalSetalready binds one on the entity side, so the concept exists in the tree and not in the grammar. Whichever is built first pays for that machinery.4. And a fourth wants the opposite of a grammar tweak. #286 and #495 both need implicit multiplication disabled for an identifier in call position, so
f(x)is an application rather thanf * x. Measured:f(x) * g(x)currently parses asf * x * g * xand differentiates to2 * f * g * x— a silent wrong answer. Also measured:f(x, y)does not parse at all today, so the multi-argument form is free to claim.What I think this adds up to. The question is usually posed as ANTLR versus Yoakke, i.e. a library choice. The measurements say the live cost is not which generator but that the generated output is committed and regeneration is a manual, risky step — which makes every grammar change expensive regardless of the tool. Four issues are queued behind grammar changes right now.
So the cheaper experiment before any rewrite: wire the existing grammar into the build (
Antlr4BuildTasks), so a grammar edit is an edit rather than an operation. If that works, the four blocked issues get much cheaper without changing parser technology at all, and the Yoakke question can be decided later on its merits rather than under the pressure of an unusable toolchain. If it does not work, that failure is itself the strongest argument for replacing it.I have not attempted the build-time integration — it is a change I would want to verify on all three TFMs, and it deserves its own issue rather than being smuggled into this one.
Correcting my comment above, and acting on the narrower thing it turned out to be.
I wrote that regeneration was "a manual, risky step". Measured on
4b2c104b: it is manual and it is not risky. A full regeneration — the vendoredantlr-4.13.1-complete.jar, thenAntlrPostProcessorReplacePublicWithInternal— reproduces every committed file byte for byte. The jar is pinned in the repo, both commands are inantlr_rerun.bat, and Java is the only external requirement.So the committed parser has not drifted from its grammar, and my case for replacing the toolchain was weaker than I made it sound.
What is real is forgetting to run it. #965 adds a CI job that regenerates and fails if the tree moves — verified both ways: it passes on a clean tree, and planting a new lexer token without regenerating makes it fail with 7 files and 443 lines of difference.
I deliberately did not do the build-time integration I suggested earlier. It would put a Java runtime between a contributor and
dotnet build, and it solves a problem the measurement says does not exist.None of this answers the question this issue actually asks — ANTLR or something else. It removes one of the reasons grammar changes feel expensive, which matters mainly because four issues are queued behind grammar work (#225, #248, #495's syntax half, #286). If the parser is eventually replaced, the check goes with the generated files it guards.
- added a commit that references this issue
on Aug 16, 2026 Recommendation: keep ANTLR, and take the performance complaint seriously as a separate, much
smaller change.What the objection actually was
#565 was not about build friction. It was measured performance, and against the rewrite:
"ParseEasy 6,815,103 ns vs 38,910 ns … (huge performance regression)", with the standing objection
"Unless there is proof of it becoming efficient, this rewrite is still out of the question."That proof never arrived. Yoakke's performance PR closed unmerged in March 2025, its last commit is
September 2025, and NuGet carries only nightlies, the most recent from 2022. Replacing a working
4,952-line generated parser with a dependency in that state is not a trade worth making.The grammar is ambiguous, and that is where the cost is
AngouriMath.gis 557 lines — 22 parser rules, 102 function-call alternatives, 155 token types.
WithLL_EXACT_AMBIG_DETECTIONover the test corpus, ANTLR reports exact ambiguities in
power_list(2,029),unary_expression(671) andnegate_expression(243). Each forces
full-context prediction, which the DFA cannot cache.Those counts are structural and do not depend on the machine.
The number I could not verify
A two-stage strategy — SLL first, falling back to LL only on a bail — is the standard answer to
exactly that, and is about ten lines inParser.cswith no regeneration. There is a reported 43×
improvement onParseHardfrom it.I could not confirm that, and I am not going to quote it as if I had. My attempt to A/B it ran
while a five-release benchmark was saturating the machine, and the baseline alone swung between
9.7 µs and 61.8 µs per parse on identical code. Under that noise the comparison says nothing. What
was stable across every run is the magnitude of the uncached hard parse: 1.5–1.9 ms, against
about 2 µs when the parse cache answers, so a caller parsing many distinct expressions pays it.The measurement to make, on a quiet machine, both arms in one process:
MathS.FromString(hard, useCache: false)before and after two-stage;- the same for the easy input, because that is the common case and the one a regression would hurt;
StringizeRoundTripTest's 168 cases and the full suite, since a mispredicting SLL pass reports a
syntax error rather than parsing slowly.
What I would not do on this issue
The other three items — round trip, error messages, a generated
Syntax.md— are real, but none of
them is an argument for a hand-written parser on its own, and the error-message one is where a
replacement would actually earn its keep. If that is ever prioritised, decide it then, on a prototype
that passesStringizeRoundTripTestand beats the two-stage numbers above rather than the numbers in
#565.
Splitting the intent of #565 out of the branch, which is being closed: it proposed replacing the ANTLR-generated parser with one built on Yoakke, and the question it raises is live even though the branch is not.
What the branch did
Deleted the whole of
Core/Antlr— the grammar, the four generated files, the.interp/.tokensartefacts and the bundled
antlr-4.8-complete.jar— and replaced it with a hand-assembled parser:+740 / −7185 across 26 files.
Why ANTLR is worth replacing
Not on parser-theory grounds. On the day-to-day cost of it:
AngouriMath.g, runantlr_rerun.bat, then run a post-processing pass whose only job is to rewritepublictointernalon the generated classes.Docs/Contributing/ImproveParser.mddocuments it,and the documentation has to warn you to regenerate the unmodified grammar first and check the diff
is empty, so that a toolchain version difference is not mistaken for your change.
SDK.
MissingOperatorParseExceptionand friends exist totranslate them, and why
Docs/Usage/Syntax.mdhad to be written by hand as a separate statement ofwhat the parser accepts.
Why it is not obviously worth doing
StringizeRoundTripTestholds printing to beingparsing's inverse across every node type, and 2.0 fixed several node shapes that did not round-trip.
A rewrite starts that guarantee from zero.
Syntax.mddescribes it, butAngouriMath.gis the thing thatdecides, and it is 480-odd lines of readable declarative rules. A hand-written recursive-descent
parser is more code and less obviously equivalent to a reader.
jar for a run-time dependency on a young library, which for a package with 313k downloads is a
different kind of risk rather than less risk.
says — precedence, notation — not about how it is produced.
What would make this decidable
StringizeRoundTripTestover every node type is the acceptance, andit did not exist when Yoakke parser rewrite #565 was written.
replacement's messages are no better than the translated ANTLR ones, the main benefit is only the
build simplification.
FromString, which caches bystring precisely because it is not free.
Syntax.mdstay true, and can it be generated from the new parser rather than maintainedbeside it?
I have no recommendation. The four costs above are real and so are the four objections, and this is a
maintainer's call about what the library wants to own. Filed so that the reasoning survives the branch
rather than being rediscovered in another four years.