Repository navigation
The property harnesses run in 8.5 minutes and CI runs none of them #1256
Description
Activity
A coverage gap in
crashcheckthat turned up while refreshing the reports against2dbeedf7, added here because it bears on what "adopt the harness" would actually be adopting.crashcheckbuilds one instance of every concreteEntitytype by reflection, and reports the types it could not build. That list grew from 22 to 24 between6a97c071and2dbeedf7, and the two additions areArgmaxfandArgminf— the binder forms added by #1222.Why they fail
The harness takes the shortest constructor of at most three parameters that are all
Entity, fills it from a cycling leaf list (x,2,1/2, …), and then requires the instance to surviveStringize()→ToEntity(), since text is how the case reaches the child process. ForArgmaxf(Expression, Var, Over)that isArgmaxf(x, 2, 1/2), which prints asargmax(x, 2 in 1/2)— and the parser refuses it, becauseargmaxrequires its second argument to be a membership whose element is a variable:parse(argmax(x, 2 in 1/2)) => InvalidArgumentParseException: argmax expects its second argument to say which variable ranges over which setSo the type is honestly reported as unreachable. Fine, and fixable with a written shape —
argmax(x ^ 2, x in [0; 1])parses and round-trips.The part worth flagging
MaximumfandMinimumfare declared identically —(Entity Expression, Entity Var, Entity Over), sameBinding.Of(Var).In(Over)initialiser — and they are not on the uncovered list. Not because they are covered. Because themax(grammar rule has a binary fall-through thatargmax(does not:'max(' … $args.list.Count == 2 && $args.list[1] is Inf { Element: Variable } ? MathS.Maximum(…) // the binder : $args.list.Aggregate((a, b) => MathS.Max(a, b)) // the binary nodeso
max(x, 2 in 1/2)does not throw — it parses, as the binaryMaxf. The round trip succeeds and the harness recordsMaximumfas built, while the expression that actually reaches the child process is a different node type carrying a nonsense2 in 1/2as an operand.So the true count of unreached binder nodes is four, and only two of them say so.
ArgmaxfandArgminfare visible precisely because their parser rule is stricter;MaximumfandMinimumfare invisible because theirs is more forgiving. A harness that reports its own gaps is the right design, and this is a hole in the gap-reporting rather than in the coverage.What it implies for this proposal
Two things, both small:
- Written shapes for the four binder forms, which is a one-line-each addition.
- "Built" has to assert the round trip returns the same node type, not merely that it parses. That check is what would have caught
Maximumf, and it costs one comparison.
It is also a concrete instance of the argument above for gating on the list rather than the count. Nothing here is a regression and no number went the wrong way —
crashcheckstill reports 0 crashed and 0 did-not-finish across 1890 cases. What carried the information was two names appearing in a list, which is exactly the signal a count would have hidden: the total went up, 1834 → 1890, at the same time as coverage of these four went nowhere.- added 6 commits that reference this issue
on Oct 1, 2026 Done. CI now runs ten of these harnesses on every change to the library, from
Sources/Tests/Harnesses, through.github/workflows/Harnesses.yml:PR harnesses fails on #1650 BoundCheck, RootCheck, CasBench a disagreement at a boundary, a missing or false root, a wrong answer #1651 PropCheck a property that does not hold #1652 SimpSweep, RuleCheck, CrashCheck a changed value, a set that does not settle or breaks its declaration, a crash #1653 CanonCheck, Confluence a change to the findings committed beside them ( HARNESS_UPDATE_BASELINE=1records one)#1656 DocSamples a wiki or website sample that does not compile, throws, or prints something its page doesn't say None of them fails on a timeout. Every report names the commit it measured and is uploaded with the run. DocSamples' first run found three wiki outputs that #1635 and a union change had made stale, and the wiki is corrected.
Three stay outside CI, each for its own reason:
intbenchneeds Rubi's suite, which carries no licence statement, so it can't be vendored;libcompareis a comparison with other libraries, and nothing in it is a defect;egraphmeasures memory cost, which is a number rather than a verdict.
The workspace copies of the ten are retired.
#746's standing condition "correctness coverage grows with the surface" is marked Met, with one caveat attached: the in-CI corpus is small and a set of property harnesses exists that CI never runs. This issue is about that caveat, and it is filed because measuring it showed the obstacle is not what I assumed.
The measurement
Ten self-contained harnesses, run one at a time against
2dbeedf7on one desktop, library already built:confluenceboundcheckSimplifykeep the value at the boundary — across a branch cut, off the real line, beside a polecanoncheckcasbenchrootcheckpropcheckegraphsimpsweeprulecheckcrashcheckAll ten exited 0. I had assumed this was a nightly-sized job and proposed it as one; it is a test-job-sized job.
crashcheckalone is half of it, and it is half of it because it spawns 1890 child processes — that cost would grow on a shared runner and much more so on Windows.Not included, and each for its own reason:
docsamplesneeds the wiki and website checked out;intbenchneeds the Rubi suite, which is downloaded rather than vendored because it carries no licence statement;libcomparepullsMathNet.SymbolicsandSymbolismfrom NuGet and was not timed.sympyparityis a special case —SymPyParity.ymlalready runs a scheduled in-repo watch, so that one is partly answered already.Why this is not just "add a workflow"
The harnesses live in a separate private workspace, not in this repository, so nothing here can run them today. That makes this a question for maintainers rather than a patch: would you want them in-repo? The code is mine to contribute if so.
There is already a shape to copy.
AotSmokeTestandDotnetBenchmarkare console projects underSources/Tests/that are in the solution, are driven by a workflow, and are not picked up bydotnet test. ASources/Tests/Harnesses/folder would sit alongside them. (One practical note:dotnet sln addrewrites the whole solution file — a single project added came out as several hundred changed lines, line endings and extra platform configurations included. The three entries are better inserted by hand.)The part that needs deciding, and it is not the runtime
These harnesses generate reports; they do not assert. Wiring them to CI means choosing what fails the build, and the naive choice is wrong: a count moving is usually not a regression.
boundcheckwent from 41 rewritten shapes to 35 over a fortnight with 0 disagreements throughout, which reads likeSimplifylosing capability and was nothing of the kind — a parser fix had stopped1/3arriving as a division, so the fold that used to happen no longer had anything to fold.So the gate should be the list, not the number — and this repository already has that pattern twice:
Corpus/corpus-baseline.tsv— a committed baseline, one line per problem, regenerated rather than hand-edited, diffed per commit.PerformanceGate— reads a committed baseline and fails on allocation moving more than 3%, while reporting time rather than failing on it, because a shared runner's wall clock belongs to whoever else is on the host.The same for a harness: commit the list of shapes it flags, fail when the list changes, and let the counts move freely.
Two constraints that follow from the harnesses' own semantics:
casbenchandcrashcheckreport a timeout as a verdict, so they must not share a runner with competing load or CI turns into a finding. That is already why the workspace's own runner script runs them alone and first.PerformanceGatealready documents.Suggested first step
One workflow, the six harnesses under 10s (
confluence,boundcheck,canoncheck,casbench,rootcheck,propcheck— 33s together), each gated on a committed list. That is small enough to run on every push rather than nightly, and it answers the caveat for the properties a corpus structurally cannot check: completeness of a root set, behaviour at a branch cut, idempotence of a normal form. The four slower ones can follow on a schedule once the baseline mechanism has proven itself.Part of #746 — the "correctness coverage grows with the surface" standing condition.