Repository navigation
Keep the corpus numbers after the job that produced them, and diff them against a record (#529, #500) - #1013
Merged
Merged
Conversation
The gate has reported solved / unsolved / wrong / error / timeout on every commit since #945, and every one of those numbers has existed only in a job log that expires. Item 36 of #746 asks for the result "as a CI artefact with per-commit history"; neither half existed. The artefact half is corpus-report.tsv, written to the test output directory before anything is asserted -- so a run that then fails still leaves the numbers behind -- and uploaded from all three operating systems with if: always(), the shape Benchmark.yml already uses for the same reason. It carries the measured verdict, the note that explains it and the answer the library gave, plus the commit, the machine and how long the corpus took. The history half is corpus-baseline.tsv, committed. An artefact cannot supply it: artefacts expire, cannot be diffed, and are not attributable to the commit that moved a number. A file in the repository is all three through git log -p and git blame. The two destination repositories #500 names for the equivalent job on the benchmark side, AngouriMathLab/performance-reports and -tools, were checked again and still do not exist. The baseline is generated from Corpus.All rather than maintained beside it. The per-problem expectation already lives in the corpus as Problem.Expect, next to the problem it is about; a second hand-kept copy would be a second thing to get wrong. This is a projection, so keeping it current is mechanical -- re-run with AM_UPDATE_CORPUS_BASELINE=1 -- and it states the two things Expect does not state anywhere: the totals, and the membership as a flat list, so a corpus that quietly shrinks shows up in the diff. The convention is PublicApiSurfaceTest's, followed rather than reinvented. The failure message now separates the four ways the corpus can move, because they ask the reader to do different things: WRONG an answer that is not right, which fails whatever the record says WORSE the library got worse here; fix it BETTER nothing is broken -- update the record, and here is how CHANGED Error and Timeout share a rank, since there is no defensible ordering between throwing and hanging, so a move between them is neither Rank(): Solved beats everything, Wrong loses to everything, Unsolved beats Error and Timeout because declining in bounded time is a legitimate answer. CSharpTest.yml gains no path filters and gains a comment saying why: it is the suite, and the leading "./" that silenced the benchmark for ninety-eight merged PRs (#775) is a trap for whoever adds the first one. Corpus gate: 3.5/3.9/3.8 s before, 3.1/3.2/3.1 s after, three runs each on one machine; the corpus itself takes 1.4 s of that and the added work is one 3.4 kB file. Suite 7534 passed, 0 failed, 14 skipped -- master's 7533 plus the new baseline test. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011emtnRT6EWTrxXNtqDVK3e
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The corpus gate reports solved / unsolved / wrong / error / timeout on every commit, on three operating systems. Those numbers existed only in the job log, which expires — so there was no way to see whether a figure had moved without re-running an old commit by hand. #746 item 36 names two things, and this is both of them.
Two files, because the item names two things one file cannot be
corpus-report.tsv— the artefact. Written to the test output directory (never into the source tree; a test must not dirty the working copy) before any assertion, so a run that then fails still leaves numbers behind. It carries the measured verdict, the note explaining it, the answer the library actually gave, plus commit, OS, runtime and how long the corpus took. Uploaded from all three matrix legs,if: always()— a verdict that differs between operating systems is itself a finding.corpus-baseline.tsv— the per-commit history, committed. Chosen over accumulating artefacts on evidence rather than taste: artefacts expire (90 days is the cap), cannot be diffed, and cannot be attributed to the commit that moved a number. A file in the repository is all three, throughgit log -pandgit blame, and it does not expire. #500's two named destinations,AngouriMathLab/performance-reportsand-tools, were re-checked and both still 404 — the org 404s on the REST listing — so pushing history elsewhere was not an option to weigh.The baseline is generated from
Corpus.All, not maintained beside it. The per-problem expectation already exists asProblem.Expect, next to the problem it is about; a second hand-kept copy is a second thing to get wrong. So the file is a projection — regenerating it is mechanical (AM_UPDATE_CORPUS_BASELINE=1), never a judgement — and it states the two thingsExpectdoes not: the totals, and membership as a flat sorted list, so a corpus that quietly shrinks shows up in the diff.That shape — committed record,
AM_UPDATE_*, and a failure that spells out both directions — is already this repository's convention inCommon/PublicApiSurfaceTest. Followed rather than reinvented.Both failure modes, demonstrated
Three problems temporarily altered, then reverted:
A wrong answer can never be recorded as expected —
NoEntryExpectsAWrongAnswerrejects it — because a record that says "this one is wrong and that is fine" is how a wrong answer becomes permanent.Verdict ranking, used only for better-versus-worse:
Solved(0) < Unsolved(1) < Timeout(2) = Error(2) < Wrong(3). Solved beats everything and Wrong loses to everything, perAGENTS.md; Unsolved beats the other two because declining in bounded time is a legitimate answer. Between threw and hung there is no defensible ordering, so they share a rank and a move between them reports asCHANGEDrather than as progress in either direction.Path filters
CSharpTest.ymlhas nopaths:filter and so cannot carry the defect that silenced the benchmark for ninety-eight merged PRs (#775). None was added — it is the suite, and every change is one it can break — with a comment saying so, so that the first person to reach for a filter checks first.The same defect is still live in
CPPBuild.yml, lines 8 and 13:"./Sources/Wrappers/AngouriMath.CPP.*/**". Its last run was 2026-01-02T14:24 — the same dateBenchmark.ymlstopped, from the same commit. That file belongs to another branch in flight and is being fixed there.Measured
--no-build, same machine). The corpus itself is 1.4 s of that; the added work is rendering and writing one 3.4 kB file.TheBaselineIsTheOneOnRecord. Build clean, the same 29 warnings as before.TheBaselineIsTheOneOnRecordfails withNo corpus baseline at …/corpus-baseline.tsv. Re-run the suite with AM_UPDATE_CORPUS_BASELINE=1 to write …, then commit it.One caveat worth stating plainly
The corpus is 40 problems, all drawn from this project's own tracker, nothing vendored — and the baseline says
Solved=39, Unsolved=1, so the gate has one problem of headroom in the improvement direction. Everything that finds new defects lives outside CI:intbenchagainst Rubi moved 536 → 604 solved in the analysis workspace, an order of magnitude more cases than the in-CI gate holds in total.This makes the gate a reliable regression detector. It does not make it the measurement the roadmap leans on, and closing that gap is separate work (#746 items 39 and 40). Not expanded here, deliberately.