Skip to content

Keep the corpus numbers after the job that produced them, and diff them against a record (#529, #500) - #1013

Merged
Rafael-SOWNet merged 1 commit into
masterfrom
feat/corpus-report-artefact
Aug 23, 2026
Merged

Rafael-SOWNet merged 1 commit into
masterfrom
feat/corpus-report-artefact

Conversation

@Rafael-SOWNet

Copy link
Copy Markdown
Member

The corpus gate reports solved / unsolved / wrong / error / timeout on every commit, on three operating systems. Those numbers existed only in the job log, which expires — so there was no way to see whether a figure had moved without re-running an old commit by hand. #746 item 36 names two things, and this is both of them.

Two files, because the item names two things one file cannot be

corpus-report.tsv — the artefact. Written to the test output directory (never into the source tree; a test must not dirty the working copy) before any assertion, so a run that then fails still leaves numbers behind. It carries the measured verdict, the note explaining it, the answer the library actually gave, plus commit, OS, runtime and how long the corpus took. Uploaded from all three matrix legs, if: always() — a verdict that differs between operating systems is itself a finding.

# generated	2026-08-23T02:32:33Z
# os	Ubuntu 26.04 LTS
# runtime	.NET 10.0.10
# elapsed_ms	1449
# totals	Solved=39	Unsolved=1	Wrong=0	Error=0	Timeout=0	Total=40
name	op	expect	verdict	status	note	answer
int:arctan-form	Integrate	Solved	Solved	ok	differentiates back	2 * arctan(2 * x / 2) / 2 + C
int:hard	Integrate	Unsolved	Unsolved	ok	integral remains	integral(sqrt(tan(x)), x)
solve:trig	Solve	Solved	Solved	ok	8 root(s) verified	{ 2 * pi * n_1, pi + 2 * pi * n_1 }

corpus-baseline.tsv — the per-commit history, committed. Chosen over accumulating artefacts on evidence rather than taste: artefacts expire (90 days is the cap), cannot be diffed, and cannot be attributed to the commit that moved a number. A file in the repository is all three, through git log -p and git blame, and it does not expire. #500's two named destinations, AngouriMathLab/performance-reports and -tools, were re-checked and both still 404 — the org 404s on the REST listing — so pushing history elsewhere was not an option to weigh.

The baseline is generated from Corpus.All, not maintained beside it. The per-problem expectation already exists as Problem.Expect, next to the problem it is about; a second hand-kept copy is a second thing to get wrong. So the file is a projection — regenerating it is mechanical (AM_UPDATE_CORPUS_BASELINE=1), never a judgement — and it states the two things Expect does not: the totals, and membership as a flat sorted list, so a corpus that quietly shrinks shows up in the diff.

That shape — committed record, AM_UPDATE_*, and a failure that spells out both directions — is already this repository's convention in Common/PublicApiSurfaceTest. Followed rather than reinvented.

Both failure modes, demonstrated

Three problems temporarily altered, then reverted:

The corpus moved: 1 wrong, 1 worse, 1 better.

WRONG    lim:ratio: (x ^ 2 + 1) / (x ^ 2 - 1) -> 1 (expected 2). An answer that is not right is
         worse than no answer, so this fails whatever the record says. Do not record it:
         NoEntryExpectsAWrongAnswer rejects Verdict.Wrong as an expectation.
WORSE    int:power: recorded Solved, measured Unsolved (integral remains). THE LIBRARY GOT WORSE
         HERE. Fix it. Only lower the record if the old verdict is shown to have been wrong, and
         say why in the commit.
BETTER   int:trig: recorded Unsolved, measured Solved (differentiates back). NOTHING IS BROKEN --
         the library got better and the record does not say so. Set Expect = Verdict.Solved on
         this problem in Corpus.cs, re-run with AM_UPDATE_CORPUS_BASELINE=1 to refresh
         corpus-baseline.tsv, and commit both.

A wrong answer can never be recorded as expected — NoEntryExpectsAWrongAnswer rejects it — because a record that says "this one is wrong and that is fine" is how a wrong answer becomes permanent.

Verdict ranking, used only for better-versus-worse: Solved(0) < Unsolved(1) < Timeout(2) = Error(2) < Wrong(3). Solved beats everything and Wrong loses to everything, per AGENTS.md; Unsolved beats the other two because declining in bounded time is a legitimate answer. Between threw and hung there is no defensible ordering, so they share a rank and a move between them reports as CHANGED rather than as progress in either direction.

Path filters

CSharpTest.yml has no paths: filter and so cannot carry the defect that silenced the benchmark for ninety-eight merged PRs (#775). None was added — it is the suite, and every change is one it can break — with a comment saying so, so that the first person to reach for a filter checks first.

The same defect is still live in CPPBuild.yml, lines 8 and 13: "./Sources/Wrappers/AngouriMath.CPP.*/**". Its last run was 2026-01-02T14:24 — the same date Benchmark.yml stopped, from the same commit. That file belongs to another branch in flight and is being fixed there.

Measured

  • Gate duration: 3.51 / 3.92 / 3.79 s before, 3.08 / 3.17 / 3.12 s after (three runs each, --no-build, same machine). The corpus itself is 1.4 s of that; the added work is rendering and writing one 3.4 kB file.
  • Suite: 7534 passed, 0 failed, 14 skipped — master's 7533 plus TheBaselineIsTheOneOnRecord. Build clean, the same 29 warnings as before.
  • The failing test was written first: with no baseline present, TheBaselineIsTheOneOnRecord fails with No corpus baseline at …/corpus-baseline.tsv. Re-run the suite with AM_UPDATE_CORPUS_BASELINE=1 to write …, then commit it.

One caveat worth stating plainly

The corpus is 40 problems, all drawn from this project's own tracker, nothing vendored — and the baseline says Solved=39, Unsolved=1, so the gate has one problem of headroom in the improvement direction. Everything that finds new defects lives outside CI: intbench against Rubi moved 536 → 604 solved in the analysis workspace, an order of magnitude more cases than the in-CI gate holds in total.

This makes the gate a reliable regression detector. It does not make it the measurement the roadmap leans on, and closing that gap is separate work (#746 items 39 and 40). Not expanded here, deliberately.

The gate has reported solved / unsolved / wrong / error / timeout on every commit
since #945, and every one of those numbers has existed only in a job log that
expires. Item 36 of #746 asks
for the result "as a CI artefact with per-commit history"; neither half existed.

The artefact half is corpus-report.tsv, written to the test output directory
before anything is asserted -- so a run that then fails still leaves the numbers
behind -- and uploaded from all three operating systems with if: always(), the
shape Benchmark.yml already uses for the same reason. It carries the measured
verdict, the note that explains it and the answer the library gave, plus the
commit, the machine and how long the corpus took.

The history half is corpus-baseline.tsv, committed. An artefact cannot supply it:
artefacts expire, cannot be diffed, and are not attributable to the commit that
moved a number. A file in the repository is all three through git log -p and git
blame. The two destination repositories
#500 names for the equivalent
job on the benchmark side, AngouriMathLab/performance-reports and -tools, were
checked again and still do not exist.

The baseline is generated from Corpus.All rather than maintained beside it. The
per-problem expectation already lives in the corpus as Problem.Expect, next to
the problem it is about; a second hand-kept copy would be a second thing to get
wrong. This is a projection, so keeping it current is mechanical -- re-run with
AM_UPDATE_CORPUS_BASELINE=1 -- and it states the two things Expect does not state
anywhere: the totals, and the membership as a flat list, so a corpus that quietly
shrinks shows up in the diff. The convention is PublicApiSurfaceTest's, followed
rather than reinvented.

The failure message now separates the four ways the corpus can move, because they
ask the reader to do different things:

  WRONG    an answer that is not right, which fails whatever the record says
  WORSE    the library got worse here; fix it
  BETTER   nothing is broken -- update the record, and here is how
  CHANGED  Error and Timeout share a rank, since there is no defensible ordering
           between throwing and hanging, so a move between them is neither

Rank(): Solved beats everything, Wrong loses to everything, Unsolved beats Error
and Timeout because declining in bounded time is a legitimate answer.

CSharpTest.yml gains no path filters and gains a comment saying why: it is the
suite, and the leading "./" that silenced the benchmark for ninety-eight merged
PRs (#775) is a trap for
whoever adds the first one.

Corpus gate: 3.5/3.9/3.8 s before, 3.1/3.2/3.1 s after, three runs each on one
machine; the corpus itself takes 1.4 s of that and the added work is one 3.4 kB
file. Suite 7534 passed, 0 failed, 14 skipped -- master's 7533 plus the new
baseline test.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011emtnRT6EWTrxXNtqDVK3e
@Rafael-SOWNet
Rafael-SOWNet merged commit 8f71597 into master Aug 23, 2026
24 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant