Repository navigation
Inter Version CPU Benchmark must run on all key commits #529
Description
Activity
- addedAcceptedFor proposals, which were approved and will be implementedFor proposals, which were approved and will be implemented
on Nov 3, 2021 Two measured things, on
masterata45a7256, and the second may make this issue much smaller than the plan in the body.1.
CommonFunctionsInterVersiondoes not compare versionsDespite the name, it benchmarks one version — the local build:
<!-- Sources/Tests/DotnetBenchmark/DotnetBenchmark.csproj --> <ProjectReference Include="..\..\AngouriMath\AngouriMath.csproj" />
There is no second version anywhere in the project. The "inter version" comparison is done by running it on different commits at different times and comparing the recorded numbers afterwards — which is exactly the thing this issue says is unreliable, because the environment drifts between those runs. So the name describes the intent and the plan, not the mechanism.
2. BenchmarkDotNet can do this natively, in the version already pinned
Job.WithNuGet(package, version)exists in 0.13.0, the version inDotnetBenchmark.csproj:BenchmarkDotNet.Jobs.JobExtensions.WithNuGet(Job, String, String, Uri, Boolean) BenchmarkDotNet.Jobs.JobExtensions.WithNuGet(Job, NuGetReferenceList)One benchmark class with one job per version, all run in a single process on a single machine on a single day, is the comparability this issue is asking for. It needs no daily schedule, no cloning of key commits, and no replacing
Program.cs— the benchmark source stays one file and only the library under it changes.The obstacle, and it is the whole of the work:
WithNuGetgenerates a project per job that references the package, so the benchmark cannot also carry aProjectReferenceto the local kernel. Either the versioned benchmark becomes its own class referencing AngouriMath by package only, or the local build is included as one more "version" via a locally-packed nupkg and a source. The first is simpler and probably right — comparing releases is a different question from benchmarking the working tree.I have not prototyped it, so treat the API as verified and the design as proposed. If you want it, I would rather build it against a couple of released versions and show the numbers than argue for it further.
Note on scope
#500 recently gained a step in this direction — the benchmark's CSVs are now retained as CI artifacts for 90 days rather than living only in an expiring job log (PR #962). That gives the history this issue wants to plot; what it does not give is comparability across those runs, which is precisely what the
WithNuGetroute would fix at the source rather than by scheduling.- added a commit that references this issue
on Aug 23, 2026 - added a commit that references this issue
on Sep 1, 2026 The measurement half of this landed in #1140:
Sources/Utils/benchmark_key_commits.shmeasures a list of commits with one benchmark on one machine,combine_key_commits.pywrites the table,key-commits.txtsays which commits, and.github/workflows/BenchmarkKeyCommits.ymlruns it.Leaving this open, because two of the steps here were deliberately not done and both are yours to decide:
Weekly and on demand, not daily. A run is about eight minutes per commit, so the five-entry list is most of an hour of runner time, and there is nothing to see on a day when nothing merged.
workflow_dispatchis how it gets used while a regression is being hunted. Easy to change tocron: dailyif you would rather.Nothing is pushed to a website. The last steps here name
AngouriMathLab/performance-reportsandAngouriMathLab/performance-reports-tools; neither repository exists, asBenchmark.ymlalready records. The table is a run artifact and a step summary until they do.One limitation worth knowing about
It reaches back three releases, not five. Measured rather than assumed: v2.3.0 and later build, v2.2.0 and earlier do not, because the benchmark's
MatchingEnginecase reachesMatchedRulesandPatterns, which were internal until v2.3.0.Going further back means compiling only a subset of the benchmark per commit — which means conditioning the benchmark project on which commit it is pointed at, and the reason these columns compare at all is that the benchmark does not vary. So the bound is recorded in
key-commits.txtand the table names the rows it could not measure rather than omitting them.The first run
Across v2.3.0 → v2.4.0 → master, bytes allocated:
benchmark v2.3.0 v2.4.0 master CompileEasy16,376 11,307 10,970 CompileHard38,193 20,434 20,433 SimplifyHard3,620,669,800 3,632,015,736 3,632,014,824 SolveHard1,431,858,544 1,446,798,528 1,446,794,072 A third less than two releases ago on the two compile cases; everything else flat to within a per cent.
Allocation and time are printed as two tables rather than two columns of one, because they are not the same kind of evidence: allocation is deterministic, while the Kernel Benchmark run twice on one unchanged commit has come back 30-57% apart on this project's own runners, with allocation over the same pair agreeing to 0.03%. The generated document says which to read and why.
- added a commit that references this issue
on Sep 5, 2026
Current situation
CI runs the benchmark on the latest commit, and we add its results to the overall perf report, collected over time.
Problem
The environment we have might change, e. g.:
And those all factors distort the result.
Solution
Program.csUnresolved questions
Such a thing would take quite a lot of time to run when there are a lot of key commits. How can we optimize it? Either by caching the results of the same env configuration (machine + .net), or by reducing the number of samples ran by BDN. Or something else...?
Related: #500