Skip to content

Inter Version CPU Benchmark must run on all key commits #529

Description

@WhiteBlackGoose

Current situation

CI runs the benchmark on the latest commit, and we add its results to the overall perf report, collected over time.

Problem

The environment we have might change, e. g.:

  • The max CPU usage
  • The max RAM
  • The default settings for .NET installation
  • The .NET runtime itself

And those all factors distort the result.

Solution

  • The CI to run this benchmark will run every day
    • On each run it clones the repo by key commits
      • On each clone it replaces the Program.cs
      • Then it runs the benchmark and saves the results to a file
    • The collected files are combined into one table and displayed
    • Then this table gets converted into an html table
    • By this table, a Plotly.NET visualization is built
    • This visualization in embedded html format is pushed to the website

Unresolved questions

Such a thing would take quite a lot of time to run when there are a lot of key commits. How can we optimize it? Either by caching the results of the same env configuration (machine + .net), or by reducing the number of samples ran by BDN. Or something else...?

Related: #500

Activity

  1. added
    AcceptedFor proposals, which were approved and will be implemented
    on Nov 3, 2021
  2. Rafael-SOWNet commented on Aug 16, 2026

    @Rafael-SOWNet
    Member

    Two measured things, on master at a45a7256, and the second may make this issue much smaller than the plan in the body.

    1. CommonFunctionsInterVersion does not compare versions

    Despite the name, it benchmarks one version — the local build:

    <!-- Sources/Tests/DotnetBenchmark/DotnetBenchmark.csproj -->
    <ProjectReference Include="..\..\AngouriMath\AngouriMath.csproj" />

    There is no second version anywhere in the project. The "inter version" comparison is done by running it on different commits at different times and comparing the recorded numbers afterwards — which is exactly the thing this issue says is unreliable, because the environment drifts between those runs. So the name describes the intent and the plan, not the mechanism.

    2. BenchmarkDotNet can do this natively, in the version already pinned

    Job.WithNuGet(package, version) exists in 0.13.0, the version in DotnetBenchmark.csproj:

    BenchmarkDotNet.Jobs.JobExtensions.WithNuGet(Job, String, String, Uri, Boolean)
    BenchmarkDotNet.Jobs.JobExtensions.WithNuGet(Job, NuGetReferenceList)
    

    One benchmark class with one job per version, all run in a single process on a single machine on a single day, is the comparability this issue is asking for. It needs no daily schedule, no cloning of key commits, and no replacing Program.cs — the benchmark source stays one file and only the library under it changes.

    The obstacle, and it is the whole of the work: WithNuGet generates a project per job that references the package, so the benchmark cannot also carry a ProjectReference to the local kernel. Either the versioned benchmark becomes its own class referencing AngouriMath by package only, or the local build is included as one more "version" via a locally-packed nupkg and a source. The first is simpler and probably right — comparing releases is a different question from benchmarking the working tree.

    I have not prototyped it, so treat the API as verified and the design as proposed. If you want it, I would rather build it against a couple of released versions and show the numbers than argue for it further.

    Note on scope

    #500 recently gained a step in this direction — the benchmark's CSVs are now retained as CI artifacts for 90 days rather than living only in an expiring job log (PR #962). That gives the history this issue wants to plot; what it does not give is comparability across those runs, which is precisely what the WithNuGet route would fix at the source rather than by scheduling.

  3. Rafael-SOWNet commented on Sep 1, 2026

    @Rafael-SOWNet
    Member

    The measurement half of this landed in #1140: Sources/Utils/benchmark_key_commits.sh measures a list of commits with one benchmark on one machine, combine_key_commits.py writes the table, key-commits.txt says which commits, and .github/workflows/BenchmarkKeyCommits.yml runs it.

    Leaving this open, because two of the steps here were deliberately not done and both are yours to decide:

    Weekly and on demand, not daily. A run is about eight minutes per commit, so the five-entry list is most of an hour of runner time, and there is nothing to see on a day when nothing merged. workflow_dispatch is how it gets used while a regression is being hunted. Easy to change to cron: daily if you would rather.

    Nothing is pushed to a website. The last steps here name AngouriMathLab/performance-reports and AngouriMathLab/performance-reports-tools; neither repository exists, as Benchmark.yml already records. The table is a run artifact and a step summary until they do.

    One limitation worth knowing about

    It reaches back three releases, not five. Measured rather than assumed: v2.3.0 and later build, v2.2.0 and earlier do not, because the benchmark's MatchingEngine case reaches MatchedRules and Patterns, which were internal until v2.3.0.

    Going further back means compiling only a subset of the benchmark per commit — which means conditioning the benchmark project on which commit it is pointed at, and the reason these columns compare at all is that the benchmark does not vary. So the bound is recorded in key-commits.txt and the table names the rows it could not measure rather than omitting them.

    The first run

    Across v2.3.0 → v2.4.0 → master, bytes allocated:

    benchmark v2.3.0 v2.4.0 master
    CompileEasy 16,376 11,307 10,970
    CompileHard 38,193 20,434 20,433
    SimplifyHard 3,620,669,800 3,632,015,736 3,632,014,824
    SolveHard 1,431,858,544 1,446,798,528 1,446,794,072

    A third less than two releases ago on the two compile cases; everything else flat to within a per cent.

    Allocation and time are printed as two tables rather than two columns of one, because they are not the same kind of evidence: allocation is deterministic, while the Kernel Benchmark run twice on one unchanged commit has come back 30-57% apart on this project's own runners, with allocation over the same pair agreeing to 0.03%. The generated document says which to read and why.

  4. added a commit that references this issue on Sep 5, 2026
    e513bf6
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    AcceptedFor proposals, which were approved and will be implemented

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions