Repository navigation
We have no test coverage of building the repo with the runtime we produce #122998
Description
Activity
- addeduntriagedNew issue has not been triaged by the area ownerNew issue has not been triaged by the area owner
on Jan 8, 2026 - addedneeds-area-labelAn area label is needed to ensure this gets routed to the appropriate area ownersAn area label is needed to ensure this gets routed to the appropriate area owners
on Jan 8, 2026 Can we do this by adding a stage 1 + stage 2 VMR validation build to replace our source-build validation legs we have today? Then we'd get full VMR validation without having to own tooling that manually patches an SDK while also validating the VMR better.
We could, but that might take more time and introduce more points of failure. I'm not sure the workflow would lend itself to easily reproducing locally. The nice thing about just building runtime is that devs should also be able to do that locally. I don't think we would need to patch the SDK, perhaps just running with some environment variable.
Codeflow into the VMR requires the entire product to be buildable with the runtime we produce.
I think we have this coverage in dotnet/dotnet as part of the two stage build validation. What am I missing?
We should add a job to https://github.com/dotnet/runtime/blob/main/eng/pipelines/runtime.yml that uses the runtime that was just built, to build the product.
This won't work when intra-release breaking changes occur. It is not unusual situation.
I think we have this coverage in dotnet/dotnet as part of the two stage build validation. What am I missing?
Right, but when we hit the problem there, it stops codeflow and accumulates debt for each day that changes are not merged. It's a double-edged sword that we have this coverage in stage2. It's great to have the coverage - but by putting it in a codeflow process that's meant to be routine and detached from the developer SMEs we increase the failure rate of that process. I think anyone working in runtime can agree that they'd much rather work in a runtime PR than an codeflow PR into VMR (no matter how much we improve the dev experience on that).
The goal of having earlier coverage would be to get the knowledge of the break further upstream and into a workflow where devs are more familiar and can leverage their normal inner-loops to efficiently investigate and fix the issue. It would make folks aware of the issue early and let them fix it as a non-fire-drill.
This won't work when intra-release breaking changes occur. It is not unusual situation.
Yes, I suppose this is true. We wouldn't be able to react to any intentional breaking change, and we can't have the policy of none ever. I guess this would have to be an optional leg if we had it, which degrades the value.
My motivation behind this issue is we are trying to make runtime (product repos in general) be a superset of testing done by VMR. We permitted very minimal testing in the VMR because we could rely on product repos doing that testing. The whole "stage2" build is one giant complex functional test which made me realize we had test gaps here. I'm trying to figure out how we find those earlier and off the critical path.
- added and removedneeds-area-labelAn area label is needed to ensure this gets routed to the appropriate area ownersAn area label is needed to ensure this gets routed to the appropriate area owners
on Jan 8, 2026 dotnet-policy-service commented
on Jan 8, 2026 ContributorMore actionsTagging subscribers to this area: @dotnet/runtime-infrastructure
See info in area-owners.md if you want to be subscribed.Right, but when we hit the problem there, it stops codeflow and accumulates debt for each day that changes are not merged.
When we hit breaks in codeflow PRs and we are not able to figure out how to fix it quickly, our first course of action should be to figure out what the revert in the contributing repo to get the code flowing again.
I think the problem with dotnet/dotnet#3873 that motivated this suggestion was poor operations. We have tried to fix the breaks in the codeflow PR that blocked the codeflow for close to a month. Instead, we should have reverted the commit that introduced the breaks so that the codeflow is not blocked, and work on figuring out the fixes for the breaks without blocking the codeflow.
We have tiered testing strategy. dotnet/runtime CI does not and cannot not test everything all the time. It is by design that breaks that block other processes get in occasionally. I have been keeping an eye on dotnet/runtime health, unapologetically reverting changes that introduced unexpected bad breaks, and strongly encouraged everybody to do the same.
My motivation behind this issue is we are trying to make runtime (product repos in general) be a superset of testing done by VMR.
I do not think this should be the goal. To deliver on this goal, runtime (and every repo in general) would have to include the VMR CI in their CI. It is complicated and expensive.
It is fine to keep track of the breaks that sneak in and try to plug the holes through which breaks sneak in frequently in targeted way or when it is cheap to do so.
Reacted by Andy GockeI think the problem with dotnet/dotnet#3873 that motivated this suggestion was poor operations. We have tried to fix the breaks in the codeflow PR that blocked the codeflow for close to a month. Instead, we should have reverted the commit that introduced the breaks so that the codeflow is not blocked, and work on figuring out the fixes for the breaks without blocking the codeflow.
The month it took was not due to the regression. It was due to the retargeting effort (and the holidays). That cannot be simply reverted. It's the necessary work to solve all the integration issues. I've shared plenty of other ideas as part of that effort, and I'm sure folks involved have even more to do this better next time. I don't see this issue as "solving" that. This was just one thing that was different about everything we hit in retargeting.
We happened to discover the product regression at the very end of fixing everything else with retargeting. We time-boxed a solution. The break wasn't "in production" yet and we had engagement from @jkoritzinsky who did a great job root causing and fixing the issue. If we hadn't had engagement or it was difficult to root cause we could have reverted - but reverting once the code is already in the VMR PR is not so simple. We might want a "playbook" for that moving forward if we expect to do it often. Do we revert in runtime first and let it flow, revert in both runtime and VMR PR, revert only in VMR PR and let it backflow?
To deliver on this goal, runtime (and every repo in general) would have to include the VMR CI in their CI. It is complicated and expensive.
No, that's not at all what I'm suggesting. I'm not trying to find everything that might be found in VMR CI - I'm trying to front load functional testing of using the product in a build workload before we actually try to do it in production builds. I'm calling out that we're doing something rather new with VMR + source-build. We're running a the portion of product we just built with a very large workload before we can commit that code and before we can publish those bits for others to test. We aren't doing this in "tests" that we can selectively ignore and asynchronously fix. We're doing it before we have bits published. Right now in preview 1 this is the first time we're doing that with substantially destabilizing features. This is a huge functional test for us. Any place our test coverage comes up short will make us have to investigate in VMR ingestion PRs or worse - VMR official builds if it happens to be platform specific. Given that VMR is a known sub-par development experience right now, I think we should do some investment to provide more coverage in runtime. I've also suggested that we need to improve the diagnosability of those VMR builds - since they're now running more "unproven" code.
Maybe the proposal of building the entire repo is too much - given the repo will be on a newer SDK and we might have build-to-build breaks between that. Running a sufficiently complex "build" workload might be a reasonable test. Our goal should be to catch these problems in runtime as much as possible since that's where folks are familiar working and that's where we'll provide the most coverage of fixes that they might need to make.
It was due to the retargeting effort (and the holidays). That cannot be simply reverted.
I understand that version bumps take always a lot of work. There was a commit (or several commits) with the version bump in the codeflow payload. I think those commits should have been reverted to unblock the codeflow, and the breaks worked through without blocking the codeflow.
We might want a "playbook" for that moving forward if we expect to do it often. Do we revert in runtime first and let it flow, revert in both runtime and VMR PR, revert only in VMR PR and let it backflow?
I have done reverts to unblock codeflow number of times, for example dotnet/dotnet#544 (comment) . It is easy (once you identify the offender) and typically the most expedient way (several hours) to unblock the codeflow. I think it is best to revert in the repo and let the revert flow since it minimizes risk of unexpected interactions with other changes.
I've also suggested that we need to improve the diagnosability of those VMR builds
+1
Our goal should be to catch these problems in runtime as much as possible
There needs to be balance between cost of the extra coverage (both machine cost and the cost of maintaining it) and the benefit of the coverage.
I think Jeremy's suggestion (#122998 (comment)) is striking a reasonable balance.
I understand that version bumps take always a lot of work. There was a commit (or several commits) with the version bump in the codeflow payload. I think those commits should have been reverted to unblock the codeflow, and the breaks worked through without blocking the codeflow.
There are a number of proposals floating around how to better do this next time: feature branch, readiness exercises, top-down. Reverting these branding/versioning changes is tricky due to conflicts and follow-on commits that depend on changes made. For better or worse we've gone the route of "drive for forward progress and fix what breaks" approach on this the past few times - the goal this time around was to do this over the holidays and have things done by the time the team got back. We must have this change in to ship in P1 so reverting doesn't really help towards that goal and if anything will delay the result and all the work that follows. It's not like reverting will change what the versioning change looks like in runtime, it just lets other code flow while folks further workout the cross-cutting changes - and to what end if those codeflow changes don't ship? Calling it a failure of operations is exceptionally harsh given all the folks involved intentionally prioritized retargeting over codeflow.
Regardless, I'll reiterate that this particular issue was not meant to be about the retargeting work but instead be a recognition of a significant test gap we have in producing builds that did not exist before. I think we need to do something to fill that gap.
There needs to be balance between cost of the extra coverage (both machine cost and the cost of maintaining it) and the benefit of the coverage.
Add to that the ability for developers to reproduce it locally and fit into their existing workflows. I would wager that no one really runs the source-build job locally today, nor runs it to validate a fix for an issue that it finds. That's a problem. I would hope we could have something that more naturally fits a developers workflow. I don't disagree with @jkoritzinsky's suggestion, but I would challenge us all to think of something that we'd be willing to do before a PR is raised, something we'd be willing to debug and run locally.
build --with-tool-runtime <path>would be be a nice workflow to have.I'll put out a PR creating the YAML for a new pipeline to do the VMR validation (the tooling from there is structured to work as a separate pipeline). Once we agree on the approach, I'll get that PR merged and have the infra folks configure the pipeline.
Actually, I might be able to put this in another pipeline, so I'll try that.
and to what end if those codeflow changes let through don't ship?
- To ensure that we are not going to accumulate multiple entangled breaks that will be hard to entangle once the code flow resumes. (I have seen similar effect in dotnet/rutime CI - if the dotnet/runtime CI is in bad shape for more like a day, we are guaranteed to pick up more breaks that creates more required work to get to a good shape again.)
- To unblock other work that depends on codeflow. For example, Fix ILVerify IndexOutOfRangeException for malformed exception handling clause bounds #122056 has been blocked on VMR codeflow. I have checked on it multiple times over the last month whether it is unblocked, only to find that it is still blocked.
Calling it a failure of operations is exceptionally harsh given all the folks involved intentionally prioritized retargeting over codeflow.
I really appreciate the work that folks put into making it work this time. I am sorry - I did not mean to offend them. At the same time, I think that blocking the codeflow for multiple weeks is unacceptable for any type of change, including version bumps. We are going to have more complicated changes going in this release (runtime consolidation, runtime async enablement). We need to execute them such that they do block the codeflow for weeks.
recognition of a significant test gap we have in producing builds that did not exist before. I think we need to do something to fill that gap.
I do not see that as a significant test gap. As I have said, we have tiered validation system. The VMR build is one of the many tiers. It is by design that failures sneak in. Our operations manual need to assume that and be prepared to handle it. We just need to ensure that it does not happen too often.
It is important to plug the holes only when we see certain class of failures to sneak in too often. I have not seen evidence that the breaks caught by this extra validation happen often enough.
I would wager that no one really runs the source-build job locally today
Yes, I think it is a good thing.
I would challenge us all to think of something that we'd be willing to do before a PR is raised
I am heavily encouraging folks to depend on the CI system for testing their changes, and only debug locally once they see something failing in the CI. For PRs created with GitHub Copilot that we encouraged to use, folks do not even have the changes locally most of the time.
- removeduntriagedNew issue has not been triaged by the area ownerNew issue has not been triaged by the area owner
on Jan 21, 2026 - locked and limited conversation to collaborators
on Feb 21, 2026
Metadata
Metadata
Assignees
Labels
Type
Projects
- StatusShow more project fieldsDone
Codeflow into the VMR requires the entire product to be buildable with the runtime we produce. If the product is not buildable, then the PR validation will fail, blocking merge. This is challenging because fixing the product at that point is more difficult - due to the size/complexity of VMR and distance from the original change/developer.
Today we have nothing in runtime that validates this. Only our unit test suite that gives good coverage, but not complete.
We should add a job to https://github.com/dotnet/runtime/blob/main/eng/pipelines/runtime.yml that uses the runtime that was just built, to build the product.
It would need to download the tar-ball, extract it, then set some environment to force the SDK to use that runtime instead of the one it bundles.
This can help us catch regressions that are only covered by build.