test: run every integration environment in the parallel batch - #981
Merged
Merged
Conversation
One fake GCS server starts in TestMain, so the datastore tests no longer call t.Setenv and no longer need NoParallel. TestMigrate becomes parallel too. Harness waits are widened because four environments now boot at once.
Thirteen tests never submit a Soroban transaction, so the raised resource limits do nothing for them. They now skip the upgrade. The upgrade itself polls Core's /sorobaninfo instead of sleeping a fixed five seconds twice. Also fixes a dead branch: the testnet case compared against the formatted file name, so it could never match.
… setup being slow The 30s window only ever passed because the settings upgrade slept 10s first. Waiting for the network to close about 16 more ledgers needs a window sized for that, not for whatever setup happened to cost.
There was a problem hiding this comment.
Pull request overview
Improves integration-test parallelism and reduces unnecessary setup time.
Changes:
- Shares one isolated fake GCS server across datastore tests.
- Parallelizes migration tests and skips unnecessary limit upgrades.
- Replaces fixed sleeps with polling and increases startup timeouts.
Reviewed changes
Copilot reviewed 14 out of 14 changed files in this pull request and generated 1 comment.
Show a summary per file
| File | Description |
|---|---|
transaction_test.go |
Skips limits for classic transaction cases. |
migrate_test.go |
Parallelizes migration environments. |
metrics_test.go |
Skips limits setup. |
main_test.go |
Adds shared GCS setup and limits helper. |
infrastructure/test.go |
Adds polling and longer startup timeouts. |
health_test.go |
Skips limits setup. |
get_version_info_test.go |
Skips limits setup. |
get_transactions_test.go |
Skips limits setup. |
get_network_test.go |
Skips limits setup. |
get_ledgers_test.go |
Shares GCS and improves waiting diagnostics. |
get_ledger_entries_test.go |
Skips limits setup. |
cors_test.go |
Skips limits setup. |
builtin_methods_disabled_test.go |
Skips limits setup. |
backfill_test.go |
Shares isolated GCS buckets and enables parallelism. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
- the health wait no longer reads what its condition goroutine writes, and the condition no longer logs, so a timeout cannot race or panic - the /sorobaninfo poll waits for a number the second upgrade file sets and enable.xdr does not; 65536 was already present and proved nothing - the number must stand alone, so 3500000 does not match 35000000 - waitForCheckpoint, waitForCoreAtLedger and the backfill waits get the same busy-machine budgets their siblings already got - TestGetLedgers waits for 15 ledgers, since 5 is exactly the first page and leaves nothing for the cursor to return - the fake GCS server only starts when integration tests are enabled
Shaptic
approved these changes
Sep 9, 2026
The wait covers 66 ledgers closing at one per second, so the budget is already the floor it needs. Nothing had failed on it, and the suite was green in CI without the extra two minutes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Makes the integration suite use all four CPUs and stop doing setup work no test asked for.
Integration test time
Note
These are not the "Integration tests (P27)" / "(P28)" job durations. They are strictly the
go testtime for the integration package, the number on this line of the job log:The job duration is a much bigger number and it is not comparable between runs:
setup-go, which builds librocksdb and libzstdgo.sumchangesmake build-libsrustup updateOne run in this PR's own history shows the problem: the job took 20m 51s while the tests took 7m 15s, because the merge changed
go.sumand every cache missed. Quoting thego testline keeps the comparison about the tests.0 failed tests, 0
TRY_AGAIN_LATERretries and 0-racereports in both jobs.Per-job links for the run after this PR:
What changed
t.Setenv("STORAGE_EMULATOR_HOST", …). An environment variable belongs to the whole program, so Go forbidst.Setenvin a parallel test. They had to setNoParallel: true.TestMigratenever calledt.Parallel(), and it does not finish until its own parallel subtests finish, so it blocked the batch twice over.TestMain, which sets the variable before any test is running to disturb.newGCSBucket(t)gives each test its own bucket, named after the test, so no two tests see each other's objects.t.SetenvandNoParallelflags deleted.TestMigrateand its subtests callt.Parallel().go test -runfiltering is unchanged.TestMainis Go's normal package entry point andm.Run()runs only what the filter selected. Verified:-run 'TestHealth$'runsTestHealthalone, 17.0s.upgradeLimitsWithFileended withtime.Sleep(5 * time.Second)before checking whether Core had applied the upgrade./sorobaninfoevery 500ms until the upgrade's value appears, capped at 30s.time.Sleepleft ininfrastructure/test.go.ApplyLimits: skipLimitsUpgrade()and skip the upgrade.upgradeLimitshadswitch limitFile { case "testnet": … }.limitFilewas alreadytestnet.p28.xdr, so the case never matched.fmt.Sprintf.Waits that had to grow
Four environments now boot at the same time instead of one. Every wait below was sized for the old schedule and expired during verification. Each returns the moment its condition holds, so a fast run pays nothing for the bigger number.
fillContainerPortsdocker compose portreporting a published portwaitForCore, both polls/info, then reaching syncwaitForRPCTestGetLedgersFromDatastoreThe same wait's failure message was also wrong:
require.Eventuallyformats its message arguments before polling starts, so it printedlast health: {LatestLedger:0 …}on every failure. It now usesassert.Eventuallyplust.Fatalfand prints the real last response.Not in this PR
35 of the 45 environments are byte-identical
NewTest(t, nil)and each still pays about 37s of setup. Sharing one network between them is a separate change.