Skip to content

[Experiment] Split native deflate hash allocation on Win64 - #134853

Closed
alinpahontu2912 wants to merge 1 commit into
dotnet:mainfrom
alinpahontu2912:fix/windows-deflate-native-allocation
Closed

alinpahontu2912 wants to merge 1 commit into
dotnet:mainfrom
alinpahontu2912:fix/windows-deflate-native-allocation

Conversation

@alinpahontu2912

Copy link
Copy Markdown
Member

Experimental native-only alternative

Draft for discussion, not merge-ready: this change improves warm parallel compression in local measurements but makes the cold workload slower. It is not a demonstrated fix for the first-operation regression in #134700.

This is a separate, native-only experiment related to #134700 and the managed-pooling experiment in #134790. It is based on upstream main at 9499094 and contains none of the managed pooling changes. PR #134790 is unchanged.

Change

On _WIN64, allocate zlib-ng's 128 KiB hash-head table separately from the consolidated deflate state:

  • Keep the window and previous-position table adjacent, preserving existing layout assumptions.
  • Preserve 64-byte hash-table alignment and initialization behavior.
  • Free the owner allocation if the second allocation fails; free the separate table before its owning metadata on teardown.
  • Reuse the existing allocation/free callbacks and deflateCopy path.
  • Keep the non-Win64 allocation path unchanged and record the vendored modification.

The patch adds no pool, retained-state policy, managed changes, public API, or native export. The motivation is to test smaller allocation blocks while preserving the dedicated Windows heap. A specific Windows decommit threshold has not been established.

Correctness validation

Local Windows x64 Release validation:

  • Native library build passed; all 80 exports match the unmodified control.
  • A session-local C contract harness passed 162 configurations covering first/second allocation failure during initialization and copying, allocation guards/free counts, 64-byte alignment, window/previous-table adjacency, copy independence after source disposal, reset without allocation, and round trips.
  • Six additional pending-output and destruction-order cases passed.
  • Unpooled native-branch System.IO.Compression tests: 2,652 passed.
  • System.IO.Compression.ZipFile tests: 367 passed.
  • Default suites exclude OuterLoop and failing categories.
  • Source review found no actionable ownership/copy/alignment issues.

The contract harness is an experimental validation artifact, not a new checked-in regression test. _WIN64 also includes Windows ARM64; ARM64 was not built or executed locally.

Performance: benefit and blocker

The reproduction uses #134700's seeded 5,000-byte, half-random payload. Control and candidate run in separate private CoreRun hosts with identical runtime and unpooled managed compression binaries; only System.IO.Compression.Native.dll differs. Output hashes and timed output-byte totals matched.

These are native-patch before/after measurements, not a direct comparison between .NET 8 and .NET 10.

Three fresh-process samples per condition. Cold runs measure 3,000 operations; warm runs measure 20,000 operations after warm-up. Median elapsed time:

Workload DOP Cold control Cold candidate Warm control Warm candidate
Fixed configuration 1 144.2 ms 143.4 ms 885.2 ms 886.5 ms
Fixed configuration 4 117.9 ms 163.5 ms 283.0 ms 291.2 ms
Fixed configuration 8 103.0 ms 138.6 ms 288.5 ms 228.2 ms
Fixed configuration 12 98.7 ms 122.4 ms 277.7 ms 198.4 ms
Mixed configurations 4 97.2 ms 159.3 ms 256.3 ms 212.4 ms

Four additional alternating-order DOP-8 pairs confirmed the tradeoff:

  • Cold: 134.2 to 175.7 ms, approximately 31% slower.
  • Warm: 359.8 to 264.1 ms, approximately 27% faster.

Shared-host noise was retained in the results. Cold page faults per operation increased; warm faults decreased. Warm improvement therefore does not establish that this addresses the original cold/first-operation regression.

No further allocator variants or additional policy machinery are proposed here. This draft preserves the small experiment and its negative result for discussion; it does not propose replacing the managed-pooling PR or claim to resolve #134700.

Note

This patch, validation summary, and PR description were prepared with GitHub Copilot.

Allocate the hash-head table separately while preserving window adjacency, alignment, allocation-failure cleanup, and copy semantics. Document the vendored change. This experiment improves warm throughput but does not fix the measured cold-start regression.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: d488efd3-0804-456f-9286-2129a5cf5c6e
@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 4 pipeline(s).
12 pipeline(s) were filtered out due to trigger conditions.
There may be pipelines that require an authorized user to comment /azp run to run.

@dotnet-policy-service

Copy link
Copy Markdown
Contributor

Tagging subscribers to this area: @karelz, @dotnet/area-system-io-compression
See info in area-owners.md if you want to be subscribed.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Parallel short-lived DeflateStream compressors: page-fault storms on .NET 10 (Windows x64), none on .NET 8

1 participant