Skip to content

perf(core): write one sort-shuffle file per task instead of one per input - #2316

Merged
Dandandan merged 2 commits into
apache:mainfrom
Dandandan:perf/consolidate-shuffle-files-per-task
Aug 21, 2026
Merged

Dandandan merged 2 commits into
apache:mainfrom
Dandandan:perf/consolidate-shuffle-files-per-task

Conversation

@Dandandan

@Dandandan Dandandan commented Aug 16, 2026 •

Copy link
Copy Markdown
Contributor

Rebased onto main now that #2315 has merged; this PR is a single commit.
The change only matters when a task owns several input partitions, which is
what #2315 made the default.

Summary

A task that owns several input partitions wrote one data.arrow plus index
per input partition. Every downstream reader fetching partition k
therefore opened one file per input partition, and a stage left M files
behind, where M is the stage's input partition count.

The writer now emits one file per task.

This matters more since ballista.scheduler.max_partitions_per_task defaults
to 0 (#2315): a task now routinely owns several input partitions, so M
files per stage became M/P.

Design

Each input partition still buckets and spills concurrently, and still encodes
its own buckets to IPC bytes on its own task, so the interleave, framing
and compression stay parallel across inputs. The coordinator — which already
awaited every input before responding — then concatenates the finished buffers
into one file and writes one index.

The file is laid out partition-major: the schema header, then output
partition 0's bytes from every input in turn, then partition 1's, and so on.

data.arrow:       [ bucket 0 ][ bucket 1 ][ bucket 2 ]
                    \_ in0,in1 _/
data.arrow.index: { 0 -> off0, 1 -> off1, 2 -> off2 }

Keeping each output partition contiguous is what lets the index stay one
offset per partition, which is why the reader is unchanged:
create_shuffle_path already resolves a sort-shuffle summary to
{stage_id}/{file_id}/data.arrow, and MultiStreamPartitionStream already
crosses concatenated IPC streams inside a byte range.

Spill directories move from {stage}/{file_id}/spill to
{stage}/{task_id}/spill-{input}, so a task owns exactly one directory under
the stage and cleanup no longer strands an empty directory per input partition.

@github-actions github-actions Bot added the documentation Improvements or additions to documentation label Aug 16, 2026
@Dandandan
Dandandan force-pushed the perf/consolidate-shuffle-files-per-task branch from 71269a0 to b6bd788 Compare August 16, 2026 13:53
…nput

A task that owns several input partitions wrote one `data.arrow` plus index
per input partition. Every downstream reader fetching partition k therefore
opened one file per input partition, and a stage left M files behind where M
is the stage's input partition count.

The writer now emits a single file per task. Each input partition still
buckets and spills concurrently, and — importantly — still encodes its own
buckets to IPC bytes on its own task, so the interleave, framing and
compression stay parallel. The coordinator, which already awaited every input
before responding, then concatenates the finished buffers into one file and
writes one index.

The file is laid out partition-major: the schema header, then output
partition 0's bytes from every input in turn, then partition 1's, and so on.
Keeping each output partition contiguous is what lets the index stay one
offset per partition, so the reader is unchanged — `create_shuffle_path`
already resolves a sort-shuffle summary to `{stage_id}/{file_id}/data.arrow`,
and `MultiStreamPartitionStream` already crosses concatenated IPC streams
inside a byte range.

An earlier revision of this change did the encoding in the coordinator, which
collapsed write parallelism from P to 1 and cost 29% on TPC-H SF10. Encoding
per input is what makes the file-count reduction free.

Spill directories move from `{stage}/{file_id}/spill` to
`{stage}/{task_id}/spill-{input}`, so a task owns exactly one directory under
the stage and cleanup no longer strands an empty directory per input.

TPC-H SF10, 2 executors x 4 vcores, `--partitions 16`,
`max_partitions_per_task=0`. Old and new binaries alternate within each round
so drift is shared, 8 runs each:

| | median | sort-shuffle files |
|---|---|---|
| one file per input partition | 21.37 s | 1762 |
| one file per task | 20.05 s | 442 |

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@Dandandan
Dandandan force-pushed the perf/consolidate-shuffle-files-per-task branch from b6bd788 to 52a511f Compare August 17, 2026 08:17
@github-actions github-actions Bot removed the documentation Improvements or additions to documentation label Aug 17, 2026
@Dandandan
Dandandan marked this pull request as ready for review August 18, 2026 05:44
@avantgardnerio

Copy link
Copy Markdown
Contributor

Benchmarked this at SF1000 on EKS. Short version: it's a solid win, and I think the mechanism is more interesting than the headline number.

Setup

TPC-H SF1000 (ZSTD parquet, 32 MiB row groups) on S3, 32 executor pods x 8 vcores = 256 vcores across 4x r6i.24xlarge, 48 GiB memory pool per pod, target_partitions=256, max_partitions_per_task=0 so every task owns 8 input partitions, lz4 shuffle compression, AQE on, 1 iteration.

Baseline is main at this PR's merge-base, so the only difference between the two builds is the two files in this PR. I crossed that with ballista.shuffle.sort_based.memory_limit_per_task_bytes at 0 (what benchmarking.md uses) and at 268435456 (the shipping default), because it turns out that setting interacts with this change.

Seven queries: the four heaviest sort-shuffle writers, plus Q1 and Q6 as controls. Q6 plans ShuffleWriterExec rather than SortShuffleWriterExec, so nothing in this PR can affect it, which makes it a useful noise gauge.

Wall clock, seconds

Config Q6 Q1 Q3 Q10 Q18 Q21 Q9 Total
main, spill off 7.64 12.60 23.50 27.73 46.58 67.25 81.69 267.00
main, spill off (repeat) 7.90 13.07 23.72 29.63 48.76 69.08 80.00 272.16
main, spill 256 MiB 7.83 12.78 26.90 30.95 49.15 69.83 61.00 258.43
this PR, spill off 7.31 12.39 21.08 23.52 46.93 63.97 64.37 239.58
this PR, spill 256 MiB 8.45 13.21 20.91 23.89 48.19 60.26 62.83 237.73

At the shipping default the PR is 8.0% faster overall, with Q3 -22.3%, Q10 -22.8%, Q21 -13.7%.

Spill volumes come out byte-identical between the two builds at each setting (Q9: 416.9 GB over 1790 events either way), which is a nice confirmation that this changes file layout only, not bucketing or spill decisions.

Hypothesis: the win is parallel encoding, not throughput

What the gather + IPC framing + lz4 work has to cross is the end-of-task serial path. There look to be three ways it can avoid that:

  1. Spilling does it incrementally, while the input stream is still arriving.
  2. This PR does it on the 8 per-input tasks, in parallel.
  3. Nothing else does.

That predicts the PR should help most exactly where spilling wasn't already helping, and the numbers line up:

  • Q18 is the one query that spills even at memory_limit_per_task_bytes=0 (35.5 GB, via memory-pool pressure rather than the per-task budget). It's also the one query the PR doesn't move: -1.6%, inside noise.
  • Q9 is where spilling helps main most on its own (81.7s -> 61.0s). It's the other query the PR doesn't improve on (+3.0%, also inside noise).
  • Q3, Q10 and Q21 never spill at 0, and they're where the PR wins double digits.

The part I find most useful is what this does to the spill budget as a tuning knob. On main it's a real trade: turning it on gains 25% on Q9 but costs 15% on Q3 and 12% on Q10. With this PR the setting stops mattering much (totals 239.58 vs 237.73, under 1%). Operators currently have to pick a side; after this they mostly don't. That reads to me as a better argument for the change than the 8%.

Two caveats on reading the above

One iteration per cell. Q6 ranged 7.31-8.45s across the five runs, so the noise floor is around 8%. The double-digit per-query deltas clear that comfortably; the 8.0% total does not, and neither Q9's +3.0% nor Q18's -1.6% should be read as anything. Happy to rerun with 3 iterations if that's worth having.

Also worth flagging because I nearly quoted it: the write_time counter shows a 96% drop here and that figure is meaningless. On main the timer wraps finalize_output, which includes the interleave, IPC framing and compression. Here it wraps write_task_consolidated, which is copies only, since the encode moved into encode_buffered_partitions and isn't timed. On top of that main accumulates the counter once per input partition (8 per task) versus once per task here, so the aggregates aren't comparable either. Timing encode_buffered_partitions under the same counter would keep before/after comparisons meaningful for whoever benchmarks the shuffle writer next.

@avantgardnerio avantgardnerio left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nice work!

I originally tried to encode to a single file when implementing MPTs, but it ran much slower because I had encoding on the serial path. This PR keeps the encoding parallel, regardless of spill, and that's a big win (see perf numbers below).

I gave it a thorough review, and the only nit is that it would be worth revisiting the metrics at some point, but that doesn't have to be in this PR.

@Dandandan
Dandandan merged commit 5f2491e into apache:main Aug 21, 2026
24 checks passed
@Dandandan

Copy link
Copy Markdown
Contributor Author

Nice work!

I originally tried to encode to a single file when implementing MPTs, but it ran much slower because I had encoding on the serial path. This PR keeps the encoding parallel, regardless of spill, and that's a big win (see perf numbers below).

I gave it a thorough review, and the only nit is that it would be worth revisiting the metrics at some point, but that doesn't have to be in this PR.

Thanks for the review! 🚀

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants