Repository navigation
Remove the task_map table, folding its length into xcom.mapped_length - #73005
Conversation
2018749 to
ef67497
Compare
c8c2bf1 to
547e2c4
Compare
kaxil
left a comment
There was a problem hiding this comment.
LGTM.
Nothing tests the backfill though, and it's the one spot where an existing DagRun's expansion length either survives the upgrade or quietly disappears. airflow-core/tests/unit/migrations/ already has seven (test_0131 is the preceding 3.4.0 data migration), and test_0117 shows the pattern, load the module by path and run the real statements against an isolated table. Worth seeding a TI with a return_value row plus a second key, where only return_value should pick up the length, and a return_value row with no task_map row, which should stay NULL through the EXISTS.
You'll also want a rebase before this goes in: main picked up 0134_3_4_0_add_dag_draining_state.py with down_revision = "f8c2a1d94e03", the same parent as this one, so post-merge there'd be two alembic heads and db migrate stops. Filenames differ, so git won't flag it as a conflict.
547e2c4 to
bb507b6
Compare
bb507b6 to
0fd38a2
Compare
The task_map table was created (by me and TP) years ago to store the number of results (length) that a task upstream of a mapped task created. This is then used by the scheduler to know how many mapped tasks to expand to. The `keys` column has never been read from (and only got written to in 2.x, but not in any 3.x), it was probably to let people iterate over tasks that returned dictionaries, but well, that turned out to not be required. This is a pre-cursor tidy up to make the PRs for task loops (AIP-111) and after that dynamic sub-graphs (AIP-113) simpler. The wire protocol for both public API and Exec API remain unchanged, this storage on the DB table was purely a server-side storage decision.
0fd38a2 to
0fb08ca
Compare
…removal Main folded the task_map table into xcom.mapped_length (apache#73005). The runtime batch size was read from that table, so the scheduler-side lookup now reads the mapped_length of the size task's return_value row, the same query main uses for a plain task's mapped length, and the tests seed and expand through main's helpers. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The task_map table is gone from main (apache#73005) and an unmapped upstream's length now lives on its map_index -1 XCom row, where a NULL is as unresolved as a missing row used to be.
The task_map table was created just to store the number of results (length) that a task upstream of a mapped task created. This is then used by the scheduler to know how many mapped tasks to expand to.
The
keyscolumn has never been written to, it was probably to let people iterate over tasks that returned dictionaries, but well, that turned out to not be required.This is a pre-cursor tidy up to make the PRs for task loops (AIP-111) and after that dynamic sub-graphs (AIP-113) simpler.
The wire protocol for both public API and Exec API remain unchanged, this storage on the DB table was purely a server-side storage decision.