Repository navigation
*: honor task collation across DXF encoding and expression paths - #69734
Conversation
|
Skipping CI for Draft Pull Request. |
|
Note Reviews pausedIt looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the Use the following commands to manage reviews:
Use the checkboxes below for quick actions:
📝 WalkthroughWalkthroughThe change moves new-collation selection into expression contexts. It updates expression builtins, table indexes, partition encoding, DDL backfilling, importer plans, and Lightning codecs to use context-specific collation settings. Tests cover legacy and new-collation behavior. ChangesCollation context and expression operations
Table and task integration
Validation
Estimated code review effort: 4 (Complex) | ~60 minutes Sequence Diagram(s)sequenceDiagram
participant Task
participant ExpressionContext
participant TableEncoder
participant IndexOrPartition
Task->>ExpressionContext: configure target collation mode
ExpressionContext->>TableEncoder: expose NewCollationEnabled
TableEncoder->>IndexOrPartition: encode values with configured collator
IndexOrPartition-->>Task: return encoded index or partition result
Possibly related PRs
Suggested labels: Suggested reviewers: Poem
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Warning Review ran into problems🔥 ProblemsGit: Failed to clone repository. Please run the Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
/retest |
| astNodeStack []ast.Node | ||
|
|
||
| planCtx *exprRewriterPlanCtx | ||
| useNewCollate bool |
There was a problem hiding this comment.
Reverted, now this value is obtained from context.
qw4990
left a comment
There was a problem hiding this comment.
LGTM for the optimizer part
| matchEnumSetElementsAsBinary := (ft.GetType() == mysql.TypeEnum || ft.GetType() == mysql.TypeSet) && !ctx.NewCollationEnabled() | ||
| if matchEnumSetElementsAsBinary { | ||
| // Legacy ENUM/SET element matching is binary even when the field metadata has a non-binary collation. | ||
| ft = ft.Clone() |
There was a problem hiding this comment.
This introduce a per datum's clone, which I think we might run into some efficiency issues
There was a problem hiding this comment.
Done in 801e25b
The background here is: some users migrating to premium retain new-collation=false in the user keyspace, while DXF IMPORT INTO and ADD INDEX tasks execute from the SYSTEM keyspace (tidb-worker) with the process-global setting set to true. The temporary FieldType is only needed for that mismatch, so the condition now also requires collate.NewCollationEnabled() to be true. Normal user execution, where the context and process-global setting agree, keeps using the original FieldType.
Besides, FieldType.Clone() is a shallow struct copy, which is inlined and the temporary value does not escape to the heap, so the remaining mismatched DXF ENUM/SET path has no per-datum heap allocation.
There was a problem hiding this comment.
it looks to me that it is the code path that only for some corner cases
| useNewCollate bool | ||
| // collatorPinned prevents an explicit collator from following later | ||
| // collation metadata changes. | ||
| collatorPinned bool |
There was a problem hiding this comment.
it looks to me useNewCollate/ctor can be inconsistant after introduce this collatorPinned, for example, if useNewCollate is set to true, and later setPinnedCollator is called to set a non-new collation collator. I think the current implementation is prone to introducing inconsistency bugs.
There was a problem hiding this comment.
Done in 801e25b . Removed the generic pinned-collator state, so the base builtin now has one invariant: its collator is derived from useNewCollate and its collation metadata. The two exceptional functions now own their behavior locally: ILIKE recomputes its binary counterpart in its concrete SetCharsetAndCollation, while WEIGHT_STRING stores the collator derived from its first argument.
|
/cherry-pick release-202603 |
|
@joechenrh: once the present PR merges, I will cherry-pick it on top of release-202603 in the new PR and assign it to you. DetailsIn response to this:
Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the ti-community-infra/tichi repository. |
|
[APPROVALNOTIFIER] This PR is APPROVED This pull-request has been approved by: D3Hunter, qw4990, windtalker, YangKeao The full list of commands accepted by this bot can be found here. The pull request process is described here DetailsNeeds approval from an approver in each of these files:
Approvers can indicate their approval by writing |
|
/retest |
|
@joechenrh: cannot checkout DetailsIn response to this:
Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the ti-community-infra/tichi repository. |
|
/cherry-pick release-nextgen-202603 |
|
@joechenrh: new pull request created to branch DetailsIn response to this:
Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the ti-community-infra/tichi repository. |
What problem does this PR solve?
Issue Number: close #69563
Problem Summary:
DXF executes tasks in the SYSTEM keyspace. If its new-collation setting differs from the submitting user keyspace, collation-sensitive encoding and expression paths can use the worker setting and produce incompatible data.
What changed and how does it work?
TableandIndexown the encoder used for comparable table and index keys, restored-data decisions, partition routing, and partial-index evaluation.BuildContexttake charges of the expression evaluation, it also owns the mode used to construct collation-sensitive scalar expressions for DDL reorganization and IMPORT generated-column or assignment evaluation.WithCollateAPIs. This keeps one task snapshot consistent across key encoding and expression evaluation.Encoderpropagation from row/value encoding. New collation changes comparable string sort keys, while row values, old-row values, and genericHashCodeserialization use non-comparable encoding and produce identical bytes in either mode. Their original APIs therefore do not need this state.Expression scope: this PR covers scalar expression evaluation used by DXF. Vectorized builtin implementations and the historical
INSTRevaluation remain unchanged.Check List
Tests
The local NextGen cluster used
new_collations_enabled_on_first_bootstrap = falsein the user keyspace andtruein the SYSTEM keyspace. Every case ranADMIN CHECK TABLE, checked index/table results where applicable, performed INSERT/UPDATE/DELETE, and ranADMIN CHECK TABLEagain.ADD INDEX
PRIMARY KEY(id) CLUSTERED,id/fk VARCHARALTER TABLE t ADD INDEX idx_fk(fk)PRIMARY KEY(id1,id2) CLUSTERED,fk INTALTER TABLE t ADD INDEX idx_fk(fk)LOWER(raw),UPPER(raw),CONCAT(id,':',raw),SUBSTR(raw,1,2)generated columnsid/raw VARCHAR, clustered VARCHAR PKLOWER,UPPER,CONCAT, andSUBSTRid VARCHAR COLLATE utf8mb4_general_ci,PARTITION BY LIST COLUMNS(id)ALTER TABLE t ADD INDEX idx_fk(fk)id VARCHAR COLLATE utf8mb4_general_ci,PARTITION BY KEY(id) PARTITIONS 4ALTER TABLE t ADD INDEX idx_fk(fk)id VARCHAR COLLATE utf8mb4_general_ci,PARTITION BY RANGE COLUMNS(id)ALTER TABLE t ADD INDEX idx_fk(fk)raw VARCHAR COLLATE utf8mb4_general_ciALTER TABLE t ADD INDEX idx_partial(fk) WHERE raw='A'=,IN,LIKE,IF,CASE,STRCMP,LOCATE, andGREATESTADMIN CHECK TABLEENUM('A','a','B'),SET('A','a','B')withutf8mb4_general_ciIMPORT INTO
The table omits storage URLs; each operation is
IMPORT INTO ... FROM <CSV>.PRIMARY KEY(id) CLUSTERED,KEY(fk)IMPORT INTO t(@1,id,fk)PRIMARY KEY(id) CLUSTERED,fk INT,KEY(fk)IMPORT INTO t(fk,id,@3)PRIMARY KEY(id1,id2) CLUSTERED,KEY(fk)IMPORT INTO t(id2,fk,id1)PRIMARY KEY(id1,id2) CLUSTERED,fk VARCHAR,KEY(fk)IMPORT INTO t(id1,id2,fk)PRIMARY KEY(id1,id2) CLUSTERED,id1/id2 CHAR,KEY(fk)IMPORT INTO t(fk,id1,id2)KEY(fk(2))IMPORT INTO t(@1,id,fk)IMPORT INTO t(@1,id,fk)LOWER,UPPER,CONCAT, andSUBSTR, all indexedIMPORT INTO t(@1,id,raw)LOWER,UPPER,CONCAT, andSUBSTRresultsIMPORT INTO t(@1,@2,@3) SET ...IMPORT INTO t(id,fk,payload)PARTITION BY LIST COLUMNS(id)IMPORT INTO t(id,fk)PARTITION BY RANGE COLUMNS(id)IMPORT INTO t(id,fk)PARTITION BY KEY(id) PARTITIONS 4IMPORT INTO t(id,fk)KEY idx_partial(fk) WHERE raw='A'IMPORT INTO t(@id,@raw) SET fk=CONCAT('v',@id)=,IN,LIKE,IF,CASE,STRCMP,LOCATE,GREATESTIMPORT INTO t(id,raw)=,IN,LIKE,IF,CASE,STRCMP,LOCATE,GREATESTIMPORT INTO t(@id,@raw) SET ...ENUM('A','a','B'),SET('A','a','B'), both indexedIMPORT INTO t(id,e,s)<=>,!=,<,>=,ILIKE,REGEXP,FIELD,LEAST,WEIGHT_STRINGIMPORT INTO t(@id,@raw) SET ...Latest upstream and this PR were tested with the same cluster and input files:
Side effects
Documentation
Release note
Please refer to Release Notes Language Style Guide to write a quality release note.
Summary by CodeRabbit
Bug Fixes
Tests