Coerce ClassLabel.names to plain str to fix invalid YAML in README.md - #8454
RudrenduPaul wants to merge 1 commit into
Conversation
numpy.str_ (and similar numpy-backed) entries slip into ClassLabel.names when built from numpy.unique(...) or a pandas column. PyYAML's default Dumper only matches representers by exact type, so these fall back to unsafe !!python/object/apply tags when the dataset card metadata is dumped, producing a README.md that fails to parse back with yaml.safe_load. Fixes huggingface#6919
shashvat-singham
left a comment
There was a problem hiding this comment.
I reproduced the motivating problem on main and it is worse than the PR description suggests — worth pasting, because the failure is invisible until someone tries to read the card back:
f = Features({"label": ClassLabel(names=list(np.array(["negative", "positive"])))})
yaml.dump(f._to_yaml_list()) '0': !!python/object/apply:numpy._core.multiarray.scalar
- !!python/object/apply:numpy.dtype
args: [U8, false, true]
state: !!python/tuple [3, <, null, null, null, 32, 4, 8]
- !!binary |
cAAAAG8AAABzAAAAaQAAAHQAAABpAAAAdgAAAGUAAAA=safe_load FAILED: ConstructorError could not determine a constructor for the tag
'tag:yaml.org,2002:python/object/apply:numpy._core.multiarray.scalar'
So the label names are emitted as base64-encoded numpy scalars and the resulting README.md metadata block cannot be parsed by yaml.safe_load at all. With this branch it dumps as plain negative / positive and round-trips. The fix is right.
Worth adding to the PR description why this slips through every existing guard: numpy.str_ is a subclass of str, so isinstance(name, str) passes everywhere and nothing upstream flags it. PyYAML dispatches on exact type(), not isinstance, which is the only place the difference shows up. That is the whole bug in one sentence and it is currently not stated anywhere in the PR.
Two behaviour changes the coercion also makes, neither obviously wrong but neither mentioned:
Non-string names are now stringified. ClassLabel(names=[0, 1, 2]) gives [0, 1, 2] on main and ['0', '1', '2'] here. str2int(0) already failed on both (Values 0 should be a string or an Iterable), so the ints were not usable as labels anyway and this arguably repairs them — but int2str(0) now returns '0' where it returned 0 before. If that is intended it is worth a line in the docstring, since ClassLabel does not currently say names must be strings.
Tuples become lists. ClassLabel(names=("a", "b")) keeps a tuple on main and becomes ['a', 'b'] here, because the comprehension rebuilds it. That looks like a genuine improvement — names is typed as list[str] and two ClassLabels built from a tuple and a list now compare equal where they did not before — just noting it is a real change beyond the stated scope.
One small thing on the test: assert type(names[0]) is not str as a sanity check will silently stop testing anything if a future numpy returns real str there. assert type(names[0]) is np.str_ states the precondition you actually depend on.
Tested on Windows 11 / Python 3.11.9 / numpy 2.4.6 / PyYAML, main vs pr-8454.
Coerce ClassLabel.names to plain str to fix invalid YAML in README.md
Fixes #6919
Problem
push_to_hubcan fail with:This happens whenever a
ClassLabel'snameslist contains elements thatare not exactly
str— most commonlynumpy.str_, which shows up whenevernamesis built fromnumpy.unique(...)or read out of a pandas column(a very common pattern for NER/token-classification labels, exactly as
described in the issue).
Root cause
ClassLabel.__post_init__already normalizes the internal_int2strmapping with
str(name), but it left the publicnameslist untouched.When the dataset card YAML metadata is generated
(
Features._to_yaml_list->yaml.dump), PyYAML's defaultDumperselects a representer by exact type match, not
isinstance. Sincenumpy.str_is a subclass ofstrbut notstritself, it fallsthrough to the generic object representer, which serializes it with
unsafe
!!python/object/apply:numpy...tags (including base64-encodedbinary state). The resulting README.md is technically written, but it
can no longer be parsed back with the strict
yaml.safe_loadused tovalidate dataset card metadata, hence the "unknown tag" error.
Fix
Coerce every element of
ClassLabel.namesto a plainstrin__post_init__, mirroring what already happens for_int2str. Thisguarantees only native Python strings ever reach the YAML dumper,
regardless of where
namesoriginally came from.Testing
(
test_class_label_names_are_coerced_to_strintests/features/test_features.py) that builds aClassLabelfromnumpy.str_values, dumps its YAML representation, and asserts itround-trips through
yaml.safe_load— this reproduces the exactfailure from the issue and fails without the fix.
tests/features/test_features.py,tests/test_info.py,and
tests/test_metadata_util.pysuites locally: 226 passed, 3skipped, no regressions.
derived via
numpy.unique, cast toSequence(ClassLabel(names=list(labels))), then dumping the resultingDatasetInfoto YAML) — the dump now round-trips throughyaml.safe_loadcleanly instead of raisingConstructorError.Scope
Only
src/datasets/features/features.py(the fix) andtests/features/test_features.py(the regression test) are touched. Nounrelated code or docs were changed.
Note: this PR was prepared with AI assistance (Claude Code) under my
direction — I reviewed the root-cause analysis, the fix, and the test
before submitting.