Skip to content

feat(modelfile): support TF SavedModel, ONNX external_data, and online-learning trees - #530

Merged
bergwolf merged 9 commits into
mainfrom
worktree-for-small-model
Oct 10, 2026
Merged

bergwolf merged 9 commits into
mainfrom
worktree-for-small-model

Conversation

@aftersnow

Copy link
Copy Markdown
Contributor

Summary

  • Add literal-name patterns feature_map and checkpoint to ModelFilePatterns so TF SavedModel artifacts no longer have these extension-less files misclassified as CODE.
  • New pkg/modelfile/onnx.go: parses ONNX protobuf with google.golang.org/protobuf/encoding/protowire to extract every external_data.location tensor path. Generator post-processes these and reclassifies them as MODEL deterministically — independent of size.
  • Online-learning directories (parent dir with base/<ts>/ + delta/<ts>/ subtrees, each holding a TF SavedModel layout) work via existing recursive walk; new tests use the real timestamp directory names from a production tree.

No changes to build / push / pull / attach / upload / codec / storage code paths. Those layers are byte-level and remain model-type-agnostic.

Why

Today InferFileType falls back to a 128MB size heuristic for extension-less files, which misclassifies:

  • feature_map, variables/checkpoint, etc. (small TF literals) → CODE
  • ONNX external tensor files with arbitrary names like tower_deep_layer_0_kernel_read__448_1 → CODE when small / MODEL when large (inconsistent)

The fix is local to the generator — once the Modelfile globs are correct, the rest of the pipeline already works because layers are encoded by media type chosen at build time and reconstructed via modelspec.AnnotationFilepath on pull.

Test plan

  • go test ./pkg/modelfile/... — all new and existing tests pass
  • go test ./... — no regression in any downstream package
  • go vet ./... — clean
  • New scenario tests use real filenames from production directory listings:
    • TestNewModelfileByWorkspace/tensorflow_saved_model_directory
    • TestNewModelfileByWorkspace/online-learning_base_plus_delta_tree
    • TestNewModelfileByWorkspace_ONNXExternalData (12 real ONNX external_data filenames)
  • TestExtractONNXExternalDataPaths_* covers empty, single, multiple, dedup, missing-file, and corrupted-bytes cases

Out of scope

  • DFS / remote source provider (user confirmed all input is local)
  • base/delta-aware incremental push or partial pull
  • New media types or a high-level "model type" field in the manifest
  • config.json handling — TF/ONNX directories don't have one and the existing skip logic is correct; users wanting metadata in the manifest can pass --family/--arch/--format to modctl modelfile generate.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 4f303daf59

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread pkg/modelfile/modelfile.go Outdated
Comment thread pkg/modelfile/modelfile.go

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces logic to correctly classify model-related files for TensorFlow SavedModels and ONNX models. It adds literal name matching for specific TensorFlow files and implements a custom ONNX parser to identify external data files referenced within .onnx files, ensuring they are categorized as model weights regardless of size or extension. A security concern was identified regarding potential path traversal vulnerabilities when processing external data paths extracted from ONNX files, and a validation check was suggested to ensure paths do not escape the workspace.

Comment thread pkg/modelfile/modelfile.go Outdated
aftersnow added a commit that referenced this pull request Apr 28, 2026
Three follow-up fixes from independent review of PR #530:

P1 (correctness/security): reclassifyONNXExternalData now only moves paths
already collected by the workspace walker, naturally inheriting all of the
walker's filtering: ExcludePatterns, isSkippable directories, file count /
size limits, and the workspace boundary. Adversarial or malformed ONNX
referencing `../escape` is silently dropped instead of being added to
mf.model.

P2 (coverage): the ONNX parser now recurses into NodeProto.attribute so
external_data references attached via Constant ops (attribute.t / tensors /
sparse_tensor / sparse_tensors) and inside If / Loop / Scan subgraphs
(attribute.g / graphs) are discovered. Subgraph recursion is bounded at 32
levels to defend against pathological inputs.

P3 (visibility): ONNX parse failures now print a clearly prefixed WARNING
line that names the offending file and explains the fallback ("external
tensor files will keep walker-assigned classification"). Generate does NOT
abort -- a single bad .onnx degrades to the pre-fix walker classification
without killing the whole pass.

New tests: NodeAttributeTensor, Subgraph (P2 unit); PathTraversalIgnored,
RespectsExclude, ParseFailureFallsBack (P1/P3 integration). go vet, go test
-race ./pkg/modelfile/..., and go test ./... all clean.

Signed-off-by: Zhao Chen <winters.zc@antgroup.com>

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 660ae9b82f

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread pkg/modelfile/onnx.go Outdated
aftersnow added a commit that referenced this pull request Apr 28, 2026
Codex review on PR #530 noted a remaining coverage gap: ExtractONNXExternal-
DataPaths only walked ModelProto.graph (field 7), missing external_data
references in training-enabled ONNX (IR v7+) where weights/states can also
appear in ModelProto.training_info[*].initialization (field 1) and
.algorithm (field 2) GraphProtos.

Refactor ExtractONNXExternalDataPaths to iterate ModelProto fields directly
via forEachField rather than the now-removed readSubMessage helper. Both
graph and training_info subtrees route into the same walkGraph entry, so
all the existing recursion (initializer / sparse_initializer / NodeProto
attribute / subgraph) applies uniformly to inference and training graphs.

Tests: TrainingInfoGraphs covers a training-only model;
InferenceAndTrainingCombined verifies both surfaces are merged.

Signed-off-by: Zhao Chen <winters.zc@antgroup.com>
aftersnow added a commit that referenced this pull request Apr 28, 2026
…ING text, fix doc

Three follow-up findings from a second codex review pass on PR #530:

1. Reject absolute external_data.location values: filepath.Join(onnxDir, /decoy)
   silently strips the leading separator and produces a relative decoy path,
   which would reclassify an unrelated workspace file. Add filepath.IsAbs(ext)
   guard before Join.

2. Lock the WARNING contract: ParseFailureFallsBack now captures os.Stderr and
   asserts the warning prefix, file path, and fallback note are printed.

3. Update ExtractONNXExternalDataPaths godoc to enumerate the real error
   sources: I/O, size cap, and malformed protobuf wire data.

go vet, go test -race ./pkg/modelfile/..., and go test ./... all clean.

Signed-off-by: Zhao Chen <winters.zc@antgroup.com>
@aftersnow

Copy link
Copy Markdown
Contributor Author

@codex review

Adds best-effort Config.Format inference to fill the previously-empty manifest field. Conservative signals only:

  • saved_model.pb / saved_model.pbtxt → tensorflow
  • *.onnx → onnx
  • *.gguf → gguf
  • *.safetensors → safetensors

Generic .bin/.pt/.pth deliberately skipped (false-positive risk).

Contracts:

  • CLI --format always wins over inference
  • Failure (no signal / panic on malformed hashset value) leaves mf.format empty with a stderr WARNING: and never aborts generation
  • Pure metadata path; build/push/pull/codec untouched

New commit: 876d157. Tests: 4 dedicated cases (CLI flag wins, no-signal stays empty, priority order saved_model.pb > .safetensors, gguf), plus updated 4 existing table cases. go vet ./... and full go test ./... pass.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 876d1571ab

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread pkg/modelfile/modelfile.go Outdated
@aftersnow

Copy link
Copy Markdown
Contributor Author

@codex review

Addresses the P2 saved_model.pbtxt false-negative finding from the previous review (commit 4c24aa3). inferFormat now scans all four walker buckets (model+config+code+doc) so signals not in ModelFilePatterns (saved_model.pbtxt) are still detected without disturbing the walker's classification. New regression: TestNewModelfileByWorkspace_InferFormatSavedModelPbtxt.

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. 🚀

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@aftersnow

Copy link
Copy Markdown
Contributor Author

E2E verification

Cross-compiled the binary from commit 4c24aa3 to a Linux x86_64 host and ran modelfile generate → build → inspect against three synthetic fixtures (and a round-trip extract on the most complex one).

Fixture 1 — TF SavedModel
Files: saved_model.pb, feature_map, tf_signature.txt, alps.meta, ais-sync.done, variables/{checkpoint, variables.index, variables.data-N-of-M}, sibling checkpoint/model.ckpt-NNN.{data-*, index, meta}.

  • Modelfile: FORMAT tensorflow
  • Manifest config.format: tensorflow
  • Layer mediaType: feature_map, variables/checkpoint, checkpoint/checkpoint → weight.v1.raw (pre-fix these would have been code.v1.raw)
  • *.meta → weight.config.v1.raw, tf_signature.txt → doc.v1.raw

Fixture 2 — ONNX + external_data
Files: synthesized model.onnx (protowire-built ModelProto with 3 external_data.location entries) plus 3 zero-byte extension-less tensor files.

  • Modelfile: FORMAT onnx
  • Manifest config.format: onnx
  • All 3 external tensor files (weights_a, weights_b, external_data_for_resource_handle) → weight.v1.raw (pre-fix all 3 would have been code.v1.raw by the size heuristic)

Fixture 3 — online-learning base + delta tree
Files: base/<ts>/... + delta/<ts>/..., each holding a full SavedModel layout.

  • Modelfile: FORMAT tensorflow
  • Manifest config.format: tensorflow (set-based scan deduplicates multiple saved_model.pb instances)
  • Round-trip: build → extract → diff -rq vs the original tree shows only Modelfile differs (expected — Modelfile is build context, not part of the artifact); the 12 model files are byte-equal.

The three behavior changes shipped in this PR are confirmed end-to-end (not only in unit tests):

  1. feature_map / literal checkpoint are lifted to MODEL via ModelFilePatterns.
  2. ONNX external_data references promote extension-less tensor files to MODEL regardless of size.
  3. inferFormat writes Format into the manifest config blob (previously empty for these layouts).

@aftersnow
aftersnow enabled auto-merge (squash) April 30, 2026 04:08
…e-learning trees

Generator now correctly classifies extension-less TF SavedModel literals
(feature_map, checkpoint) as MODEL via new ModelFilePatterns entries, and
deterministically reclassifies all ONNX external_data tensor files as MODEL by
parsing the .onnx protobuf with google.golang.org/protobuf/encoding/protowire.

This removes the size-heuristic dependency that previously misclassified small
external tensor files as CODE. Online-learning directories (a parent dir
containing base/<ts> and delta/<ts> subtrees) are handled by existing recursive
walk; new test cases mirror the real workflow_10062365 layout.

No changes to build/push/pull/attach/upload/codec/storage paths --- those remain
byte-level and model-type-agnostic.

Signed-off-by: Zhao Chen <winters.zc@antgroup.com>
Three follow-up fixes from independent review of PR #530:

P1 (correctness/security): reclassifyONNXExternalData now only moves paths
already collected by the workspace walker, naturally inheriting all of the
walker's filtering: ExcludePatterns, isSkippable directories, file count /
size limits, and the workspace boundary. Adversarial or malformed ONNX
referencing `../escape` is silently dropped instead of being added to
mf.model.

P2 (coverage): the ONNX parser now recurses into NodeProto.attribute so
external_data references attached via Constant ops (attribute.t / tensors /
sparse_tensor / sparse_tensors) and inside If / Loop / Scan subgraphs
(attribute.g / graphs) are discovered. Subgraph recursion is bounded at 32
levels to defend against pathological inputs.

P3 (visibility): ONNX parse failures now print a clearly prefixed WARNING
line that names the offending file and explains the fallback ("external
tensor files will keep walker-assigned classification"). Generate does NOT
abort -- a single bad .onnx degrades to the pre-fix walker classification
without killing the whole pass.

New tests: NodeAttributeTensor, Subgraph (P2 unit); PathTraversalIgnored,
RespectsExclude, ParseFailureFallsBack (P1/P3 integration). go vet, go test
-race ./pkg/modelfile/..., and go test ./... all clean.

Signed-off-by: Zhao Chen <winters.zc@antgroup.com>
Codex review on PR #530 noted a remaining coverage gap: ExtractONNXExternal-
DataPaths only walked ModelProto.graph (field 7), missing external_data
references in training-enabled ONNX (IR v7+) where weights/states can also
appear in ModelProto.training_info[*].initialization (field 1) and
.algorithm (field 2) GraphProtos.

Refactor ExtractONNXExternalDataPaths to iterate ModelProto fields directly
via forEachField rather than the now-removed readSubMessage helper. Both
graph and training_info subtrees route into the same walkGraph entry, so
all the existing recursion (initializer / sparse_initializer / NodeProto
attribute / subgraph) applies uniformly to inference and training graphs.

Tests: TrainingInfoGraphs covers a training-only model;
InferenceAndTrainingCombined verifies both surfaces are merged.

Signed-off-by: Zhao Chen <winters.zc@antgroup.com>
…ING text, fix doc

Three follow-up findings from a second codex review pass on PR #530:

1. Reject absolute external_data.location values: filepath.Join(onnxDir, /decoy)
   silently strips the leading separator and produces a relative decoy path,
   which would reclassify an unrelated workspace file. Add filepath.IsAbs(ext)
   guard before Join.

2. Lock the WARNING contract: ParseFailureFallsBack now captures os.Stderr and
   asserts the warning prefix, file path, and fallback note are printed.

3. Update ExtractONNXExternalDataPaths godoc to enumerate the real error
   sources: I/O, size cap, and malformed protobuf wire data.

go vet, go test -race ./pkg/modelfile/..., and go test ./... all clean.

Signed-off-by: Zhao Chen <winters.zc@antgroup.com>
Best-effort fill of mf.format when --format is not passed. Inference
keys off uniquely diagnostic signals only:

  saved_model.pb / saved_model.pbtxt -> tensorflow
  *.onnx                              -> onnx
  *.gguf                              -> gguf
  *.safetensors                       -> safetensors

Generic .bin/.pt/.pth are intentionally NOT signals: they appear in
many formats and would produce false positives on existing repos.

Failure (no signal, panic on a malformed hashset value, etc.) leaves
mf.format empty and degrades silently with a stderr WARNING. CLI
--format always wins over inference. Format remains best-effort
metadata; downstream build/push/pull paths already handle it being
blank, so this never blocks generation.

Signed-off-by: Zhao Chen <winters.zc@antgroup.com>
Codex review (P2) caught that saved_model.pbtxt is not in
ModelFilePatterns -- the walker lands it in CODE/DOC, so the prior
inferFormat (which only scanned mf.model) silently missed the textual
SavedModel variant. A workspace whose only TF signal was a .pbtxt
would yield format="" instead of "tensorflow".

Fix: extract the per-file scan into a closure and apply it to all four
walker buckets (model + config + code + doc). Set-based: duplicates
across buckets are harmless. Walker classification is unchanged.

Adds TestNewModelfileByWorkspace_InferFormatSavedModelPbtxt as
regression coverage and updates the doc comment to spell out the
multi-bucket scan and the rationale.

Signed-off-by: Zhao Chen <winters.zc@antgroup.com>
@aftersnow
aftersnow force-pushed the worktree-for-small-model branch from 4c24aa3 to 5c87b50 Compare September 7, 2026 14:50
bergwolf
bergwolf previously approved these changes Oct 10, 2026

@bergwolf bergwolf left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm!

@bergwolf

Copy link
Copy Markdown
Member

@copilot resolve the merge conflicts in this pull request

# Conflicts:
#	go.mod

Co-authored-by: bergwolf <88972+bergwolf@users.noreply.github.com>
auto-merge was automatically disabled October 10, 2026 14:18

Head branch was pushed to by a user without write access

Co-authored-by: bergwolf <88972+bergwolf@users.noreply.github.com>

Copilot AI commented Oct 10, 2026

Copy link
Copy Markdown
Contributor

@copilot resolve the merge conflicts in this pull request

Resolved by merging main and reconciling go.mod; merge commit 9909b70. I also updated gRPC to patched v1.83.2 in 95238ee.

Copilot AI requested a review from bergwolf October 10, 2026 14:20
The merge of main in 9909b70 brought in the recordingStrategy helper from
#592, which uses io.Reader and io.Copy, but the io import was not added.
golangci-lint and go vet fail with "undefined: io".

Signed-off-by: Zhao Chen <zhaochen.zju@gmail.com>

@bergwolf bergwolf left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm!

@bergwolf
bergwolf merged commit f241cec into main Oct 10, 2026
5 checks passed
@bergwolf
bergwolf deleted the worktree-for-small-model branch October 10, 2026 14:50
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants