Skip to content

Read a zip archive as one workbook (csv and xlsx entries as sheets) #19

Description

@vaceslav

Read a zip archive as if it were one workbook: every readable entry contributes its sheets to one flat list, and each sheet carries its Source (the entry's path) so a UI can group sheets by it.

Use cases:

  1. One large csv, zipped to shrink the upload (a 572 MB export is 60–80 MB zipped), read straight from the archive without unpacking it to disk.
  2. Several csv files of one export in one archive.
  3. A mixed archive with workbooks and csv files side by side. Rare, but it comes almost for free with 1 and 2.

Design: docs/superpowers/specs/2026-09-25-archive-as-workbook-design.md, part 2.

Prerequisite (part 1, before v0.1, done):

  • per-sheet Format / Source / Dialect / Diagnostics;
  • MappingPlan.SheetName / SheetSource with structure.sheet-changed.

Scope of this issue (additive, after v0.1):

  • Detection by bytes:
    • [Content_Types].xml means a workbook;
    • an OpenDocument spreadsheet is TabularFormat.Ods (done in Read OpenDocument spreadsheets (.ods) #13); another OpenDocument type stays format.unsupported;
    • anything else is TabularFormat.Zip.
  • Entries:
    • Skipped silently: directories, __MACOSX/, hidden files.
    • Every other entry is sniffed: an xlsx or ods workbook contributes its sheets, text becomes a csv sheet.
    • Nested archives, other OpenDocument types, OLE2, binary, XML documents and encrypted entries are skipped, each with a reason in FileProfile.SkippedEntries.
  • Order and names: sheets are ordered by entry path, ordinally. A csv sheet is named after its file without the extension. Source + Name identify a sheet.
  • Csv entries: read as streams. The dialect probe is buffered and rejoined to the front, because an entry stream cannot seek.
  • Workbook entries: buffered in memory under MaxEmbeddedWorkbookBytes (default 256 MB, configurable past 2 GB through a chunked buffer), at most one at a time.
  • Options: ArchiveCursorOptions (Csv, Xlsx, Ods, MaxEntries, MaxUncompressedBytes, MaxEmbeddedWorkbookBytes); TabularOpenOptions.Archive.
  • Tests: built from raw bytes. They cover a zip bomb, junk entries, a workbook renamed to .zip, sheet order that does not depend on entry order, and moving back to an earlier csv sheet.
  • Performance: the 5M-row csv zipped vs. unpacked (time, peak memory).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions