You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Read a zip archive as if it were one workbook: every readable entry contributes its sheets to one flat list, and each sheet carries its Source (the entry's path) so a UI can group sheets by it.
Use cases:
One large csv, zipped to shrink the upload (a 572 MB export is 60–80 MB zipped), read straight from the archive without unpacking it to disk.
Several csv files of one export in one archive.
A mixed archive with workbooks and csv files side by side. Rare, but it comes almost for free with 1 and 2.
Design: docs/superpowers/specs/2026-09-25-archive-as-workbook-design.md, part 2.
Prerequisite (part 1, before v0.1, done):
per-sheet Format / Source / Dialect / Diagnostics;
MappingPlan.SheetName / SheetSource with structure.sheet-changed.
Every other entry is sniffed: an xlsx or ods workbook contributes its sheets, text becomes a csv sheet.
Nested archives, other OpenDocument types, OLE2, binary, XML documents and encrypted entries are skipped, each with a reason in FileProfile.SkippedEntries.
Order and names: sheets are ordered by entry path, ordinally. A csv sheet is named after its file without the extension. Source + Name identify a sheet.
Csv entries: read as streams. The dialect probe is buffered and rejoined to the front, because an entry stream cannot seek.
Workbook entries: buffered in memory under MaxEmbeddedWorkbookBytes (default 256 MB, configurable past 2 GB through a chunked buffer), at most one at a time.
Tests: built from raw bytes. They cover a zip bomb, junk entries, a workbook renamed to .zip, sheet order that does not depend on entry order, and moving back to an earlier csv sheet.
Performance: the 5M-row csv zipped vs. unpacked (time, peak memory).
Read a zip archive as if it were one workbook: every readable entry contributes its sheets to one flat list, and each sheet carries its
Source(the entry's path) so a UI can group sheets by it.Use cases:
Design:
docs/superpowers/specs/2026-09-25-archive-as-workbook-design.md, part 2.Prerequisite (part 1, before v0.1, done):
Format/Source/Dialect/Diagnostics;MappingPlan.SheetName/SheetSourcewithstructure.sheet-changed.Scope of this issue (additive, after v0.1):
[Content_Types].xmlmeans a workbook;TabularFormat.Ods(done in Read OpenDocument spreadsheets (.ods) #13); another OpenDocument type staysformat.unsupported;TabularFormat.Zip.__MACOSX/, hidden files.FileProfile.SkippedEntries.Source+Nameidentify a sheet.MaxEmbeddedWorkbookBytes(default 256 MB, configurable past 2 GB through a chunked buffer), at most one at a time.ArchiveCursorOptions(Csv,Xlsx,Ods,MaxEntries,MaxUncompressedBytes,MaxEmbeddedWorkbookBytes);TabularOpenOptions.Archive..zip, sheet order that does not depend on entry order, and moving back to an earlier csv sheet.