Summary
Add a multi-source bibliographic lookup client as a plugin: query external library catalogues (Library of Congress, BnF, DNB, K10plus and other SRU/Z39.50 endpoints) for existing records, and feed the results into the metadata pipeline Pinakes already has, in the formats it already ingests (MARCXML / MARC21 / JSON). Inspired by Shelvd's "32 library sources", but built on the interop layer Pinakes already ships.
Why
Today the scraping pipeline (ScrapingService, ScrapeController, BulkEnrichController) resolves metadata mostly from commercial/ISBN sources. For serious cataloguing — and for older or non-commercial titles that ISBN lookups miss — authoritative national-library records via SRU/Z39.50 are the right source. Pinakes already speaks the output side of this world (Z39.50 server, OAI-PMH, MARCXML export); this closes the loop with a client that reads those catalogues.
Scope
- Plugin, opt-in, following the bundled-plugin pattern (
ensureSchema() called from both onActivate() and onInstall(); no doAction()/applyFilters() in onActivate()).
- Query configurable SRU (and where needed Z39.50) endpoints. Ship a small default set (LoC, BnF, DNB, K10plus) and let the admin add more.
- Parse the returned MARCXML/MARC21 into the internal book representation, reusing the existing MARC handling rather than a parallel parser.
- Register as a metadata provider in the existing scraping/enrich flow, so a lookup result flows through
BookDataMerger / the enrich UI exactly like the current providers — not a separate one-off screen. Should compose with the proposed field-by-field enrich mode.
- Per-source toggle, timeout and rate-limit handling; graceful degradation when a source is down (never block the scrape on one dead endpoint).
Alignment
Fits the existing interop roadmap (ISNI/VIAF/NCIP/UNIMARC/MAG) and reuses the MARC/authority code already in tree (viaf-authority, z39-server, oai-pmh-server, bibframe-linked-data). No core version bump implied — it lives as a plugin.
Summary
Add a multi-source bibliographic lookup client as a plugin: query external library catalogues (Library of Congress, BnF, DNB, K10plus and other SRU/Z39.50 endpoints) for existing records, and feed the results into the metadata pipeline Pinakes already has, in the formats it already ingests (MARCXML / MARC21 / JSON). Inspired by Shelvd's "32 library sources", but built on the interop layer Pinakes already ships.
Why
Today the scraping pipeline (
ScrapingService,ScrapeController,BulkEnrichController) resolves metadata mostly from commercial/ISBN sources. For serious cataloguing — and for older or non-commercial titles that ISBN lookups miss — authoritative national-library records via SRU/Z39.50 are the right source. Pinakes already speaks the output side of this world (Z39.50 server, OAI-PMH, MARCXML export); this closes the loop with a client that reads those catalogues.Scope
ensureSchema()called from bothonActivate()andonInstall(); nodoAction()/applyFilters()inonActivate()).BookDataMerger/ the enrich UI exactly like the current providers — not a separate one-off screen. Should compose with the proposed field-by-field enrich mode.Alignment
Fits the existing interop roadmap (ISNI/VIAF/NCIP/UNIMARC/MAG) and reuses the MARC/authority code already in tree (
viaf-authority,z39-server,oai-pmh-server,bibframe-linked-data). No core version bump implied — it lives as a plugin.