CLI
Lead:
Operator:
Inspecting chain state / query node data:
It's not obvious how much of this, if anything should be part of the CLI (alternatively this can be checked using query node playground and a tool for inspecting substrate chain state like polkadot-js/apps, but this will be less convenient):
- checking current distribution policies for dynamic bags (chain / possibly query-node)
- checking current distribution buckets per bag limit (chain / possibly query-node)
- inspecting current buckets metadata (query node)
- inspecting current buckets status (acceptingNewBags / distributing) (query-node / chain)
- inspecing bags in bucket (query-node / chain)
- inspecting data objects in bag (query-node / chain)
- etc.
Protobuf
Query node
Distributor Node
API:
State / cache / storage
Startup, cleanup, data integrity:
Logging
Initial sync
initial sync of joining nodes - use an external source to determine populary of the content in order to initialize data objects cache (Orion? other nodes?)
What if a node already has a full cache and is is tasked with storing a completely new bucket? Should it still try to guess and pre-fetch some potentially popular content then?
Other features to consider
fetching data between distributor nodes
Testing
CLI
Lead:
acceptsNewBagsflag) - shouldn't this be done by worker?distributingflag)Operator:
Inspecting chain state / query node data:
It's not obvious how much of this, if anything should be part of the CLI (alternatively this can be checked using query node playground and a tool for inspecting substrate chain state like
polkadot-js/apps, but this will be less convenient):Protobuf
protobufjs)Query node
storage-v2andcontent-directorymappings oncemasteris merged tostorage-v2branchDataObject.typeDistributor Node
API:
Range)'all'buckets config option - distributing all buckets that are part of runtime worker assignmentHEADrequests on data object endpoint (for the purpose of checking headers without triggering LRU cache state update and potential data object fetching)/bucketsendpoint exposing supported bucketsState / cache / storage
LRU-SPcache policyIt's not yet clear when should it happen - on request? (probably simplest approach). In a separate job set up via
setIntervaland performed every X seconds?dataObjectIdinstead of hashes (?)Startup, cleanup, data integrity:
resuming pending downloads / downloading missing data objectsLogging
Initial sync
initial sync of joining nodes - use an external source to determine populary of the content in order to initialize data objects cache (Orion? other nodes?)What if a node already has a full cache and is is tasked with storing a completely new bucket? Should it still try to guess and pre-fetch some potentially popular content then?
Other features to consider
fetching data between distributor nodesTesting