Skip to content

Storage v2 distribution - progress tracking #2543

Description

@Lezek123

CLI

  • General accounts/api setup

Lead:

  • create bucket family
  • delete bucket family
  • create bucket
  • update distribution buckets for bag
  • update distribution bucket status (acceptsNewBags flag) - shouldn't this be done by worker?
  • delete distribution bucket
  • update distribution buckets per bag limit
  • update distribution bucket mode (distributing flag)
  • update families in dynamic bag creation policy (number of buckets per family that should store given dynamic bag)
  • invite distribution bucket operator
  • cancel distribution bucket operator invitation
  • remove distribution bucket operator
  • set bucket family metadata

Operator:

  • accept distribution bucket invitation
  • set distribution operator metadata

Inspecting chain state / query node data:

It's not obvious how much of this, if anything should be part of the CLI (alternatively this can be checked using query node playground and a tool for inspecting substrate chain state like polkadot-js/apps, but this will be less convenient):

  • checking current distribution policies for dynamic bags (chain / possibly query-node)
  • checking current distribution buckets per bag limit (chain / possibly query-node)
  • inspecting current buckets metadata (query node)
  • inspecting current buckets status (acceptingNewBags / distributing) (query-node / chain)
  • inspecing bags in bucket (query-node / chain)
  • inspecting data objects in bag (query-node / chain)
  • etc.

Protobuf

  • Storage bucket metadata format
  • Distributor bucket metadata format
  • Distributor bucket family metadata format
  • Merge with master's protobuf lib (implies switch to protobufjs)

Query node

Distributor Node

API:

  • data object request endpoint
  • download requested data objects from storage node when missing (picking first-response node and creating a queue of alternative endpoints to try on failure)
  • limit number of initial requests made when object is missing (currently "check" requests are made to every storage provider that's supposed to store given object)
  • bandwidth / simultaneous connections management (limits, prioritizing)
  • proxying requests while downloading data objects or streaming from file (depending on Range)
  • controlling buckets served on a given instance via an authorized endpoint
  • allowing 'all' buckets config option - distributing all buckets that are part of runtime worker assignment
  • simple status endpoint
  • support HEAD requests on data object endpoint (for the purpose of checking headers without triggering LRU cache state update and potential data object fetching)
  • /buckets endpoint exposing supported buckets

State / cache / storage

  • persistent object -> file type / extension association
  • LRU-SP cache policy
  • dropping data objects when no longer supported (no longer part of distributed buckets)
    It's not yet clear when should it happen - on request? (probably simplest approach). In a separate job set up via setInterval and performed every X seconds?
  • storing objects in a directory structure that would faciliate faster access
  • simplification: storing objects by dataObjectId instead of hashes (?)

Startup, cleanup, data integrity:

  • recreating object -> file type association on startup when missing
  • waiting for pending downloads to finish on exit
  • resuming pending downloads / downloading missing data objects
  • a tool to run full data integrity check & fix any potential issues (checking data objects hashes etc.)

Logging

  • adjust logging levels
  • support for custom elasticsearch endpoint

Initial sync

  • initial sync of joining nodes - use an external source to determine populary of the content in order to initialize data objects cache (Orion? other nodes?)
    What if a node already has a full cache and is is tasked with storing a completely new bucket? Should it still try to guess and pre-fetch some potentially popular content then?

Other features to consider

  • fetching data between distributor nodes

Testing

  • handling massive spikes of simultaneous requests
  • handling random node shutdowns
  • handling data corruption
  • ...

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions