What kind
A new connector (source, sink, or deserializer)
Which components
S3 source (spate-s3)
What you are trying to do, and what stops you
Part of #749. A backfill from Google Cloud Storage has no source: spate-s3 enables only the object_store aws feature (crates/spate-s3/Cargo.toml:27), so a gs:// URL fails when the store is built.
spate-gcs is a store crate over the backfill engine extracted in #750, with the gcp client built by the shared builder crate from #685. It carries:
GcsSource and GcsSourceConfig, with the same fields as S3SourceConfig and gs:// URLs;
- the
spate_gcs_source_* metric namespace;
- a
gcs feature on the spate facade;
- a docs page under
docs/user-guide/04-connectors/sources/.
Client hardening. spate-s3 builds its S3 client from the environment and applies hardened client defaults (crates/spate-s3/src/source.rs:403). The gcp builder also reads the environment, so the GCS client needs the same review of which variables it honours and which defaults it hardens, with tests.
Rejected credentials. spate-s3 recognises a listing rejected with 401 or 403 by matching the S3 client's error text (crates/spate-s3/src/error.rs:68), because object_store reports it as Generic. How the gcp client reports a rejected listing and a rejected read is unverified. A failing test against a local server pins both before the classification is written.
Split identity. Each store crate passes its own domain tag and fingerprint prefix to the engine, so a GCS job's persisted ids are never confused with an S3 job's.
This pull request adds the ADR for the one-crate-per-store layout, since it is the first time two store crates use the engine.
Blocked by #750 and by #685.
Line numbers are on a58ebf94.
Semver impact, as best you can tell
Additive — nothing existing changes
Does this touch any of these?
The crate registers a new spate_gcs_source_* namespace with the same families as spate_s3_source_*.
Would you want to implement it?
Yes, if the design is agreed
What kind
A new connector (source, sink, or deserializer)
Which components
S3 source (spate-s3)
What you are trying to do, and what stops you
Part of #749. A backfill from Google Cloud Storage has no source:
spate-s3enables only theobject_storeawsfeature (crates/spate-s3/Cargo.toml:27), so ags://URL fails when the store is built.spate-gcsis a store crate over the backfill engine extracted in #750, with thegcpclient built by the shared builder crate from #685. It carries:GcsSourceandGcsSourceConfig, with the same fields asS3SourceConfigandgs://URLs;spate_gcs_source_*metric namespace;gcsfeature on thespatefacade;docs/user-guide/04-connectors/sources/.Client hardening.
spate-s3builds its S3 client from the environment and applies hardened client defaults (crates/spate-s3/src/source.rs:403). Thegcpbuilder also reads the environment, so the GCS client needs the same review of which variables it honours and which defaults it hardens, with tests.Rejected credentials.
spate-s3recognises a listing rejected with 401 or 403 by matching the S3 client's error text (crates/spate-s3/src/error.rs:68), becauseobject_storereports it asGeneric. How thegcpclient reports a rejected listing and a rejected read is unverified. A failing test against a local server pins both before the classification is written.Split identity. Each store crate passes its own domain tag and fingerprint prefix to the engine, so a GCS job's persisted ids are never confused with an S3 job's.
This pull request adds the ADR for the one-crate-per-store layout, since it is the first time two store crates use the engine.
Blocked by #750 and by #685.
Line numbers are on
a58ebf94.Semver impact, as best you can tell
Additive — nothing existing changes
Does this touch any of these?
select!spate-coreThe crate registers a new
spate_gcs_source_*namespace with the same families asspate_s3_source_*.Would you want to implement it?
Yes, if the design is agreed