Skip to content

[arm] the Spate arm lags the framework and will report the wrong wire format #68

Description

@MarcusKainth

Which system

spate

What is configured wrong

The Spate arm is pinned to a published spate-clickhouse 0.2.0 and is behind the framework in two steps.

entrants/spate/src/main.rs:232 builds the sink with from_component_config(...) alone. Since spate-etl/spate#407 that returns a builder, and a runnable sink needs .with_row::<Owned<Row>>(), which derives the insert column list from the row struct rather than from a columns: list in YAML. The arm does not compile against the current framework.

The second step is the one that reaches a published number. spate-etl/spate#413 changes format: rowbinary to send INSERT … FORMAT RowBinaryWithNamesAndTypes: the schema is always fetched and its column names and types open every request body, so the server rejects a body that no longer describes the table. entrants/spate/entrant.toml:121 still reports:

reports  = { wire_format = "rowbinary" }

harness/src/ceiling.rs:549 gates each arm against the ceiling whose format matches that string, and environments/ceilings/c8gd-metal-24xl-ec2-docker.json carries rowbinary and rowbinary_nt as separate entries. Once the arm writes the headed format, it is measured against the ceiling for the other one.

The prose around the arm also predates it. entrant.toml:107-114 and entrants/spate/README.md:94-99 both say the rowbinary variant is the one to read against Flink because the official Flink connector can only write RowBinaryWithNamesAndTypes — written when that was an approximation. It stops being one.

What it should be instead

reports.wire_format = "rowbinary_nt" for the rowbinary variant, and src/main.rs updated for with_row and for ClickHouseEncoder::with_schema(sink.schema()), which is now that encoder's only constructor. The variant id and label can stay: the config value is still format: rowbinary, and only the wire keyword changed.

The two fairness paragraphs get stronger rather than weaker, and should say what is now true — both arms write the same wire format — rather than that one approximates the other.

What you expect that to do to the number, and why

No material change, and the interest is in confirming that rather than in the direction.

Measured on the framework side, six interleaved 1M-row inserts per format against ClickHouse 26.3 on a laptop that was not quiet: 87 ms average (78–96) for RowBinary against 90 ms (84–101) for RowBinaryWithNamesAndTypes, at 73 and 77 ms of server CPU. The ranges overlap. The header is one 78-byte prefix per request body, not per chunk or per row, and at this arm's batch sizes it is far below the noise floor. Client-side encode is untouched: the header is written by the sink's writer, not its encoder.

So the expectation is that rowbinary lands where it lands today, and that the change visible in the results is the reported format name. If re-measurement disagrees by more than run-to-run spread, that is worth knowing and this issue should carry the finding.

The two rules

  • This is configuration, or a switch to a better API the system already ships — not a hand-written replacement for the system's own internals.
  • The system still gives at-least-once delivery with a comparable durability interval. Turning fault tolerance off to go faster is not permitted.

The wire format is not selectable independently of correctness in the framework any more: it fetches the schema and sends the header, and that is the only thing format: rowbinary does now.

Your relationship to this system, if any

I maintain it, and I made the framework change that puts the arm out of date.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions