The Flink fairness review found documentation and validation rules that no longer describe the measured configuration. This issue tracks their correction; it does not claim a measured performance improvement.
Evidence at benchmark commit 9b4cab3:
entrants/flink/README.md calculates 800,000 buffered/in-flight rows. The descriptor now uses 32 subtasks, 100,000 buffered rows and two 50,000-row requests: 3,200,000 buffered plus up to 3,200,000 in flight.
- The README says task off-heap memory is 128 MiB;
config.yaml sets 512 MiB.
- The approximately 49-second drain and Docker Desktop explanation describe older experiments. The archived Linux/Graviton4 runs use 40 million messages and take approximately 988 seconds for Flink.
- The statements that POJO fields never use Kryo and that no serialization occurs need qualification. Chaining/object reuse avoid inter-operator copies; Avro decoding, ClickHouse encoding and checkpoint serialization remain. Nested field type extraction must be verified independently of the outer POJO type.
- Equal parallelism is only one chaining condition. Explicitly setting equal per-operator parallelism does not itself break chaining. Fewer consumers than partitions does not imply a universal half-rate penalty.
- The claims that larger heaps necessarily worsen GC and that the live set does not grow need run provenance. Current records show approximately 17.5 GiB post-GC occupancy against a similarly sized heap.
flink_jvm_sizing_fits_its_declared_container enforces historical tuning choices as a minimum and a compressed-reference boundary as a maximum. These are not fairness requirements; legitimate smaller heaps and larger heaps must remain testable within the resource envelope.
- Spate's README still describes eight-shard aggregate concurrency and approximately 99.7% batch-cap fill. Its current RowBinary descriptor uses 32 shards, and current actual rows per INSERT need to replace historical batch-fill claims.
Acceptance:
- Correct configuration-derived statements and distinguish configured capacity from observed occupancy/batch size.
- Link empirical claims to exact run records, hardware, corpus and protocol; remove unsupported generalizations.
- Explain the difference between Flink POJO typing and the ClickHouse connector's typed mode.
- Document the actual checkpoint and connector serialization paths, effective chain and resolved field serializers.
- Replace historical heap restrictions with resource-envelope validation.
- Link the completed review and configuration changes here. Keep tuning records separate from published results.
This review is requested by the benchmark maintainer. All candidate changes preserve the contract's public-API requirement and five-second at-least-once durability.
The Flink fairness review found documentation and validation rules that no longer describe the measured configuration. This issue tracks their correction; it does not claim a measured performance improvement.
Evidence at benchmark commit
9b4cab3:entrants/flink/README.mdcalculates 800,000 buffered/in-flight rows. The descriptor now uses 32 subtasks, 100,000 buffered rows and two 50,000-row requests: 3,200,000 buffered plus up to 3,200,000 in flight.config.yamlsets 512 MiB.flink_jvm_sizing_fits_its_declared_containerenforces historical tuning choices as a minimum and a compressed-reference boundary as a maximum. These are not fairness requirements; legitimate smaller heaps and larger heaps must remain testable within the resource envelope.Acceptance:
This review is requested by the benchmark maintainer. All candidate changes preserve the contract's public-API requirement and five-second at-least-once durability.