Skip to content

Phase 5: Add distributed operator contracts and A0-A5 profiles #31

Description

@inaniloquentee

Background

Split from #6 Phase 5: Distributed Operator Contracts.

This issue owns distributed operator contracts and cross-configuration drift attribution. It starts only after the corresponding single-card operator comparison is available, so distributed drift can be separated from ordinary single-card arithmetic drift.

Scope

  • Add distributed contract tests and sweeps only for operators whose corresponding single-card Phase 4 comparisons pass or have a documented reference path.
  • Add TP matmul projection tests, including the cases where sharded partial sums require an explicit reduction contract.
  • Add SP RMSNorm boundary tests, including activation shard placement and reduction/downcast semantics.
  • Add CP attention tests for cache/position metadata and LSE-merge semantics.
  • Generalize the existing linear_logp TP test surface into a cross-configuration drift harness rather than a one-off TP check.
  • Add distributed batch-invariance sweeps across TP/SP/CP degree, microbatch partitioning, rank placement, active-token density, and packed-batch layouts.
  • Add A0-A5 named profiles over the same result path:
    • A0 fully aligned reference
    • A1 arithmetic-only mismatch
    • A2 reduction/topology-only mismatch
    • A3 representation-only mismatch
    • A4 pairwise mismatches
    • A5 production mismatch
  • Add or formalize contract-preserving distributed abstractions where strict consistency needs them, such as explicit placement, sharding, reduction, and fixed-order reduce helpers.
  • Record requested-vs-actual distributed provenance, including shard layout, rank placement, collective backend, custom all-reduce status, accumulation dtype, reduction order, and fallback behavior.

Non-Goals

  • Do not add missing single-card second/reference implementations; that belongs to Phase 4.
  • Do not duplicate the Phase 3 audit/replay harness or result-cube implementation.
  • Do not treat general profiling and fast-path expansion as part of this issue; that belongs to later roadmap work.
  • Do not claim final Blackwell/B-card operator replacement from this issue alone. [Experiment] Build Blackwell train-inference mismatch baseline for VIME post-training #24 can consume the attribution results for baseline and candidate A/B experiments, but hardware-specific CUDA kernel fixes should be scoped separately once the drifting boundary is identified.
  • Do not silently fall back to full-replica or native distributed paths in strict cases without recording that fallback as the declared axis under test.

Acceptance Criteria

  • Distributed tests run only after the corresponding single-card comparison passes or a strict reference path is documented.
  • TP/SP/CP drift can be attributed to sharding, reduction order, dtype, custom all-reduce, metadata mismatch, or fallback behavior.
  • Strict distributed runs keep same-sample dlogp within tolerance across the supported batch-layout and distributed-layout variants.
  • The report names the smallest A0-A5 alignment set that passes instead of requiring full reference execution by default.
  • Strict consistency has a documented distributed reference path.
  • Requested-vs-actual distributed provenance is visible enough to reject undeclared fallback in strict runs.

Dependencies

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    consistencyRollout-training consistencyphase-5Phase 5 workstreamrl-kernelRL-Kernel integrationtelemetryStructured telemetry and observability

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions