Skip to content

Shared Device maxConcurrentClaims is unreachable through the daemon: one claimant per Device row #520

Description

@vicondoa

What we found

While closing the §36 "shared-backend limit" row we proved the ceiling is enforced at the authority layer (d2b-core-controller host_global_hardware_matrix_cannot_be_bypassed_by_zone_or_private_class: two holders admitted, third refused AuthorityCapacityExceeded; the systemd and host-network limits have their own refusal tests), but the maxConcurrentClaims ceiling cannot be reached through the daemon's shared-Device path:

  1. The plumbing is real: packages/d2bd/src/shared_provider_effects.rs reads the spec's /maxConcurrentClaims (default 1), builds GpuAuthorityAdmission with it, and passes max_holders into AuthorityRequest::gpu_from_core, which the daemon admits.
  2. The daemon derives exactly one claimant per Device row (one port per reconciled row) and keys the authority by that row's backing digest (which mixes the row's uid and generation). A repeat reconcile of the same row is AuthorityError::DuplicateActiveReservation - not a ceiling refusal - and two different Device rows never share an authority key, so each row gets its own holder ceiling.
  3. The authority's capacity arm requires ≥2 distinct owner proofs on one key, which the daemon never submits.
  4. The per-claimant machinery that could enforce a shared ceiling - GpuArbitrator / GpuAuthorityIndex::reserve (packages/d2b-provider-device-gpu/src/{arbitration.rs,authority.rs}, refusal GpuClaimError::MaxClaimsExceeded, wire code device-claim-max-exceeded) - has no caller under packages/d2bd.

Net effect: a Device row that declares shared arbitration and a maxConcurrentClaims > 1 has that ceiling applied only to itself, so the declared limit is inert across rows. It is also untestable end to end today: the daemon's admission needs a live controller-session generation and a published Zone/Resource plane, so a hermetic test cannot drive it either.

Why it matters

Either the shared-device arbitration is a real requirement - in which case the daemon path must submit per-claimant proofs to a shared key (or be routed through the provider's existing GpuArbitrator/GpuAuthorityIndex) and the ceiling refusal becomes an end-to-end test - or shared Device access is deliberately per-row and the spec field should say so (and the unused machinery should be retired rather than kept as a decoy).

Suggested resolution

Decide which; both are small. The evidence chain with file:line is in the plan's U15 status block, residual (4): docs/plans/2026-09-09-001-refactor-v3-resource-runtime-rewrite-plan.md.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions