Skip to content

ABW candidate cache fills during a long block and every share is refused until the next disclosure #19

Description

@paulscode

A user mining to Convoy with ABW enabled reported that all of his shares had stopped being
submitted. His gateway log shows BLAKE2b anti-withholding candidate cache is full repeating
with no successful submissions, then recovering the moment the pool disclosed a key:

14:44:35 .. 14:46:41   ERROR: BLAKE2b anti-withholding candidate cache is full   (repeating)
14:46:38.872  NEW NETWORK BLOCK: ...973010
14:46:41.709  datum_protocol_abw_reveal: DATUM server retired BLAKE2b assignment slot 0
14:46:42.341  Share accepted: NONCE: 4e54fb67 ...

That accepted share is the first in the capture. The last cache error is at 14:46:41.043 and
the capture runs to 14:47:59, so nothing failed again after the disclosure. The shares were
not bad either: they passed every local check and were refused at the submit step, so the work
was done and then dropped.

I have not reproduced this on a test rig. Everything below is from reading the code against
that log, so I may have it wrong.

mining.abw_verify_all_shares_on_disclosure defaults to true (datum_conf.c:132). With it on,
the drain on an accepted share does not run:

// datum_protocol.c:2052
if (data[0] == DATUM_POW_SHARE_RESPONSE_ACCEPTED &&
    !datum_config.mining_abw_verify_all_shares_on_disclosure)
	datum_protocol_abw_forget_exact(data[10] + 1, data + 11);

Proofs are then held until disclosure, which is what the setting is for.

The limit that seems to bite first is DATUM_ABW_TEMPLATE_CACHE rather than the pending cache.
It is 256 (datum_protocol.c:609), a slot is keyed by source generation (datum_protocol.c:777),
next_template_generation increments on every GBT refresh, and a slot is freed only once the
last pending referencing it clears, which happens at disclosure.

If that reading is right, the ceiling is roughly 256 * work_update_seconds of mining without
a disclosure. At the default 40 that is about two and a half hours. This operator runs 5,
which our packaging sets rather than anything he chose, so about twenty minutes. His capture
has one block arriving 18 seconds after the one before it and another running long enough to
fill the cache, which is the spread this chain has been showing.

One thing I could not settle from the log: whether Convoy answers ACCEPTED or
ACCEPTED_TENTATIVELY for these. The drain above only fires on plain ACCEPTED, so if it is the
tentative code then turning the setting off would not help either, and I would be giving the
operator bad advice. If you know which it is, that saves me guessing at it.

Options I can see:

  • size the template cache from work_update_seconds instead of a fixed 256
  • evict the oldest generation when full instead of refusing the share
  • submit without the retained proof and log that audit coverage was lost for that share

I would argue for the third if it were mine to pick, since refusing costs the miner the whole
share while losing withholding detection on one share costs less. That is a policy call about
what ABW is for though, and it is yours.

I can write whichever you prefer as a PR.

Separately: we added per-reason counters on the local reject path downstream, which is how
this got noticed instead of showing up as an unexplained reject rate. I can send that up as
well if you want it.

Activity

  1. jasonsopko commented on Sep 20, 2026

    @jasonsopko

    Your reading matches the code on master (b9ea7dc). A template slot is matched on the exact source_generation (datum_protocol_abw_template_matches_source), so every generation that receives a share takes a slot, and datum_protocol_abw_pending_clear only frees the slot when the last pending that references it goes. With abw_verify_all_shares_on_disclosure on, nothing clears a pending before the reveal: the share response skips the drain, the candidate receipt only sets pool_handled, and the candidate release does nothing. When datum_protocol_abw_cache_candidate returns false the share is dropped with that error and never reaches datum_queue_add_item, so it is refused after passing every local check, as you saw.

    On the response code, from the gateway side: with the setting off, plain ACCEPTED (0x50) drains on the response and ACCEPTED_TENTATIVELY (0x55) does not, but the candidate receipt (0xA5) and the candidate release both call datum_protocol_abw_forget_exact when the setting is off. So turning it off helps in either case, as long as the pool follows a tentative accept with a receipt or a release. Under the default the code makes no difference at all. Which code Convoy sends I cannot tell you; nothing here mines to it at the moment.

    How often the ceiling is reached on this chain, from a node's last 2000 blocks (971249 to 973248): mean interval 6.4 minutes, median 4.3, longest 50. 100 intervals, five percent, ran past 21.3 minutes, which is the ceiling at work_update_seconds 5; none ran past 171 minutes, the ceiling at the default 40. About 9 of those 215 hours were spent past the 21-minute mark. A slot is only taken when a share lands on that generation, so a small miner fills more slowly than that; anything producing a share every five seconds hits it on schedule. The pending cache is the other bound, 65536 entries, so a farm submitting ten shares a second runs out of that in under two hours, with the same dependence on disclosure.

    Of the three, sizing the template cache from work_update_seconds only moves the ceiling. Submitting without the retained proof and logging it keeps the miner paid and costs one share of audit coverage, which is what I would take too.

  2. paulscode commented on Oct 1, 2026

    @paulscode
    Author

    This looks fixed by 27c16ff ("protocol: Submit shares even if the ABW disclosure buffer has some issue"), merged 26 September. When the candidate can't be retained, datum_protocol_pow_submit no longer returns early. It clears the ABW health latch, logs "Could not retain anti-withholding candidate; non-disclosure detection is compromised", and submits the share. That is the option @jasonsopko and I both preferred: the miner is paid, at the cost of audit coverage for that share.

    Our gateway package has carried 27c16ff since 30 September. Closing unless anyone sees the cache-full refusals again on a build that includes it.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions