Repository navigation
ABW candidate cache fills during a long block and every share is refused until the next disclosure #19
Description
Activity
Your reading matches the code on master (b9ea7dc). A template slot is matched on the exact
source_generation(datum_protocol_abw_template_matches_source), so every generation that receives a share takes a slot, anddatum_protocol_abw_pending_clearonly frees the slot when the last pending that references it goes. Withabw_verify_all_shares_on_disclosureon, nothing clears a pending before the reveal: the share response skips the drain, the candidate receipt only setspool_handled, and the candidate release does nothing. Whendatum_protocol_abw_cache_candidatereturns false the share is dropped with that error and never reachesdatum_queue_add_item, so it is refused after passing every local check, as you saw.On the response code, from the gateway side: with the setting off, plain ACCEPTED (0x50) drains on the response and ACCEPTED_TENTATIVELY (0x55) does not, but the candidate receipt (0xA5) and the candidate release both call
datum_protocol_abw_forget_exactwhen the setting is off. So turning it off helps in either case, as long as the pool follows a tentative accept with a receipt or a release. Under the default the code makes no difference at all. Which code Convoy sends I cannot tell you; nothing here mines to it at the moment.How often the ceiling is reached on this chain, from a node's last 2000 blocks (971249 to 973248): mean interval 6.4 minutes, median 4.3, longest 50. 100 intervals, five percent, ran past 21.3 minutes, which is the ceiling at
work_update_seconds5; none ran past 171 minutes, the ceiling at the default 40. About 9 of those 215 hours were spent past the 21-minute mark. A slot is only taken when a share lands on that generation, so a small miner fills more slowly than that; anything producing a share every five seconds hits it on schedule. The pending cache is the other bound, 65536 entries, so a farm submitting ten shares a second runs out of that in under two hours, with the same dependence on disclosure.Of the three, sizing the template cache from
work_update_secondsonly moves the ceiling. Submitting without the retained proof and logging it keeps the miner paid and costs one share of audit coverage, which is what I would take too.This looks fixed by 27c16ff ("protocol: Submit shares even if the ABW disclosure buffer has some issue"), merged 26 September. When the candidate can't be retained,
datum_protocol_pow_submitno longer returns early. It clears the ABW health latch, logs "Could not retain anti-withholding candidate; non-disclosure detection is compromised", and submits the share. That is the option @jasonsopko and I both preferred: the miner is paid, at the cost of audit coverage for that share.Our gateway package has carried 27c16ff since 30 September. Closing unless anyone sees the cache-full refusals again on a build that includes it.
A user mining to Convoy with ABW enabled reported that all of his shares had stopped being
submitted. His gateway log shows
BLAKE2b anti-withholding candidate cache is fullrepeatingwith no successful submissions, then recovering the moment the pool disclosed a key:
That accepted share is the first in the capture. The last cache error is at 14:46:41.043 and
the capture runs to 14:47:59, so nothing failed again after the disclosure. The shares were
not bad either: they passed every local check and were refused at the submit step, so the work
was done and then dropped.
I have not reproduced this on a test rig. Everything below is from reading the code against
that log, so I may have it wrong.
mining.abw_verify_all_shares_on_disclosuredefaults to true (datum_conf.c:132). With it on,the drain on an accepted share does not run:
Proofs are then held until disclosure, which is what the setting is for.
The limit that seems to bite first is DATUM_ABW_TEMPLATE_CACHE rather than the pending cache.
It is 256 (datum_protocol.c:609), a slot is keyed by source generation (datum_protocol.c:777),
next_template_generation increments on every GBT refresh, and a slot is freed only once the
last pending referencing it clears, which happens at disclosure.
If that reading is right, the ceiling is roughly
256 * work_update_secondsof mining withouta disclosure. At the default 40 that is about two and a half hours. This operator runs 5,
which our packaging sets rather than anything he chose, so about twenty minutes. His capture
has one block arriving 18 seconds after the one before it and another running long enough to
fill the cache, which is the spread this chain has been showing.
One thing I could not settle from the log: whether Convoy answers ACCEPTED or
ACCEPTED_TENTATIVELY for these. The drain above only fires on plain ACCEPTED, so if it is the
tentative code then turning the setting off would not help either, and I would be giving the
operator bad advice. If you know which it is, that saves me guessing at it.
Options I can see:
I would argue for the third if it were mine to pick, since refusing costs the miner the whole
share while losing withholding detection on one share costs less. That is a policy call about
what ABW is for though, and it is yours.
I can write whichever you prefer as a PR.
Separately: we added per-reason counters on the local reject path downstream, which is how
this got noticed instead of showing up as an unexplained reject rate. I can send that up as
well if you want it.