Skip to content

Bugfix: stratum: Process the current job in newly started stratum threads - #21

Open
alphaminetech wants to merge 3 commits into
CONVOYMining:masterfrom
alphaminetech:new-thread-current-job
Open

alphaminetech wants to merge 3 commits into
CONVOYMining:masterfrom
alphaminetech:new-thread-current-job

Conversation

@alphaminetech

Copy link
Copy Markdown

A stratum thread is started for each new connection until stratum.max_threads threads exist (8 by default), and threads never stop, so this concerns the first stratum.max_threads connections after startup. A new thread records the job that is current when it starts as already seen, so its loop never processes that job:

  • full_coinbase_ready stays false, and every client that subscribes on the thread gets the class 0 coinbase instead of the full one. In pooled mining that is the coinbase that pays only the pool; in solo mining all classes are the same coinbase. It lasts until the thread's next job reaches the client: the next work update, at most bitcoind.work_update_seconds after the current job (or a new block, if one comes first), which is then paced across the thread's clients like any regular update. Only clients of a newly started thread are affected, and only during that thread's first job.
  • If the thread starts in the few milliseconds while empty work for a new block is going out, it counts towards the threads that have to send it but never does. "Empty work send completed" never comes, and the full work for that block waits for the template thread's ~4 s timeout.

Commits:

  1. Bugfix: stratum: Check for the coinbaser before applying its timeout. stratum_job_coinbaser_ready checked the 5 s fallback first, so a thread that first looks at a job more than 5 s after it was created fell back to class 0 even though the coinbaser was there. That is the case for most new threads (a job is older than 5 s for most of its life), and for a coinbaser that lands in the last loop interval (up to ~10 ms) before the 5 s mark. The timeout itself is unchanged: with no coinbaser after 5 s, the thread still goes ahead without it.

    One side effect: when the fetch gets no usable reply, it builds the job's coinbases with 0 outputs (every class is then the pool-only coinbase) and clears need_coinbaser. A thread that first looks after that now sends this coinbase under the miner's class (4) instead of class 0, as master already does when the fetch gives up before the 5 s mark. The coinbase itself is the same.

  2. Bugfix: stratum: Process the current job in newly started stratum threads. Start the thread without a job, so its first loop iteration picks up the current one like any other job update. A thread always runs its loop once before it handles any client command, so the job is processed before a client on it can subscribe. This needs commit 1, since most new threads start on a job that is older than 5 s.

  3. QA: stratum_tests: Check that new stratum threads process the current job. Three tests: stratum_job_coinbaser_ready with a coinbaser that is there after the 5 s mark, a client subscribing on a new thread gets class 4, and a thread started during an empty-work send takes part in it. On master 4 asserts fail; without commit 1, 3 fail; without commit 2, 3 fail.

--test passes in the four CI configurations (API with gcc and with clang, Umbrel define, API off) with ASan+UBSan and leak detection, and with plain -Wall -Werror gcc and clang, on Debian 12 (GCC 12.2, clang 14) and Fedora 44 (GCC 16.2, clang 22).

To reproduce on master: start a gateway in pooled mode and connect a miner once the first job's coinbaser is in. The job id of its first mining.notify ends in 00 (class 0), and only its thread's next job brings the miner's class (04). With this branch the first notify ends in 04.

End to end, against a local mock node and a mock DATUM pool. First mining.notify per client, stratum.max_threads 3, bitcoind.work_update_seconds 15, coinbaser answered at once:

client thread job age at connect master this branch
A new 6.5 s class 0 (class 4 after 9.1 s, with the next job) class 4
B new 7.5 s class 0 (class 4 after 16.1 s: next job, then pacing) class 4
C new 0.6 s class 0 class 4
D existing 2.1 s class 4 class 4

Other runs ("split" = the coinbaser's outputs):

scenario master this branch
new block, then a connection that starts a new thread 0-12 ms later (12 blocks) full work 4.28-4.29 s after the block in 10 of 12 full work 6-16 ms after the block in all 12
coinbaser reply 4.985-5.000 s after the request (35 jobs, client on an existing thread) split used for 8; class 0 although the coinbaser was in before the 5 s mark: 7 split used for 16; such class 0: none
reply 0.05-4.98 s (35 jobs) split used for all 35 split used for all 35
reply 5.0-5.5 s, the fetch gives up (25 jobs) class 0: 25 class 0: 15; class 4 label on the same pool-only coinbase: 10

Related PRs:

🤖 Generated with Claude Code

alphaminetech and others added 3 commits September 30, 2026 23:50
stratum_job_coinbaser_ready applied the 5 second timeout before looking
at whether the job's coinbaser had arrived. A thread that first looks at
a job more than 5 seconds after it was created therefore treated it as
timed out even when the coinbaser was ready, and sent the class 0
coinbase instead of the full one until the next job.

Check the coinbaser first, and only fall back after 5 seconds if it is
still missing.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…eads

Stratum threads are started as clients connect, up to stratum.max_threads.
A new thread recorded the current job as already seen, so its loop never
processed that job:

- full_coinbase_ready stayed false, and every client on the thread was
  sent the class 0 coinbase (in pooled mining it pays only the pool)
  until the next job, up to bitcoind.work_update_seconds later.
- A thread started while empty work for a new block was being sent counted
  towards the threads that have to send it, but never did, so the full work
  for that block waited for the template thread's 4 second timeout.

Start the thread without a job, so the first loop iteration picks up the
current one like any other job update.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… job

Cover stratum_job_coinbaser_ready using a coinbaser that is already there
even after the 5 second fallback, a client subscribing on a newly started
stratum thread getting the current job with the full coinbase (class 4)
rather than the pool-only one (class 0), and a thread started while empty
work is being sent taking part in that send.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant