Multi-venue crypto order book reconstruction with continuous exchange-checksum verification.
A header-only C++20 library that rebuilds L2 order books from exchange websocket feeds and proves the result is correct against the exchange's own arithmetic, continuously, on live data.
Most order book implementations are tested against themselves. This one is tested against Kraken's CRC32 checksum of the top 10 levels, published on every single update — so "the book is correct" is a measurement with a number attached, not an assertion.
211,733 checksums verified against live Kraken across three instruments and three price scales. Zero mismatches. (how to reproduce)
Crypto venues hand you a correctness oracle for free, and the handlers that do
collect it treat it as optional. cryptofeed
validates across six venues behind a checksum_validation=True flag,
NautilusTrader
does Kraken L3 by default but needs credentials for it, and
ccapi — the closest C++ analog —
implements checksums for OKX and Bitfinex, defaults them off, and has none for
Kraken at all. cryptofeed's own docs explain why they are opt-in: 10 to 100
microseconds per update.
So the claim here is not that verification is novel. It is that verification should not be a flag, and it only stops being one when it costs a few hundred nanoseconds instead of tens of microseconds — which is a consequence of the fixed-point decision below, not of being clever. What each venue hands you:
| Venue | Ground truth | Public? |
|---|---|---|
Kraken L2 book |
CRC32 of the top 10 levels, on every update | Yes, no credentials |
| Binance spot diff-depth | U/u sequence bounds + REST snapshot reconciliation |
Yes, no credentials |
| Binance futures diff-depth | pu continuity field |
Yes, no credentials |
Coinbase full |
per-message sequence for gap detection |
Requires auth |
In equities the equivalent data can't be redistributed, which is why public ITCH
order book repositories generally ship without runnable data and ask to be
believed. Here, crossbook's correctness claim is reproducible by anyone who
clones it — no API key, no paid feed, no sample file to trust. The committed
capture and the match rate it produces are in
The measurement below.
Kraken computes its checksum by taking each level's price and quantity as they
appear on the wire, removing the decimal point, stripping leading zeros, and
concatenating. From their own documented example: price 45285.2 becomes
452852, quantity 0.00100000 becomes 100000.
Both are exactly the decimal digits of the value's fixed-point mantissa at the instrument's scale.
So storing prices as scaled int64 mantissas — rather than double — turns
checksum generation into an integer-to-ASCII with no formatting step and no
rounding to get wrong. A double-based book has to format its way back to
decimal to reproduce the checksum: a step that costs more than the comparison it
enables, and that has to be exactly right at every instrument's scale. The
integer path does not have the step.
Fixed-point here isn't a style preference. The venue's own verification algorithm requires it. That's why the verifier runs on every update in the hot path instead of being sampled.
The identity has one precondition — the venue must spell values canonically at the instrument's scale — and that precondition is documented, tested, and guarded at ingest rather than assumed.
The distinction matters, and calling all five "independent" would blur exactly the point this README opens with. Only the first compares the book against something outside this repository. The other four are consistency checks — they are worth having, they catch real bugs, and they cannot tell you that your reading of the venue's spec was wrong. That is the checksum's job, and it is why the checksum is the one that runs on live data.
- Exchange checksum — the external oracle. Kraken's CRC32, recomputed locally on every update. The only mechanism here that can find a bug in what we believe about Kraken, and it did.
- Differential testing. Two independent book implementations — a
std::mapreference and a tick-indexed array — driven through identical event streams, with full state compared after every single update. Plus a differential fuzzer doing the same with coverage-guided adversarial input. - Sequence continuity. Per-venue gap detection with the correct contract for each. Spot and futures genuinely differ, and a test asserts they disagree on the same event.
- Determinism. Identical input produces a bit-identical state hash on every platform. No floating point anywhere in the book, so there is nothing left that could vary by compiler or CPU. Checked across Linux, macOS, and Windows in CI.
- Recovery. A book that is known to be wrong is never served and never
updated further until it has been rebuilt from a snapshot. The failure that
costs money is not the dropped message — it is applying the next one anyway.
feed.hppexists to make that impossible.
Divergences are never summarised away. Every mismatch is recorded with a cause, because a match rate without an enumerated remainder isn't evidence.
crossbook_verify connects to Kraken, rebuilds the book, and recomputes the
exchange's CRC32 over local state on every update. No API key, no account.
cmake --preset release && cmake --build build/release
./build/release/tools/crossbook_verify --venue kraken --symbol BTC/USD --seconds 180A three-minute run on BTC/USD, 2026-08-01:
frames 2939
applied 2759
checksums verified 2759
checksum mismatches 0
match rate 100.000000% (2759 of 2759)
state hash 7648057f6909c67a
The exit status is the point: any mismatch, any decode failure, any resync and it exits non-zero.
And you can check this without taking my word for it. A recorded minute of
that feed is committed at tests/fixtures/kraken_btcusd_l2.cbcap — 72 KB
of verbatim Kraken bytes — and replays offline, deterministically, on every
platform:
./build/release/tools/crossbook_verify --replay tests/fixtures/kraken_btcusd_l2.cbcap
# 301 of 301 checksums matched, state hash e9613af632f40653CI runs exactly that on Linux, macOS and Windows on every push, and asserts the state hash is bit-identical across all three. That is why the number above is a regression test rather than an anecdote. In equities the equivalent data is licensed and cannot be redistributed, which is why every public ITCH order book repository ships without runnable data and asks to be believed.
Recording your own is one command, and works for Binance too:
./build/release/tools/crossbook_capture --venue kraken --symbol ETH/USD \
--seconds 60 --out eth.cbcap
./build/release/tools/crossbook_verify --replay eth.cbcapcrossbook_capture deliberately does not decode anything. Recording and
interpreting are separate jobs, and keeping them separate is what makes a
capture evidence rather than output: change the book implementation and the
capture is still the bytes the exchange sent, so the new implementation can be
held to them.
Worth stating plainly, because it is the reason the verifier exists.
The first live run reported 98.66% — 4 of 298 updates mismatched — and the book held 20 bid levels for a subscription that asked for 10.
The cause is a gap in the depth-limited contract that unit tests do not reach. Kraken reports cancellations, so a reader that handles those looks correct. It never reports that a level fell out of the top ten because a better level arrived — from the venue's side there is nothing to say. Those orphaned levels sit below the checksummed depth doing no harm, until enough removals near the touch promote one back into view, and then the checksum fails on an update that was itself perfectly fine. The divergence is minutes away from its cause.
The fix is BasicL2Book::trim, and the reason it
is trustworthy is the same reason the bug was found: replaying the committed
capture with trimming disabled still fails, and
a test asserts that it does.
No amount of testing the book against itself would have surfaced this. The exchange's checksum did, in sixty seconds.
Measured, with the methodology stated, because a number without one is noise.
Book update — the path every message takes:
| Spread from touch | std::map |
tick-indexed array | overflow |
|---|---|---|---|
| 8 ticks (tight book) | 58.5 ns | 8.23 ns | 0 |
| 200 ticks | 68.7 ns | 8.02 ns | 0 |
| 5000 ticks (wide) | 115 ns | 8.55 ns | 0 |
| 40000 ticks (window defeated) | 148 ns | 17.9 ns | 3167 |
The last row is the one worth having. Everything below 65,536 ticks fits the
price window and is the array's home turf; past that the book spills into an
overflow container and re-anchors. That row used to read 25,000 ns — a
rebuild with a heap allocation on every update, 184x slower than the
std::map it exists to beat, sitting immediately past the widest case anything
benchmarked. It is in the table now precisely because the table is what let it
hide.
Reads — where the array does not win:
Per call, on a book of 200 levels per side. gap is the tick distance between
adjacent levels: real Kraken BTC/USD at a 0.1 tick has levels every $0.50-$10,
so gap 1 is not a realistic book and is shown only as the best case.
| Operation | std::map |
tick-indexed array | |
|---|---|---|---|
| best bid + best ask | gap 1 | 1.33 ns | 1.36 ns |
| gap 100 | 1.30 ns | 1.43 ns | |
| top 10 levels | gap 1 | 61 ns | 59 ns |
| gap 100 | 61 ns | 610 ns | |
| Kraken checksum | gap 1 | 451 ns | 411 ns |
| gap 100 | 459 ns | 1326 ns |
std::map is completely insensitive to sparsity. The array is not: it scans
empty slots, so at a realistic gap it is 3.2x slower on the checksum and 8.2x
slower on top(10). Earlier versions of this table published only the gap-1
row, which is the array's best case presented as its typical one. (The
checksum rows predate the slice-by-8 CRC kernel; making the CRC cheaper grew
the iteration's share, so the sparsity penalty measured after that change is
larger — 4.1x — not smaller. The direction of the conclusion is unchanged:
sparse reads are the array's weakness and the map's strength.)
The touch read is a genuine tie. The previous table claimed 1.33 ns against
2.67 ns and conceded a 2x read loss — that was two different units compared
against each other, one of them measured with the work hoisted out of the loop
by an incorrect DoNotOptimize. Fixing the benchmark removed the loss.
The whole frame, end to end. Feed::handle — JSON decode, RFC 8259 scalar
validation, symbol routing, canonical-spelling guards, timestamp parse,
sequence check, book update, depth trim, and CRC32 verification of the result
against the exchange — measured per frame on a 200-level book:
| per frame | frames/s (single thread) | |
|---|---|---|
| 1-level update (the common case) | ~2.1 us | ~480,000 |
| 10-level update | ~4.5-10.6 us | ~100,000-230,000 |
Every one of those frames has its checksum recomputed and compared. The handlers cited at the top of this README make that verification an opt-in flag because it costs them 10-100 us per update; here the entire frame — decode included — costs a fifth of their verification step alone.
Where the frame's cost lives now (one quiet-machine run, per-family):
| Stage | ns | share |
|---|---|---|
| JSON decode + ingest guards | ~1900 | ~75% |
| Kraken checksum (iterate + CRC) | ~600 | ~25% |
| book update | ~8 | 0.3% |
Decode used to be ~2400 ns and 89% of the frame. The fix was not removing the
guards — the RFC 8259 scalar validation, symbol routing, timestamp parse, and
canonical-spelling byte-compare are the checks that make the rest of this
README true, and they stayed. The fix was json::for_each_member: the old
decoder restarted a find() from the front of the frame for every field it
read, and an audit measured one Kraken frame being re-walked 9.3x —
checksum and timestamp are spelled after the level arrays on the wire,
so each of those lookups walked both arrays to reach its key. One member walk,
with the walk itself carrying the structural validation a separate
well_formed pass used to pay for, removed the redundancy without touching
what is checked. The CRC kernel also went from byte-at-a-time to slice-by-8
over one buffered payload (SSE4.2's crc32 instruction is CRC32C and cannot
compute Kraken's IEEE polynomial — the header documents the trap).
Measured properly — both versions rebuilt and run back to back in the same machine state, 9 repetitions each, because this desktop drifts between fast and slow states by 2x across a morning:
| median, same-state A/B | before | after | |
|---|---|---|---|
| decode, 1-level frame | 10.7 us | 4.9 us | 2.2x |
| decode, 100-level snapshot | 532 us | 247 us | 2.2x |
Feed::handle, 1-level frame |
13.5-14.6 us | 5.5-5.9 us | 2.4-2.7x |
Feed::handle, 50-level frame |
163-176 us | 64-73 us | 2.2-2.7x |
(Those absolutes are from a slow machine state; the ~2.1 us figure above is the same benchmark in a quiet state. The ratios are the claim, and they hold in both states.)
So swapping the tick-indexed array back for a std::map still costs a few
percent end to end, not 8.5x. The array's speedup is not what makes the
library fast — decode still dominates. What the speedup buys is headroom:
it is why verifying every update is affordable rather than sampled, which
is the trade this library exists to make and the one the handlers cited at the
top declined.
Methodology. Median of 7+ repetitions, Google Benchmark, MSVC 19.50 /O2,
Windows 11, 16 logical cores @ 2995 MHz, on an untuned desktop that was not
otherwise idle. Read the ratios within a table, not the absolute nanoseconds
across tables; before/after comparisons are only ever taken from back-to-back
runs in the same machine state. Reproduce with
cmake --preset bench && ./build/bench/bench/crossbook_bench --benchmark_repetitions=7.
One methodology note that cost real time to learn: do not run the whole suite in one process. The 40,000-tick row does sustained window rebuilds for tens of seconds and thermally throttles everything scheduled after it — the same binary measured the touch read at 12.0 ns in-suite and 3.1 ns run alone. Every figure above comes from a per-family run.
What these numbers are not. These are throughput microbenchmarks — mean time per operation, warm cache, tight loop. They are not latency measurements. Nothing is pinned to a core, turbo and C-states are untouched, and the input is synthetic. A mean is the wrong statistic for latency anyway, because the distribution is heavily right-tailed and the mean hides the tail you would actually be paid to fix.
Tail latency needs a different harness, and
it now exists: replay_open_loop paces frames at
their recorded inter-arrival times regardless of whether the consumer keeps up,
and measures each one from the instant it was supposed to be processed rather
than the instant work began. A stall therefore lands on every message queued
behind it, exactly as production would experience — no correction step needed,
because nothing was omitted.
That distinction is not academic. A closed loop stops issuing work while it is stalled, so a 100ms hiccup contributes one slow sample instead of the ten thousand messages actually delayed behind it, and the reported p99.9 describes the harness rather than the system. Gil Tene named this coordinated omission; a test asserts the harness does not commit it.
Latency is now measured against a real captured Kraken feed rather than a synthetic generator — see below — but still on an untuned desktop. Pinned cores, disabled turbo, and disabled C-states would be needed before any of it should be compared against a production system.
The central claim is not a design argument. It is a measurement, reproducible with two commands:
./build/release/tools/crossbook_capture --venue kraken --symbol BTC/USD \
--depth 25 --seconds 900 --out cap.cbcap
./build/release/tools/crossbook_verify --replay cap.cbcap --depth 25| Instrument | Price scale | Duration | Checksums verified | Mismatches |
|---|---|---|---|---|
| BTC/USD | 1 decimal | 9.5 min | 13,937 | 0 |
| ETH/USD | 2 decimals | 40 min | 105,334 | 0 |
| XRP/USD | 5 decimals | 40 min | 92,462 | 0 |
| 211,733 | 0 |
Every one of those updates had its CRC32 recomputed locally from the reconstructed book and compared to the value Kraken published in that same message.
The three scales are the point, not decoration. XRP trades near $0.50 quoted to five decimals, so its mantissas land in the same numeric range as BTC's at $63,000 quoted to one. A scale bug that a single instrument would hide has nowhere to go across this set — and a test proves the wrong scale is caught rather than silently producing a plausible book.
A ~250 KiB slice of each capture is committed as a fixture and re-verified on every CI run across Linux, macOS, and Windows, so this is a continuous claim rather than a snapshot.
Binance would have been a fourth venue, but its websocket returns HTTP 451 from this location. The Binance decoder is therefore tested against its documented message shapes and fuzzed, but not verified against live traffic, and this README does not claim otherwise.
The depth-trim bug described in Live verification found a real bug was found twice, independently, by two separate live runs against this data — once at 99.47% and once at 98.66%. Same cause both times, and only observable by comparing against the exchange: the library passed every spec-derived test while getting it wrong.
Replaying the BTC/USD fixture open-loop at true 1x pacing, measuring each frame from the instant it was supposed to be processed — full JSON decode, ingest validation, book update, and CRC32 verification per frame:
| p50 | p99 | p99.9 | max | |
|---|---|---|---|---|
Median of 5 runs, --realtime |
8.7 us | 62.2 us | 176.9 us | 209.0 us |
Two waypoints behind that number, both worth keeping. An early version of this pipeline measured p50 at 2.4 us — before the scalar validation, symbol routing, and canonical-spelling checks that make the verification claims above true were on the measured path. Adding them moved p50 to 15.1 us, and that was not a regression to apologise for. The single-pass decoder then bought most of it back: same guards, same checks, half the walking. Market data arrives in bursts — frames captured microseconds apart queue behind each other, and open-loop measurement charges that queueing to every frame it delays, by construction — so a cheaper handler clears the burst faster, which is what the p50 improvement is mostly measuring.
Pinning a core made this consistently worse, which was not the expected
result. --pin on a hybrid CPU produced p99 figures two to four orders of
magnitude higher across every core tried; raising scheduler priority alone gave
the numbers above. The tool supports both flags and prints which were actually
applied, because a latency figure whose measurement conditions are unstated is
not a measurement — and one whose stated conditions silently failed to apply is
worse.
The tool warns that it fell behind on some events, and that warning is generated by the harness rather than added afterwards. Do not compare these to a production system: turbo and C-states are untouched, the capture is read from disk, and the run-to-run spread on the tail is roughly 5x. The tail is measuring this desktop.
Earlier versions of this section reported a 3.4 ms max. Most of that was a metric bug, not the machine — see below.
A single latency figure at a single rate is how throughput numbers become
marketing: in an open-arrival system, latency sits flat while there is
headroom and diverges as the offered rate approaches saturation, so the honest
artifact is the curve — swept over a recorded real feed, which is the same
shape as STAC-M1, the industry's feed-handler benchmark. ReplayOptions::speed
existed from the start; an audit pointed out that nothing ever swept it, which
made it a control that could not fail. Now the tool does:
./build/release/tools/crossbook_verify --replay tests/fixtures/kraken_btcusd_l2.cbcap --sweepOne run on this desktop, replaying the committed BTC/USD capture tiled to at least 10,000 samples per rung:
| offered (msg/s) | p50 | p99 | n | zero late events |
|---|---|---|---|---|
| 5 (recorded pace) | 27.0 us | 69.1 us | 36 | yes |
| 121 | 24.0 us | 139.0 us | 1,192 | no |
| 614 | 20.4 us | 138.9 us | 6,124 | no |
| 1,251 | 23.2 us | 257.9 us | 10,161 | no |
| 3,128 | 25.6 us | 640.5 us | 10,161 | no |
| 6,257 | 39.4 us | 911.9 us | 10,161 | no |
The shape is the point: p50 is flat within noise across three orders of magnitude of offered rate — the handler is nowhere near compute-bound, as the ~2 us service time predicts — while the tail grows with rate because every scheduling hiccup lands on more queued frames. The strict conclusion the tool prints ("sustained 5 msg/s") uses a zero-late-events definition of sustained, and on a general-purpose OS that mostly measures the scheduler; the p99 column is the informative one, and rungs whose sample count cannot resolve a p99.9 are footnoted rather than reported. Reading a saturation knee for the library off this desktop would be dishonest either way — what the sweep demonstrates is that the measurement exists, fails when it should, and travels with its conditions. Run it on tuned hardware and the knee means something.
A [timing] test failed intermittently and was diagnosed as machine
contention. It was not.
behind_schedule incremented whenever an event arrived past its deadline — by
any amount. But a spin loop can only observe that a deadline has passed after
it passes, so every wait overshoots by at least one clock read. On an idle
machine, a 5 ms schedule reported one event "behind" by 100 nanoseconds, and
kept_pace() — including the saturation warning printed next to every latency
figure above — called that a system that could not keep up.
Overshoot below 1 us is now not lateness, with tests pinning both directions: noise is not reported, and a 20 ms stall against a 1 ms schedule still is. The contention guard stayed, because CI runners really are shared, but it is no longer covering for a metric that was over-sensitive by three orders of magnitude.
The lesson is the uncomfortable one: a flaky test was blamed on the environment for several rounds before the environment turned out to be innocent.
Worth stating plainly, because it is the reason the benchmarks exist.
The first implementation of the tick-indexed array scanned the window from its
edge to find the best price. Writes were 8x faster than the tree, exactly as
designed. Reads were ~15,000x slower — 67 µs to answer "what is the best
bid", versus 4 ns for std::map — because every read walked ~32,000 empty slots.
Correct, fully passing its equivalence tests, and completely useless: every quoting decision reads the touch. The fix was a maintained best-index hint, and the differential oracle confirmed the optimisation changed no behaviour. Numbers in the table above are post-fix.
That one happened before the first commit, so unlike the depth-trim bug above you cannot dig it out of the history — take it as an anecdote about why the benchmarks exist, not as evidence. The checksum bug is the one with the audit trail: a capture you can replay, a fix you can diff, and a test that fails without it.
Requires CMake 3.24+ and a C++20 compiler (GCC 11+, Clang 14+, MSVC 19.30+).
git clone https://github.com/jdardash/crossbook.git
cd crossbook
cmake --preset release
cmake --build build/release
ctest --preset releaseHeader-only, so consuming it is just an include path:
include(FetchContent)
FetchContent_Declare(crossbook
GIT_REPOSITORY https://github.com/jdardash/crossbook.git
GIT_TAG v0.3.0
GIT_SHALLOW TRUE
SYSTEM)
FetchContent_MakeAvailable(crossbook)
target_link_libraries(your_target PRIVATE crossbook::crossbook)Or vendor it. There is no generated header, no configure step, and no dependency outside the standard library, so copying the tree is a complete install — and for a lot of desks that is the honest answer:
cp -r include/crossbook third_party/include/ is the entire library. find_package(crossbook CONFIG REQUIRED)
works too, against an installed prefix.
#include "crossbook/book.hpp"
#include "crossbook/checksum.hpp"
using namespace crossbook;
// BTC/USD on Kraken: 1 decimal of price, 8 of quantity.
ArrayBook book(InstrumentSpec{"BTC/USD", 1, 8});
// Parse wire values exactly — no double, ever.
const auto price = parse_fixed("45285.2", book.spec().price_scale);
const auto qty = parse_fixed("0.00100000", book.spec().qty_scale);
book.apply(Side::kAsk, Price{price.mantissa}, Qty{qty.mantissa});
// Verify against what the exchange said.
if (kraken_checksum(book) != message_checksum) {
// Enumerate it — never just count it.
log.record({DivergenceKind::kChecksumMismatch, "kraken", "BTC/USD",
ts, seq, message_checksum, kraken_checksum(book),
kraken_checksum_payload(book)});
}Or drive the whole pipeline — decode, verify, recover — from raw frames:
#include "crossbook/feed.hpp"
#include "crossbook/venues/kraken.hpp"
using namespace crossbook;
using Kraken = Feed<venues::KrakenBookDecoder, ArrayBook>;
Kraken feed("kraken", venues::KrakenBookDecoder(InstrumentSpec{"BTC/USD", 1, 8}),
SequencePolicy::kStrictIncrement);
for (std::string_view frame : frames_from_your_transport) {
switch (feed.handle(frame)) {
case FeedStatus::kApplied: break; // Verified against Kraken's CRC32.
case FeedStatus::kIgnored: break; // Heartbeat, ack, stale duplicate.
case FeedStatus::kRejected: break; // Logged; book untouched.
case FeedStatus::kNeedsSnapshot:
// The book is known wrong. Resubscribe. Do NOT read it until then.
resubscribe();
break;
}
}
// feed.synced() must be true before anyone reads the book.
// feed.match_rate() is only evidence if feed.divergences().verified() > 0.v0.3 — the roadmap is done, except the socket.
- Exact fixed-point decimal, refuses to round rather than silently rounding
- L2 book, two implementations, differentially tested against each other
- Kraken CRC32 checksum, allocation-free, matching the documented algorithm
- Sequence continuity for Binance spot, Binance futures, Coinbase
- Divergence log with cause classification
- No-allocation hot path under a walking touch, enforced by a test
that hooks global
operator new - Determinism via state hashing
- Zero-dependency JSON scanner returning raw wire tokens
- Kraken v2
bookand Binance spot/futures depth decoders - Feed handler with resnapshot recovery and staleness detection
- HDR histogram with coordinated-omission correction
- Open-loop replay harness measuring against the schedule
- Rate-vs-latency sweep over a recorded capture (
--sweep), with honest sample counts and a stated p99 bound - Single-pass venue decoders over a one-walk JSON member cursor
- Slice-by-8 CRC32 over a buffered checksum payload
- L3 order-by-order book: arena-pooled intrusive queues, open-addressed id lookup, and queue position
-
executable_sizeandcost_to_trade— what you can actually trade - Consolidated cross-venue book: fee-adjusted, staleness-filtered
- Websocket transport: RFC 6455 framing, TLS via Schannel and OpenSSL
- Depth-limited book trimming — found by live verification, not by a test
- Capture and byte-exact offline replay, with recorded captures committed at three price scales
-
-Werror, ASan + UBSan, and a differential fuzzer per subsystem - Automatic Binance REST snapshot reconciliation in the tool (v0.3)
The correctness core still decodes, verifies, and recovers without opening a
socket: the transport is a separate, optional target, and consuming
crossbook::crossbook pulls in no TLS stack. -DCROSSBOOK_BUILD_TOOLS=OFF
drops it entirely. That boundary is what keeps the core testable offline —
which is also how the whole stack gets tested, since CI verifies a recorded
capture rather than a live venue.
The library has none: standard library only. The JSON reader, the RFC 6455 codec, SHA-1 and base64 are written here rather than pulled in, and two of those are load-bearing rather than stylistic.
The JSON reader returns the untouched wire token for every value, because
Kraken's checksum is computed over the digits as the venue spelled them — a
parser that hands back a double has already destroyed the information needed
to verify the book. And SHA-1 is here so that Sec-WebSocket-Accept is actually
verified rather than assumed; that check is what proves the peer parsed the
upgrade request rather than merely answering 101, and it is the step most
hand-rolled clients skip.
TLS is the one thing that cannot reasonably be written here, so each platform's
own is used: Schannel on Windows, which ships with the OS, and OpenSSL
elsewhere. cmake --build therefore produces a working client on a stock
Windows machine with nothing installed.
The touch is not a size. "Best ask 45283.6" says nothing about whether you can buy one coin there or fifty, and sizing a position off it is the most common way a spread that looked profitable turns out not to be.
#include "crossbook/execution.hpp"
// How much can I buy within 5 bps of the touch?
const Execution e = executable_size(book, Side::kAsk, from_bps(5));
e.qty; // Total available inside the limit
e.vwap; // What you would actually pay, size-weighted
e.slippage; // Cost against the touch, in 0.01 bps units
e.depth_exhausted; // Ran out of book vs stopped by the limit — these differ
// What does one coin cost?
const Execution c = cost_to_trade(book, Side::kAsk, Qty{100'000'000});cost_to_trade reports a partial fill rather than extrapolating a price for
size that is not in the book, because inventing depth is how a backtest produces
returns a live account cannot.
This is also the concrete reason a correct book matters rather than an approximately correct one. The touch is refreshed constantly and self-corrects; a level ten deep can sit there wrong for hours. Depth is exactly what a reconstruction bug corrupts silently, and it is exactly what these functions read.
Equities have Reg NMS, a SIP, and a legally defined national best bid and offer.
Crypto has none of it — no authority on the best price, no shared clock, no
obligation for venues to agree. ConsolidatedBook makes each judgement call
explicit instead of burying it:
- Fees, not quotes. Taker fees are large relative to crypto spreads, so the venue with the best headline price is frequently not the cheapest to trade. A test pins a case where a 26 bps venue quoting 3 ticks better loses to a 1 bp venue.
- Size changes the answer. A venue can be best on one coin and worst on
fifty.
best_executiontherefore takes a quantity and walks each venue's real depth — a size-free "best venue" is not a well-defined question. - Staleness excludes, and fails closed. A quiet venue looks exactly like a
stable one, so entries past a configured age are dropped rather than quoted.
An unstamped quote is unusable rather than immortally fresh, and a
timestamp in the future is treated as a clock fault — the failure modes here
all had to be inverted, because every one of them originally failed open.
StalenessPolicy::kDisabledis explicit rather than a magic zero. - Local clocks only. Venue timestamps are never compared to each other.
- Scales must match, and it is checked. Two venues can quote the same
instrument at different price and quantity scales, and comparing their raw
mantissas is meaningless — Kraken at
price_scale=1against Binance at2makes one venue win every bid and lose every ask, forever, while mismatched quantity scales silently corrupt the VWAP weights with no visible symptom.VenueQuotetherefore carries itsInstrumentSpec,update()returnsfalseand refuses a quote that disagrees, andbest_executiondrops any book whose spec does not match. - Rounding costs the trader, never flatters them. Integer division truncates toward zero, which understates what an ask costs and overstates what a bid pays — enough to invert the venue ranking outright. Fees round away from zero and VWAP rounds against the taker, so a venue is never reported cheaper than it is.
- Ties break deterministically, on venue name rather than arrival order. Two processes reading the same market must route identically, and rounding onto integer prices manufactures exact ties often enough for it to matter.
The questions a feed-handler engineer asks within a minute of reading "tick-indexed array", answered plainly rather than left to be inferred.
Memory. ArraySide defaults to 65,536 slots of int64, so 512 KiB per
side and 1 MiB per ArrayBook, allocated up front whether the book is full or
empty. At Kraken BTC/USD's 0.1 tick that window spans about $6.5k. Two
hundred instruments is therefore ~200 MiB of windows, which is the number to
budget against. MapBook has no floor and costs roughly 64 B per live level;
Feed is templated on the book type, so Feed<Decoder, MapBook> is a
one-word change when footprint matters more than update cost. The slot count is
a constructor parameter — 8,192 slots (64 KiB/side) is ample for a depth-10 or
depth-100 subscription and lets far more books stay resident in L2.
Prices outside the window. They are not dropped. They go to an overflow
container and stay correct there, and the window re-anchors around the touch
once enough have accumulated. That path is the one most likely to harbour a bug,
which is why the differential test hammers it — but it is also genuinely more
expensive than the array fast path, so a book whose live span persistently
exceeds the window is one that should be using MapBook.
Threading. The library is single-threaded and contains no std::mutex,
std::atomic, or std::thread by design. One Feed belongs to one thread; if
another thread reads the book, you supply the synchronisation. Nothing here
publishes a consistent snapshot for you, and Feed::handle is synchronous with
no internal queue, so backpressure between your socket and the handler is also
yours to design. This is a deliberate boundary — the same one that keeps the
correctness core testable offline — not an oversight, but it is your problem
and the README should say so.
- Not a trading system. No signals, no strategy, no positions, no PnL. There is no type in this library that represents a position.
- Not a matching engine. Exchanges do not need theirs rebuilt.
- Not competitive with colocated production systems. It is not tuned for a latency budget, and no benchmark here is run on tuned hardware.
It is a correctness-first feed handler core, built so that the claim "this book is right" is something you can check rather than something you have to believe.
fixed.hpp— why prices are integers, and the one precondition the checksum fast path rests onbook.hpp— two side containers and why both staychecksum.hpp— Kraken's algorithm, and why it is cheap heresequence.hpp— three venue contracts, one state machinejson.hpp— why the scanner is hand-written, and whywell_formedhas to run firstfeed.hpp— the recovery state machinehistogram.hpp— HDR bucketing and coordinated-omission correctionreplay.hpp— open-loop pacing, and why the spin threshold is 20msl3.hpp— arena, intrusive queues, and why a size reduction keeps priority while an increase does notexecution.hpp— why the unit is 0.01 bpsconsolidated.hpp— the four judgement calls behind a crypto "best price"bench_book.cpp— what the numbers mean and don't
- Kraken WebSocket v2 book checksum
- Kraken
level3channel - Binance spot WebSocket streams
- Binance futures: managing a local order book
- Coinbase Exchange websocket channels
- Gil Tene, How NOT to Measure Latency — the coordinated omission problem
Issues and PRs welcome. Two rules, both non-negotiable and both mechanically checked:
- No floating point in the book. Prices and quantities are integer
mantissas. This is what makes checksums and determinism possible. A CI job
greps the book path for
float/double— reporting code may use them, because a percentile is a ratio; the book may not. (This rule was described as mechanically checked for a while before anything checked it. It is now.) - No allocation on the hot path.
tests/test_no_alloc.cpphooks globaloperator newand will fail the build if you add one. It covers a book whose touch walks, not merely one that sits still — the earlier version drove a fixed touch and so never entered the only branch that could allocate. Beyond the price window the book degrades rather than allocating per update, anddegraded()reports when that has happened.
MIT — see LICENSE.