Skip to content

Repository files navigation

Quorum

deletion test tests first paid claim fee proven drains site

Prove it, or drop it.

Quorum finds exploitable bugs in EVM contracts and publishes only the ones it can prove by running the exploit on a fork — and it publishes how often it is wrong on hacks it has never seen.

A finding is a claim until an exploit settles it. Quorum hunts with the strongest reader available (Pashov's twelve-agent solidity-auditor), scores every tool on a blind benchmark of hacks nobody tuned for, and then, for each candidate, has a model write a Foundry exploit on a fork at the block the bug was live. The scaffold measures the attacker's real balance change and publishes only what actually pays. Everything else is kept as a candidate, never asserted.

Point it at your own contract. The prove stage is not tied to the benchmark — give it a chain, a verified address, the block, and a one-line hypothesis and it settles the claim by execution: python prove/tool.py prove base 0x… --block 27000000 --function swapV3 --attack "free-mint moves the pool price". It answers PROVEN with the exploit that pays, or unproven with the reason it does not. Method: prove/PROVE.md.

Or let it hunt and prove in one shot. python prove/hunt.py base 0x… --block latest reads the verified source, proposes candidate bugs, and settles the top few on a fork — so it finds and proves without you writing the hypothesis. It dedupes and ranks the finder's output and proves only the top handful, which is the answer to reader noise: the prover, not a person, decides which candidates were real. Method: prove/PROVE.md.

This is a rebuild. Quorum began as the lens-swarm described below (a Sibyl Labs hackathon entry). Measuring it honestly showed the regex lenses are the weakest reader — 1 of 9 real 2026 hacks, 0 of 2 on a blind pair — so they were demoted to a cheap baseline and the engine was rebuilt around proof. The lens-swarm sections that follow are kept as the origin story and still carry their real measured numbers.

What it does today, measured:

  • Four real 2026 hacks drained by a generated exploit, each reproducing to the wei with the model gone, no faked state: Unistreet (+0.0072 WETH), Sandbox (+1,000,000 SAND), Squid (+712 USDC), and — through the whole pipeline blind, on a hack the operator never read — GaslessReservoirEnabler (+0.836 WETH). Proofs: prove/proofs/, method: prove/PROVE.md.
  • A benchmark that cannot be tuned for. Only hacks reproduced after a freeze date count, scored before anyone reads the cause, order enforced in code: bench/STREAM.md.
  • The one number nobody else publishes: how often it is wrong, on hacks it has never seen. That is the whole differentiator — a trust moat, re-derivable by anyone.

Read the hacks, proven ↗ · The prove stage ↗ · The blind benchmark ↗

Telegram ↗ · Live site ↗ · Watch the demo ↗ · Verify it yourself ↗ · The paid claim ↗

Built for the Sibyl Labs Hackathon. Named one of fifteen consolation winners among 92 submissions when the results came out on 16 September 2026 (the announcement).


▶ Demo

The memory switch on the live front page: memory on, 2 confirmed; memory off, 0 confirmed; the terminal below crossfades between the two runs

The switch on runquorum.site, recorded live: flip the memory off and the same swarm on the same contracts confirms nothing.

Every terminal line in the demo is a real run: an empty memory, the lenses of that build on two teaching contracts, two findings published and five held back, three processes sharing one memory and doing twenty-four units of work once each, the reentrancy idiom recognised inside Friend.tech's live contract from a single sighting, the same swarm with memory removed confirming nothing, and the first claim on Base verified against memory. The video predates the fee: the paid claim on Robinhood Chain is on the registry page, not in the video.

quorum-demo-v2.mp4 (release asset, 12 MB) · the demo is also live: the deletion test switch on the front page and the in-browser claim verifier.

Turn the memory off. It finds nothing. Same files. Same checkers. One difference. No burn. No memory.
Turn the memory off. It finds nothing. Same files, same checkers, one difference: 2 confirmed with memory, 0 without No burn, no memory: 100,000 QUORUM per claim

Table of contents

The engine today

The origin (hackathon lens swarm), kept with its measured numbers


The problem the lens swarm was built for

This is the problem the origin architecture set out to solve. The problem Quorum solves today — telling a real finding from a plausible one — is answered by proof, not by memory; see the top.

A single detector that reports everything it sees is noise. Ten of them are ten times the noise.

  • A lens on its own cannot tell a finding from a sighting. Nobody checks whether a second, independent reading agrees.
  • Agents that share nothing duplicate everything. Without a shared record of who claimed what, every process scans every unit.
  • Nothing learned survives the session. A pattern confirmed today is re-derived from scratch tomorrow, on the same contract.
  • Human corrections evaporate. Retire a false positive and the next run reports it again.

Every one of those is a memory problem, not a detection problem. Quorum is the coordination and memory layer; the lenses are the honest minimum needed to have something real to coordinate about.

The origin: the lens swarm

This section describes the hackathon architecture Quorum grew out of. It is the weakest reader in the current benchmark (see the rebuild note up top), kept because its coordination and memory ideas are real and its numbers are honestly measured. The engine that does the work today is the prove-it pipeline.

Ten regex-and-brace-matching lenses over Solidity source, coordinated through one Sibyl Memory file and nothing else. The loop:

SCAN → CORROBORATE → REMEMBER → RECOGNISE → CLAIM → VERIFY

  1. Scan. Ten lenses read the source, two to four per risk, each reasoning from different evidence. A lens records what it saw and reads nothing about what its peers saw.
  2. Corroborate. A finding becomes real only when two lenses that work from different evidence arrive at the same conclusion. The tally of who agreed lives on the finding in memory (WARM tier), not in any agent's head. Disagreement is kept as a candidate and never published.
  3. Remember. A confirmed idiom is promoted to permanent swarm knowledge (REFERENCE tier). Every sighting, promotion, suppression and on-chain claim is appended to the COLD journal.
  4. Recognise. In a later session, on a contract the swarm has never read, a confirmed idiom is matched on sight from a single sighting. No quorum needed the second time.
  5. Claim. quorum attest calls the claim registry on Robinhood Chain, which pulls the fee (100,000 QUORUM), burns it through the token's own burn, and records the claim digest, all in one transaction. A record cannot exist without its fee. The finding itself never leaves the machine.
  6. Verify. quorum verify reads the registry's Claimed log back, checks the fee left the claimant and the supply in that same transaction, and recomputes the digest from memory. A claim whose fee was never burned cannot exist.

Verify it yourself in 60 seconds

No key, no wallet, no GPU. Every line below was run on a fresh memory file before it was written here; the expected results are in the comments.

git clone https://github.com/Yonkoo11/quorum && cd quorum
python3 -m venv .venv && .venv/bin/pip install -e .          # or: uv venv --python 3.12 .venv && uv pip install --python .venv/bin/python -e .

.venv/bin/python -m pytest tests -q                           # → 98 passed
.venv/bin/quorum --db fresh.db run --targets fixtures/*.sol   # → confirmed 2 | recalled 1 | candidates 5
.venv/bin/quorum --db fresh.db run --no-memory --targets fixtures/*.sol
                                                              # → confirmed 0 | recalled 0: without memory the swarm
                                                              #   cannot corroborate, recognise or forget.
.venv/bin/quorum --db fresh.db import 0x7556ec748f8ffb9e2ca5809c4383e407f4affb9b847124c0d33205281e905f32
                                                              # → ok digest matches the claim
                                                              #   ok revealed by the claim's signer
                                                              #   ok claim fee of 100,000 QUORUM burned
                                                              #   imported as a hint: still needs two local lenses

The import reads Robinhood Chain and refuses unless the revealed fields hash to the claim's digest, the reveal came from the claim's signer, and the claim's fee was burned. quorum verify <tx> on your machine checks the same chain half and then looks for the finding in your memory; on a fresh file it reports, honestly, that no finding there reproduces the digest, because the digest commits to the exact finding the publishing swarm held. The zero-install path is the in-browser verifier, which reads the transaction from a public node in your browser and compares it with the digest printed on the page.


The headline result

$ quorum run --targets fixtures/VulnerableVault.sol fixtures/OpenFeeSetter.sol
  QUORUM    VulnerableVault.sol:withdraw   reentrancy   corroborated by callorder-lens, guard-lens
  QUORUM    OpenFeeSetter.sol:setFeeRate   unguarded-state-write   corroborated by modifier-lens, sender-lens
scanned 20 lens-units | confirmed 2 | recalled 1 | candidates 5

# new process, new day, contracts it has never seen. Real verified Base mainnet source
$ quorum run --targets targets/*.sol
recalled before reading any code: 2 confirmed pattern(s)
  RECALLED  FriendtechSharesV1.sol:buyShares   reentrancy
            first confirmed on VulnerableVault.sol   (1 sighting was enough)
scanned 40 lens-units | confirmed 0 | recalled 1 | candidates 25

# same swarm, same contracts, memory removed
$ quorum run --no-memory --targets targets/*.sol
scanned 40 lens-units | confirmed 0 | recalled 0 | candidates 26
nothing was confirmed, recalled or suppressed: without memory the swarm
cannot corroborate, recognise or forget.
$ quorum recall

confirmed patterns (REFERENCE tier)
  reentrancy             reentrancy:e7801161b97ad175
    first confirmed on VulnerableVault.sol by callorder-lens, guard-lens
    recognised since on FriendtechSharesV1.sol  (1 sighting each, no quorum needed)

The pattern was learned on a teaching fixture and recognised in production code deployed on Base. msg.sender.call{value: amount}("") and protocolFeeDestination.call{value: protocolFee}("") are the same idiom, so they hash to the same signature. weth.deposit{value: amountETH}() is a different idiom and does not.

Run Quorum against audited production contracts and it mostly holds its tongue. On Aerodrome's Router, WETH9 and a Compound proxy it confirms nothing and files twenty candidates. That is the intended behaviour, not a failure to find bugs.


Architecture

flowchart TB
  L["ten lenses, separate processes, no messages between them<br/>callorder · guard · modifier · sender · consistency · forward · wrap · bound · ledger · payout"]
  subgraph M["one Sibyl Memory file: quorum/memory.py"]
    direction TB
    HOT["HOT · state/ · who claimed which unit"]
    WARM["WARM · entities/ · sightings, and who agreed"]
    REF["REFERENCE · reference/ · confirmed idioms"]
    ARCH["ARCHIVE · archive/ · retired findings"]
    COLD["COLD · journal · every event, auditable after the fact"]
  end
  RH[("Robinhood Chain<br/>QUORUM2 claim · QUORUM3 reveal")]
  L -- "claim_work" --> HOT
  L -- "record_sighting" --> WARM
  WARM -- "two lenses, different evidence: promote" --> REF
  REF -- "known_pattern: recognised on sight" --> L
  ARCH -- "is_retired" --> L
  REF -- "quorum attest: burn 100,000 QUORUM, then write the digest" --> RH
  RH -- "quorum verify · quorum import" --> REF
Loading

Every Sibyl Memory read and write in this project is in one file, quorum/memory.py. The chain layer, quorum/chain.py, is the only module that knows the token exists; tests/test_token.py asserts that the scanner modules never touch it.

Where memory is load-bearing

Four call sites carry the whole product.

What breaks without it Written at Read at Tier
Agents duplicate each other's work. A lens claims a (contract, lens) unit; a peer that finds it claimed does not scan it. This is the only coordination mechanism in the system. claim_work same call HOT state/
Quorum can never be reached. Corroboration is accumulated across lenses, processes and sessions on the finding entity. A lens has no idea who else agreed with it; memory does. record_sighting swarm.py:74 WARM entities/
Nothing is ever recognised again. A confirmed idiom becomes permanent swarm knowledge and is matched on sight in later sessions, on contracts the swarm has never read. promote known_pattern → swarm.py:79 REFERENCE reference/
Human corrections evaporate. Retire a finding once and no later session reports that shape again. retire is_retired → swarm.py:64 ARCHIVE archive/

Every sighting, promotion, suppression and on-chain claim is also appended to the COLD journal (log), which is what makes a published claim auditable after the fact.

The deletion test, as a command

The rules ask what happens if you delete the memory layer. Rather than assert an answer, Quorum ships it as a runtime flag. NoMemory implements the identical interface and forgets everything the instant it is written:

$ quorum run --no-memory

Every claim succeeds, so agents duplicate work. Every sighting looks like the first, so corroboration never accumulates and quorum is never reached. Nothing is recognised from an earlier session. Retirements do not stick. Confirmed findings: 0. Recalled: 0. The core function is gone, and the same behaviour is asserted in tests/test_quorum.py.


Coordination without a message bus

Quorum's agents are separate operating-system processes. They share one memory file and nothing else: no queue, no broker, no RPC between them.

$ quorum swarm --workers 3

3 agent processes, one memory, no message bus
  agent-1   scanned  26  stood down on  64 units a peer had already claimed
  agent-2   scanned  27  stood down on  63 units a peer had already claimed
  agent-3   scanned  37  stood down on  53 units a peer had already claimed

  90 units scanned in total, 180 skipped, in 15.3s
  no agent sent a message to any other agent. The HOT tier decided who did what.

Three processes, ninety units of work (ten lenses over nine files), each done exactly once. Claiming is optimistic because a read-then-write across processes is not atomic: an agent writes its own id into the claim, waits out the window in which a peer could be writing too, then reads the claim back and stands down unless it sees itself (claim_work). Take the HOT tier away and all three agents do all ninety units.

Why a quorum

Quorum pairs its lenses by risk, two of them for three risks and four for the access risk, and each lens reasons from different evidence:

Risk Lens A Lens B
reentrancy callorder-lens: an external call precedes a state write in the same function guard-lens: the function moves value out and carries no reentrancy guard
unguarded-state-write modifier-lens: externally callable, writes storage, carries no modifier at all sender-lens: writes a privileged-looking variable with no msg.sender check anywhere on the path.
consistency-lens: writes state that the contract's other writers of it only write behind a guard this function does not carry
forward-lens: anyone can call it, and it hands the caller's own bytes to another contract as itself (built from 2026 hacks, bench/HACKS-RECALL.md)
unsafe-math wrap-lens: the compiler lets this storage arithmetic wrap (a pre-0.8 pragma, or an unchecked block) bound-lens: nothing in the function bounds the operands before the write
accounting-mismatch ledger-lens: this balance has no way down anywhere in the contract, and this function relies on it payout-lens: value leaves this function against a balance it reads and never reduces

Each pair is two readings of one bug, never two bugs. The first arithmetic pair broke that rule (one lens for unchecked blocks, one for division before multiplication) and could never agree with itself; the benchmark showed it, and it was rebuilt. Agreement is signal. Disagreement is kept as a candidate and never published.


Claims on Robinhood Chain

When a finding reaches quorum it stops being a private opinion. quorum attest writes the claim digest to Robinhood Chain as a self-addressed 0-value transaction. The first claim was written to Base before the token existed, with QUORUM1-prefixed calldata:

$ quorum attest
signer 0xf9946775891a24462cD4ec885d0D4E2675C84355  balance 0.000500 ETH
  claimed on Base FriendtechSharesV1.sol:buyShares:reentrancy
    https://basescan.org/tx/0xa648821d91093df770b72c60be56834d069c9355c785e6195183e911f00bf713

A live claim, block 51138878. It can be read back and checked against the evidence that produced it:

$ quorum verify 0xa648821d91093df770b72c60be56834d069c9355c785e6195183e911f00bf713

claim on Base  block 51138878  2026-09-10T19:05:03+00:00
  published by 0xf9946775891a24462cD4ec885d0D4E2675C84355
  digest       0xfd7d5ec6e350aa28f160c2d3cf60d3faf4b52d8f7280bfe9f665ce1809042721

  the evidence for this claim is still in memory
    FriendtechSharesV1.sol:buyShares:reentrancy
    corroborated by guard-lens  (recall)
  digest recomputed from memory matches the chain

The digest commits to the risk, the idiom signature, the contract, the function and the exact set of lenses that corroborated it (chain.py). The finding itself never leaves the machine. Sibyl Memory is local-first and so is this. What goes on chain is a timestamped, verifiable claim that this swarm knew this shape at this block, which is what a disclosure timeline actually needs. The transaction hash is written back onto the finding entity in memory, so the claim and its evidence stay joined.

The signing key is read from the process environment at call time. It is never logged, printed or written to disk.

Publishing costs. Scanning does not.

Since 18 September 2026 the claim goes through a contract, the ClaimRegistry at 0xDeA0792cEc959CE6893C24dEeFc6FE9B047a3Ea3 on Robinhood Chain, deployed at block 66593107 by the claim wallet (tx). claim(digest) pulls the fee, burns it, and only then records claimedAt[digest][claimant], so a record cannot exist without its own fee and one fee cannot back two records. No owner, no pause, no upgrade, no ether, nothing to sweep; the fee and the token are immutable, so a different fee would be a different registry. Its source is verified on the explorer as an exact match of the build in this repository; the same source, 18 Foundry tests, a fuzz and an invariant are in contracts/. The first claim through it, made from this machine the same day and read back:

$ quorum verify 0x2f22250d80352b2633da7e75111bd06c0a811219675bb56626da606d1b5ffc3d

claim on Robinhood Chain  block 66594959  2026-09-18T22:49:36+00:00
  published by 0xf9946775891a24462cD4ec885d0D4E2675C84355
  digest       0x87f31bdf56ab844d8709d8bfef7725819f29a8df856134ad901d144393c545b4
  fee burned   100,000 QUORUM, pulled and burned by the registry 0xDeA0792cEc959CE6893C24dEeFc6FE9B047a3Ea3 in this same transaction

  the evidence for this claim is still in memory
    OpenFeeSetter.sol:setTreasury:unguarded-state-write
    corroborated by modifier-lens, sender-lens  (recall)
    evidence: treasury = newTreasury;

  digest recomputed from memory matches the chain

Its reveal is tx 0x99195f17…, imported into an empty memory as a hint the same evening. Claims made before the registry existed (the QUORUM2 shape below) still read and verify against their separate burn; a self-addressed claim mined after the registry existed is refused.

The claims above form a public registry of "this swarm knew this bug shape at this block". A public registry that is free to write to fills with junk, so writing to it has a cost, and the cost is destroyed rather than paid to anyone: each claim burns 100,000 QUORUM through the token contract's own burn(uint256) on Robinhood Chain before the claim is written to the same chain, and the claim's calldata (QUORUM2 shape) carries the burn's transaction hash. quorum verify checks both halves: the digest on chain, and that the burn it points at is a real burn of at least the fee by the same signer. Burn and claim share one chain, so one RPC verifies both. A claim whose fee was never burned does not verify.

One chain, on purpose. The token pays for one thing: publishing a claim. Claims live on Robinhood Chain because that is where the token is. The tool reads verified code from Ethereum, Base, Arbitrum, Optimism, Polygon and BNB Smart Chain, and that is the part of those chains Quorum cares about. The token goes to a second chain only when three things are true at once: a claim registry exists on that chain, a canonical or audited bridge path exists for the token, and someone on that chain wants to publish claims. None of the three is true today. There is no date, and we will not give one.

The first claim at the current fee, 2026-09-12. (The launch claim at the 1,000 fee, claim 0xacd123…efb0 at block 60748972 with burn 0x76da2f…d5d7, came 22 minutes earlier and still verifies.)

$ quorum attest --limit 1
signer 0xf994...4355  0.002358 ETH on Robinhood Chain  100,830 QUORUM on Robinhood Chain
each claim burns 100,000 QUORUM before it is written. Scanning is free; publishing is not.
  burned 100,000 QUORUM
    https://robinhoodchain.blockscout.com/tx/0x2005d3b84a3286980ff134d0641d120e7859e13568abf003e393300ac3c9b34a
  claimed on Robinhood Chain FriendtechSharesV1.sol:buyShares:reentrancy
    https://robinhoodchain.blockscout.com/tx/0xb999d218981ad9985b587da6c4017ae7dc8557ef702e27c9bbc9ca4f68bf1655

$ quorum verify 0xb999d218981ad9985b587da6c4017ae7dc8557ef702e27c9bbc9ca4f68bf1655
claim on Robinhood Chain  block 60762176  2026-09-12T02:46:55+00:00
  published by 0xf9946775891a24462cD4ec885d0D4E2675C84355
  digest       0xfd7d5ec6e350aa28f160c2d3cf60d3faf4b52d8f7280bfe9f665ce1809042721
  fee burned   100,000 QUORUM on Robinhood Chain, block 60762150, by the same signer
  ...
  digest recomputed from memory matches the chain

The burn is saved to memory the moment it lands, before the claim is sent, so if the claim transaction fails the next quorum attest reuses that burn instead of paying a second fee.

Once a claim is paid, its owner can reveal the pattern behind it, and any other swarm can import it. Run live the same day, the import into a memory database that had never seen anything:

$ quorum reveal FriendtechSharesV1.sol:buyShares:reentrancy      # discloses the claimed fields on chain (QUORUM3)
  revealed on Robinhood Chain FriendtechSharesV1.sol:buyShares:reentrancy
    https://robinhoodchain.blockscout.com/tx/0x7556ec748f8ffb9e2ca5809c4383e407f4affb9b847124c0d33205281e905f32

$ quorum --db fresh.db import 0x7556ec748f8ffb9e2ca5809c4383e407f4affb9b847124c0d33205281e905f32   # someone else's machine
reveal reentrancy  reentrancy:97d18d17cfa4494a  by 0xf9946775891a24462cD4ec885d0D4E2675C84355
  ok  digest matches the claim
  ok  revealed by the claim's signer
  ok  claim fee of 100,000 QUORUM burned
  imported as a hint: a sighting of this idiom is flagged, and still needs two local lenses to confirm

An import checks three things and refuses if any fails: the revealed fields hash to the claim's digest, the reveal came from the claim's signer, and the claim's fee was burned. So a pattern nobody paid to publish never enters anyone's memory. And a pattern somebody did pay for still cannot confirm anything on its own: it enters as a hint, a sighting that matches it is shown as a candidate with a note, and the finding needs two local lenses like any other. Only when this swarm reaches quorum on the same idiom does the imported pattern earn the trust of one it confirmed itself. One fee burn admits one import per memory. That is the whole job of the token: it is the cost of being heard by other people's swarms, not the right to be believed. Everything else, run, swarm, recall, retire, the memory tiers, the --no-memory test, has no token in it, and tests/test_token.py asserts that the scanner modules never touch it.

Token: QUORUM on Robinhood Chain (chain id 4663), contract 0xa6452Fd7134218f62056a304eaf501F8714A26b9. The first claim (Base, block 51138878) predates the fee and the move to Robinhood Chain; the first paid claim is the one above; quorum verify looks on Robinhood Chain first, then Base, and reads it back as a v1 claim with no burn to check. QUORUM_CHAIN_ID can point new claims at any chain in chain.CHAINS; the fee burn is always on Robinhood Chain, where the token is. The fee was 1,000 QUORUM at launch and is 100,000 from Robinhood Chain block 60761164 (chain.FEE_SCHEDULE); a burn is judged against the fee in force at its own block, so claims paid at the launch rate keep verifying.


What's real, and what we deliberately did not claim

Capability Status
The deletion test Real, and a runtime flag, not a paragraph. --no-memory runs the identical swarm through NoMemory: confirmed 0, recalled 0, asserted in tests/test_quorum.py.
Cross-session recognition Real. Learned on a teaching fixture, recognised in verified Base mainnet source in a new process from one sighting. The signature hashes the idiom on a line, not the identifiers on it.
Coordination without a message bus Real. Three OS processes, one memory file, 24 units each done exactly once; the HOT tier decided who did what. Take it away and all three do all 24.
Paid claims on chain Real. Fee burned through the token's own burn(uint256), claim written with the burn hash in its calldata, both live on Robinhood Chain (block 60762176). Reveal and import ran live the same day.
Measured against labelled bugs Three corpora, one harsh rule: a finding counts only if it names a labelled function. bench/BENCHMARK.md, SmartBugs-curated (143 files from 2017, 73 targets): two-witness 64% recall at 52% precision, reentrancy 90% / 72%, arithmetic 81% / 40%, after ten rounds of lens changes made on that corpus, so tuned; the accounting pair's five confirmations there count as false because the corpus labels none of its files for that bug (bench/ACCOUNTING.md, each of the five read by hand). The eighth round came from reading every confirmation of a run over 631 public repos (bench/ROBINHOOD.md): 3 of 281 were true. That run was later redone over the whole ecosystem with the ninth lens: of its 339 confirmations, the 115 the new lens takes part in were all read by hand, and two are real bugs on live code the audits had not touched, a Uniswap V4 hook whose fee claim anyone can redirect and a circuit-breaker registry anyone can seize. A separate question, asked of real hacks rather than audit reports, is in bench/HACKS-2026.md: of the 53 hacks of 2026 that DeFiHackLabs reproduces with a root-cause note, read one at a time, 23 are one of the four shapes these lenses read for, access control is the joint-largest cause, and reentrancy is down to 2. All 53 calls are published with their reasons in bench/labels/hacks-2026.json, the table is printed from that file rather than typed, and python3 bench/hacks.py --check <clone> confirms no file carrying a root-cause note was left unlabelled; an earlier pass said 21 but never recorded its per-file calls, and the file says so. That is an upper bound on what could be caught and not a hit rate, and the hit rate has since been measured against the same hacks: bench/HACKS-RECALL.md fetches the deployed source of every one of those victims whose code can still be read, nine of the 23, and scores the lenses on the function each root-cause note blames. The two-witness rule confirmed the labelled bug in one of the nine, at 12% precision. It named the right function in three, calling two access-control bugs reentrancy. A tenth lens, forward-lens, was then built from those two misses (a function anyone can call hands the caller's bytes to another contract as itself), and with it the rule confirms three of nine at 30%. That is fitted to the same nine and is not an out-of-sample number; the one of nine is. It moved nothing on SmartBugs or the contest corpus and cost one false confirmation on the held-out set. The accounting pair scored 0 of 3 and no lens sighted any of them. Fourteen could not be scanned at all, six because the victim never published source, which is a permanent limit on a source reader rather than a gap to close. By dollars the picture inverts, because most 2026 losses came from stolen keys and operational failure, which a source reader cannot see at all; the Bitget breach is the clearest case and Quorum could not have helped with it. The ninth came from bench/MODERN.md, the first measurement against recent audit contests, which is the number that matters most and is the worst one here: of 131 High and Medium findings across 16 contests, 7 are shapes these lenses look for at all, and of those 7 the two-witness rule confirmed 0 until a ninth lens was built for exactly that gap, and now confirms 2. All 29 of its confirmations have been read by hand; the other 27 are false. The lens was developed against that corpus, so its 29% recall there is no longer a measurement on unseen code, and the file says so. bench/HELDOUT.md, DeFiVulnLabs (57 modern files, 10 targets hand-labelled before any lens was changed, and no lens was adjusted against it): 50% recall at 62% precision, down from 71% when the tenth lens confirmed the arbitrary call in UnsafeCall.sol, a file the labels put outside the scored risks before that lens existed. bench/SLITHER.md and bench/SLITHER-HELDOUT.md score Slither the same way: 48% / 41% on SmartBugs (reentrancy 90% / 62%, and far better on access control) and 40% / 20% on the held-out set. Slither has no overflow detector, so it scores 0% on arithmetic on 2017 code; the rebuilt pair scores 81% at 35%. bench/README.md has the tables, the history and the caveats.
Restraint on production code Measured, not asserted. On Aerodrome's Router, WETH9 and a Compound proxy, fresh memory, nine lenses: 0 confirmed, 20 candidates held back. On the wrapped token from each of five chains: 0 confirmed, 29 held back (bench/MULTICHAIN.md). On 1,164 files from twelve audited codebases the accounting pair confirmed nothing (bench/ACCOUNTING.md). The first cut of the rebuilt arithmetic pair confirmed WETH9's deposit (balanceOf[msg.sender] += msg.value), which is why bound-lens now counts what the chain itself bounds as bounded.
Tests 89, run in CI on every push. They cover the idiom signature matching across contracts, that one lens never confirms, that two lenses reach quorum, that the deletion test really confirms nothing, the claim and reveal calldata shapes, that a burn is only valid for the fee on the token, that attest burns before it claims, that the scanner modules never touch the token, that burn and claim share one chain by default, and that the first Base claim still reads after the move, that the pre-0.5 call idiom reaches quorum, that an unnamed 0.4 fallback is parsed as a function, that the call-order lens follows a storage alias, that the SARIF export names both witnesses and leaves candidates out, that the arithmetic pair reads one bug from two sides (bounded arithmetic and checked arithmetic each get one witness only), that an imported pattern never confirms alone and is upgraded by local quorum, that one burn admits one import, that reverted, unpaid and mis-addressed claims are refused, that malformed reveals are refused rather than crashed, that the v2 signature keeps the member name, that a 2300-gas transfer is not an external call while a token transfer is, that braces inside comments are ignored and lines are counted from the brace, that 0.4 constructors and constant functions are skipped, that bound-lens reads msg.value as a dotted name and accepts an equality bound, and that an explorer's answer can only become a file under targets/<chain>/ or be refused (multi-file joins, 404s, unverified or oversized bodies, a name that tries to leave the folder), that the accounting pair reads one bug from two sides and confirms only together (the fixed twin gets neither lens; a counter read by a setter and a payout against a balance lowered elsewhere each get one), that one file name on two chains stays two targets, that every variable a line mentions is reported and in a fixed order (it reported one, chosen from a set, so the same file confirmed an accounting finding on one run and not the next), that a struct's fields are not contract state so a local sharing a field's name is not a state write, and that the third reading of the access risk fires on a function whose siblings agree on a guard it does not carry while staying quiet on a one-shot initializer, a function that tests its own caller, a caller writing their own row, and a write the caller cannot reach.
Findings in the Security tab Live: three alerts on the fixtures in this repository's Security tab. run --sarif writes SARIF 2.1.0; tests/test_quorum.py asserts that the fixture run yields three results (two by quorum, one recalled), that each names both lenses, that candidates are left out, that the fingerprint is the idiom signature, that a second run on the same memory keeps the findings in the log, and that quoted source cannot carry a link and is bounded. The first version of the action installed from the main branch at run time and interpolated its inputs into a shell line; both were found in the 2026-09-13 review and fixed the same day (install from the pinned ref, inputs through the environment, actions pinned to commits, upload split from the scan).
Attacked, 2026-09-13 Four ways to poison the registry were found by reviewing the tool itself. Fixed: an imported pattern was recalled from one sighting exactly like a locally confirmed one (now a hint until local quorum); one fee burn could admit unlimited imports into a memory (now one per burn); the idiom signature dropped the member name, so x.delegatecall(y) and t.approve(s) hashed the same and one retirement silenced both (signature v2 keeps it); a claim was accepted even if reverted, not self-addressed, or unpaid after the fee existed (all refused). Not fixed: globally, one burn can still back more than one claim, because nothing on chain ties a burn to a claim. That needs a contract or an indexer and is written here instead of pretended.
The lenses Deliberately simple: regex-and-brace-matching heuristics over source text, not a compiler front end. They read lines rather than a call graph: a call inside a modifier, a write reached through an internal call, and a guard in the caller are invisible to them, which is why Slither finds far more access-control bugs on old code, and wrap-lens reads the pragma rather than the compiler, so a file with no pragma is treated as checked, and they key a finding by file, function and risk, so two contracts in one file can share a key. The accounting pair sees a balance that only ever grows; it cannot see a sibling function that forgot one debit another function has, because there the balance does go down, on the other path. The point of this project is the coordination and memory layer.
Vulnerability claims None. Quorum publishes corroborated idioms worth review, not confirmed vulnerabilities. A quorum means two independent lenses agreed on a shape, nothing more. The Friend.tech recall above is a pattern match on a call idiom, not an allegation about that contract.
Against a modern audit Measured, not guessed: the four risks are 5% of what 16 recent contests reported at High and Medium, and on the 7 findings in those shapes the rule confirmed 0. Four of the seven misses come from one cause, that this tool reads a file at a time while modern protocols spread state and logic across files. bench/MODERN.md names every miss and its reason.
The registry Deployed 2026-09-18, source verified on the explorer as an exact match, one real claim made, verified, revealed and imported the same day. The fork test against the real token never completed on a public node; the on-chain claim stands in its place.
The fixtures fixtures/ are vulnerable on purpose and are not deployed anywhere.
The first claim On Base, block 51138878, before the fee and the move. It reads back as a v1 claim with no burn to check, and it is accepted only because it predates the fee: an unpaid claim mined anywhere after the fee existed is refused by verify and import, as is a claim whose transaction reverted or was not self-addressed.
Exploits, proofs of concept, severity Not claimed, anywhere in this repository.

Tech stack

  • Language: Python 3.10 to 3.13. No framework; the CLI is argparse.
  • Memory: Sibyl Memory, all five tiers, load-bearing. Every read and write in one file.
  • Chain: web3.py against Robinhood Chain (chain id 4663) for the token, the fee burn, claims, reveals and imports; Base mainnet for the first claim. Verified target source comes from the Blockscout instances of Ethereum, Base, Arbitrum, Optimism and Polygon, and from Sourcify for BNB Smart Chain, which has no Blockscout instance and whose own explorer wants a key. No API key either way (bench/MULTICHAIN.md).
  • Tests: pytest, 89 tests, no chain access needed (the chain is mocked where it matters).
  • Site: static HTML, CSS and JavaScript in docs/, served by GitHub Pages at runquorum.site; the in-browser verifier reads the chain through public JSON-RPC nodes.
  • Demo: the terminal recording lives in demo/ and the video assembly in video/.

Project layout

prove/           # THE ENGINE TODAY — a hypothesis is a finding only when its exploit runs and pays
  harness.py     # fork at the block the bug was live, a model writes the exploit, measure real profit
  scaffold/      # ForkPoCBase: [PROOF] fires only on a measured attacker gain; the safety gates
  proofs/        # the exploits that pay, reproducible with `forge test` (four real 2026 hacks)
  PROVE.md       # the results and the method
bench/
  stream.py      # the blind benchmark: intake, contestants, labeller kept apart, order enforced in code
  STREAM.md      # the scored table and the blind end-to-end validation
learn/           # the education arm: each hack card backed by a running proof, plus defensive-habit guides
bot/             # the QUORUM buy bot (v4 pool on Robinhood Chain), token read from env, never in the repo
quorum/
  agents.py      # the ten lenses, two to four per risk (the origin baseline; the weakest reader now)
  swarm.py       # the run: claim units, record sightings, promote, recall, retire
  memory.py      # every Sibyl Memory read and write (HOT, WARM, REFERENCE, ARCHIVE, COLD) and NoMemory
  chain.py       # claim digest, QUORUM1/2/3 calldata, fee schedule, burn check, attest, verify, reveal, import
  targets.py     # quorum fetch: verified source from six chains, Blockscout and Sourcify, no key
  sarif.py       # confirmed findings as SARIF 2.1.0: both witnesses, what each read, the idiom as the fingerprint
  cli.py         # the command line
tests/           # 119 tests across the lenses, the token boundary, the stream harness, the prove gates and the buy bot
fixtures/        # two teaching contracts, vulnerable on purpose
docs/            # the site (runquorum.site): five pages, one stylesheet, one script, self-hosted fonts
brand/           # the cards, marks and fonts the site and the posts are built from
  card-maker.html  # one file, no install: type the words, save the card as a PNG (see MAKING-CARDS.md)
                 # bench/ also holds the lenses and Slither scored on three labelled corpora (the origin baseline)
contracts/       # the ClaimRegistry (Foundry): source, 18 tests, a fuzz and an invariant, the deploy script
diary/           # the Telegram diary: what the repo did, in plain words, every two hours, silent when nothing happened
action.yml       # `uses: Yonkoo11/quorum@v0.4.2`: scan, write the page to the job summary, write SARIF, upload to the Security tab
demo/            # the recorded terminal session and its beats
video/           # the demo video pipeline

Run it

Python 3.10 to 3.13. If python3 -m venv fails on your machine, the uv path below avoids it entirely.

git clone https://github.com/Yonkoo11/quorum && cd quorum

python3 -m venv .venv && .venv/bin/pip install -e .   # or:
uv venv --python 3.12 .venv && uv pip install --python .venv/bin/python -e .

# pull real verified source, no API key: --chain ethereum | base | arbitrum | optimism | polygon (default base)
.venv/bin/quorum fetch 0xCF205808Ed36593aa40a44F10c7f7C2F67d4A4d4 \
                       0xcF77a3Ba9A5CA399B7c97c74d54e5b1Beb874E43 \
                       0x4200000000000000000000000000000000000006
.venv/bin/quorum fetch --chain ethereum 0xC02aaA39b223FE8D0A0e5C4F27eAD9083C756Cc2   # lands in targets/ethereum/

.venv/bin/quorum run --targets fixtures/*.sol   # the swarm learns
.venv/bin/quorum run                            # a fresh session recognises
.venv/bin/quorum recall                         # what it knows, and how it knows it
.venv/bin/quorum swarm --workers 3              # three processes, one memory
.venv/bin/quorum verify <tx>                    # check a claim against memory
.venv/bin/quorum recall --since 2026-09-10T00:00:00+00:00   # what it learned since
.venv/bin/quorum run --no-memory                # the deletion test
.venv/bin/quorum run --sarif quorum.sarif       # the same run, findings written for the GitHub Security tab
.venv/bin/python -m pytest tests -q             # 89 tests

quorum attest additionally needs DEPLOYER_PRIVATE_KEY in the environment, gas on Robinhood Chain, and the claim fee in QUORUM. QUORUM_RPC overrides the public Robinhood Chain endpoint; BASE_RPC overrides the public Base endpoint used only to read the first claim.

Commands: fetch, run [--sarif PATH] [--summary PATH], swarm, recall [--since], retire <key> --reason, attest, verify <tx>, reveal <key>, import <tx>, status.

In a GitHub workflow

Findings go where reviewers already look. The action writes the run as a page in the job summary, which a reviewer reaches from the pull request's checks: each confirmed finding with the two lenses that agreed, the line each one read, and the count of candidates held back. Each Security-tab alert carries the same two witnesses and the idiom signature it will be recognised by next time; a candidate seen by one lens is not written at all. A small repository never opens the Security tab, so the page also lands as one comment on the pull request when the workflow adds the comment job shown in .github/workflows/quorum.yml: it downloads the page as an artifact and posts it from a job that runs no repository code. The fingerprint is the idiom signature, so the same shape is one alert across runs and renames, and a finding confirmed on an earlier run stays in the log while the code still has it. Pin the action to a tag or a commit: it installs the code at the ref you pinned and fetches nothing from a branch at run time. Source lines quoted in an alert are escaped and cut at 160 characters, so a comment in a contract cannot plant a link in your Security tab.

permissions:
  security-events: write
steps:
  - uses: actions/checkout@v4
  - uses: Yonkoo11/quorum@v0.4.2
    with:
      targets: contracts/**/*.sol

.github/workflows/quorum.yml runs it on this repository's own fixtures on every push, so the two teaching findings are this repository's own alerts. That workflow keeps the scan, the upload and the comment in separate jobs on purpose: the scan runs repository code with a read-only token; the upload and the comment need write-scoped tokens and run no repository code; the upload never runs on a pull request and the comment only runs on one from this repository. Every third-party action is pinned to a commit.

Tests

.venv/bin/python -m pytest tests -q             # 52 passed

tests/test_quorum.py drives the swarm end to end on the fixtures: one lens never confirms, two lenses from different evidence do, the signature matches across contracts, the deletion test confirms nothing, a retirement sticks. tests/test_token.py pins the calldata shapes, the digest a reveal must reproduce, the burn rules a claim must satisfy, that attest burns before it claims and reuses a saved burn rather than paying twice, that the scanner modules never import the chain, and that the first Base claim still reads after the move to Robinhood Chain. The same suite runs in CI on every push.

The diary

Every two hours a workflow reads what changed in this repo (commits, merged pull requests, releases, failed runs), hands it to Claude with the rules in diary/diary.py, and posts a short entry in plain words to the Telegram group. If nothing happened, or only noise happened, it posts nothing. The entry never mentions price, the market or the token, never invents, says when something broke, and uses the weakest of designed, built, tested or proven that the evidence supports. Links, keys, hashes and addresses are scrubbed before the text leaves the process. Each run covers the time since the previous scheduled run's window closed, so nothing is posted twice and a missed run is covered by the next.

python diary/diary.py --dry --hours 48    # print what it would say about the last two days, post nothing

Site and docs

runquorum.site · The lenses · The memory · The registry, with the in-browser verifier · Run it

MIT licensed. Memory is a local file; nothing is uploaded.

About

A swarm of Solidity security lenses that coordinates only through Sibyl Memory: two lenses must agree before a finding exists, memory carries it across sessions, and publishing a claim burns QUORUM on Robinhood Chain. Built for the Sibyl Labs Hackathon.

Topics

Resources

Security policy

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages