Skip to content

About

Provenance-enforced open dataset mapping US chartered sponsor banks to the fintech programs that ride on them: sponsorships, corporate control, middleware, enforcement actions and custodial structure. Every record carries a resolvable public citation.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Sponsor Bank / Fintech Program Map

validate license CDLA-Permissive-2.0 records 554

A public, open-source dataset mapping US chartered sponsor banks to the fintech programs that ride on them, together with the corporate, middleware and custodial structure around those relationships.

Routing-number → bank-of-record is a solved lookup, but only for authorized users: the Federal Reserve's bulk routing directory has not been publicly available since 2018 and current copies are licensed. Bank-of-record → which fintech program is open ground, and that mapping is this project's contribution.

This is an independent research project published under CrimsonVector. It is not conducted on behalf of, or endorsed by, any employer of the maintainer.


The provenance rule

Every record carries a public, resolvable citation. A record without one cannot exist in this repository.

This is enforced mechanically, not by discipline. tools/validate.py fails the build on any record missing a citation, missing or misusing its confidence field, carrying an unnamespaced ID, violating its schema, or citing a malformed URL — and on graph, supersession and ABA-checksum violations besides. CI runs it on every push; a pre-commit hook runs it before changes reach history. How strongly each record is held is recorded in two fields, described under the data model.


The data model

Eleven entity types and twelve edge types. Rectangles hold records; stadium outlines are types the schema defines but the dataset does not yet populate. Solid arrows are populated edges, dotted arrows defined-but-empty ones — the empty ones are shown deliberately, because what the model supports is part of what a reader needs to know.

flowchart LR
  CharteredBank["CharteredBank<br/>45"]
  BankHoldingCompany["BankHoldingCompany<br/>10"]
  FintechProgram["FintechProgram<br/>23"]
  ServiceProvider["ServiceProvider<br/>117"]
  RoutingNumber["RoutingNumber<br/>14"]
  AccountSchema(["AccountSchema<br/>0"])
  BIN(["BIN<br/>0"])
  CryptoRail(["CryptoRail<br/>0"])
  EnforcementAction["EnforcementAction<br/>17"]
  Advisory["Advisory<br/>6"]
  Typology["Typology<br/>3"]

  CharteredBank -->|"SPONSORS 39"| FintechProgram
  BankHoldingCompany -->|"CONTROLS 10"| CharteredBank
  ServiceProvider -->|"MANAGES_PROGRAM_AT 127"| CharteredBank
  FintechProgram -.->|"ROUTES_THROUGH* 0"| ServiceProvider
  CharteredBank -->|"OWNS_RTN 14"| RoutingNumber
  CharteredBank -.->|"ISSUES_BIN 0"| BIN
  ServiceProvider -->|"HOLDS_FBO_AT* 3"| CharteredBank
  FintechProgram -.->|"REFERS_KYC_TO* 0"| ServiceProvider
  ServiceProvider -.->|"SETTLES_ON* 0"| CryptoRail
  CharteredBank -->|"NAMED_IN* 23"| EnforcementAction
  EnforcementAction -->|"EXHIBITS* 2"| Typology
  FintechProgram -.->|"OFFRAMPS_VIA 0"| CryptoRail

  subgraph Legend
    direction LR
    P["has records"] -->|"populated"| Q["has records"]
    R(["defined, empty"]) -.->|"defined, empty"| S(["defined, empty"])
  end
Loading

* marks an edge type that accepts more than one endpoint type; the diagram draws one representative line each. Full rules, enforced by tools/validate.py:

Edge Records From To
SPONSORS 39 CharteredBank FintechProgram
CONTROLS 10 BankHoldingCompany CharteredBank
MANAGES_PROGRAM_AT 127 ServiceProvider CharteredBank
ROUTES_THROUGH 0 FintechProgram CharteredBank or ServiceProvider
OWNS_RTN 14 CharteredBank RoutingNumber
ISSUES_BIN 0 CharteredBank BIN
HOLDS_FBO_AT 3 FintechProgram or ServiceProvider CharteredBank
REFERS_KYC_TO 0 FintechProgram CharteredBank or ServiceProvider
SETTLES_ON 0 FintechProgram or ServiceProvider CryptoRail
NAMED_IN 23 BankHoldingCompany or CharteredBank or FintechProgram or ServiceProvider Advisory or EnforcementAction
EXHIBITS 2 CharteredBank or EnforcementAction or FintechProgram or RoutingNumber or ServiceProvider AccountSchema or Typology
OFFRAMPS_VIA 0 FintechProgram CryptoRail

Both diagrams are generated from the records by tools/gen_diagrams.py; a test fails if they drift out of step with the data.

Two units are used throughout. 239 is the number of record files; 554 is the number of records the validator checks — larger because three collection files carry 90, 98 and 127 members each. Bulk sources that refresh wholesale are stored as collections so a refresh is one diff rather than thousands.

Confidence and verification. Every record carries both:

  • confidence — confirmed (primary document, or original derivation cited to METHODOLOGY.md), reported (credible secondary), inferred (pattern-derived, basis stated).
  • verification — direct_fetch (the page was fetched and read), supplied_snapshot (read from a snapshot supplied to the maintainer), or search_excerpt (only a search-engine excerpt was seen). A confirmed record must rest on a direct read; a primary-but-unread page is reported.

Per record file: 237 confirmed, 2 reported; 227 direct_fetch, 10 supplied_snapshot, 2 search_excerpt. The two reported/search_excerpt records are the same pair — a BaaS provider whose bank disclosure is served only as JavaScript.

Model notes. A company gets one ServiceProvider record and its role is carried by the edge — Marqeta, Galileo, Lithic, Highnote and Alviere each act as program manager in one arrangement and platform or processor in another, and two records per company would make "what does this firm touch" return a fraction of the truth. Holding companies are modelled separately from their banks (CONTROLS), because a parent's statements are not the subsidiary's.


What's in it

Entity Count Edge Count
CharteredBank 45 SPONSORS 39
ServiceProvider 117 MANAGES_PROGRAM_AT 127
FintechProgram 23 NAMED_IN 23
EnforcementAction 17 OWNS_RTN 14
RoutingNumber 14 CONTROLS 10
BankHoldingCompany 10 HOLDS_FBO_AT 3
Advisory 6 EXHIBITS 2
Typology 3

Plus a 98-row manifest of the prepaid agreements sampled for the custodial finding, so its percentages can be recomputed rather than trusted.


Measured findings

Each carries its denominator. Methods and caveats are in METHODOLOGY.md.

1. Programs name their bank far more often than they publish its routing number. Of the 23 programs in the dataset, 5 do not name a sponsor bank (22%), but 21 do not publish the routing number their own customers use (91%) — only Brex and Wise do. A program is roughly four times more likely to withhold the identifier that would let an outsider resolve a payment than the name of the bank holding the money.

2. Enforcement actions are about fintech partners but almost never name them. Of 17 enforcement actions, 15 carry third-party or fintech-partner findings — and 2 name an individual partner. Those two are the same matter seen twice: a state and a federal supervisor, a day apart, both naming the same program manager. Every other action in the corpus describes partner-related failures without identifying the partner. This is a property of the documents themselves, measured over every action in the corpus, so it is not affected by how the corpus was assembled. Sutton Bank's consent order is the clearest illustration: it compels the bank to compile "a complete inventory of third-party relationships" — establishing that a definite partner list exists and that the supervisor expects to see it — while the public order names none of them.

The same measurement one layer deeper: naming a partner is not the same as describing what happened. Matching the corpus against typologies the federal government has already published gives a third step down. Of the same 17 actions, 15 carry third-party findings, 2 name a partner, and exactly 1 matter describes realized harm specifically enough to match a published federal typology — the Metropolitan Commercial Bank orders, which map to FinCEN's unemployment insurance fraud pattern and account for both EXHIBITS edges in the dataset. So the corpus narrows at every step: most actions are about partners, few name one, and one describes the consequence. 2 of the 3 recorded typologies carry no EXHIBITS edge at all. They are in the dataset because a cited advisory names the pattern, not because anything here was found to exhibit it, and their presence should not be read as evidence that it occurred.

Enforcement material can nonetheless surface relationships no disclosure page would carry — demonstrated once. The Federal Reserve's 2023 order against Metropolitan Commercial Bank describes MovoCash, Inc. as "a former third-party program manager" for the bank's Global Payments Group. That is the dataset's only sponsor relationship obtainable from enforcement material alone, and no disclosure page could have produced it: the program was already defunct, and defunct programs do not maintain "which bank" pages. Treat it as an existence proof of the channel, not as a rate.

Not a finding: the discovery_source split. 37 of 39 SPONSORS edges came from program disclosures, 1 from a bank's disclosure, and 1 from an enforcement order. That ratio is confounded by collection effort — 23 programs were swept systematically while 17 enforcement actions were collected opportunistically — so it describes this dataset's own composition and collection history, not the relative yield of the sources. It is recorded on every edge for provenance and reproducibility, and should not be read as a measurement of where sponsor relationships are discoverable in general.

3. Prepaid agreements condition FDIC insurance on identifying the cardholder. Across 98 agreements from 19 issuers, 43% (68% of issuers) condition insurance on registration or identification, and 36% (58%) name the bank holding the funds. One clause recurs near-verbatim across 12 of 19 issuers, varying only in the bank name — funds are insured "if specific deposit insurance requirements are met and your card is registered". We call it the registered-card condition. Under 12 CFR 330.5 and 330.7 pass-through insurance depends on ownership being ascertainable from records, so the clause is the consumer-facing form of a documented requirement — the same dependency the enforcement corpus addresses from the supervisory side.

4. A quarter of the CFPB prepaid registry is unreadable in the format the rule contemplates. The Prepaid Accounts rule expects text-searchable, digitally-created PDFs, yet 34 of 132 sampled agreements (26%) contained no PDF at all in the archive the CFPB serves. Every archive that did contain one yielded usable text, so the gap is in what was filed and published. Neither number is derivable from build/dataset.duckdb — the sampling frame and the skipped archives are recorded in the collection file, because a skipped archive produced no record for the artifact to hold.

Worked example: Revolut US

One program exercises nearly the whole model — four banks scoped by product, two dated supersessions, a holding company, two enforcement actions against a former sponsor, and the typology they both exhibit. Every element below is read from the records.

flowchart TD
  PROG["Revolut (US)<br/>FintechProgram"]
  cross_river_bank["Cross River Bank<br/>CharteredBank"]
  lead_bank["Lead Bank<br/>CharteredBank"]
  metropolitan_commercial_bank["Metropolitan Commercial Bank<br/>CharteredBank"]
  sutton_bank["Sutton Bank<br/>CharteredBank"]
  HC["Metropolitan Bank Holding Corp.<br/>BankHoldingCompany"]
  ACT1["Consent Order (NYSDFS)<br/>2023-10-18"]
  ACT2["Order to Cease and Desist and Order of Assessment of a Civil Money Penalty (BGFRS)<br/>2023-10-19"]
  TYP1["Unemployment insurance fraud<br/>Typology"]

  cross_river_bank -->|"SPONSORS · savings vaults (created after 2025-07-29)<br/>from 2025-07-29"| PROG
  cross_river_bank -->|"SPONSORS · Revolut Visa Credit Card<br/>from 2026-08-08"| PROG
  lead_bank -->|"SPONSORS · prepaid card account and Cardholder account<br/>from 2024-11-12"| PROG
  metropolitan_commercial_bank -.->|"SPONSORS · prepaid card account (Revolut and Revolut Bu…<br/>2024-08-26 → 2024-11-12 (ended)"| PROG
  sutton_bank -.->|"SPONSORS · savings vaults (created on or before 2025-07…<br/>2024-08-26 → 2025-07-29 (ended)"| PROG

  sutton_bank ==>|"superseded 2025-07-29"| cross_river_bank
  metropolitan_commercial_bank ==>|"superseded 2024-11-12"| lead_bank
  HC -->|"CONTROLS"| metropolitan_commercial_bank
  metropolitan_commercial_bank -->|"NAMED_IN"| ACT1
  metropolitan_commercial_bank -->|"NAMED_IN"| ACT2
  ACT1 -->|"EXHIBITS"| TYP1
  ACT2 -->|"EXHIBITS"| TYP1

  subgraph Legend2 [Legend]
    direction LR
    A1["current relationship"] -->|"solid"| A2["·"]
    B1["ended relationship"] -.->|"dotted"| B2["·"]
    C1["predecessor"] ==>|"supersession"| C2["successor"]
  end
Loading

Read it as: the prepaid card account moved from Metropolitan Commercial Bank to Lead Bank on 2024-11-12, and savings vaults from Sutton Bank to Cross River Bank on 2025-07-29 — two independent transitions affecting different products of the same program, each recorded as a superseded edge with its own effective date rather than as an overwrite.

The right-hand side is the typology layer on the same bank: two orders a day apart, from a state and a federal supervisor, both linked to the pattern FinCEN names. The EXHIBITS edges attach to the orders, not to the bank — the claim is that these documents describe that pattern, and nothing more.

Withdrawn: the retail-RTN hypothesis

One earlier explanation — that program-focused sponsor banks avoid publishing retail routing numbers — was tested against prepaid-program density and is not supported: program-heavy banks publish at 38%, banks with no prepaid programs at 40%. It is withdrawn. This rate is not derivable from the artifact: it is computed over the 31 banks whose sites could be fetched, a subset the artifact does not record. The cross-tabulation, the named confound and the small-n caveat are in METHODOLOGY.md.


Why the opacity matters

The four findings above all measure the same thing from different angles: how much of a sponsored program's structure is visible from outside it. They do not say why that visibility is worth measuring. The Advisory and Typology layers answer that using only documents the government has already published.

Supervisors describe the gap in their own words. The 2023 interagency guidance from the Federal Reserve, FDIC and OCC states that "A banking organization's use of third parties does not diminish its responsibility to meet these requirements to the same extent as if its activities were performed by the banking organization in-house" — which is why this dataset records a bank of record for every program, whoever faces the customer. The FDIC's 2024 custodial recordkeeping proposal describes the failure mode directly: consumers "have been unable to access their funds at IDIs for an extended period of time while the IDIs attempt to determine ownership of the funds deposited by fintechs."

Typologies are recorded, never coined. A Typology record exists only where a government publication names the pattern, and it carries that document's own indicator text rather than a paraphrase. The collector asserts every indicator string appears verbatim in the fetched source, so an invented indicator fails the build rather than passing review. Three are recorded: money mule schemes and unemployment insurance fraud (FinCEN FIN-2020-A003 and FIN-2020-A007, with FBI/IC3 PSA I-120419-PSA), and third-party payment processor risk (FIN-2012-A010).

Links are document-to-document. An EXHIBITS edge runs from an enforcement action to a typology, and it means one thing only: the order describes a pattern and a cited advisory names that pattern. Both endpoints are primary. No edge is drawn from a bank or a program to a typology, because that would be an assertion about an institution rather than a comparison of two documents — see scope rule 6 in CLAUDE.md.

The result is two edges, and the shortfall is the point. Across 17 enforcement actions, only the Metropolitan Commercial Bank matter — the New York DFS consent order of 2023-10-18 and the Federal Reserve order of the following day — describes a pattern specific enough to match a published typology. The DFS order finds that "fraud actors opened these prepaid card accounts using another individual's identity and directed payments, including direct deposit payroll payments and government benefits, onto the fraudulently opened cards," and records that the bank's delay "helped facilitate more than $300 million in pandemic unemployment benefits to be misdirected." FIN-2020-A007 names that pattern. The 15 remaining actions describe partner-oversight failures without describing what was realized through them, so no honest edge can be drawn — the same opacity the findings measure, showing up again one layer down.

Two candidate links were considered and rejected on reading the documents: an order mentioning unemployment benefits does so about legitimate customers being blocked from their own funds, and the only order mentioning a payment processor attributes an operational data-migration error to one. Neither is the published pattern, and keyword overlap is not a finding.


Scope rules

  • Volume cutoff for bulk sources. From the CFPB prepaid agreements database, issuers with ten or more active, most-recent agreements are ingested; below that, only issuers already present for other reasons. 22 of 85 active issuers clear the line, covering 2,230 of 2,351 agreements — 95% of the volume in 26% of the issuers. The tail is single-card credit unions that These are not derivable from the artifact either, being measured over the whole CFPB metadata snapshot rather than over the records ingested from it. would swamp the sponsor signal. This is a scope decision, not a claim about importance.
  • Excluded source: the Federal Reserve E-Payments Routing Directory, in any form. Its terms restrict it to financial institutions and authorized users and bar re-licensing, and bulk access has required a FedLine solution since December 2018 — incompatible with this repository's licence. Routing numbers here come only from institutions' own published pages. tools/docs_check.py fails the build if an excluded source reappears outside its exclusion note.
  • No named private individuals. Bulk registries contain free-text fields where filers sometimes enter a person's name instead of a firm. Those rows are skipped and logged, never recorded.
  • Derived values are computed, never transcribed. Any hash, count or extracted date must be computed from the artifact in the same step that writes it — a fabricated value of the correct shape passes every pattern check.

Known limitations

  • The enforcement corpus is a discovery method with a known bias, not a census. It surfaces banks that had supervisory problems, not all banks operating sponsor programs. Counting actions per bank measures supervisory attention, not program volume.
  • Routing numbers identify banks, not programs — for now. No bank in the dataset has both a verified retail RTN and a known program RTN, so the distinct-RTN ratio has zero comparable pairs and is not reported. Treat RTNs here as a bank-identification capability, not a program-detection one.
  • No account-number fingerprints exist here, deliberately. METHODOLOGY.md explains what determines account-number structure (the 17-position NACHA ceiling, the absence of a US format standard, virtual vs core numbering) but establishes no institution's actual format. That needs per-bank ground truth which has not been collected, and no public source substitutes for it.
  • Coverage is bounded by disclosure, and the two disclosure rates above measure how tight that bound is.
  • valid_from is first-observed, not inception. Sponsor relationships are routinely older than the page documenting them.
  • Some sources are unreachable from the collection environment; see COLLECTION_QUEUE.md, which records each with its reason.

Corrections made during development

Two provenance defects reached commits and were corrected: records marked confirmed whose pages had only been read via search excerpt, and parent-company statements attributed to subsidiary banks. Both were caught by tooling built for other reasons, both are visible in the log, and both are written up in full — what went wrong, how it surfaced, what changed — in METHODOLOGY.md. A third class was caught before committing and never entered history.


Repository structure

data/          the dataset: records as YAML, one file per record
  entities/      nodes (banks, programs, providers, actions, routing numbers)
  edges/         relationships (sponsors, controls, named-in, owns-rtn, …)
  collections/   bulk sources that refresh wholesale, many members per file
schema/        JSON Schema (2020-12) — the contract every record must satisfy
tools/         validate.py, build.py, docs_check.py, scope_check.py, pdfdoc.py
collectors/    re-runnable source collectors, each citing what it emits
tests/         validator tests, including negative cases proving uncited
               records are rejected
snapshots/     out-of-band source snapshots (contents gitignored)

Quickstart

python -m venv .venv
source .venv/bin/activate        # Windows: .venv\Scripts\activate
pip install -r requirements-dev.txt

python tools/validate.py         # provenance gate: must exit 0
pytest                           # includes negative tests proving uncited records fail
python tools/docs_check.py       # excluded sources must not reappear
python tools/build.py            # build/dataset.duckdb + build/csv/*.csv

tools/build.py refuses to run on a tree that fails validation, so the artifact is gated by the provenance rule and not merely the source files.

Querying the artifact

-- duckdb build/dataset.duckdb

-- which banks sponsor which programs, and how each was found
SELECT b.name AS bank, p.name AS program, e.product_scope, e.discovery_source
FROM edge_sponsors e
JOIN entity_chartered_bank b ON b.id = e.from_id
JOIN entity_fintech_program p ON p.id = e.to_id
ORDER BY 1, 2;

-- banks by number of distinct prepaid program managers (CFPB filings)
SELECT b.name, count(DISTINCT m.from_id) AS managers
FROM edge_manages_program_at m
JOIN entity_chartered_bank b ON b.id = m.to_id
GROUP BY 1 ORDER BY managers DESC;

-- recompute a published percentage from committed data
SELECT count(*) FILTER (WHERE fdic_conditioned_on_id) AS conditioned,
       count(*) AS total
FROM manifest_prepaid_agreement;

Every table carries provenance columns (citation_url, citation_retrieved, confidence, valid_from/valid_to), so any row can be traced to its source without leaving SQL.

Contributing

pre-commit install
pre-commit install --hook-type commit-msg

Hooks validate records and screen staged changes and commit messages against CLAUDE.md's scope rules. Organization-name screening uses a local-only denylist that is never committed — create ~/.config/sponsor-map/scope-denylist.txt (or set SCOPE_DENYLIST_FILE), one term per line. It lives in your home directory rather than the repository, so it is created once per machine and survives cloning; a fresh clone of this repository does not need a new one. The hook warns conspicuously whenever it is absent, and falls back to generic pattern checks only.

Collectors set a declared User-Agent with contact details from SPONSOR_MAP_USER_AGENT; SEC EDGAR and some other hosts require one.

Pull requests that fail validation are not mergeable. There is no "cite it later".

Citing this work

Cite a fact to its own source, not to this dataset. Every record carries a public_citation naming the primary document it came from — a regulator's order, an institution's disclosure, a filing. If you are relying on a specific sponsor relationship, routing number or finding, cite that document. This dataset is the path you took to it, not the authority for it.

Cite the dataset itself when referring to the collection, its method, or a measurement derived from it (the disclosure rates, the custodial-structure prevalences, the registered-card condition):

Parra, D. (2026). Sponsor Bank / Fintech Program Map: a provenance-enforced dataset of US sponsor-bank and fintech-program relationships. CrimsonVector. https://github.com/crimsonvector/sponsor-map — contact info@crimsonvector.com

@misc{sponsor_map_2026,
  author       = {Parra, Diego},
  title        = {Sponsor Bank / Fintech Program Map: a provenance-enforced
                  dataset of US sponsor-bank and fintech-program relationships},
  year         = {2026},
  organization = {CrimsonVector},
  howpublished = {\url{https://github.com/crimsonvector/sponsor-map}},
  note         = {Contact: info@crimsonvector.com}
}

Replace the URL with the repository's actual location if it differs. Records are snapshots of sources that change: cite the retrieved date on the record you rely on, and where a record carries a document_sha256, quote it so a reader can confirm they hold the same document.

License

Dataset and repository: Community Data License Agreement – Permissive – Version 2.0 (see LICENSE). CDLA-Permissive-2.0 imposes no share-alike obligation on derived data or results, so the dataset can be vendored and extended downstream, publicly or privately.

About

Provenance-enforced open dataset mapping US chartered sponsor banks to the fintech programs that ride on them: sponsorships, corporate control, middleware, enforcement actions and custodial structure. Every record carries a resolvable public citation.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages