A crypto market database built to eliminate survivorship bias and selection bias from downstream backtests, plus the tools to build it from scratch. This repo produces marketdata.db, a SQLite file containing historical market cap rankings, token metadata, and daily OHLCV bars going back to 2009 for the oldest tokens.
This is NOT a trading tool. It is a data pipeline. Other repos like a RSPS System read from this database to run backtests and generate signals. This repo is the source of that data: it records which coins were actually in the top 100 by market cap on each historical date (including coins that later crashed, delisted, or disappeared), so the RSPS can use it to eliminate survivorship bias and selection bias from its own backtests.
Shoutout to:
- Ghe (Kung Fu Panda): for maintaining the
dietmarb01/tvdatafeedfork, a code-corrected version of the upstreamrongardF/tvdatafeedthat makes reliable OHLCV fetching possible. See his post on the topic.
If you backtest a momentum strategy on today's top 20 coins, the results will look amazing. That is because those coins are in the top 20 precisely because they went up. You are testing on winners and calling it a strategy.
This is called survivorship bias. The fix is to test on the coins that were actually in the top 100 on each historical date, including coins that have since crashed, been delisted, or disappeared entirely. A coin that was #5 in 2019 but no longer exists today still appears in this database for 2019.
The close cousin is selection bias: how you pick which coins to test on in the first place. A hand-rolled list of coins you know and like biases the sample toward your own intuition, not toward any neutral definition of "what was eligible at the time". The same fix works here: let the market pick the sample. Whoever was in the top 100 by market cap on a given date is in, no cherry picking.
That is what this pipeline builds. It pulls daily top 100 snapshots from CoinMarketCap going back to 2018, checks which of those tokens have tradeable data on TradingView, and downloads their full OHLCV history. Using this database as the sample source is what lets the RSPS eliminate both biases from its own backtests.
The database has 4 tables:
| Table | What it stores |
|---|---|
daily_top |
Daily snapshot of the top 100 coins by market cap (from CoinMarketCap), going back to 2018. Think of it as a daily leaderboard: which coins were #1, #2, ... #100 on each day. |
symbol_metadata |
Which coins are available on TradingView and what exchange they trade on (CRYPTO: vs INDEX:). |
ohlcv |
Daily OHLCV bars (Open, High, Low, Close, Volume) for every coin that ever appeared in the top 100. Same data you would see on a TradingView daily chart. |
market_totals |
Daily OHLCV for CRYPTOCAP:TOTALES, the total crypto market cap. Same as the TOTALES chart on TradingView. Populated by tradingview/totals_sync.py (part of the pipeline). |
We only include tokens available on the CRYPTO: exchange (plus BTCUSD and ETHUSD on INDEX:). CRYPTO: is TradingView's aggregated feed. Tokens only appear there if they are listed on enough major exchanges with sufficient liquidity to aggregate meaningfully. If a token only shows up as BINANCE:XYZUSDT but not CRYPTO:XYZUSD, it tells you:
- Narrow exchange presence: essentially only traded on one or very few venues
- Lower liquidity: not enough cross exchange volume to warrant aggregation
- Higher manipulation risk: thin order books, one dominant exchange controlling the price
- Higher delisting risk: these tokens come and go, CRYPTO: tokens tend to be more established
This acts as a built in quality filter for the database.
Don't worry about understanding every file. This section is just a reference. To run the whole pipeline at once, run setup.bat (Windows) or setup.sh (macOS / Linux). To run it step by step, see the "How to Rebuild the Database from Scratch" section further down. The [step N] markers in the tree tell you where each script fits in that flow.
marketdata/
│
├── marketdata.db The database itself (gitignored, build it locally via setup or rebuild)
├── setup.bat [RUN on Windows] One-shot setup (env + packages + optional full DB rebuild)
├── setup.sh [RUN on macOS / Linux] Same as setup.bat
├── update.bat [RUN on Windows] Daily incremental update (wrapper for update.py)
├── update.sh [RUN on macOS / Linux] Daily incremental update (wrapper for update.py)
├── update.py Runs the 4 pipeline steps in order to bring the DB to today
├── filters.py Rules for excluding stablecoins, wrapped tokens, etc.
├── rebrands_list.py Maps old ticker names to new ones (e.g. VEN to VET)
├── requirements.txt Python package list
│
├── setup/
│ └── init_db.py [step 1] Creates or verifies the 4 database tables
│
├── cmc_snapshots/ Daily top 100 ingestion
│ ├── cmc_config.py CoinMarketCap API settings (imported, not run directly)
│ ├── cmc_fetcher.py [step 2, option B] Fetches daily top 100 from the CoinMarketCap API
│ ├── cg_fetcher.py [step 2, option C] Same idea from CoinGecko (alternative source)
│ ├── check_tv_availability.py [step 3] Checks which coins exist on TradingView
│ └── cmc_historical_snapshots.zip Bundled historical CMC snapshot CSV (extract in place before use)
│
├── tradingview/
│ ├── ohlcv_backfill.py [step 4] Downloads full price history from TradingView
│ └── totals_sync.py [step 5] Syncs CRYPTOCAP:TOTALES into market_totals
│
├── historical_csv_data/ TradingView CSV exports for tokens beyond the 5000-bar limit (gitignored)
│
└── tools/
├── import_cmc_snapshots.py [step 2, option A] Imports the bundled CMC CSV into the database
├── import_csv_backfill.py [step 6, optional] Imports TradingView CSV exports for capped tokens
└── daily_universe.py Quick check: "what were the top 20 coins on date X?"
From the project root, run the one that matches your OS:
# Windows
setup.bat# macOS / Linux
chmod +x setup.sh # first time only
./setup.shBoth scripts do the same thing. They're idempotent and safe to re-run any time: every step either does nothing (if the work is already done) or picks up where the last run left off. Pass /yes (Windows) or --yes (macOS / Linux) to auto-approve every prompt, or /help / --help for the full flag list.
Both scripts run these 5 stages in order:
- Environment. If a conda env or venv is already active, it uses that. Otherwise it shows an interactive menu to pick conda, venv, or system Python (or auto-creates a
.venvunder/yes). You can skip the menu entirely by passing/create-conda <name>to go straight to creating that conda env, or/create-venv <path>to go straight to creating a venv at that path. - Python interpreter. Reports the resolved Python version, absolute interpreter path, and the active env. Sanity check that the previous stage produced a usable Python.
- Packages. Runs
pip install -r requirements.txt. Pip skips already-satisfied packages, so subsequent runs are fast. - Database. If
marketdata.dbexists, runssetup/init_db.pyto verify schemas. If it's missing, prints a summary of the 4-step ingestion pipeline, warns that steps 3 and 4 are rate-limited and can take many hours (both resumable), and asks to confirm. Onyit: creates the empty DB file, runsinit_db.pyto build the schema, extractscmc_historical_snapshots.zipif only the archive is present, fillsdaily_topfrom the CSV (or falls back to the CoinMarketCap API), probes TradingView for every symbol, and downloads daily OHLCV for each confirmed symbol. - Smoke test. Imports
filters.should_excludeandrebrands_list.REBRANDSto confirm the interpreter, packages, and top-level modules all load cleanly.
Finally it prints a cheat sheet of the individual script commands you'd run for updates or partial re-ingestion later.
If you prefer to do it manually, follow the two steps below (and then the Rebuild section further down, if marketdata.db is missing).
Python 3.10 or higher is needed. Both conda and venv are supported (pick whichever you already use). The goal is just to have an isolated environment active before installing packages.
Option A: conda (Anaconda or Miniconda)
conda create -n quant python=3.11
conda activate quantThis creates an isolated conda env named quant. Reactivate it (conda activate quant) every new terminal session before running anything.
Option B: venv (Python standard library, no conda needed)
# Windows
python -m venv .venv
.\.venv\Scripts\activate# macOS / Linux
python3 -m venv .venv
source .venv/bin/activateEither option leaves you with an isolated environment. setup.bat / setup.sh pick venv by default under /yes if you prefer to skip this step and let the script handle it.
From the project root folder, run:
pip install -r requirements.txtThis installs all the libraries the pipeline needs (pandas, requests, python-dotenv, tvDatafeed).
If you see errors, make sure you activated your conda env or venv first.
Note on tvDatafeed. The TradingView data scripts (cmc_snapshots/check_tv_availability.py and tradingview/ohlcv_backfill.py) use the tvDatafeed package, a third party library for downloading historical OHLCV from TradingView's websocket API (the fork used here is provided by Ghe (Kung Fu Panda), as mentioned above). This is already included in requirements.txt, so if you ran the command above it is already installed. If you need to install it separately for any reason, the command is:
pip install git+https://github.com/dietmarb01/tvdatafeed.git
If marketdata.db gets deleted or corrupted, follow these steps in order. setup.bat / setup.sh automate all of them when the DB file is missing.
init_db.py is intentionally paranoid: it refuses to run unless marketdata.db already exists. That safety keeps it from silently creating a phantom DB in the wrong directory. So start by creating an empty file and then run the script to build the schema:
type nul > marketdata.db
python setup/init_db.pyThis leaves you with marketdata.db containing all 4 empty tables (ohlcv, market_totals, symbol_metadata, daily_top).
This step populates the daily_top table. You have three options depending on what you have available.
Option A: From the bundled CSV (fastest)
The repo ships with cmc_snapshots/cmc_historical_snapshots.zip. Extract it in place (so the CSV lands right next to the zip, not in a subfolder), then run the importer:
cd cmc_snapshots
powershell Expand-Archive -Path cmc_historical_snapshots.zip -DestinationPath .
cd ..
python tools/import_cmc_snapshots.pyThis bulk loads every daily snapshot bundled in the CSV from 2018 onward into the daily_top table (top 50 per day, the legacy CMC export size). After this finishes, run cmc_fetcher.py (option B) to gap fill each day up to the current target of 100 entries. The CSV must sit at cmc_snapshots/cmc_historical_snapshots.csv, import_cmc_snapshots.py hardcodes that path.
Option B: From the CoinMarketCap API (slower, no CSV needed)
python cmc_snapshots/cmc_fetcher.pyThis fetches one day at a time from the CoinMarketCap API. It will take a while if you are fetching years of data, but it is resumable. You can stop and restart it.
Option C: From CoinGecko (alternative)
python cmc_snapshots/cg_fetcher.pySame idea, different data source. Requires a CoinGecko API key in a .env file.
python cmc_snapshots/check_tv_availability.pyThis goes through every coin that ever appeared in the top 100 and checks if TradingView has price data for it. Results are saved in the symbol_metadata table.
This is resumable: if you stop it, it picks up where it left off (tracked in tv_check_queue.txt). Delete that file to start over.
python tradingview/ohlcv_backfill.pyThis downloads daily OHLCV bars from TradingView for every coin confirmed in Step 3. It fetches up to 5000 bars per coin (roughly 13+ years of daily data).
This is also resumable: progress is saved in backfill_queue.txt. Delete that file to start over.
This step takes the longest. It fetches one coin at a time with delays between requests to avoid getting rate limited by TradingView.
python tradingview/totals_sync.pyPulls the latest CRYPTOCAP:TOTALES daily bars from TradingView into market_totals. This table holds the total crypto market cap and is what downstream strategies use as a market regime signal (e.g. trend filters in the RSPS).
This is incremental: each run picks up where the last one left off, so it stays cheap to re-run.
Some tokens have more than 5000 days of history, which is TradingView's API limit. The backfill script logs these to tradingview/max_bars_tokens.txt and creates a historical_csv_data/ folder.
To fill in the older data:
- Open each token's chart on TradingView (daily timeframe)
- Export the full CSV using the chart export button
- Drop the CSV files into
historical_csv_data/ - Run:
python tools/import_csv_backfill.pyTokens with a matching CSV are imported and removed from the txt file. Tokens without a CSV are left for later. When all tokens have been imported, the txt file is deleted.
After these steps, marketdata.db is fully populated and ready for use by other repos.
Once marketdata.db is built, run update.py (or its update.bat / update.sh wrappers) to bring it forward to today's date:
# Windows
update.bat# macOS / Linux
chmod +x update.sh # first time only
./update.shIt runs the four pipeline steps in order:
cmc_snapshots/cmc_fetcher.pyfills missing days indaily_top(top 100 per day, gap fills any day with fewer than 100 rows).cmc_snapshots/check_tv_availability.pyprobes TradingView for any new symbols and writes them tosymbol_metadata.tradingview/ohlcv_backfill.pydownloads daily OHLCV for every confirmed symbol up to the current date.tradingview/totals_sync.pypulls the latest CRYPTOCAP:TOTALES bars intomarket_totals.
Each step is idempotent and resumable, so re-running is safe. A failure in one step prints its traceback and the next step still runs, so a transient TradingView outage does not block the whole pipeline. The exit code is non-zero if any step failed.
Every run of update.py will show roughly 7 historical dates being re-fetched in step 1, for example 2021-09-28, 2023-11-23, 2023-12-06, and a few others. This is expected and safe to ignore. The CoinMarketCap historical API does not return the full 100 tokens for those dates (2021-09-28 only has 81 rows, the others are usually missing one specific rank). Our daily_top table therefore stays under 100 rows for those days, and cmc_fetcher.py re-queues any date with fewer than 100 rows. CMC has no more data to give, so nothing changes in the table. You can either leave it as is, or maintain a skip list of known incomplete dates that the fetcher ignores as long as the table already has rows for them.
Two files in the root handle data cleanup. They are meant to be imported by your RSPS system (or any downstream strategy repo), so the same exclusion and rebrand rules apply everywhere that consumes this database.
filters.py: Defines which coins to exclude from the universe. This covers several categories:
| Category | Examples | Why excluded |
|---|---|---|
| Stablecoins | USDT, USDC, DAI, BUSD | Pegged to fiat, no independent price action |
| Wrapped BTC/ETH | WBTC, STETH, CBETH | Mirror the underlying asset, not independent |
| Liquid staked tokens | MSOL, JITOSOL, SAVAX | Same as above, staking derivatives |
| Exchange tokens | BNB, OKB, FTT, LEO | Platform tokens, not pure crypto assets |
| Wrapped L1 tokens | WMATIC, WAVAX, WBNB | Wrapped versions of existing L1 tokens |
| Gold pegged | XAUT, PAXG | Commodity pegged, not crypto |
You would not want any of these in a momentum strategy. They either track something else or have artificial price stability.
rebrands_list.py: Some coins changed their ticker over the years. For example, VeChain went from VEN to VET, IOTA went from MIOTA to IOTA, and Nano went from NANO to XNO. This mapping ensures historical data lines up correctly so a coin is not double counted under its old and new name.
DM or tag shshs21 in IM💎 section.