Skip to content

Latest commit

 

History

26 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

trading-channel-dataset

A research dataset of crypto trading signals published by the public Telegram channel WallstreetQueenOfficial, together with the tooling that scraped the posts and converted them into structured JSON with an LLM.

The dataset covers 7 May 2021 – 30 September 2026 (65 monthly folders):

Raw channel posts 9,189
LLM verdicts 921 (521 long, 294 short, 106 wait)
Tickers with full 1m-candle coverage 82

What it is for

The dataset is the input for backtesting a signal-following strategy and for calibrating its exits. The idea, described in detail in ARCHITECTURE.md, is to take from the channel author only the direction and the moment of entry, and to replace the author's own targets, stop-loss and leverage with exits calibrated on how the channel's signals actually behaved:

  1. Every signal opens a position that is held untouched for 24 hours, and a per-minute net PnL trajectory is recorded.
  2. Trajectories are split into winners and losers, and their separability is checked (go/no-go).
  3. Exit thresholds — hard stop, per-minute PnL floor, profit lock, trailing take — are fitted as low-parameter curves and validated by replay against a "hold until the horizon, with liquidation" baseline.

The backtest and calibration themselves run in a separate project built on backtest-kit. This repository holds only the data, the extraction pipeline and the design document.

Repository layout

ARCHITECTURE.md            Exit-calibration plan: model, procedures, worked numeric example
assets/target_tickers.txt  82 *USDT symbols with gap-free Binance spot candles over the whole range
content/<mon>_<year>/      One folder per month, may_2021 … sep_2026
  assets/messages.jsonl      Raw channel posts for that month
  assets/signals.jsonl       LLM-extracted signals for that month
  prompt.mustache            Extraction prompt, tuned to the post format of that period
  generate-signals.mjs       messages.jsonl -> signals.jsonl generator
scripts/                   Scraping, validation and batch-run helpers

The channel's post format changed over the years, so each month carries its own prompt.mustache with examples and formatting rules for that period.

Data format

Both files are JSON Lines and are stored in Git LFS (git lfs pull after cloning; content/ is about 420 MB).

messages.jsonl

One line per channel post:

Field Description
id Telegram message id
content Post text
channel Channel name
date Publication time, ISO 8601 UTC
photo Attached image as base64 JPEG, when present

Because of the embedded photos, read these files line by line as a stream rather than loading them whole.

signals.jsonl

One line per (post, ticker) pair that passed the pre-filter:

{
  "messageId": 6590,
  "symbol": "CFXUSDT",
  "entry": {
    "id": 6590,
    "symbol": "CFXUSDT",
    "position": "short",
    "entryRange": { "from": 0.468, "to": 0.475 },
    "targets": [0.455, 0.445, 0.435, 0.425, 0.405, 0.385, 0.368, 0.35],
    "stoploss": 0.484,
    "reasoning": "The message provides a clear Short Set-Up for CFX/USDT ...",
    "_context": { "model": "gemma4:31b-cloud" }
  },
  "url": "https://t.me/WallstreetQueenOfficial/6590",
  "message": { "id": 6590, "content": "...", "channel": "WallstreetQueenOfficial", "date": "2024-04-01T09:52:46.000Z" }
}
  • entry.position is long, short or wait. All verdicts are stored, including wait (result reports, news, analysis), so consumers must filter on entry.position.
  • Missing values are encoded as zeros: entryRange of 0/0 means a market entry with no stated range, stoploss: 0 means no stop was given.
  • message is the source post without the photo.

jan_2023 has no signals.jsonl: no post in that month passed the signal filter.

How signals are generated

generate-signals.mjs in each month folder:

  1. Streams assets/messages.jsonl, dropping photos.
  2. Keeps posts that look like a signal: an entry keyword (entry, buy zone, market price, …) and a levels keyword (target, take profit, stop).
  3. Extracts candidate tickers from hashtags (#ADA, #ADA/USDT, #ADAUSDT → ADAUSDT).
  4. For each (post, ticker) pair, asks the LLM (gemma4:31b-cloud via Ollama, through json-inference) to fill a strict JSON schema, using the month's prompt.mustache.
  5. Appends the verdict to assets/signals.jsonl.

The script is resumable: pairs already present in signals.jsonl are skipped.

Usage

Requires Node.js, Git LFS and access to an Ollama endpoint serving the model.

git lfs pull
npm install

# one month
node content/apr_2024/generate-signals.mjs

# all 65 months, chronologically
./scripts/linux/generate_test_all.sh

# or in 10 shards that can run in parallel
./scripts/linux/generate_test_1.sh   # … generate_test_10.sh

Windows equivalents are in scripts/win/*.bat.

Helper scripts

Script Purpose
scripts/fetch_messages.mjs Scrape the channel day by day into content/<month>/assets/messages.jsonl (needs a Telegram session for telegram-reader)
scripts/read_posts.mjs Print recent posts or specific post ids, to study the author's format
scripts/fetch_photos.mjs Save post photos as JPEG files for inspection
scripts/check_filter.mjs Find signal-like posts that the keyword pre-filter drops
scripts/validate_prompt.mjs Run the prompt and schema against known posts (signal, result report, news)
scripts/read_months.mjs Collect all tickers mentioned in signals and check them against Binance spot
scripts/check_candles.mjs Check candle coverage of each ticker over the whole backtest range
scripts/read_cache.mjs Inspect the LLM cache that backtest-kit keeps in MongoDB

Several helpers were carried over from the parent backtest project and expect its environment: ccxt and mongoose are not listed in package.json, check_candles.mjs reads config/loader.config.ts, and the Telegram scripts need a session.txt.

About

A research dataset of crypto trading signals published by the public Telegram channel WallstreetQueenOfficial, together with the tooling that scraped the posts and converted them into structured JSON with an LLM.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages