A research dataset of crypto trading signals published by the public Telegram channel WallstreetQueenOfficial, together with the tooling that scraped the posts and converted them into structured JSON with an LLM.
The dataset covers 7 May 2021 – 30 September 2026 (65 monthly folders):
| Raw channel posts | 9,189 |
| LLM verdicts | 921 (521 long, 294 short, 106 wait) |
| Tickers with full 1m-candle coverage | 82 |
The dataset is the input for backtesting a signal-following strategy and for calibrating its exits. The idea, described in detail in ARCHITECTURE.md, is to take from the channel author only the direction and the moment of entry, and to replace the author's own targets, stop-loss and leverage with exits calibrated on how the channel's signals actually behaved:
- Every signal opens a position that is held untouched for 24 hours, and a per-minute net PnL trajectory is recorded.
- Trajectories are split into winners and losers, and their separability is checked (go/no-go).
- Exit thresholds — hard stop, per-minute PnL floor, profit lock, trailing take — are fitted as low-parameter curves and validated by replay against a "hold until the horizon, with liquidation" baseline.
The backtest and calibration themselves run in a separate project built on backtest-kit. This repository holds only the data, the extraction pipeline and the design document.
ARCHITECTURE.md Exit-calibration plan: model, procedures, worked numeric example
assets/target_tickers.txt 82 *USDT symbols with gap-free Binance spot candles over the whole range
content/<mon>_<year>/ One folder per month, may_2021 … sep_2026
assets/messages.jsonl Raw channel posts for that month
assets/signals.jsonl LLM-extracted signals for that month
prompt.mustache Extraction prompt, tuned to the post format of that period
generate-signals.mjs messages.jsonl -> signals.jsonl generator
scripts/ Scraping, validation and batch-run helpers
The channel's post format changed over the years, so each month carries its own prompt.mustache with examples and formatting rules for that period.
Both files are JSON Lines and are stored in Git LFS (git lfs pull after cloning; content/ is about 420 MB).
One line per channel post:
| Field | Description |
|---|---|
id |
Telegram message id |
content |
Post text |
channel |
Channel name |
date |
Publication time, ISO 8601 UTC |
photo |
Attached image as base64 JPEG, when present |
Because of the embedded photos, read these files line by line as a stream rather than loading them whole.
One line per (post, ticker) pair that passed the pre-filter:
{
"messageId": 6590,
"symbol": "CFXUSDT",
"entry": {
"id": 6590,
"symbol": "CFXUSDT",
"position": "short",
"entryRange": { "from": 0.468, "to": 0.475 },
"targets": [0.455, 0.445, 0.435, 0.425, 0.405, 0.385, 0.368, 0.35],
"stoploss": 0.484,
"reasoning": "The message provides a clear Short Set-Up for CFX/USDT ...",
"_context": { "model": "gemma4:31b-cloud" }
},
"url": "https://t.me/WallstreetQueenOfficial/6590",
"message": { "id": 6590, "content": "...", "channel": "WallstreetQueenOfficial", "date": "2024-04-01T09:52:46.000Z" }
}entry.positionislong,shortorwait. All verdicts are stored, includingwait(result reports, news, analysis), so consumers must filter onentry.position.- Missing values are encoded as zeros:
entryRangeof0/0means a market entry with no stated range,stoploss: 0means no stop was given. messageis the source post without the photo.
jan_2023 has no signals.jsonl: no post in that month passed the signal filter.
generate-signals.mjs in each month folder:
- Streams
assets/messages.jsonl, dropping photos. - Keeps posts that look like a signal: an entry keyword (
entry,buy zone,market price, …) and a levels keyword (target,take profit,stop). - Extracts candidate tickers from hashtags (
#ADA,#ADA/USDT,#ADAUSDT→ADAUSDT). - For each (post, ticker) pair, asks the LLM (
gemma4:31b-cloudvia Ollama, throughjson-inference) to fill a strict JSON schema, using the month'sprompt.mustache. - Appends the verdict to
assets/signals.jsonl.
The script is resumable: pairs already present in signals.jsonl are skipped.
Requires Node.js, Git LFS and access to an Ollama endpoint serving the model.
git lfs pull
npm install
# one month
node content/apr_2024/generate-signals.mjs
# all 65 months, chronologically
./scripts/linux/generate_test_all.sh
# or in 10 shards that can run in parallel
./scripts/linux/generate_test_1.sh # … generate_test_10.shWindows equivalents are in scripts/win/*.bat.
| Script | Purpose |
|---|---|
scripts/fetch_messages.mjs |
Scrape the channel day by day into content/<month>/assets/messages.jsonl (needs a Telegram session for telegram-reader) |
scripts/read_posts.mjs |
Print recent posts or specific post ids, to study the author's format |
scripts/fetch_photos.mjs |
Save post photos as JPEG files for inspection |
scripts/check_filter.mjs |
Find signal-like posts that the keyword pre-filter drops |
scripts/validate_prompt.mjs |
Run the prompt and schema against known posts (signal, result report, news) |
scripts/read_months.mjs |
Collect all tickers mentioned in signals and check them against Binance spot |
scripts/check_candles.mjs |
Check candle coverage of each ticker over the whole backtest range |
scripts/read_cache.mjs |
Inspect the LLM cache that backtest-kit keeps in MongoDB |
Several helpers were carried over from the parent backtest project and expect its environment: ccxt and mongoose are not listed in package.json, check_candles.mjs reads config/loader.config.ts, and the Telegram scripts need a session.txt.