Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
158 changes: 136 additions & 22 deletions README.md

Large diffs are not rendered by default.

158 changes: 136 additions & 22 deletions README_CN.md

Large diffs are not rendered by default.

92 changes: 92 additions & 0 deletions RESOURCES.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,92 @@
# RESOURCES — What This Project Needs, Priced

Two things stand between the current repository and a validated instrument:
**cleaned data** and **training compute**. This file prices both. It is
simultaneously the project's internal shopping list and — for a potential
partner or investor — the use-of-funds statement. Experiment-level detail
for the data lives in [DATA_REQUIREMENTS.md](DATA_REQUIREMENTS.md); the
Phase 2 architecture that consumes the compute lives in
[docs/PHASE2_NEURAL_GAME.md](docs/PHASE2_NEURAL_GAME.md).

---

## 1 · Pillar one: cleaned, *typed* data

The framework does not want "more data" — it wants **each agent type's
behavior observed at the level that type actually acts**. That is a
different shopping list from a factor shop's, organized by which term of
the model each stream feeds:

| Agent type / model term | Data | Source | Cost tier |
|---|---|---|---|
| Physical vs behavioral noise (Thm 1) | Intraday trades & quotes, 500 equities × 10 yr | WRDS TAQ / Polygon | academic access / ~$2k·yr |
| Event operators (all 22) | Corporate actions, M&A, IPO, delistings, earnings | CRSP + SDC + I/B/E/S | academic access |
| Institutional mean field μ* (L1–L2) | Quarterly 13F holdings, futures COT positioning | EDGAR (free) + CFTC (free) | free |
| Retail cohort policies (Day 16) | **The E5 LLM query atlas** — stratified audit of consumer AI investment advice; the dataset exists nowhere and we create it | consumer LLM APIs | ~$200 |
| Funding / stress channel (Λₜ) | FRED spreads, FINRA margin, OFR indices, CBOE | public | free |
| Cross-market L0 | TIC flows, dollar indices | public | free |
| Alt-data layer (later) | News embeddings, positioning surveys | vendor | deferred |

**The punchline stays the same as DATA_REQUIREMENTS.md:** experiments
E1 + E2 + E4 + E5 — enough for the first real-data paper — are executable
in one quarter by one person inside any research group with WRDS access
plus about **$200** of API budget. The scarce resource is access, not money.

### E7 — the real-data denoised-price validation (To-C line)

The synthetic concept demo ([`demo/denoised_price_2026.py`](demo/denoised_price_2026.py))
graduates to a real-data experiment:

> **E7.** Reconstruct the denoised equilibrium track P^eq for the memory
> sector through the July 2026 unwind (SOX −19%, worst month since 2008)
> using only point-in-time data: prices (CRSP/Polygon), institutional
> positioning (13F + COT), retail-flow proxies, and the E5 retail-AI
> kernel. Measure: did divergence D_t cross threshold, with the
> institutional field rotating out, materially before July 24? Deliverable:
> a walk-forward figure exactly like the 2008 hindcast — same honesty
> rules, no look-ahead.

Data unlock: same as E1–E3 (nothing new to buy). Priority: **P0 for the
product line** — this is the first exhibit any retail-facing partner will
ask for.

## 2 · Pillar two: training compute

Phase 1 (everything in this repo) runs on a laptop; that was the point.
Phase 2 — the [Neural Network Game Structure](docs/PHASE2_NEURAL_GAME.md),
where every neuron is itself a small network — is where the compute bill
arrives. Costs below are honest ranges at mid-2026 cloud prices, not
precision estimates:

| Stage | What runs | Hardware | Est. cost |
|---|---|---|---|
| Phase 1 (today) | HJB–FPK solvers, demos, 50 tests | laptop / free Colab | ~$0 |
| Real-data Phase 1 (E1–E7) | Encoder training on real panel, walk-forward backtests | 1× consumer GPU or A100 spot | $1–3k |
| Phase 2 pilots (E8–E10) | Architecture ablations on synthetic market, ~10² agent-networks | 1× H100/H200 | $5–15k |
| Phase 2 full train (E11) | L1+L2 NNGS: ~10³ neuron-networks × 10⁵–10⁶ params, adversarial co-training + reflexivity loop | **8× H200 node, weeks-scale runs** | $100–250k per campaign |
| Type 2 sandbox (Horizon 3+) | Agent-level "capitalism simulator" disciplined by the Type 1 equilibrium | multi-node cluster | deferred until E11 says it's earned |

Why it squares: each unit's learning signal depends on every other unit's
current policy (non-stationary co-training), and the units themselves are
models — see the complexity-wall section of the
[Phase 2 design doc](docs/PHASE2_NEURAL_GAME.md#5--the-complexity-wall--and-the-four-tools-against-it)
for the four reductions (mean-field factorization, latent embeddings,
hierarchy-as-curriculum, sparse strategic attention) that keep the bill in
five figures for pilots rather than seven.

## 3 · Use of funds, by scenario

| Scenario | Budget | Buys |
|---|---|---|
| **Bootstrap** (status quo) | ~$500 | E5 atlas ($200) + Polygon starter + misc. Everything else free/academic. |
| **Seed research grant** | ~$25k | All of the above + real-data Phase 1 (E1–E7) + Phase 2 pilots (E8–E10) on rented H100/H200. Output: the NeurIPS-workshop paper *and* the E7 product exhibit. |
| **Partner / pre-seed** | ~$300k | One full NNGS training campaign (E11) + one year of data subscriptions + paper-trading infrastructure live (Airflow + Alpaca, already scaffolded in [`online/`](online/)). |

No headcount is priced in: the maintainer cost of this project is one
student who refuses to stop.

---

*Offering access, compute, or capital: see
[Partnerships & Contact](README.md#partnerships--contact) — or open an
issue.*
205 changes: 205 additions & 0 deletions demo/denoised_price_2026.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,205 @@
"""
Denoised Equilibrium Price · 2026 — the retail product line, as a concept demo
on SYNTHETIC data, qualitatively calibrated to the July 2026 memory-sector
unwind (SOX −19% in July 2026, its worst month since 2008).

⚠️ THIS IS A SYNTHETIC SCENARIO, NOT A HINDCAST ON REAL DATA AND NOT
INVESTMENT ADVICE. Every path below is generated by the dual-noise +
mean-field machinery of this repo. The July 2026 episode supplies the
*narrative shape* (institutional crowding → LLM-homogenized retail
chase → fundamental re-rating → forced unwind with retail capitulating
at the bottom while institutions re-enter); the numbers are simulated.
The real-data version of this demo is experiment E7 in RESOURCES.md.

The product thesis (To-C line):
A retail investor cannot out-trade the behavioral noise ν_η — Theorem 1
(dual Cramér-Rao) says nobody can. But the *equilibrium* component of
the price — what the asset is worth once every agent has played its
rational strategy — is exactly what the MFG solves for. Publish that
denoised track and the divergence D_t = P_t/P_t^eq − 1 becomes a
mid/long-horizon positioning signal: when the market trades 25% above
its own equilibrium while the institutional mean field is already
rotating out, you do not need to call the day of the crash — you need
to not be the one buying the top.

Mechanics (same objects as the rest of the repo):
P_t^eq — equilibrium price: the MFG-consistent fundamental track.
Steps down once, on the guidance shock (a Mode-I event
operator: genuine information, Theorem 3 irreversibility).
κ_t — institutional crowding (Level-1/2 mean-field concentration).
f_t — retail flow: momentum-chasing, LLM-homogenized (Day 16).
w_t — the behavioral wedge: dw = a·(κ + b·f⁺) dt − c·w dt + jumps.
Market price P_t = P_t^eq · (1 + w_t).
Signal — U_t fires when D_t > θ_D while the smoothed κ-trend has
turned negative for ≥3 consecutive sessions: the wedge is
large AND the agents holding it up are leaving.

Run: python demo/denoised_price_2026.py → figures/denoised_price_demo.png
"""
import os
import numpy as np
import pandas as pd
import matplotlib
matplotlib.use("Agg")
import matplotlib.pyplot as plt

HERE = os.path.dirname(os.path.abspath(__file__))
FIGS = os.path.join(os.path.dirname(HERE), "figures")

rng = np.random.default_rng(20260729)

# ── calendar ───────────────────────────────────────────────────────────────────
dates = pd.bdate_range(end="2026-07-29", periods=252)
n = len(dates)
t = np.arange(n)

GUIDANCE = pd.Timestamp("2026-07-13") # HBM4-delay / guidance shock (Mode-I event)
UNWIND = pd.Timestamp("2026-07-24") # unwind accelerates (Korea selloff spillover)
CAPITUL = pd.Timestamp("2026-07-28") # capitulation session
i_guid = dates.get_loc(GUIDANCE)
i_cap = dates.get_loc(CAPITUL)

# ── equilibrium (denoised) price ───────────────────────────────────────────────
# Fundamental drift of an AI-memory franchise: strong but not manic.
mu_eq = 0.00055
eq_shock = np.zeros(n)
eq_shock[i_guid:] = np.log(1 - 0.06) # −6% genuine re-rating at guidance
log_eq = np.log(100.0) + mu_eq * t + eq_shock \
+ np.cumsum(rng.normal(0, 0.0035, n)) # physical noise σ_τ only
P_eq = np.exp(log_eq)

# ── institutional crowding κ_t (mean-field concentration, Level 1–2) ───────────
# Logistic build-up through the AI-memory trade; institutions begin rotating
# out a few sessions BEFORE the guidance shock (they see the channel checks),
# and de-crowd hard after it.
kappa = 0.15 + 0.75 / (1 + np.exp(-(t - n * 0.55) / 18))
i_rot = i_guid - 4
decay = np.zeros(n)
decay[i_rot:] = np.linspace(0, 1, n - i_rot) ** 1.35
kappa = kappa - 0.55 * decay * kappa
kappa = np.clip(kappa + rng.normal(0, 0.008, n), 0.05, 0.95)
# Post-capitulation: institutions buy the dip (the AlphaGBM observation)
kappa[i_cap:] += np.linspace(0, 0.05, n - i_cap)

# ── retail flow f_t (LLM-homogenized momentum chase, Day 16) ───────────────────
mom = pd.Series(log_eq).diff(20).fillna(0).to_numpy()
h = 0.85 # homogenization coefficient
f = h * np.tanh(35 * mom) + rng.normal(0, 0.06, n)
f[i_guid:i_cap] += 0.25 # retail buys the first dip…
f[i_cap:] = -0.9 + rng.normal(0, 0.05, n - i_cap) # …then capitulates at the low

# ── behavioral wedge w_t ───────────────────────────────────────────────────────
w = np.zeros(n)
a, b, c = 0.0030, 0.55, 0.0105
for k in range(1, n):
drive = a * (kappa[k] + b * max(f[k], 0.0))
jump = 0.0
if rng.uniform() < 0.04 + 0.10 * w[k - 1]: # jumps cluster with crowding
jump = rng.choice([-1, 1], p=[0.35, 0.65]) * rng.exponential(0.006)
if k >= i_cap - 2: # forced unwind: 4 sessions
drive, c_eff = -0.055, 0.28
jump = -abs(rng.normal(0.012, 0.006))
else:
c_eff = c
w[k] = w[k - 1] + drive - c_eff * w[k - 1] + jump
w = np.clip(w, -0.08, 0.40)

P = P_eq * (1 + w)
D = P / P_eq - 1 # divergence

# ── the signal U_t ─────────────────────────────────────────────────────────────
THETA_D = 0.20
kap_trend = pd.Series(kappa).ewm(span=10).mean().diff()
falling = (kap_trend < 0).rolling(3).sum() == 3
armed = (pd.Series(D) > THETA_D) & falling.fillna(False)
sig_idx = armed[armed & (armed.index >= i_guid - 10)].index.min()
sig_date = dates[sig_idx]
lead_unwind = int(((dates > sig_date) & (dates <= UNWIND)).sum())
lead = int(((dates > sig_date) & (dates <= CAPITUL)).sum())

peak = P.max()
trough = P[i_cap - 2:].min()

print("=" * 70)
print("SYNTHETIC CONCEPT DEMO — denoised equilibrium price (To-C product line)")
print(f"scenario calendar : {dates[0].date()} → {dates[-1].date()} ({n} sessions)")
print(f"guidance shock : {GUIDANCE.date()} (Mode-I operator, eq −6%)")
print(f"signal U_t fires : {sig_date.date()} — D = {D[sig_idx]*100:+.1f}% above "
f"equilibrium, κ-trend negative 3 sessions")
print(f"lead time : {lead_unwind} trading days (≈2 weeks) before the "
f"{UNWIND.date()} unwind acceleration,")
print(f" {lead} trading days before the {CAPITUL.date()} capitulation")
print(f"peak → trough : {(trough/peak - 1)*100:.0f}% "
f"(market returns to its denoised track)")
print(f"at the trough : retail flow f = {f[i_cap+1]:+.2f} (panic sell) | "
f"institutional κ rising (buying the dip)")
print("=" * 70)

# ── figure ─────────────────────────────────────────────────────────────────────
OK = dict(blue="#0072B2", orange="#E69F00", red="#D55E00", ink="#111827",
mute="#6B7280", grid="#E5E7EB")
plt.rcParams.update({"font.family": "sans-serif", "font.size": 10,
"axes.spines.top": False, "axes.spines.right": False,
"figure.facecolor": "white"})
fig, axes = plt.subplots(2, 1, figsize=(13.2, 7.6), sharex=True,
gridspec_kw=dict(height_ratios=[1.15, 1.0], hspace=0.10))
lo, hi = dates[0], dates[-1]

ax = axes[0]
ax.plot(dates, P, color=OK["ink"], lw=1.5, label="market price $P_t$ (equilibrium + behavioral wedge)")
ax.plot(dates, P_eq, color=OK["blue"], lw=1.7,
label="denoised equilibrium price $P_t^{eq}$ (the product)")
ax.fill_between(dates, P_eq, P, where=P > P_eq, color=OK["orange"], alpha=0.25,
label="behavioral wedge $w_t$ (crowding, Thm 1)")
ax.axvline(sig_date, color=OK["orange"], lw=1.5, ls="--")
ax.axvline(UNWIND, color=OK["red"], lw=1.0, ls=":")
ax.axvline(CAPITUL, color=OK["red"], lw=1.4, ls="--")
ax.annotate(f"{sig_date.strftime('%b %d')} — signal: price {D[sig_idx]*100:+.0f}% above\n"
f"equilibrium while institutions rotate out.\n"
f"{lead_unwind} trading days (≈2 weeks) before the\n"
f"unwind accelerates on Jul 24.",
xy=(sig_date, P[sig_idx]), xytext=(dates[int(n*0.35)], peak * 0.915),
fontsize=9.5, fontweight="bold", color="#B45309",
arrowprops=dict(arrowstyle="->", color="#B45309", lw=1.2))
ax.annotate("Jul 28 — capitulation:\nretail panic-sells at the low,\n"
"institutions buy; price re-joins\nits denoised track.",
xy=(CAPITUL, trough), xytext=(dates[int(n*0.46)], P_eq[0] * 0.985),
fontsize=9.5, fontweight="bold", color=OK["red"],
arrowprops=dict(arrowstyle="->", color=OK["red"], lw=1.2,
connectionstyle="arc3,rad=0.15"))
ax.axvline(GUIDANCE, color=OK["mute"], lw=0.9, ls=":")
ax.text(GUIDANCE - pd.Timedelta(days=3), 119.5, "guidance shock\nJul 13 (Mode-I)",
fontsize=7.8, color=OK["mute"], ha="right")
ax.set_ylabel("price (indexed)")
ax.legend(loc="upper left", fontsize=8.5, frameon=False)
ax.set_title("Denoised equilibrium price — SYNTHETIC concept demo of the To-C product line "
"(scenario shaped on the July 2026 memory unwind)",
fontsize=12.5, fontweight="bold", loc="left", pad=10)
ax.grid(alpha=0.3)

ax = axes[1]
ax.plot(dates, D * 100, color=OK["ink"], lw=1.5, label="divergence $D_t = P_t/P_t^{eq}-1$")
ax.axhline(THETA_D * 100, color=OK["red"], ls=":", lw=1.3)
ax.text(dates[int(n*0.42)], THETA_D * 100 + 1.2, f"θ_D = {THETA_D*100:.0f}%",
fontsize=8.5, color=OK["red"])
ax.fill_between(dates, THETA_D * 100, np.clip(D * 100, THETA_D * 100, None),
color=OK["orange"], alpha=0.45)
ax2 = ax.twinx()
ax2.plot(dates, kappa, color=OK["blue"], lw=1.3, alpha=0.85)
ax2.set_ylabel("institutional crowding κ_t", color=OK["blue"])
ax2.tick_params(axis="y", labelcolor=OK["blue"])
ax2.spines.top.set_visible(False)
ax.axvline(sig_date, color=OK["orange"], lw=1.5, ls="--")
ax.axvline(CAPITUL, color=OK["red"], lw=1.4, ls="--")
ax.set_ylabel("divergence (%)")
ax.set_xlim(lo, hi)
ax.legend(loc="upper left", fontsize=8.5, frameon=False)
ax.grid(alpha=0.3)

fig.text(0.01, 0.005,
"SYNTHETIC DATA — a concept demo of the denoised-price product, not investment advice and not "
"a real-data hindcast. Signal U_t: D_t > θ_D while the EMA(10) κ-trend is negative 3 consecutive "
"sessions. Real-data version: experiment E7 (RESOURCES.md).",
fontsize=7.5, color=OK["mute"])
fig.savefig(os.path.join(FIGS, "denoised_price_demo.png"), dpi=200, bbox_inches="tight")
print("✓ figures/denoised_price_demo.png")
Loading
Loading