Skip to content

finharness transplants: point-in-time integrity, honest parsing, sealed ledger and evaluation stack from FinanceHarness, StockAgent, TradingAgents and five papers - #5

Open
linroger wants to merge 458 commits into
mainfrom
feat/finharness-transplants
Open

linroger wants to merge 458 commits into
mainfrom
feat/finharness-transplants

Conversation

@linroger

@linroger linroger commented Oct 2, 2026

Copy link
Copy Markdown
Owner

What this is

This PR brings into DeepResearchForecast (DRF) what three codebases and five research papers do well. The codebases are FinanceHarness 0.1.0, StockAgent and TradingAgents 0.5.1. The papers are F²Agent, the FinanceHarness paper, Nexus, TIEM and the StockAgent paper. Every idea was checked against DRF's code before it was built.

The full write-up is docs/research/2026-09-29-finharness-essence.md. It covers the essence of each source and the roadmap. It also has the implementation record (wave by wave, with knobs and gate results), the open issues and the appendices. Start with its synthesis chapter.

How it was built

  • Study. Agents read 281 of the 282 source files and every page of the five papers. That produced 83 transplant candidates. Each candidate was mapped to the DRF file and function it would change, then checked three ways:
    • does the source really do this;
    • does DRF really lack it;
    • would it improve DRF's forecasts?
  • Plan. The plan has 100 work packages: 81 to implement and 19 deferred with recorded evidence.
  • Implementation. All 81 planned packages are merged, plus 12 follow-up packages (FU-1..FU-12) from triaging the open issues the merged packages reported.
    • Each package was built in its own worktree and reviewed adversarially until approved. Most needed one to three review rounds; REPORT-13 needed seven.
    • Packages were merged in waves, with the full offline suite run after each batch.

What changed, by area

  • Time and evidence integrity.
    • Hindcast admission and a pinned as_of policy.
    • Point-in-time evidence gates and a citation wall.
    • Source publication dates.
    • Market end-date hygiene and price-time provenance.
    • Official-data tools: FRED/ALFRED vintages and SEC EDGAR as-filed statements.
  • Research.
    • Verbatim evidence spans.
    • Typed quantitative rows with page verification.
    • A question spec.
    • Typed absence markers.
    • DERIVED findings, computed by a hardened Decimal evaluator.
    • A forecast-prompt context packer.
    • Fetch-layer shell detection.
  • Simulation.
    • Roster-bound decision validation.
    • Scheduled-event provenance.
    • Truthful period context.
    • A hollow-run gate.
    • A prior-echo diagnostic.
  • Report.
    • Honest probability parsing: an unreadable value becomes needs_review and is never coerced.
    • Narrative sync.
    • A labelled verified-figures block, with a shadow figure check.
    • A logic/number slot audit.
    • A tool-call argument boundary.
    • Structured binary targets.
    • A counter-case pass (off by default).
  • Evaluation.
    • Sealed, idempotent forecast-ledger commits and deterministic settlement.
    • Eval statistics, a golden set and contamination probes.
    • A per-stage scorecard and a cost card.
    • An eval bundle, a market-relative skill scorer and a label-free value-add study.
  • Platform.
    • Hermetic tests: an ambient-env scrub, a no-egress guard and FakeLLMClient parity.
    • LLM transport normalisation.
    • A strict config audit.
    • Identifier containment and local API hardening.
    • An AST import fence for egress.
    • JSON parsers that never raise on hostile model replies.

Review notes

  • Knobs and defaults. Every new behaviour sits behind a Config knob documented in .env.example: 145 new entries, plus 2 former ghost knobs now declared. At its default a knob leaves output byte-identical, unless its .env.example entry says the default is on. Most default-on knobs are fail-closed honesty fixes; a few are provenance or shadow diagnostics.
  • Opt-in features that change forecasts stay off. Examples: RESEARCH_DERIVED_FINDINGS, RESEARCH_V3_FORECAST_INPUTS and REPORT_COUNTER_CASE. Each needs a live check before its default flips.
  • No live runs. Nothing was run live: agents made no LLM or network calls, and every test uses FakeLLMClient.
  • Licensing.
    • No FinanceHarness or StockAgent code was copied; neither ships a licence.
    • Code adapted from TradingAgents (Apache-2.0) is attributed in NOTICE, with the licence in LICENSES/Apache-2.0.txt. It covers the FRED/EDGAR readers in deerflow_bridge/data_tools.py and the EDGAR test fixtures.
  • Docs. DRF_ARCHITECTURE.md §20 maps every mechanism to its module, and the README gains trust section 6. A final fact-check checked these docs against the code. It confirmed 30 errors, and all are corrected.
  • Where the diff is. The large diffs are in deerflow_bridge/linear_research.py, backend/app/services/report_agent.py and pipeline_orchestrator.py. About 86k of the ~148k added lines are tests.
  • Base. The branch is based on main at c20ec60 (the current main).

Testing

  • Final full gate (head dbc4435; the commits after it touch docs only):
    • backend python -m pytest -q -p no:cacheprovider: 12,410 passed, 0 failed, 53 skipped, 11 xfailed;
    • frontend node --test: 79/79.
  • Baseline before the program: 4,276 backend tests passed.
  • Gates along the way. Every merge batch was gated. Each failure was either fixed on the branch or traced to the environment; the gate table in the research document records each one.
  • Document anchors: 1,066 code anchors and 514 paper anchors validated.

Follow-ups (not in this PR)

The research document lists these under "Open issues after implementation". The main ones:

  • live checks before default flips, including one run of the official-data tools with real credentials;
  • one owner decision: the order between TIME-13's contradiction drop and RESEARCH-8's derivation;
  • some older defects, each needing its own package.

Roger Lin added 30 commits October 1, 2026 03:17
- linear_research._quant_sanity: reconcile_quantitative now runs once per
  scope (_quant_scopes / _quant_scope: the end of period_end, else the year
  of as_of_date; the canonical region or geography; reported vs projected by
  epistemic_class when typed, else value_type).  v3 rows keep their period,
  geography and value type outside the metric, so grouping by (metric, unit)
  alone turned forecast trajectories, series over time and different regions
  into contested.json claims (6 of 6 reconciled claims on the four real v3
  handoffs were such; scoped: 0, while the legacy-shaped poll disagreements
  still reconcile).  Reconciled claims and unit warnings name their scope
  (claim suffix, warning "scope"); caps are applied after concatenation, and
  reconcile stops at its first failed or malformed result (one
  analytics_errors entry).  The bridge helper and the legacy engine are
  unchanged.
- The future-dated bound for flag_implausible_quant is a pinned run's as-of
  itself (RESEARCH_AS_OF: no midnight to cross, a later publication is a
  leak); live runs keep the day after the plan's as-of.
- meta.quant_sanity_truncated records the pre-cap totals of any capped
  sanity list (unit warnings, reconciled claims, implausible flags).
- The extract-only salvage of a v3 handoff drops the v3 quant sanity keys
  (_SALVAGE_VOLATILE_META_KEYS), which describe the files it rewrites.
- config.py / .env.example text updated to match.

Tests: test_research_quant_parity_v3.py (scoped v3 run: trajectory, series,
two regions and actual-vs-forecast add nothing while the same-quarter 1000x
pair reconciles; _quant_scope keys and labels; per-scope stub caps and
truncation totals; reconcile failure mid-scopes; pinned hindcast and
pinned-today bounds; salvage drops the sanity keys).
Numeric guard parser: negated events (not/n't/never/fails to/no longer/
cannot/without/unless; 未/不会/没有/从未/不再...) invert the comparator and a
negated touch only infers the violated side; ASCII '~' is a range separator
only between tight numbers; a bare year never opens a range with a marked
amount (except a tight rising dash) and 'from A to B' / '从A至B' are
trajectories; decline verbs suppress only size-of-change readings, never
below/under/between levels; suffix comparators (以上/以下/or more/or less)
and at-or-above/below, 达到或超过; time spans are windows (touch), not levels;
number-letter identifiers (28nm) and scientific notation are read correctly.

Binding: research rows use the repo vocabulary (quant_typing.quant_class,
observed/value_kind), source_ref falls back to source; as_of/period_end
count as known only once the stated period ends, and hindcasts require a
readable as_of (an unreadable pin as_of records error instead of using
today).

Shadow binary draws get 64 extra max_tokens per row (min 10 rows); off keeps
4096. numeric_guard_mode joins the EVAL-18 config fingerprint. Knob docs
say the prompt addition can shift drafting. Rebased onto the integration
branch (INFRA-8 conflicts resolved keeping both sides).
- Post-hoc parity with the in-band dead-round carry (low #1): new pure
  decision_channel._carry_unreplayed_events moves the events of rounds absent
  from the action log into the next replayed round (own events first, copies
  labelled carried_from_round = source round; events past the last replayed
  round reach nothing, as in-band). run_decision_channel applies it when the
  new keyword carry_unreplayed_events is true; the runner's
  _posthoc_decision_events passes it only while SIM_PERIOD_CONTEXT_V2 is on
  (SIM-6's own gate). Default False keeps every existing call unchanged.
- In-band rounds delivered with an empty roster (low #5): _step_round keeps
  their decision-context events (labelled via _mark_carried) for the next
  elicitation, V2-gated; the digest summary is not repeated.
- Events renderer failure is logged before degrading (low #2).
- Carried label tests `is not None`, like world_delta (low #3).
- _event_pair collapses whitespace runs so dedupe, digest and rendering share
  one identity; _events_digest documents why order/labels stay out of the
  key, per spec (low #4).
- Knob docs in config.py and .env.example mention the carries.

Tests: whitespace dedupe + carried_from_round=0 label, logged renderer
failure, digest identity, pure carry helper, post-hoc carry in
run_decision_channel, runner helper kwargs under V2 on/off, post-hoc
dead-round parity via main's kwargs, in-band empty-roster carry.
…c binary guard [REPORT-11]

What and why (candidate C20, re-scoped to "measure first"): DRF pushes binary
sets hard away from 0.5 and runs an always-on red-team critique on the scenario
spine, but neither effect was measured and the uncritiqued spine was never
persisted. This change measures both and ships the symmetric guard held behind
a default-off knob.

- New pure module app/services/probability_shape.py: probability_shape()
  (prob-shape/v1) computes the scenario peak (max_probability, same name as the
  gate's quality.max_probability), normalized entropy, TV distance from uniform
  and the post-minus-pre critique delta; for binaries the 0.40-0.60 midband
  share (the scorecard's band), 0.45-0.55 near-half share, extreme share,
  mean |p-0.5|, stdev, a decile histogram and move counts toward/away from 0.5
  from market_influence and pre_reconciliation_probability stamps. Never raises;
  ignores bool/None/NaN/strings/out-of-range values.
- ReportAgent snapshots the spine's scenarios before the pre-prose critique and
  before the post-hoc critique; on a successful critique it stores them as
  quality.pre_critique_scenarios (quality dict copied, the uncritiqued spine and
  every critique/premortem prompt stay unchanged). _finalize_structured_forecast
  writes quality.forecast_policy when FORECAST_BINARY_SYMMETRIC_GUARD is on
  (independent of the telemetry knob) and quality.probability_shape when
  FORECAST_PROBABILITY_SHAPE is on, after the binary block and before
  forecast.json, so the final audit fingerprint covers both. No gate reads them.
- Ledger (critic amendment): ledger_commit.commit_report passes the SEALED
  forecast's quality.probability_shape as objective_signals to
  forecast_ledger.commit_published_forecast (new optional kwarg; stored as a
  copy, never joins commit id or target key; non-JSON signals are dropped with
  objective_signals_dropped rather than failing the commit). The legacy
  append_forecast path is byte-identical. New forecast_ledger.shape_summary()
  groups production rows by policy.binary_symmetric_guard with per-group
  n / mean / median of the key stats; tolerates rows without signals.
- CLI: forecast_tools.py shape --ledger-dir DIR | --forecasts FILE...
- Symmetric guard: _BINARY_SYMMETRIC_GUARD is appended right after the
  contrarian or low-p rule in every binary _draw prompt when the knob is on
  (it avoids the 'CONTRARIAN FRAMING' / '0.05-0.35 range' literals that test
  fakes route on). capture_safety_policy_v1 records
  forecast_binary_symmetric_guard; INFRA-9 forks inherit it. Idea from the
  TradingAgents (Apache-2.0) symmetric Hold rule; own wording, no code copied.
- drf2 forecast-report SKILL.md section 2.1 and its checklist item: removed the
  "gate will fail the report otherwise" cue; the numeric targets now describe a
  healthy table instead of being aims.

Knobs (.env.example documented): FORECAST_PROBABILITY_SHAPE=true (observability
only), FORECAST_BINARY_SYMMETRIC_GUARD=false (probability-moving policy, waits
for the WP14 outcome-blind promotion gate).

Tests: new test_probability_shape.py and test_binary_symmetric_guard.py;
test_forecast_ledger.py extended (commit rows, commit_report, shape_summary).
Knob-off proofs: binary prompts byte-identical with the guard off; forecast.json,
LLM calls and the ledger commit row unchanged with the shape knob off; legacy
ledger rows unchanged.
- rejection_excerpt fails closed: cut at Final Answer / blank line, start at
  the first call opener (else the first brace, else empty), keep only
  call-syntax tokens so prose after an unterminated block never reaches the
  ungated tool_rejected row.
- The per-section tool minimum counts dispatched calls only (ReAct and native);
  charged rejections still spend the max budget, and once they exhaust it the
  minimum is waived instead of looping (report_tool_args.evidence_floor_unmet).
- Unparseable blocks skipped beside a selected call: the first 3 get their own
  row and outcome count, the rest one summary row; the observation lists at
  most 3 reasons.
- A call after leading prose braces is found at the next top-level
  {"name":/{"tool": opener (bounded attempts, nested objects never taken).
- A bare reply that starts with a live-tool call but lacks its outer brace is
  repaired, or surfaced as args_not_json, instead of dropped.
- REPORT_TOOL_MAX_REJECTED_PER_SECTION comment and .env.example line state that
  it only applies with REPORT_TOOL_ARG_REPAIR on.
…zation and v3 QA [RESEARCH-9]

Real runs strip 55-94% of report citation markers before the read-only
final audit, and one run also deleted 912 sentences and machine-added 71
citations, yet final_audit.json showed a clean document.  This package
records that surgery as deterministic, read-only telemetry (candidate C33,
re-scoped to telemetry per the verifier corrections).

Report (REPORT_FINALIZATION_TELEMETRY, default true):
- New pure module app/services/citation_finalization_telemetry.py keeps a
  per-report log and builds the report-citation-finalization/1 record.
- _stabilize_publish_markdown starts the log (markers_before measured on the
  body without References before the first finalization); _finalize_citations
  adds the dangling repair counts and the semantic remaps/strips (kept and
  unverifiable are the last pass's state, since every pass re-checks the
  surviving markers); each pass records the FIRST, repairing
  _repair_final_quantitative_grounding call (the probe that follows reports
  zeros and overwrites totals) plus overuse strips, removed quotes and passes.
- _audit_final_published_markdown writes final_audit.json pre_audit_repairs
  and forecast.json quality.citation_finalization before serialization, so
  forecast_sha256 seals it: marker_strip_ratio and
  markers_lost_with_removed_text = before + machine-added - final - stripped.
  A WARNING is logged at a strip ratio >= 0.5.  No gate, integrity-issue or
  policy-version change; telemetry failures are logged and ignored.

v3 research (RESEARCH_V3_CITATION_STATS, default true, forwarded to the v3
child through RESEARCH_CHILD_V3_KNOBS):
- renumber_citations gains an optional dropped out-list (one ledger sid per
  removed occurrence); None keeps the bytes identical.
- _run_qa writes qa.json citation_stats (v3-citation-stats/1): markers before
  QA and published, orphan markers (occurrences, distinct, sample), stale
  groups, cited fetched vs snippet sources (as sources.json publishes them,
  via the extracted _published_page helper) and the snippet marker share,
  unused fetched pages, writer bibliographies and scaffold echo lines
  detected, and prose numbers no VERIFIED/REPORTED finding, cited fetched
  page or cited snippet traces.  Detection only.
- _qa_meta and _analytics mirror it into meta.research_qa and
  meta.research_quality; a resumed qa.json without the key still finalizes.

With either flag off no new key is written; research_report.md, sources.json,
full_report.md and every gate are byte-identical in every mode.

Tests: test_report_finalization_telemetry.py (persisted record, first-call
quantitative capture and the corrected balance formula, flag-off identity and
gate, degrade-safe failures, knob) and test_research_engine_v3_citation_stats.py
(renumber out-parameter, detectors, end-to-end engine run with an orphan
marker and a writer bibliography, flag-off byte identity, resume, knob).
…tching [INFRA-11]

Actor identity rested on a lossy key (NFKC + casefold, keep only [0-9a-z]
and U+3400-U+9FFF). Kana, hangul and Cyrillic names collapsed to an empty
key, so PREPARE raised "duplicate selected actor_id in context build" (or
"missing a stable name") for Japanese/Korean questions, names such as
トヨタ自動車 / ホンダ自動車 collided on 自動車, and every hangul or
Cyrillic role contract shared actor_e3b0c44298fc1c14. Report-side name
lookups guessed instead of failing closed: opinion_shift merged any agent
whose name contained the target ('US' matched Russia, Australia, Business
Roundtable), trace_cascade resolved 2-character containments, and the
entity resolver's alias map was last-writer-wins.

Identity (unflagged; Latin ids byte-identical, pinned by literals):
- utils/actors: legacy_actor_key (moved verbatim from actor_context),
  actor_key_is_lossy (a dropped letter outside the Latin script, spacing
  modifier letters excepted), actor_identity_key and stable_actor_id
  (legacy hash when lossless, else sha256('idk1\x1f' + normalize_name)).
- actor_context.actor_id_for delegates to stable_actor_id; relationship
  endpoints match on lossless identity keys, so a non-Latin actor no
  longer collects every relationship with a non-Latin endpoint.
- actor_role_prompt: one NFKC key with actor_context; role-contract ids
  derive from the raw name exactly like actor_id_for; _matches_actor and
  the pack-dict lookup compare normalize_name first and fall back to the
  legacy key only when neither side is lossy.

ACTOR_NAME_MATCH_STRICT (default true, documented in .env.example):
- opinion_shift resolves roster exact name/alias, exact agent name,
  roster >=4-char containment, agent >=4-char containment; several
  candidates return a message naming them instead of merging.
- _resolve_entity_name consults exact roster aliases, needs >=4-char
  containment and a unique node.
- actor_alias_map drops (and logs) aliases claimed by two actors; the
  graph pruner's core set keeps every alias (actor_alias_norms).
- ZepToolsService.actor_roster, set by ReportAgent from its actors.
Flag off = legacy matching (tests pin it).

Candidate: C37 (re-scoped to the identity-key hygiene fix).
Tests: backend/tests/test_actor_identity.py (39); affected suites
(actor role/context, provenance, entity resolver, KG, report wiring,
config audit) pass.
- Symmetric guard: append it outside the contrarian block, so with
  FORECAST_BINARY_CONTRARIAN off it follows the base RULES (which also
  push away from 0.5). Knob on now means every binary draw prompt carries
  the guard, matching quality.forecast_policy and the admission pin.
  Knob-off prompts stay byte-identical (contrarian on and off).
- probability_shape: an int too large for a float is ignored instead of
  collapsing the record; the catch-all fallback keeps the policy and adds
  error: True so a failure never reads as an empty forecast.
- shape_summary: _finite_number tolerates huge ints in ledger rows.
- forecast-report SKILL.md: §2.2, the §8 heading and checklist, and §9
  no longer present spread as a gate to pass; the spread line is an
  explicit diagnostic.
Latest-actual binding (llm_field and quant_row) now takes only a value that
states one figure: a trajectory ("2.6 → 13.0", "from 2.6% to 13.0%",
"从2.6%到13%") or two figures stays unbound instead of binding the oldest
figure (report_1c312b400d33 F1 was read as 2.6 against a 13.0 actual). A year
beside the figure is a date; a lone year-like value is the figure.

Non-finite numbers ("$1e999", 400 digits, unit scale overflow) are no
quantity, so forecast.json stays strict JSON; claim/quantity raw is capped.
Comparator pairing is linear (bisect + span sweep) instead of quadratic.
A bare "not" negates only after an auxiliary/modal outside a relative clause
and "without" only within one word of the comparator; "equals/meets or
exceeds" and "equal to or above/below" are inclusive. as_of dates keep the
day of "Sep 15, 2026" and read slash dates (later reading when ambiguous).

capture_safety_policy_v1 warns when it pins an invalid NUMERIC_GUARD_MODE as
shadow. The off-mode extractor test is pinned to digests recorded on the
pre-TIME-5 base (948a792).
- report audit: a re-audit that produces no record (flag off, no log, or the
  log of another report) drops forecast.json quality.citation_finalization, so
  an earlier publish's record is never re-sealed under forecast_sha256.  Flag
  off stays byte-identical: the key never existed before this package.
- citation_finalization_telemetry: coverage values that are not finite (or
  overflow float) read as None and counts that overflow int read as 0, so the
  record always serializes with allow_nan=False.
- v3 writer_bibliographies: linear in the text.  A block's scan stops at its
  first non-list line and resumes there (every opener before that line runs to
  the same heading and fails on it); counts are unchanged, pinned against the
  literal definition on 3,000 generated texts.
- v3 untraced numbers: scenario weights trace only as percentages ("%35"), so
  a bare "35 plants" sentence is still counted.
- v3 unused_fetched_sids reads fetched as sources.json publishes it
  (_published_page): an uncited extraction shell is no unused fetched page.

Tests: stale-record re-audit (flag off and foreign log), non-finite record
values, detector equivalence and a regex-call bound on 5,000 opener lines,
shell-aware unused fetched sids, percent-only weight tracing.
- probability_shape._round and forecast_ledger._distribution normalise -0.0 to 0.0
  (a degenerate 1/0 split's entropy; a small negative mean or a -0.0 median), so
  forecast.json, ledger objective_signals and the shape CLI never print -0.0.
- The module docstring states that normalized_entropy and tv_from_uniform are taken
  over the readable probabilities renormalised to sum 1, while the peak stays raw.
- New test pins that the pre-critique stamp comes after the critique and the
  pre-mortem: with REPORT_PREMORTEM off and on, the red-team prompts are identical
  with FORECAST_PROBABILITY_SHAPE on and off and never contain pre_critique_scenarios.
opinion_shift (ACTOR_NAME_MATCH_STRICT): containment is now weighed across
the roster and the action log at once. A unique roster containment no
longer hides an agent that also contains the target: 'Japan' against the
roster actor Bank of Japan and the unrostered agent Government of Japan
is reported as two candidates instead of silently returning the Bank of
Japan trajectory. A target resolved through an alias or a containment is
named in the output ("## 「Japan」(解析为「Bank of Japan」)逐轮行为轨迹"),
also in the not-found and injected-only messages; flag off is unchanged.

Canonical names outrank another actor's alias: actor_match_candidates
returns the row whose canonical name matches before alias-only rows, so
'China' resolves to the China actor even when the CCP lists 'China' as an
alias (it used to return an ambiguity the model could not resolve).
Agents are assigned to a roster actor by exact name (canonical first; an
alias two actors share names neither) before match_actor, so the
contested alias no longer pulls the China agent into the CCP trajectory,
and trace_cascade's roster-alias expansion skips contested aliases (CCP
no longer resolves to the China node).

actor_context relevance terms use the lossless key for kana/hangul/
Cyrillic text (legacy key for everything it keeps losslessly), so such
terms are no longer dropped as empty, de-duplicated, or equated with a
name sharing only its kanji. Latin packs are byte-identical; generic
term keys are precomputed and actor_key_is_lossy has an ASCII fast path.

The ACTOR_NAME_MATCH_STRICT comment and .env.example line now state that
unique containments under 4 characters (Fed -> Federal Reserve) stop
resolving unless the short name is an exact roster name or alias.
- NO-polarity criteria: a clause stating the NO outcome ('Resolves NO if the
  rate exceeds 3%', 'NO unless ...', 'No: ...') inverts the reading once more;
  the outcome word nearest before the reading governs it.
- Negators inside the metric: no negator but 'unless' negates right after a
  relative pronoun (contractions, never, cannot, no longer, fail to included);
  'never' after an auxiliary is read with it, and a bare 'never' before a past
  participle and a prepositional phrase is a reduced relative clause.
- Split fiscal years ('2026/27', 'FY2026-27', 'Q4 2026/27') end in their end
  year, so a hindcast never binds one before it has ended.
- status_quo_contradiction findings record horizon_gap_days (diagnostic only).
…e-dispatch validation, uncharged invalid calls
… module [TIME-10]

New deerflow_bridge/data_tools.py (stdlib plus a lazily imported httpx; never
imports the backend; never raises): the official-data vendor core and the
FRED/ALFRED macro series tool. It is inert until TIME-13 binds it to the
research engine, so every run is unchanged.

- DataResult (frozen), a pluggable Transport (httpx by default; an exception
  is reported by its type name only), _DiskCache (sha256 file names, atomic
  tmp + os.replace, corrupt/expired entries are misses, only ok fetches
  stored) and a process-level _Throttle (FRED 0.5 s).
- MACRO_ALIASES (35 aliases; no SP500/NASDAQCOM/BAML) and resolve_series,
  which rejects descriptive phrases and query-smuggling text before any call.
- vendor_today_chicago / fred_pit: min(as_of, FRED's America/Chicago today)
  compared as date objects; without tz data, UTC minus six hours (never later
  than the Chicago date).
- fred_series: both requests carry realtime_start == realtime_end == pit <=
  as_of (a request that would not is never sent); conservative status mapping
  (400 "does not exist" -> not_found today / no_vintage at a historical pin,
  never an unpinned retry; other 400, 429, 5xx, transport failure ->
  unavailable). Deterministic rendering whose units DRF's page-number parser
  verifies (4.3%, 159,000 thousand persons, 29,000.5 billion USD); the header
  and every support sentence carry source, series id and vintage; YoY and
  change lines are labelled derived and use true endpoints; model_text is
  capped at 2600 characters; supports in English or Chinese. The API key never
  reaches a result, URL, cache file or log (httpx's request log line is
  redacted).
- Knobs (env, documented in .env.example): DATA_TOOLS_CACHE_DIR,
  DATA_FRED_CACHE_TTL_H (6; a closed vintage is kept 30 days),
  DATA_TOOL_TIMEOUT_S (20), DATA_FRED_WINDOW_YEARS (10, clamped 1-40).
- Deployed by setup.sh's bridge loop and _DEPLOYED_BRIDGE_MODULES; the INFRA-13
  import-fence policy admits httpx in _httpx_transport only.
- Root NOTICE with the TradingAgents Apache-2.0 attribution (portions adapted
  from tradingagents/dataflows/vendors/fred.py 0.5.1).

Candidate: C06.
Tests: backend/tests/test_data_tools_fred.py (new, 97 offline tests with a
fake transport), test_deerflow_bridge_sync_guard.py (data_tools.py deployed by
both lists), test_import_fences.py.
…losed on truncated forecast JSON [INFRA-3]

Reasoning models (MiniMax / GLM / Kimi) can spend the whole max_tokens budget on
thinking and return empty content with finish_reason=length. chat() treated this
as a generic RuntimeError (three same-cap backoff retries, then failover), the
report preflight (max_tokens=64) reported a reachable provider as down, and a
truncated spine / binary / critique JSON reply was silently bracket-repaired
into forecast content.

- app/utils/llm_recovery.py (new, pure): next_max_tokens doubles the cap with a
  1024 floor, clamped to the provider ceiling and to context window - prompt -
  reserve; None when there is no headroom.
- llm_client: chat() calls the new instance method _chat_openai_escalating in
  place of _chat_openai. On EmptyCompletion(finish_reason='length') it meters
  the rejected attempt with the exception usage, checks the run budget
  (BudgetExceeded propagates and is never retried), and re-sends at once with
  the next cap, at most LLM_MAX_ESCALATIONS times per chat() call (the
  escalation state is shared across chat()'s transient retries). Exhausted:
  escalation_exhausted=True and chat() fails over instead of retrying. The
  ceiling is PROVIDER_META[provider]['max_output_tokens'] else
  LLM_MAX_TOKENS_CEILING; the window is Config.context_window_for(provider).
  An escalated reply is cached under the original key only on finish 'stop'.
- telemetry: LLMMeter.record_recovery(kind, outcome); snapshot gains
  'recovery' {kind: {outcome: n}} and 'recovery_by_stage' when non-empty.
- pipeline_orchestrator: the report preflight treats EmptyCompletion with
  finish_reason 'length' as a reachable provider (info log) instead of failing.
- forecast_extractor (LLM_JSON_TRUNCATION_FAIL_CLOSED): _reply_truncated reads
  last_call_meta() (method or dict) for finish 'length' or
  json_truncation_repaired. Truncated spine draws are discarded (quality
  llm_truncation.spine_draws_discarded; all discarded = failed spine), a
  truncated binary draw drops its last item (binary_quality
  llm_truncation_trimmed), critique / premortem return the input unchanged,
  and the post-hoc extractor flags quality.llm_truncation.
- Knobs (config.py + .env.example): LLM_LENGTH_ESCALATION=true,
  LLM_MAX_ESCALATIONS=2, LLM_MAX_TOKENS_CEILING=32768,
  LLM_JSON_TRUNCATION_FAIL_CLOSED=true. Off: legacy behaviour.

Candidate: C15.

Tests: tests/test_llm_recovery.py (schedule, clamps, purity) and
tests/test_llm_length_escalation.py (recovery + metering + recovery snapshot,
cache rule, provider/window clamps, exhausted -> single failover without
retries, BudgetExceeded stops escalation, transient resume, non-length not
escalated, flag-off legacy path, knob parsing and .env.example docs, report
preflight pass/fail through the real state machine, and the fail-closed
forecast paths with method-style metadata fakes plus a real LLMClient on a
stubbed transport). The orchestrator test harness gains a
report_preflight_error parameter.
- opinion_shift (strict): an agent named exactly after an alias two roster
  actors share (owned by neither) is now reachable: its own trajectory is
  shown, and the header says whose shared alias the name also is. Without
  such an agent the shared alias stays ambiguous.
- opinion_shift (strict): agents that match_actor attributes to a roster
  actor under another name are listed in the header
  ("(合并 agent:Bank、Bank of Japan)", capped at 12 like the ambiguity
  list), never silently counted under the queried name. The helper now
  returns the full header note, so resolution and merge labelling live in
  one place. Flag off: legacy output unchanged (tested).
- Role-contract id: pinned by test and documented in code. For two Latin
  edge cases the role id changes once. Names over 180 characters used to
  hash the truncated display name; names replaced by the unsafe-text
  placeholder used to share one id. Both now use the unchanged context-pack
  id (actor_id_for), which the base's actor_id_for already produced.
  Persisted role manifests are validated against their own stored ids, so
  sealed artifacts stay valid.
- stable_actor_id docstring: it is the backend fallback behind
  actor_id_for, not the producer's deerflow_research.stable_actor_id, and
  explicit producer actor_ids take precedence.
… diagnostics [RESEARCH-6]

Candidate N03 (expert-consensus capture), re-scoped to its shadow core.

RESEARCH_FORECASTER_ATTRIBUTION (default false, forwarded to the v3 child):
the facts task gains one field rule (after the date and forecast-inputs
rules) asking estimate/forecast/target rows for their forecaster, a
forecaster-free metric, and the low/high range and n_forecasters the report
states. normalize_quant(with_items=True) pairs each row with its own
extracted item; the pure attribute_forecast_row then runs before the bridge
enriches the rows: the forecaster is also written to analyst (so charts split
by forecaster, not publisher; source is kept), low/high are kept only as a
pair whose numbers are all in the report's page_number_set and that bracket
the value (scale-aware), n_forecasters only as an integer >= 2 next to a
count noun, a range value becomes low/high (range_kind stated_range), and
actual rows never carry the keys. meta.forecaster_attribution counts kept and
dropped fields; a failure is recorded in analytics_errors and leaves the rows
unattributed. The changed task hash re-extracts the facts memo.

REPORT_CONSENSUS_DIAGNOSTICS (off | shadow, default off): the new pure module
app/services/consensus_evidence.py groups projected quant rows by metric
family (else forecaster-free metric), region, target year and unit, reports
groups with >= 2 forecasters (min/max/median over each forecaster's newest
vintage, max/min spread for level units, vintages, staleness from age and
later timeline events, deterministic same-forecaster revisions), lists
within-row ranges and excludes rows dated after the as-of (after_as_of); the
payload carries a canonical-JSON sha256 and never raises. In shadow mode
ReportAgent writes consensus_evidence.json atomically and a digest to
forecast.quality.consensus right before the final forecast.json write: no
prompt, probability, gate or model-call change.

Both knobs off: facts prompt, quantitative.json and forecast.json are
byte-identical. Knobs are on Config, in .env.example, the v3 forwarding
registry and the v3 test hermetic list.

Tests: backend/tests/test_consensus_diagnostics.py (32) covers the addendum,
provenance checks, stated ranges, count nouns, actual rows, analyst,
item pairing, memo re-extraction, degrade-safe attribution, the actor-context
regression, the humanoid-2030 group, the leakage guard, Dell'Oro revisions,
staleness, malformed input and sha stability, and the report stage (shadow
digest and sidecar, identical prompts and call count, off/invalid modes, a
failing writer), plus check_env_drift --strict and child forwarding.
- sim_schedule_audit: a row that makes fire_scheduled_events raise while it
  builds the due list (non-mapping, or a round int() rejects) blocks the whole
  schedule, so every otherwise-fireable row is now counted as
  blocked_by_invalid_round (unreachable == scheduled); non-finite floats in
  samples are stringified so run_summary.json stays strict JSON.
- Drift guard test runs the real fire_scheduled_events over the run and pins
  scheduled - unreachable to what it actually posts.
- forecast_extractor: a world-state block with an explicit non-valid verdict
  line is neither an allowed source label nor parsed into
  world_state_outcome / sim_adjustment (fail-closed, knob on or off), which
  restores the pre-SIM-3 outcome for those blocks.
- Provenance test now runs reconcile_forecast_contract and asserts
  sim_adjustment; new legacy_prompt tests for inconclusive blocks.
- Scope quant reconciliation by period length as well as end: a year no
  longer meets its fourth quarter, second half or December, nor a
  multi-year total its last year (or a span starting elsewhere), so no
  fabricated contested claims from annual-vs-quarterly rows.
- Scope labels name the reported/projected class (projected, unclassified)
  and an as_of_date year ('as of 2025'), so distinct scopes never share a
  claim text.
- v3 also flags claimed actuals whose period_end ends after the as-of
  (_unfinished_period_flags, the program classifier's
  future_dated_reported), under the same cap; the helper keeps rows whose
  as_of_date is already after it.
- Probable unit-scale disagreements go first under the reconcile cap; the
  report block's 15-claim limit is noted at the constant.
- Extract-only reference date is never after today and ignores a year- or
  month-only extraction.
…ed headline [EVAL-8]

Candidate P02 (cut to the S honesty fix). A golden Brier is a skill estimate
only over rows the run could not look up, so golden_eval now classifies every
matched row from the scored run's provenance and reports a headline over
prospective rows only.

- classify_tier(q, *, run_created_at, lead_tolerance_days, hindcast=None) is pure
  and compares UTC dates: prospective (run before the resolution date and no
  later than as_of + GOLDEN_PROSPECTIVE_LEAD_TOLERANCE_DAYS), late_origin,
  hindcast_retrieval_exposed (run on or after resolution: live retrieval and
  model memory see the outcome), hindcast_pit (integration adjustment: a TIME-7
  hindcast pin, whatever the dates, with TIME-9's integrity verdict) and
  unknown (no provenance; naive stamps are rejected).
- score-forecast-file gains --pipeline-dir (run.json created_at, else
  pipeline_state.json; the hindcast pin from pipeline_state options, else
  run.json as_of_enforcement), --run-created-at and --require-headline, all
  read with getattr. A forecast stamped 'hindcast' is a hindcast too. The
  report gains headline {status, tier_counts, run_created_at, source,
  lead_tolerance_days, metrics over prospective rows} and
  characterization.by_tier; 'metrics' is unchanged with metrics_scope
  all_matched_characterization. The markdown opens with HEADLINE: or
  HEADLINE WITHHELD: above the EVAL-7 banner. --require-headline exits 5 unless
  the status is ok; otherwise the exit codes are unchanged.
- --to-ledger forwards each row's tier; append_golden_result writes golden_tier
  only when given. score-ledger builds its golden headline from golden_tier
  'prospective' rows (rows without the key are unknown).
- Knobs: GOLDEN_HEADLINE_GATE=true (false: byte-identical pre-EVAL-8 reports,
  ledger rows and score-ledger output; --require-headline then exits 5 with
  status gate_disabled) and GOLDEN_PROSPECTIVE_LEAD_TOLERANCE_DAYS=7, both in
  .env.example. The cutoff registry, post_cutoff/straddles tiers,
  --reclassify, harvest-holdout, list_closed_markets and the AS_OF prompt
  block stay deferred (docstring).

The committed 2024-2025 set scored for a 2026 run is withheld: 30/30 hindcast.

Tests: new backend/tests/test_golden_tiering.py (classify_tier boundaries,
committed set withheld and exit 5, prospective headline, hindcast pin tiers,
no provenance, gate-disabled byte identity pinned to e57b26d hashes, ledger
golden_tier forwarding and the score-ledger split). Existing tests with the
default gate on were amended minimally: the EVAL-7 additive-keys hash now also
excludes the EVAL-8 keys, the committed-set markdown check expects the
HEADLINE line above the banner, and the pre-EVAL-9 byte test runs with the
gate off. Golden, ledger, evaluation-run, config audit, env drift, import fence
and hindcast suites pass.
…ORT-3]

S11 anchors on the first occurrence of each scenario's full name, so a stale
probability written against an alias passed the final audit:
report_ffe1ea6bf50d published the summary blockquote "基准情景(40%)" while
forecast.json held A:基准扩张 = 0.35 (P08 stage 1b; split from P08, candidate
id carried by REPORT-2).

New pure module services/logic_number.py (no LLM, no IO):
- derive_scenario_aliases: full name, enumerator label aliases (情景A / A情景 /
  Scenario A), core, head, and role words resolved to the unique scenario
  whose name holds the role keyword (not _is_residual_scenario_name, which
  counts 基准 as residual); aliases naming two scenarios are dropped.
- find_probability_slots: strict slots only (ALIAS(N%), ALIAS(概率 0.NN),
  ALIAS:N% 概率, N% 的概率 ALIAS, ALIAS at N% probability); REPORT-2's range,
  quantity and sum guards (plus a history context) make a finding unresolved.
  Bare role words and two-character CJK name parts count only when they stand
  free ("价格上行(10%)" is a price move).
- substitute_probability_slots rewrites only fixable number spans, keeping the
  format; audit_markdown skips fences, References, the Part-1 block and every
  blockquote but the summary blockquote after the H1.

narrative_sync.py exposes its guards as range_guarded / quantity_guarded /
sum_guarded (no behaviour change, per the critic amendment).

Wiring (report_agent.py):
- the outline summary is repaired right after plan_outline, before
  report.outline / self._outline_summary / save_outline, so outline, meta,
  quote-audit exemption and the published blockquote stay identical; a late
  resync keeps them in lockstep when the spine was not ready at planning;
- _repair_logic_number runs after language purity and before the editorial
  lint and _stabilize_publish_markdown (REPORT_NARRATIVE_SYNC on + spine
  scenarios), rewriting full_report.md;
- REPORT_LOGIC_NUMBER_GATE (off|observe|numeric, default observe): observe
  records forecast.quality.logic_number and final_audit.json logic_number
  (with the repair record) and never feeds hard issues; numeric adds alias
  mismatches to _audit_numeric_consistency and, via lint_report(...,
  alias_aware_s11=True), to check_scenario_probabilities. numeric changes a
  hard rule and needs a policy-version bump plus replay tool (owner decision);
  REPORT_FINAL_AUDIT_POLICY_VERSION stays 3.

Tests: test_logic_number.py (aliases, ffe1 slot, guards/decimal, slot forms,
weak aliases, history guard, markdown scope, caps, S11 strings, gate parsing)
and test_logic_number_repair.py (outline lockstep through generate_report,
repair before lint, late resync, ffe1 fixture end to end with the real final
audit, gate modes, observe changes no hard rule, config/.env.example).
test_report_lint_projection's lint_report signature pin now includes the new
keyword-only flag.
- A refused escalated max_tokens (a deterministic 400 past the provider's output
  cap) now ends the ladder with the last EmptyCompletion (escalation_exhausted),
  so a fallback provider never enters the 900 s deterministic cooldown and the
  report preflight still sees a reachable provider.
- _try_fallback propagates BudgetExceeded from the fallback while
  LLM_LENGTH_ESCALATION is on (flag off: legacy swallow).
- An escalated reply that is still length-cut is counted as recovery outcome
  'partial', not 'recovered'.
- Market match and divergence-revision passes drop the last item of a
  truncated reply and record binary_quality.llm_truncation_market_trimmed.
- quality.llm_truncation is kept out of the critic/pre-mortem view, and a spine
  lost to truncated draws is merged into forecast.json quality.llm_truncation.
- Secrecy: vendor text is redacted whole before anything cuts it. _fred_get
  redacts the key and api_key= values from the body text and from every string
  of the parsed JSON, and drops a trailing key fragment left by a transport's
  own 500-character cut; the default transport redacts before its cut. A key
  straddling the 160-character detail clip, the 500-character body cut or a
  metadata clip no longer leaks a prefix (regression tests assert no
  8-character piece of the key in any result field or log).
- Vintage: answers are checked, not only requests. A response or row whose
  realtime_start/realtime_end stamps exclude the pin (or are malformed) is
  no_vintage at a historical pin and unavailable at today's, never labelled the
  pinned vintage.
- Derived arithmetic runs in an 80-digit decimal context: differences of
  32-digit values are exact and extreme percent changes quantize instead of
  raising InvalidOperation (which dropped valid data as unavailable).
- A naive now is local time (Python's reading); the vendor clock never raises
  at the datetime limits.
- window_years accepts whole numbers given as floats or numeric strings; any
  other value is invalid_input with zero calls, never a silent default.
- Cache: an open vintage's entry is capped at the current DATA_FRED_CACHE_TTL_H
  at read time; provenance carries fetched_at (when FRED answered).
- Tests: observations-stage "does not exist", throttle wiring (one wait per
  sent request, none for refused or cached lookups), and the straddle cases.
- LICENSES/Apache-2.0.txt (the full licence text, from TradingAgents 0.5.1),
  referenced from NOTICE (Apache-2.0 section 4(a)).
consensus_evidence:
- Dates (as_of_date and timeline event dates) are read with quant_typing's
  copy of the engine's period parser, so free-text forms ("August 2026",
  "2026年8月", "2026年底") meet the leakage guard. A stated as_of_date no
  reading can date ("FY29") is excluded as unparsed_as_of (fail closed); a
  blank or placeholder one stays an undated vintage.
- The bridge's year is a target year only when as_of_date does not name it
  (the bridge fills year from as_of_date when a row states no period).
- The value scale word is the one right after the number value_num is read
  from (first range, else first number): "250,000 (1 million by 2035)"
  stays 250,000.
- Values, bounds, medians and spread ratios that leave the float range are
  counted as no_value / dropped / None instead of sinking the payload.
- The ranges sort key covers every field (n_forecasters, range_kind), so
  the sha256 never depends on row order.

linear_research: _magnitude returns None for numbers beyond the float range
(a JSON integer like 10**400 raised OverflowError and cost the whole run its
attribution); _forecaster_count caps ints at 9,999,999 like its text form.

Tests: the actor-context check now asserts the monotone match predicate and
pins the 32-row pack-cap displacement; knob docs note the cap caveat.
- A run resumed or regenerated in place keeps its created_at, so the run is now
  dated by its last recorded activity. With --pipeline-dir, the run date is the
  latest of created_at and every pipeline_state.json activity stamp:
  created_at, heartbeat_at, last_progress_at, options.resumed_at,
  options.force_report_regen, the ensemble_wall window and every stages.* window.
  updated_at is not used, because bookkeeping writes move it. That date drives
  both the resolution check and the as_of + tolerance check. A run created
  before resolution but resumed after it becomes hindcast_retrieval_exposed.
  The headline records run_last_activity_at and run_last_activity_source.
- The run stamps are withheld (fail closed) in these cases: pipeline_state.json
  cannot be read (run.json alone cannot rule out a resume); an activity stamp or
  its container is malformed; the state's report_id is missing or is not the
  scored forecast's report directory; or options is not an object.
- The headline records pipeline_dir and pipeline_id. A --run-created-at stamp is
  taken as given, and the headline carries a note saying so.
- run.json resolved.as_of_enforcement is always read. If either record marks a
  hindcast, the rows are hindcast_pit, and the state pin's verdict wins when both
  do. An as_of_enforcement that is not an object withholds the stamp.
- The lead tolerance is capped at 0..3650 days (MAX_LEAD_TOLERANCE_DAYS). An
  as_of + tolerance past date.max clamps instead of raising OverflowError.
- An ok headline adds HEADLINE_SCOPE_NOTE under the unchanged EVAL-7 banner in
  both markdown layouts.
- An invalid non-string series is named by its type, never stringified:
  the text of a list, bytes or dict holding the key could be clipped
  inside it, past the whole-key scrub.
- Observation values are cached and reported fixed-point: str(Decimal)
  wrote 0.0000002 as 2E-7, which the cache parse rejected (every lookup
  missed) and which disagreed with the rendered page.
- An HTTP 200 without the seriess or observations list, or with a series
  entry that is not an object, is unavailable; only a list FRED sends
  empty is no_vintage/not_found.
…inned macro series (pure module, mocked tests)
…as labelled exogenous items (P23 follow-on)
Roger Lin added 25 commits October 2, 2026 08:15
PRICE_TIME_BASES mirrors prediction_markets' PRICE_TIME_BASIS_* values (requote, observed,
snapshot; a parity test pins them), so an exact anchor dated by the row's observed_at stays
in the headline and its basis mix. The market_p caveat and the monitor legend name all three.
…AL-5]

The citation check leaves out an implied_yes_prob that is no probability (inf, 1e308, ...)
instead of raising OverflowError; such an anchor already fails market_eligibility, so only
its row is skipped. Any other item whose lookup or row raises is logged and counted as
unscored row_unreadable rather than failing the ledger-wide block.
expired_unresolved now has one stated scope: the reports in ledger.jsonl or
resolutions.jsonl plus the reports the monitor run covers. _cmd_run carries the batch's
report ids on MarketSkillReads, so every page of a run --all-recent batch counts the same
reports; market_skill.expired_unresolved_scope (ledger_reports for summary,
ledger_and_run_reports for a run) and the md line name the scope.
…d of cutting them; noun vs hedge probability rule [REPORT-13]
…ng [EVAL-5]

Review round 3 (approve) low: an admitted item whose target binary cannot be
found was counted as missing_market_price, so the monitor could not tell a
vanished target from an anchor without a price; it is now unscored under its
own reason, target_missing.
…ce the forecast saw and divergence hit rate vs a market-implied null

Reviewed in three rounds (round 3 approved: flag-off monitor output
byte-identical to the base; FU-11's 'observed' price-time basis accepted); its
one low (target_missing) was fixed before the merge.
Applied by the orchestrator from the reviewer's tested patterns:
- A period ends a sentence only before whitespace and a non-lower-case word or
  at the end, so 'U.S.', 'U.K.', 'e.g.' and 'approx.' never separate a chance
  noun from its quantity.
- Two-numeral Chinese tenths ranges (七八成, 六七成, 三四成) are percentages,
  with the single-numeral tail exclusions (三一成立 is none); Chinese decimals
  (零点七) are quantities.
- The Chinese hope-word gap holds only linking and comparison words (and
  whitespace, for the signal-threshold row), so 车企希望产能提升一倍 and
  出口机会增加两倍 pass; 一半的机会 stays a probability (pinned).
- The prompt states the 300-character trigger limit; the docstring no longer
  claims the pass only ever publishes non-probability text.
Merged feat/finharness-transplants (81a84e8) first.
Round 5's narrowed Chinese hope-word gap (orchestrator) let adverb phrasings
through that round 4 blocked (成功的机会还不到一半, 机会已超过一半, 约有一半,
把握还是三比一); the gap now takes an optional adverb and compound linking words
(the reviewer's tested patch), with single optional spaces so it stays linear
(the reviewer's \s* form was quadratic on long whitespace). The two-numeral
tenths range covers 一两成 and skips a year in Chinese numerals (二〇二五成都车展).
The docstring states that a sentence ending in a single-letter token is read
with the next one (fail-closed).
…whatever the unit [RESEARCH-8]

_result_form reads a formula that keeps its data operands' unit as a
quantity whether or not that unit is a unit class, so a difference of
counts or bare numbers ('700%' from 107 - 100 units, '2,400%' from 37 - 13)
is no longer accepted as a ratio x 100. Failure cases and derived_quant_match
rows cover it; the flag-off pin is recomputed (base 30ab072 and dc89859 agree).
… x 100 [RESEARCH-8]

derived_numbers.is_quotient tells whether a formula's top operation divides
a term holding a data operand by another; _ResultForm.ratio replaces
quantity, so products, powers and calls of operands ('48,100%' from a*b,
'608%' from sqrt(a)) and any unit-keeping result state no percentage. The
flag-off pin is recomputed (bases 30ab072 and dc89859 agree with HEAD).
…ad single-digit results [RESEARCH-8]

derived_numbers._dimension keeps the data operands' unit through + - abs min
max only when every term is in that unit, the literal 0 aside, so 'a+100',
'100-a' and 'min(a, 1000)' over GW or dollars state no figure in that unit
(100-a over percentages is still a percentage). derived_quant_match reads
the one number that may state a result (can_state), as a finding does, so a
'7 %' row matches a 7% growth. The flag-off pin is recomputed (bases 30ab072
and dc89859 agree with HEAD).
A sum or difference of percentages is in percentage points and is stated in
points only ('26 percentage points', '26 pp', '26个百分点'; '26%' for
68% - 42% is result_mismatch); any other percentage result and a ratio's
x 100 reading reject the points form ('62 percentage points' for a relative
change, '162 pp' for a ratio). _NumberOccurrence records the points form
(_POINTS_NUMBER_RE). The flag-off pin is recomputed (bases 30ab072 and
dc89859 agree with HEAD).

Round 4 in full: a unit-keeping result is never a percentage whatever its
unit (medium); only a top-level quotient of data operands reads x 100; a
unitless literal offset drops the unit; derived_quant_match reads
single-digit percentage and unit rows; percentage points as above. Open
issue 5 (TIME-13 drop before derive) is left fail-closed as decided.
…trong-tier call feeding Part-2 and dated update triggers (published probabilities untouched)

Reviewed in seven rounds (round 7 approved). The worded-probability wall is a
documented lexicon rule (cross-sentence and detached-hedge forms out of scope,
pinned by tests); over-length claims and triggers are dropped, never cut. Off by
default (REPORT_COUNTER_CASE=false).
is_quotient requires both sides of the top division to be in the data
operands' unit (as keeps_unit reads them), so a quotient whose sides differ
in unit ('a*b/b' over GW, 'a/b/c') is no ratio and '3,700%' states no 37 GW.
New derived_numbers.gives_points: over percentage operands only sums and
differences (the literal 0 aside) or a complement taking them from 100 give
percentage points; _result_is_points uses it, so 'a+100' over 68% is no
'168 percentage points'. Merged feat/finharness-transplants 44a4561; the
flag-off pin is recomputed (bases 30ab072 and 44a4561 agree with HEAD).
Low 2 (explicit *100 on a unit-keeping formula) stays a recorded follow-up.
…ation key

Gate wave27a: test_golden_tiering::test_gate_disabled_legacy_keys pins the
pre-EVAL-8 bytes of the score-ledger report with GOLDEN_HEADLINE_GATE off.
EVAL-5 (merged dc89859) adds n_unmatched_outcome to calibration_report as its
spec requires (additive; legacy numbers unchanged), which that pin did not
expect. The test now pops it, as it already does for EVAL-12's additive
contamination block, and asserts it is 0 for the fixture.
…valuator, operand-on-page verification, derived supports

Reviewed in five rounds (round 5 approved: the flag-off pin equals the merged
base); its round-5 lows were fixed before the merge (d225ee4). Off by default
(RESEARCH_DERIVED_FINDINGS=false).
… trust section

DRF_ARCHITECTURE.md gains section 20 (where each transplanted mechanism lives,
by pipeline stage, with its knob and default) and a key-file-map row; the
README trust section gains '6. Time, settlement and evaluation integrity'.
Every cited path and knob was checked against the merged branch, and every
stated default against config.py.
…code

A fact-check of every statement against the branch found:
- hindcast admission lives in PipelineOrchestrator.start, not PipelineManager;
- the TIME-8 gates live in source_dates.py and research_gateway.py;
- some default-on knobs are not honesty fixes;
- anchors carry price_time plus a basis (quoted_at and snapshot_as_of
  are on market rows);
- page verification runs under RESEARCH_VERIFIED_FACTS.

Opt-in features (counter-case, market-blend arithmetic, backbone check,
settlement fold for the report) are now marked off by default. Rejected
tool calls are free only up to the per-section cap. The scorecard and
cost card are written by every run. Follow-ups now run through FU-12.
…sues

Assembled from the program's essence chapters, roadmap, implementation
status and gate records. It covers the essence of FinanceHarness,
StockAgent, TradingAgents and five papers. It records how the 81 planned
packages and 12 follow-ups landed: knobs, gates, integration commits and
licensing. It also lists the open issues left after implementation.
Mechanically validated: 1,066 code anchors and 514 paper anchors.

Also adds the session's agent-progress entry.
Copilot AI balanced review requested due to automatic review settings October 2, 2026 16:32

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 58818ec1f1

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

if self.pit is None:
return self.records
admissible = _memoized_sid_check(self._pit_admissible)
return {kid: pit_wall_record(record, admissible, per_claim=True)[0] for kid, record in self.records.items()}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Drop mixed co-citations from hindcast fallbacks

When a gated hindcast uses a deterministic fallback and a finding ends in a mixed cluster such as [S1][S2], where only S1 is available by the cutoff, per_claim=True strips S2 and publishes the entire finding under S1. If part of that finding was supported only by the post-cutoff S2, the fallback leaks future information into the hindcast and can invalidate its report and evaluation results; use the fail-closed whole-line wall applied to the writer digest instead.

Useful? React with 👍 / 👎.

if not isinstance(stamp, str) or not stamp.strip():
return None
try:
return datetime.fromisoformat(stamp.strip()).astimezone(timezone.utc)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Preserve the timezone of report completion timestamps

When a report generated in one timezone is backfilled on a machine in another timezone, completed_at is a naive local timestamp, so astimezone() interprets it in the replay machine's timezone rather than the generation machine's timezone. This shifts the cutoff used to stamp window-ended markets and can make the same archived report produce different bytes or labels across environments; persist an aware UTC timestamp and handle legacy naive values without assigning the replay host's zone.

Useful? React with 👍 / 👎.

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-02T16:37:05.604343Z 58818ec PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Unmatched outcomes corrupt calibration, the golden question changes its threshold, and scoring and provenance helpers have correctness gaps.

Review effort: Balanced
Findings: 3 Medium severity

Open (3)
What changed in this PR

This PR adds point-in-time research integrity, safer simulation/reporting, sealed evaluation infrastructure, and platform hardening across DRF.

Changes:

  • Adds hindcast, provenance, typed evidence, honest parsing, and market-integrity controls.
  • Expands evaluation, scoring, ledger, telemetry, and cost-accounting capabilities.
  • Hardens identifiers, provider handling, configuration, testing, and deployment.
File Description
setup.sh Deploys new bridge modules.
scripts/​doctor.sh Audits honesty pins and provider requests.
README.md Documents trust guarantees.
NOTICE Adds third-party attribution.
frontend/​tests/​marketStatus.test.mjs Tests incomplete market states.
frontend/​src/​utils/​marketStatus.js Classifies market lookup failures.
frontend/​src/​components/​research/​DossierViewer.vue Displays incomplete-market messaging.
drf2/​skills/​custom/​forecast-report/​SKILL.md Reframes probability calibration guidance.
drf2/​skills/​custom/​deep-research/​SKILL.md Distinguishes failed and empty searches.
drf2/​driver/​state.py Hardens pipeline identifiers.
deerflow_bridge/​skills/​forecast-visuals/​scripts/​render.py Filters future-dated actuals.
deerflow_bridge/​research_budget.py Adds quota circuit handling.
deerflow_bridge/​market_tools.py Adds delayed HTTP retries.
backend/​tests/​test_worldstate.py Tests non-finite vote handling.
backend/​tests/​test_utils_security.py Makes URL tests hermetic.
backend/​tests/​test_utils_dates.py Tests date precision periods.
backend/​tests/​test_sim_market_priors.py Tests expired-market filtering.
backend/​tests/​test_sim_event_provenance.py Tests event provenance helpers.
backend/​tests/​test_seed_ensemble.py Pins ensemble seed derivation.
backend/​tests/​test_research_spend_flush.py Tests cached-token accounting.
backend/​tests/​test_research_gateway.py Tests shell detection and ledger rollback.
backend/​tests/​test_research_evidence_quality.py Tests future-date classification.
backend/​tests/​test_research_engine_v3.py Extends v3 question-spec fixtures.
backend/​tests/​test_research_engine_v3_text.py Updates engine test doubles.
backend/​tests/​test_research_engine_v3_round3_synth.py Updates synthesis test context.
backend/​tests/​test_research_engine_v3_round2.py Tests subprocess and model telemetry.
backend/​tests/​test_report_market_evidence.py Tests typed market absence.
backend/​tests/​test_report_lint_absence.py Tests absence-marker leakage removal.
backend/​tests/​test_quantity_scoring.py Tests quantity scoring rules.
backend/​tests/​test_prediction_markets.py Pins market clocks.
backend/​tests/​test_point_in_time.py Tests strict hindcast dates.
backend/​tests/​test_parallel_research_merge.py Tests market-status merging.
backend/​tests/​test_market_influence_boundary.py Stabilizes market anchoring tests.
backend/​tests/​test_llm_recovery.py Tests token-cap escalation.
backend/​tests/​test_forecast_diff.py Tests unscoreable probabilities.
backend/​tests/​test_foglamp_containment.py Tests pinned numeric policy.
backend/​tests/​test_env_drift_pins.py Tests honesty-critical pins.
backend/​tests/​test_ensemble_backtest.py Tests unmatched outcomes.
backend/​tests/​test_deerflow_bridge_sync_guard.py Verifies bridge deployment parity.
backend/​tests/​test_camel_context_delivery.py Tests CAMEL prompt delivery.
backend/​tests/​test_bridge_overhaul_v3.py Pins market time in bridge tests.
backend/​tests/​test_bridge_market_tools.py Tests retry and outage semantics.
backend/​tests/​test_binary_source_rule.py Tests forecast provenance prompts.
backend/​tests/​test_audit_fixes_infra.py Tests recovery and resume messaging.
backend/​tests/​test_actor_context_runtime.py Removes test-time network dependencies.
backend/​tests/​fixtures/​question_spec_golden.json Adds question-spec golden data.
backend/​tests/​eval/​rubric.md Tightens grounding criteria.
backend/​scripts/​run_twitter_simulation.py Uses guarded environment loading.
backend/​scripts/​run_reddit_simulation.py Uses guarded environment loading.
backend/​scripts/​preflight.py Centralizes provider request overrides.
backend/​scripts/​model_comparison.py Preserves null probabilities.
backend/​scripts/​cost_card.py Adds offline cost-card rebuilding.
backend/​scripts/​batch_runs.py Preserves fork policies and lineage.
backend/​run.py Separates startup validation from audits.
backend/​pyproject.toml Enables strict offline test markers.
backend/​app/​utils/​sim_timeline.py Documents question-spec horizons.
backend/​app/​utils/​security.py Adds safe identifiers and containment.
backend/​app/​utils/​provider_overrides.py Centralizes provider-specific requests.
backend/​app/​utils/​point_in_time.py Adds strict temporal parsing.
backend/​app/​utils/​oasis_llm.py Applies resolved-provider overrides.
backend/​app/​utils/​llm_recovery.py Adds token escalation policy.
backend/​app/​utils/​env_loading.py Prevents test-time .env loading.
backend/​app/​utils/​dates.py Adds precision-aware date periods.
backend/​app/​utils/​ctxpool.py Propagates context into workers.
backend/​app/​utils/​canonical_json.py Centralizes canonical hashing.
backend/​app/​utils/​atomic.py Adds secure writes and strict JSON.
backend/​app/​services/​backtest.py Tracks unmatched outcomes in calibration.
backend/​app/​services/​zep_graph_memory_updater.py Honors pinned feedback policy.
backend/​app/​services/​zep_entity_resolver.py Rejects ambiguous aliases.
backend/​app/​services/​worldstate.py Rejects non-finite decisions.
backend/​app/​services/​simulation_manager.py Contains simulation paths.
backend/​app/​services/​simulation_config_generator.py Adds horizon and market gates.
backend/​app/​services/​sim_event_provenance.py Labels injected simulation events.
backend/​app/​services/​report_visualizer.py Adds typed and validity-aware charts.
backend/​app/​services/​quantity_scoring.py Adds quantitative forecast scoring.
backend/​app/​services/​oasis_profile_generator.py Propagates context and filters markets.
backend/​app/​services/​graphiti_client/​runtime.py Contains Kuzu graph paths.
backend/​app/​services/​graph_pruner.py Protects all actor aliases.
backend/​app/​services/​graph_builder.py Contains layout paths.
backend/​app/​services/​exec_brief.py Reports research degradation.
backend/​app/​services/​ensemble.py Excludes unreadable forecast runs.
backend/​app/​services/​actor_role_prompt.py Hardens actor identity matching.
backend/​app/​models/​project.py Contains project paths.
backend/​app/​mcp/​sim_server.py Validates simulation IDs.
backend/​app/​mcp/​kg_server.py Validates graph IDs.
backend/​app/​api/​simulation.py Hardens simulation file access.
backend/​app/​api/​settings.py Reports persistence and reuses overrides.
backend/​app/​api/​research.py Adds strict admission and options.
backend/​app/​__init__.py Adds host and identifier gates.
agent-progress.txt Records implementation progress.
.gitignore Excludes secret temporary files.
.github/​workflows/​ci.yml Expands offline CI coverage.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +108 to +109
if sc.get("outcome_matched_a_scenario") is False:
n_unmatched += 1
Comment on lines +70 to +72
def covered80(q10: float, q90: float, y: float) -> bool:
"""Whether y lies in [q10, q90] (bounds inclusive)."""
return _finite(q10, "q10") <= _finite(y, "y") <= _finite(q90, "q90")
Comment on lines +5 to +6
"operational_question": "Will installed global data-centre capacity exceed 250 GW on 31 December 2027?",
"outcome_definition": "Global installed data-centre IT capacity is at least 250 GW (≥ 250 GW) on 31 December 2027.",

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants