Repository navigation
finharness transplants: point-in-time integrity, honest parsing, sealed ledger and evaluation stack from FinanceHarness, StockAgent, TradingAgents and five papers - #5
Conversation
- linear_research._quant_sanity: reconcile_quantitative now runs once per scope (_quant_scopes / _quant_scope: the end of period_end, else the year of as_of_date; the canonical region or geography; reported vs projected by epistemic_class when typed, else value_type). v3 rows keep their period, geography and value type outside the metric, so grouping by (metric, unit) alone turned forecast trajectories, series over time and different regions into contested.json claims (6 of 6 reconciled claims on the four real v3 handoffs were such; scoped: 0, while the legacy-shaped poll disagreements still reconcile). Reconciled claims and unit warnings name their scope (claim suffix, warning "scope"); caps are applied after concatenation, and reconcile stops at its first failed or malformed result (one analytics_errors entry). The bridge helper and the legacy engine are unchanged. - The future-dated bound for flag_implausible_quant is a pinned run's as-of itself (RESEARCH_AS_OF: no midnight to cross, a later publication is a leak); live runs keep the day after the plan's as-of. - meta.quant_sanity_truncated records the pre-cap totals of any capped sanity list (unit warnings, reconciled claims, implausible flags). - The extract-only salvage of a v3 handoff drops the v3 quant sanity keys (_SALVAGE_VOLATILE_META_KEYS), which describe the files it rewrites. - config.py / .env.example text updated to match. Tests: test_research_quant_parity_v3.py (scoped v3 run: trajectory, series, two regions and actual-vs-forecast add nothing while the same-quarter 1000x pair reconciles; _quant_scope keys and labels; per-scope stub caps and truncation totals; reconcile failure mid-scopes; pinned hindcast and pinned-today bounds; salvage drops the sanity keys).
Numeric guard parser: negated events (not/n't/never/fails to/no longer/ cannot/without/unless; 未/不会/没有/从未/不再...) invert the comparator and a negated touch only infers the violated side; ASCII '~' is a range separator only between tight numbers; a bare year never opens a range with a marked amount (except a tight rising dash) and 'from A to B' / '从A至B' are trajectories; decline verbs suppress only size-of-change readings, never below/under/between levels; suffix comparators (以上/以下/or more/or less) and at-or-above/below, 达到或超过; time spans are windows (touch), not levels; number-letter identifiers (28nm) and scientific notation are read correctly. Binding: research rows use the repo vocabulary (quant_typing.quant_class, observed/value_kind), source_ref falls back to source; as_of/period_end count as known only once the stated period ends, and hindcasts require a readable as_of (an unreadable pin as_of records error instead of using today). Shadow binary draws get 64 extra max_tokens per row (min 10 rows); off keeps 4096. numeric_guard_mode joins the EVAL-18 config fingerprint. Knob docs say the prompt addition can shift drafting. Rebased onto the integration branch (INFRA-8 conflicts resolved keeping both sides).
- Post-hoc parity with the in-band dead-round carry (low #1): new pure decision_channel._carry_unreplayed_events moves the events of rounds absent from the action log into the next replayed round (own events first, copies labelled carried_from_round = source round; events past the last replayed round reach nothing, as in-band). run_decision_channel applies it when the new keyword carry_unreplayed_events is true; the runner's _posthoc_decision_events passes it only while SIM_PERIOD_CONTEXT_V2 is on (SIM-6's own gate). Default False keeps every existing call unchanged. - In-band rounds delivered with an empty roster (low #5): _step_round keeps their decision-context events (labelled via _mark_carried) for the next elicitation, V2-gated; the digest summary is not repeated. - Events renderer failure is logged before degrading (low #2). - Carried label tests `is not None`, like world_delta (low #3). - _event_pair collapses whitespace runs so dedupe, digest and rendering share one identity; _events_digest documents why order/labels stay out of the key, per spec (low #4). - Knob docs in config.py and .env.example mention the carries. Tests: whitespace dedupe + carried_from_round=0 label, logged renderer failure, digest identity, pure carry helper, post-hoc carry in run_decision_channel, runner helper kwargs under V2 on/off, post-hoc dead-round parity via main's kwargs, in-band empty-roster carry.
…c binary guard [REPORT-11] What and why (candidate C20, re-scoped to "measure first"): DRF pushes binary sets hard away from 0.5 and runs an always-on red-team critique on the scenario spine, but neither effect was measured and the uncritiqued spine was never persisted. This change measures both and ships the symmetric guard held behind a default-off knob. - New pure module app/services/probability_shape.py: probability_shape() (prob-shape/v1) computes the scenario peak (max_probability, same name as the gate's quality.max_probability), normalized entropy, TV distance from uniform and the post-minus-pre critique delta; for binaries the 0.40-0.60 midband share (the scorecard's band), 0.45-0.55 near-half share, extreme share, mean |p-0.5|, stdev, a decile histogram and move counts toward/away from 0.5 from market_influence and pre_reconciliation_probability stamps. Never raises; ignores bool/None/NaN/strings/out-of-range values. - ReportAgent snapshots the spine's scenarios before the pre-prose critique and before the post-hoc critique; on a successful critique it stores them as quality.pre_critique_scenarios (quality dict copied, the uncritiqued spine and every critique/premortem prompt stay unchanged). _finalize_structured_forecast writes quality.forecast_policy when FORECAST_BINARY_SYMMETRIC_GUARD is on (independent of the telemetry knob) and quality.probability_shape when FORECAST_PROBABILITY_SHAPE is on, after the binary block and before forecast.json, so the final audit fingerprint covers both. No gate reads them. - Ledger (critic amendment): ledger_commit.commit_report passes the SEALED forecast's quality.probability_shape as objective_signals to forecast_ledger.commit_published_forecast (new optional kwarg; stored as a copy, never joins commit id or target key; non-JSON signals are dropped with objective_signals_dropped rather than failing the commit). The legacy append_forecast path is byte-identical. New forecast_ledger.shape_summary() groups production rows by policy.binary_symmetric_guard with per-group n / mean / median of the key stats; tolerates rows without signals. - CLI: forecast_tools.py shape --ledger-dir DIR | --forecasts FILE... - Symmetric guard: _BINARY_SYMMETRIC_GUARD is appended right after the contrarian or low-p rule in every binary _draw prompt when the knob is on (it avoids the 'CONTRARIAN FRAMING' / '0.05-0.35 range' literals that test fakes route on). capture_safety_policy_v1 records forecast_binary_symmetric_guard; INFRA-9 forks inherit it. Idea from the TradingAgents (Apache-2.0) symmetric Hold rule; own wording, no code copied. - drf2 forecast-report SKILL.md section 2.1 and its checklist item: removed the "gate will fail the report otherwise" cue; the numeric targets now describe a healthy table instead of being aims. Knobs (.env.example documented): FORECAST_PROBABILITY_SHAPE=true (observability only), FORECAST_BINARY_SYMMETRIC_GUARD=false (probability-moving policy, waits for the WP14 outcome-blind promotion gate). Tests: new test_probability_shape.py and test_binary_symmetric_guard.py; test_forecast_ledger.py extended (commit rows, commit_report, shape_summary). Knob-off proofs: binary prompts byte-identical with the guard off; forecast.json, LLM calls and the ledger commit row unchanged with the shape knob off; legacy ledger rows unchanged.
- rejection_excerpt fails closed: cut at Final Answer / blank line, start at
the first call opener (else the first brace, else empty), keep only
call-syntax tokens so prose after an unterminated block never reaches the
ungated tool_rejected row.
- The per-section tool minimum counts dispatched calls only (ReAct and native);
charged rejections still spend the max budget, and once they exhaust it the
minimum is waived instead of looping (report_tool_args.evidence_floor_unmet).
- Unparseable blocks skipped beside a selected call: the first 3 get their own
row and outcome count, the rest one summary row; the observation lists at
most 3 reasons.
- A call after leading prose braces is found at the next top-level
{"name":/{"tool": opener (bounded attempts, nested objects never taken).
- A bare reply that starts with a live-tool call but lacks its outer brace is
repaired, or surfaced as args_not_json, instead of dropped.
- REPORT_TOOL_MAX_REJECTED_PER_SECTION comment and .env.example line state that
it only applies with REPORT_TOOL_ARG_REPAIR on.
…zation and v3 QA [RESEARCH-9] Real runs strip 55-94% of report citation markers before the read-only final audit, and one run also deleted 912 sentences and machine-added 71 citations, yet final_audit.json showed a clean document. This package records that surgery as deterministic, read-only telemetry (candidate C33, re-scoped to telemetry per the verifier corrections). Report (REPORT_FINALIZATION_TELEMETRY, default true): - New pure module app/services/citation_finalization_telemetry.py keeps a per-report log and builds the report-citation-finalization/1 record. - _stabilize_publish_markdown starts the log (markers_before measured on the body without References before the first finalization); _finalize_citations adds the dangling repair counts and the semantic remaps/strips (kept and unverifiable are the last pass's state, since every pass re-checks the surviving markers); each pass records the FIRST, repairing _repair_final_quantitative_grounding call (the probe that follows reports zeros and overwrites totals) plus overuse strips, removed quotes and passes. - _audit_final_published_markdown writes final_audit.json pre_audit_repairs and forecast.json quality.citation_finalization before serialization, so forecast_sha256 seals it: marker_strip_ratio and markers_lost_with_removed_text = before + machine-added - final - stripped. A WARNING is logged at a strip ratio >= 0.5. No gate, integrity-issue or policy-version change; telemetry failures are logged and ignored. v3 research (RESEARCH_V3_CITATION_STATS, default true, forwarded to the v3 child through RESEARCH_CHILD_V3_KNOBS): - renumber_citations gains an optional dropped out-list (one ledger sid per removed occurrence); None keeps the bytes identical. - _run_qa writes qa.json citation_stats (v3-citation-stats/1): markers before QA and published, orphan markers (occurrences, distinct, sample), stale groups, cited fetched vs snippet sources (as sources.json publishes them, via the extracted _published_page helper) and the snippet marker share, unused fetched pages, writer bibliographies and scaffold echo lines detected, and prose numbers no VERIFIED/REPORTED finding, cited fetched page or cited snippet traces. Detection only. - _qa_meta and _analytics mirror it into meta.research_qa and meta.research_quality; a resumed qa.json without the key still finalizes. With either flag off no new key is written; research_report.md, sources.json, full_report.md and every gate are byte-identical in every mode. Tests: test_report_finalization_telemetry.py (persisted record, first-call quantitative capture and the corrected balance formula, flag-off identity and gate, degrade-safe failures, knob) and test_research_engine_v3_citation_stats.py (renumber out-parameter, detectors, end-to-end engine run with an orphan marker and a writer bibliography, flag-off byte identity, resume, knob).
…tching [INFRA-11]
Actor identity rested on a lossy key (NFKC + casefold, keep only [0-9a-z]
and U+3400-U+9FFF). Kana, hangul and Cyrillic names collapsed to an empty
key, so PREPARE raised "duplicate selected actor_id in context build" (or
"missing a stable name") for Japanese/Korean questions, names such as
トヨタ自動車 / ホンダ自動車 collided on 自動車, and every hangul or
Cyrillic role contract shared actor_e3b0c44298fc1c14. Report-side name
lookups guessed instead of failing closed: opinion_shift merged any agent
whose name contained the target ('US' matched Russia, Australia, Business
Roundtable), trace_cascade resolved 2-character containments, and the
entity resolver's alias map was last-writer-wins.
Identity (unflagged; Latin ids byte-identical, pinned by literals):
- utils/actors: legacy_actor_key (moved verbatim from actor_context),
actor_key_is_lossy (a dropped letter outside the Latin script, spacing
modifier letters excepted), actor_identity_key and stable_actor_id
(legacy hash when lossless, else sha256('idk1\x1f' + normalize_name)).
- actor_context.actor_id_for delegates to stable_actor_id; relationship
endpoints match on lossless identity keys, so a non-Latin actor no
longer collects every relationship with a non-Latin endpoint.
- actor_role_prompt: one NFKC key with actor_context; role-contract ids
derive from the raw name exactly like actor_id_for; _matches_actor and
the pack-dict lookup compare normalize_name first and fall back to the
legacy key only when neither side is lossy.
ACTOR_NAME_MATCH_STRICT (default true, documented in .env.example):
- opinion_shift resolves roster exact name/alias, exact agent name,
roster >=4-char containment, agent >=4-char containment; several
candidates return a message naming them instead of merging.
- _resolve_entity_name consults exact roster aliases, needs >=4-char
containment and a unique node.
- actor_alias_map drops (and logs) aliases claimed by two actors; the
graph pruner's core set keeps every alias (actor_alias_norms).
- ZepToolsService.actor_roster, set by ReportAgent from its actors.
Flag off = legacy matching (tests pin it).
Candidate: C37 (re-scoped to the identity-key hygiene fix).
Tests: backend/tests/test_actor_identity.py (39); affected suites
(actor role/context, provenance, entity resolver, KG, report wiring,
config audit) pass.
- Symmetric guard: append it outside the contrarian block, so with FORECAST_BINARY_CONTRARIAN off it follows the base RULES (which also push away from 0.5). Knob on now means every binary draw prompt carries the guard, matching quality.forecast_policy and the admission pin. Knob-off prompts stay byte-identical (contrarian on and off). - probability_shape: an int too large for a float is ignored instead of collapsing the record; the catch-all fallback keeps the policy and adds error: True so a failure never reads as an empty forecast. - shape_summary: _finite_number tolerates huge ints in ledger rows. - forecast-report SKILL.md: §2.2, the §8 heading and checklist, and §9 no longer present spread as a gate to pass; the spread line is an explicit diagnostic.
Latest-actual binding (llm_field and quant_row) now takes only a value that
states one figure: a trajectory ("2.6 → 13.0", "from 2.6% to 13.0%",
"从2.6%到13%") or two figures stays unbound instead of binding the oldest
figure (report_1c312b400d33 F1 was read as 2.6 against a 13.0 actual). A year
beside the figure is a date; a lone year-like value is the figure.
Non-finite numbers ("$1e999", 400 digits, unit scale overflow) are no
quantity, so forecast.json stays strict JSON; claim/quantity raw is capped.
Comparator pairing is linear (bisect + span sweep) instead of quadratic.
A bare "not" negates only after an auxiliary/modal outside a relative clause
and "without" only within one word of the comparator; "equals/meets or
exceeds" and "equal to or above/below" are inclusive. as_of dates keep the
day of "Sep 15, 2026" and read slash dates (later reading when ambiguous).
capture_safety_policy_v1 warns when it pins an invalid NUMERIC_GUARD_MODE as
shadow. The off-mode extractor test is pinned to digests recorded on the
pre-TIME-5 base (948a792).
- report audit: a re-audit that produces no record (flag off, no log, or the
log of another report) drops forecast.json quality.citation_finalization, so
an earlier publish's record is never re-sealed under forecast_sha256. Flag
off stays byte-identical: the key never existed before this package.
- citation_finalization_telemetry: coverage values that are not finite (or
overflow float) read as None and counts that overflow int read as 0, so the
record always serializes with allow_nan=False.
- v3 writer_bibliographies: linear in the text. A block's scan stops at its
first non-list line and resumes there (every opener before that line runs to
the same heading and fails on it); counts are unchanged, pinned against the
literal definition on 3,000 generated texts.
- v3 untraced numbers: scenario weights trace only as percentages ("%35"), so
a bare "35 plants" sentence is still counted.
- v3 unused_fetched_sids reads fetched as sources.json publishes it
(_published_page): an uncited extraction shell is no unused fetched page.
Tests: stale-record re-audit (flag off and foreign log), non-finite record
values, detector equivalence and a regex-call bound on 5,000 opener lines,
shell-aware unused fetched sids, percent-only weight tracing.
- probability_shape._round and forecast_ledger._distribution normalise -0.0 to 0.0 (a degenerate 1/0 split's entropy; a small negative mean or a -0.0 median), so forecast.json, ledger objective_signals and the shape CLI never print -0.0. - The module docstring states that normalized_entropy and tv_from_uniform are taken over the readable probabilities renormalised to sum 1, while the peak stays raw. - New test pins that the pre-critique stamp comes after the critique and the pre-mortem: with REPORT_PREMORTEM off and on, the red-team prompts are identical with FORECAST_PROBABILITY_SHAPE on and off and never contain pre_critique_scenarios.
opinion_shift (ACTOR_NAME_MATCH_STRICT): containment is now weighed across
the roster and the action log at once. A unique roster containment no
longer hides an agent that also contains the target: 'Japan' against the
roster actor Bank of Japan and the unrostered agent Government of Japan
is reported as two candidates instead of silently returning the Bank of
Japan trajectory. A target resolved through an alias or a containment is
named in the output ("## 「Japan」(解析为「Bank of Japan」)逐轮行为轨迹"),
also in the not-found and injected-only messages; flag off is unchanged.
Canonical names outrank another actor's alias: actor_match_candidates
returns the row whose canonical name matches before alias-only rows, so
'China' resolves to the China actor even when the CCP lists 'China' as an
alias (it used to return an ambiguity the model could not resolve).
Agents are assigned to a roster actor by exact name (canonical first; an
alias two actors share names neither) before match_actor, so the
contested alias no longer pulls the China agent into the CCP trajectory,
and trace_cascade's roster-alias expansion skips contested aliases (CCP
no longer resolves to the China node).
actor_context relevance terms use the lossless key for kana/hangul/
Cyrillic text (legacy key for everything it keeps losslessly), so such
terms are no longer dropped as empty, de-duplicated, or equated with a
name sharing only its kanji. Latin packs are byte-identical; generic
term keys are precomputed and actor_key_is_lossy has an ASCII fast path.
The ACTOR_NAME_MATCH_STRICT comment and .env.example line now state that
unique containments under 4 characters (Fed -> Federal Reserve) stop
resolving unless the short name is an exact roster name or alias.
- NO-polarity criteria: a clause stating the NO outcome ('Resolves NO if the
rate exceeds 3%', 'NO unless ...', 'No: ...') inverts the reading once more;
the outcome word nearest before the reading governs it.
- Negators inside the metric: no negator but 'unless' negates right after a
relative pronoun (contractions, never, cannot, no longer, fail to included);
'never' after an auxiliary is read with it, and a bare 'never' before a past
participle and a prepositional phrase is a reduced relative clause.
- Split fiscal years ('2026/27', 'FY2026-27', 'Q4 2026/27') end in their end
year, so a hindcast never binds one before it has ended.
- status_quo_contradiction findings record horizon_gap_days (diagnostic only).
…e-dispatch validation, uncharged invalid calls
… module [TIME-10] New deerflow_bridge/data_tools.py (stdlib plus a lazily imported httpx; never imports the backend; never raises): the official-data vendor core and the FRED/ALFRED macro series tool. It is inert until TIME-13 binds it to the research engine, so every run is unchanged. - DataResult (frozen), a pluggable Transport (httpx by default; an exception is reported by its type name only), _DiskCache (sha256 file names, atomic tmp + os.replace, corrupt/expired entries are misses, only ok fetches stored) and a process-level _Throttle (FRED 0.5 s). - MACRO_ALIASES (35 aliases; no SP500/NASDAQCOM/BAML) and resolve_series, which rejects descriptive phrases and query-smuggling text before any call. - vendor_today_chicago / fred_pit: min(as_of, FRED's America/Chicago today) compared as date objects; without tz data, UTC minus six hours (never later than the Chicago date). - fred_series: both requests carry realtime_start == realtime_end == pit <= as_of (a request that would not is never sent); conservative status mapping (400 "does not exist" -> not_found today / no_vintage at a historical pin, never an unpinned retry; other 400, 429, 5xx, transport failure -> unavailable). Deterministic rendering whose units DRF's page-number parser verifies (4.3%, 159,000 thousand persons, 29,000.5 billion USD); the header and every support sentence carry source, series id and vintage; YoY and change lines are labelled derived and use true endpoints; model_text is capped at 2600 characters; supports in English or Chinese. The API key never reaches a result, URL, cache file or log (httpx's request log line is redacted). - Knobs (env, documented in .env.example): DATA_TOOLS_CACHE_DIR, DATA_FRED_CACHE_TTL_H (6; a closed vintage is kept 30 days), DATA_TOOL_TIMEOUT_S (20), DATA_FRED_WINDOW_YEARS (10, clamped 1-40). - Deployed by setup.sh's bridge loop and _DEPLOYED_BRIDGE_MODULES; the INFRA-13 import-fence policy admits httpx in _httpx_transport only. - Root NOTICE with the TradingAgents Apache-2.0 attribution (portions adapted from tradingagents/dataflows/vendors/fred.py 0.5.1). Candidate: C06. Tests: backend/tests/test_data_tools_fred.py (new, 97 offline tests with a fake transport), test_deerflow_bridge_sync_guard.py (data_tools.py deployed by both lists), test_import_fences.py.
…losed on truncated forecast JSON [INFRA-3]
Reasoning models (MiniMax / GLM / Kimi) can spend the whole max_tokens budget on
thinking and return empty content with finish_reason=length. chat() treated this
as a generic RuntimeError (three same-cap backoff retries, then failover), the
report preflight (max_tokens=64) reported a reachable provider as down, and a
truncated spine / binary / critique JSON reply was silently bracket-repaired
into forecast content.
- app/utils/llm_recovery.py (new, pure): next_max_tokens doubles the cap with a
1024 floor, clamped to the provider ceiling and to context window - prompt -
reserve; None when there is no headroom.
- llm_client: chat() calls the new instance method _chat_openai_escalating in
place of _chat_openai. On EmptyCompletion(finish_reason='length') it meters
the rejected attempt with the exception usage, checks the run budget
(BudgetExceeded propagates and is never retried), and re-sends at once with
the next cap, at most LLM_MAX_ESCALATIONS times per chat() call (the
escalation state is shared across chat()'s transient retries). Exhausted:
escalation_exhausted=True and chat() fails over instead of retrying. The
ceiling is PROVIDER_META[provider]['max_output_tokens'] else
LLM_MAX_TOKENS_CEILING; the window is Config.context_window_for(provider).
An escalated reply is cached under the original key only on finish 'stop'.
- telemetry: LLMMeter.record_recovery(kind, outcome); snapshot gains
'recovery' {kind: {outcome: n}} and 'recovery_by_stage' when non-empty.
- pipeline_orchestrator: the report preflight treats EmptyCompletion with
finish_reason 'length' as a reachable provider (info log) instead of failing.
- forecast_extractor (LLM_JSON_TRUNCATION_FAIL_CLOSED): _reply_truncated reads
last_call_meta() (method or dict) for finish 'length' or
json_truncation_repaired. Truncated spine draws are discarded (quality
llm_truncation.spine_draws_discarded; all discarded = failed spine), a
truncated binary draw drops its last item (binary_quality
llm_truncation_trimmed), critique / premortem return the input unchanged,
and the post-hoc extractor flags quality.llm_truncation.
- Knobs (config.py + .env.example): LLM_LENGTH_ESCALATION=true,
LLM_MAX_ESCALATIONS=2, LLM_MAX_TOKENS_CEILING=32768,
LLM_JSON_TRUNCATION_FAIL_CLOSED=true. Off: legacy behaviour.
Candidate: C15.
Tests: tests/test_llm_recovery.py (schedule, clamps, purity) and
tests/test_llm_length_escalation.py (recovery + metering + recovery snapshot,
cache rule, provider/window clamps, exhausted -> single failover without
retries, BudgetExceeded stops escalation, transient resume, non-length not
escalated, flag-off legacy path, knob parsing and .env.example docs, report
preflight pass/fail through the real state machine, and the fail-closed
forecast paths with method-style metadata fakes plus a real LLMClient on a
stubbed transport). The orchestrator test harness gains a
report_preflight_error parameter.
- opinion_shift (strict): an agent named exactly after an alias two roster
actors share (owned by neither) is now reachable: its own trajectory is
shown, and the header says whose shared alias the name also is. Without
such an agent the shared alias stays ambiguous.
- opinion_shift (strict): agents that match_actor attributes to a roster
actor under another name are listed in the header
("(合并 agent:Bank、Bank of Japan)", capped at 12 like the ambiguity
list), never silently counted under the queried name. The helper now
returns the full header note, so resolution and merge labelling live in
one place. Flag off: legacy output unchanged (tested).
- Role-contract id: pinned by test and documented in code. For two Latin
edge cases the role id changes once. Names over 180 characters used to
hash the truncated display name; names replaced by the unsafe-text
placeholder used to share one id. Both now use the unchanged context-pack
id (actor_id_for), which the base's actor_id_for already produced.
Persisted role manifests are validated against their own stored ids, so
sealed artifacts stay valid.
- stable_actor_id docstring: it is the backend fallback behind
actor_id_for, not the producer's deerflow_research.stable_actor_id, and
explicit producer actor_ids take precedence.
… diagnostics [RESEARCH-6] Candidate N03 (expert-consensus capture), re-scoped to its shadow core. RESEARCH_FORECASTER_ATTRIBUTION (default false, forwarded to the v3 child): the facts task gains one field rule (after the date and forecast-inputs rules) asking estimate/forecast/target rows for their forecaster, a forecaster-free metric, and the low/high range and n_forecasters the report states. normalize_quant(with_items=True) pairs each row with its own extracted item; the pure attribute_forecast_row then runs before the bridge enriches the rows: the forecaster is also written to analyst (so charts split by forecaster, not publisher; source is kept), low/high are kept only as a pair whose numbers are all in the report's page_number_set and that bracket the value (scale-aware), n_forecasters only as an integer >= 2 next to a count noun, a range value becomes low/high (range_kind stated_range), and actual rows never carry the keys. meta.forecaster_attribution counts kept and dropped fields; a failure is recorded in analytics_errors and leaves the rows unattributed. The changed task hash re-extracts the facts memo. REPORT_CONSENSUS_DIAGNOSTICS (off | shadow, default off): the new pure module app/services/consensus_evidence.py groups projected quant rows by metric family (else forecaster-free metric), region, target year and unit, reports groups with >= 2 forecasters (min/max/median over each forecaster's newest vintage, max/min spread for level units, vintages, staleness from age and later timeline events, deterministic same-forecaster revisions), lists within-row ranges and excludes rows dated after the as-of (after_as_of); the payload carries a canonical-JSON sha256 and never raises. In shadow mode ReportAgent writes consensus_evidence.json atomically and a digest to forecast.quality.consensus right before the final forecast.json write: no prompt, probability, gate or model-call change. Both knobs off: facts prompt, quantitative.json and forecast.json are byte-identical. Knobs are on Config, in .env.example, the v3 forwarding registry and the v3 test hermetic list. Tests: backend/tests/test_consensus_diagnostics.py (32) covers the addendum, provenance checks, stated ranges, count nouns, actual rows, analyst, item pairing, memo re-extraction, degrade-safe attribution, the actor-context regression, the humanoid-2030 group, the leakage guard, Dell'Oro revisions, staleness, malformed input and sha stability, and the report stage (shadow digest and sidecar, identical prompts and call count, off/invalid modes, a failing writer), plus check_env_drift --strict and child forwarding.
- sim_schedule_audit: a row that makes fire_scheduled_events raise while it builds the due list (non-mapping, or a round int() rejects) blocks the whole schedule, so every otherwise-fireable row is now counted as blocked_by_invalid_round (unreachable == scheduled); non-finite floats in samples are stringified so run_summary.json stays strict JSON. - Drift guard test runs the real fire_scheduled_events over the run and pins scheduled - unreachable to what it actually posts. - forecast_extractor: a world-state block with an explicit non-valid verdict line is neither an allowed source label nor parsed into world_state_outcome / sim_adjustment (fail-closed, knob on or off), which restores the pre-SIM-3 outcome for those blocks. - Provenance test now runs reconcile_forecast_contract and asserts sim_adjustment; new legacy_prompt tests for inconclusive blocks.
- Scope quant reconciliation by period length as well as end: a year no
longer meets its fourth quarter, second half or December, nor a
multi-year total its last year (or a span starting elsewhere), so no
fabricated contested claims from annual-vs-quarterly rows.
- Scope labels name the reported/projected class (projected, unclassified)
and an as_of_date year ('as of 2025'), so distinct scopes never share a
claim text.
- v3 also flags claimed actuals whose period_end ends after the as-of
(_unfinished_period_flags, the program classifier's
future_dated_reported), under the same cap; the helper keeps rows whose
as_of_date is already after it.
- Probable unit-scale disagreements go first under the reconcile cap; the
report block's 15-claim limit is noted at the constant.
- Extract-only reference date is never after today and ignores a year- or
month-only extraction.
…ed headline [EVAL-8]
Candidate P02 (cut to the S honesty fix). A golden Brier is a skill estimate
only over rows the run could not look up, so golden_eval now classifies every
matched row from the scored run's provenance and reports a headline over
prospective rows only.
- classify_tier(q, *, run_created_at, lead_tolerance_days, hindcast=None) is pure
and compares UTC dates: prospective (run before the resolution date and no
later than as_of + GOLDEN_PROSPECTIVE_LEAD_TOLERANCE_DAYS), late_origin,
hindcast_retrieval_exposed (run on or after resolution: live retrieval and
model memory see the outcome), hindcast_pit (integration adjustment: a TIME-7
hindcast pin, whatever the dates, with TIME-9's integrity verdict) and
unknown (no provenance; naive stamps are rejected).
- score-forecast-file gains --pipeline-dir (run.json created_at, else
pipeline_state.json; the hindcast pin from pipeline_state options, else
run.json as_of_enforcement), --run-created-at and --require-headline, all
read with getattr. A forecast stamped 'hindcast' is a hindcast too. The
report gains headline {status, tier_counts, run_created_at, source,
lead_tolerance_days, metrics over prospective rows} and
characterization.by_tier; 'metrics' is unchanged with metrics_scope
all_matched_characterization. The markdown opens with HEADLINE: or
HEADLINE WITHHELD: above the EVAL-7 banner. --require-headline exits 5 unless
the status is ok; otherwise the exit codes are unchanged.
- --to-ledger forwards each row's tier; append_golden_result writes golden_tier
only when given. score-ledger builds its golden headline from golden_tier
'prospective' rows (rows without the key are unknown).
- Knobs: GOLDEN_HEADLINE_GATE=true (false: byte-identical pre-EVAL-8 reports,
ledger rows and score-ledger output; --require-headline then exits 5 with
status gate_disabled) and GOLDEN_PROSPECTIVE_LEAD_TOLERANCE_DAYS=7, both in
.env.example. The cutoff registry, post_cutoff/straddles tiers,
--reclassify, harvest-holdout, list_closed_markets and the AS_OF prompt
block stay deferred (docstring).
The committed 2024-2025 set scored for a 2026 run is withheld: 30/30 hindcast.
Tests: new backend/tests/test_golden_tiering.py (classify_tier boundaries,
committed set withheld and exit 5, prospective headline, hindcast pin tiers,
no provenance, gate-disabled byte identity pinned to e57b26d hashes, ledger
golden_tier forwarding and the score-ledger split). Existing tests with the
default gate on were amended minimally: the EVAL-7 additive-keys hash now also
excludes the EVAL-8 keys, the committed-set markdown check expects the
HEADLINE line above the banner, and the pre-EVAL-9 byte test runs with the
gate off. Golden, ledger, evaluation-run, config audit, env drift, import fence
and hindcast suites pass.
…ORT-3]
S11 anchors on the first occurrence of each scenario's full name, so a stale
probability written against an alias passed the final audit:
report_ffe1ea6bf50d published the summary blockquote "基准情景(40%)" while
forecast.json held A:基准扩张 = 0.35 (P08 stage 1b; split from P08, candidate
id carried by REPORT-2).
New pure module services/logic_number.py (no LLM, no IO):
- derive_scenario_aliases: full name, enumerator label aliases (情景A / A情景 /
Scenario A), core, head, and role words resolved to the unique scenario
whose name holds the role keyword (not _is_residual_scenario_name, which
counts 基准 as residual); aliases naming two scenarios are dropped.
- find_probability_slots: strict slots only (ALIAS(N%), ALIAS(概率 0.NN),
ALIAS:N% 概率, N% 的概率 ALIAS, ALIAS at N% probability); REPORT-2's range,
quantity and sum guards (plus a history context) make a finding unresolved.
Bare role words and two-character CJK name parts count only when they stand
free ("价格上行(10%)" is a price move).
- substitute_probability_slots rewrites only fixable number spans, keeping the
format; audit_markdown skips fences, References, the Part-1 block and every
blockquote but the summary blockquote after the H1.
narrative_sync.py exposes its guards as range_guarded / quantity_guarded /
sum_guarded (no behaviour change, per the critic amendment).
Wiring (report_agent.py):
- the outline summary is repaired right after plan_outline, before
report.outline / self._outline_summary / save_outline, so outline, meta,
quote-audit exemption and the published blockquote stay identical; a late
resync keeps them in lockstep when the spine was not ready at planning;
- _repair_logic_number runs after language purity and before the editorial
lint and _stabilize_publish_markdown (REPORT_NARRATIVE_SYNC on + spine
scenarios), rewriting full_report.md;
- REPORT_LOGIC_NUMBER_GATE (off|observe|numeric, default observe): observe
records forecast.quality.logic_number and final_audit.json logic_number
(with the repair record) and never feeds hard issues; numeric adds alias
mismatches to _audit_numeric_consistency and, via lint_report(...,
alias_aware_s11=True), to check_scenario_probabilities. numeric changes a
hard rule and needs a policy-version bump plus replay tool (owner decision);
REPORT_FINAL_AUDIT_POLICY_VERSION stays 3.
Tests: test_logic_number.py (aliases, ffe1 slot, guards/decimal, slot forms,
weak aliases, history guard, markdown scope, caps, S11 strings, gate parsing)
and test_logic_number_repair.py (outline lockstep through generate_report,
repair before lint, late resync, ffe1 fixture end to end with the real final
audit, gate modes, observe changes no hard rule, config/.env.example).
test_report_lint_projection's lint_report signature pin now includes the new
keyword-only flag.
- A refused escalated max_tokens (a deterministic 400 past the provider's output cap) now ends the ladder with the last EmptyCompletion (escalation_exhausted), so a fallback provider never enters the 900 s deterministic cooldown and the report preflight still sees a reachable provider. - _try_fallback propagates BudgetExceeded from the fallback while LLM_LENGTH_ESCALATION is on (flag off: legacy swallow). - An escalated reply that is still length-cut is counted as recovery outcome 'partial', not 'recovered'. - Market match and divergence-revision passes drop the last item of a truncated reply and record binary_quality.llm_truncation_market_trimmed. - quality.llm_truncation is kept out of the critic/pre-mortem view, and a spine lost to truncated draws is merged into forecast.json quality.llm_truncation.
- Secrecy: vendor text is redacted whole before anything cuts it. _fred_get redacts the key and api_key= values from the body text and from every string of the parsed JSON, and drops a trailing key fragment left by a transport's own 500-character cut; the default transport redacts before its cut. A key straddling the 160-character detail clip, the 500-character body cut or a metadata clip no longer leaks a prefix (regression tests assert no 8-character piece of the key in any result field or log). - Vintage: answers are checked, not only requests. A response or row whose realtime_start/realtime_end stamps exclude the pin (or are malformed) is no_vintage at a historical pin and unavailable at today's, never labelled the pinned vintage. - Derived arithmetic runs in an 80-digit decimal context: differences of 32-digit values are exact and extreme percent changes quantize instead of raising InvalidOperation (which dropped valid data as unavailable). - A naive now is local time (Python's reading); the vendor clock never raises at the datetime limits. - window_years accepts whole numbers given as floats or numeric strings; any other value is invalid_input with zero calls, never a silent default. - Cache: an open vintage's entry is capped at the current DATA_FRED_CACHE_TTL_H at read time; provenance carries fetched_at (when FRED answered). - Tests: observations-stage "does not exist", throttle wiring (one wait per sent request, none for refused or cached lookups), and the straddle cases. - LICENSES/Apache-2.0.txt (the full licence text, from TradingAgents 0.5.1), referenced from NOTICE (Apache-2.0 section 4(a)).
consensus_evidence:
- Dates (as_of_date and timeline event dates) are read with quant_typing's
copy of the engine's period parser, so free-text forms ("August 2026",
"2026年8月", "2026年底") meet the leakage guard. A stated as_of_date no
reading can date ("FY29") is excluded as unparsed_as_of (fail closed); a
blank or placeholder one stays an undated vintage.
- The bridge's year is a target year only when as_of_date does not name it
(the bridge fills year from as_of_date when a row states no period).
- The value scale word is the one right after the number value_num is read
from (first range, else first number): "250,000 (1 million by 2035)"
stays 250,000.
- Values, bounds, medians and spread ratios that leave the float range are
counted as no_value / dropped / None instead of sinking the payload.
- The ranges sort key covers every field (n_forecasters, range_kind), so
the sha256 never depends on row order.
linear_research: _magnitude returns None for numbers beyond the float range
(a JSON integer like 10**400 raised OverflowError and cost the whole run its
attribution); _forecaster_count caps ints at 9,999,999 like its text form.
Tests: the actor-context check now asserts the monotone match predicate and
pins the 32-row pack-cap displacement; knob docs note the cap caveat.
- A run resumed or regenerated in place keeps its created_at, so the run is now dated by its last recorded activity. With --pipeline-dir, the run date is the latest of created_at and every pipeline_state.json activity stamp: created_at, heartbeat_at, last_progress_at, options.resumed_at, options.force_report_regen, the ensemble_wall window and every stages.* window. updated_at is not used, because bookkeeping writes move it. That date drives both the resolution check and the as_of + tolerance check. A run created before resolution but resumed after it becomes hindcast_retrieval_exposed. The headline records run_last_activity_at and run_last_activity_source. - The run stamps are withheld (fail closed) in these cases: pipeline_state.json cannot be read (run.json alone cannot rule out a resume); an activity stamp or its container is malformed; the state's report_id is missing or is not the scored forecast's report directory; or options is not an object. - The headline records pipeline_dir and pipeline_id. A --run-created-at stamp is taken as given, and the headline carries a note saying so. - run.json resolved.as_of_enforcement is always read. If either record marks a hindcast, the rows are hindcast_pit, and the state pin's verdict wins when both do. An as_of_enforcement that is not an object withholds the stamp. - The lead tolerance is capped at 0..3650 days (MAX_LEAD_TOLERANCE_DAYS). An as_of + tolerance past date.max clamps instead of raising OverflowError. - An ok headline adds HEADLINE_SCOPE_NOTE under the unchanged EVAL-7 banner in both markdown layouts.
- An invalid non-string series is named by its type, never stringified: the text of a list, bytes or dict holding the key could be clipped inside it, past the whole-key scrub. - Observation values are cached and reported fixed-point: str(Decimal) wrote 0.0000002 as 2E-7, which the cache parse rejected (every lookup missed) and which disagreed with the rendered page. - An HTTP 200 without the seriess or observations list, or with a series entry that is not an object, is unavailable; only a list FRED sends empty is no_vintage/not_found.
…inned macro series (pure module, mocked tests)
…as labelled exogenous items (P23 follow-on)
PRICE_TIME_BASES mirrors prediction_markets' PRICE_TIME_BASIS_* values (requote, observed, snapshot; a parity test pins them), so an exact anchor dated by the row's observed_at stays in the headline and its basis mix. The market_p caveat and the monitor legend name all three.
…AL-5] The citation check leaves out an implied_yes_prob that is no probability (inf, 1e308, ...) instead of raising OverflowError; such an anchor already fails market_eligibility, so only its row is skipped. Any other item whose lookup or row raises is logged and counted as unscored row_unreadable rather than failing the ledger-wide block.
expired_unresolved now has one stated scope: the reports in ledger.jsonl or resolutions.jsonl plus the reports the monitor run covers. _cmd_run carries the batch's report ids on MarketSkillReads, so every page of a run --all-recent batch counts the same reports; market_skill.expired_unresolved_scope (ledger_reports for summary, ledger_and_run_reports for a run) and the md line name the scope.
…d of cutting them; noun vs hedge probability rule [REPORT-13]
…ws and documented residuals [REPORT-13]
…ng [EVAL-5] Review round 3 (approve) low: an admitted item whose target binary cannot be found was counted as missing_market_price, so the monitor could not tell a vanished target from an anchor without a price; it is now unscored under its own reason, target_missing.
…ce the forecast saw and divergence hit rate vs a market-implied null Reviewed in three rounds (round 3 approved: flag-off monitor output byte-identical to the base; FU-11's 'observed' price-time basis accepted); its one low (target_missing) was fixed before the merge.
# Conflicts: # .env.example
Applied by the orchestrator from the reviewer's tested patterns: - A period ends a sentence only before whitespace and a non-lower-case word or at the end, so 'U.S.', 'U.K.', 'e.g.' and 'approx.' never separate a chance noun from its quantity. - Two-numeral Chinese tenths ranges (七八成, 六七成, 三四成) are percentages, with the single-numeral tail exclusions (三一成立 is none); Chinese decimals (零点七) are quantities. - The Chinese hope-word gap holds only linking and comparison words (and whitespace, for the signal-threshold row), so 车企希望产能提升一倍 and 出口机会增加两倍 pass; 一半的机会 stays a probability (pinned). - The prompt states the 300-character trigger limit; the docstring no longer claims the pass only ever publishes non-probability text. Merged feat/finharness-transplants (81a84e8) first.
Round 5's narrowed Chinese hope-word gap (orchestrator) let adverb phrasings through that round 4 blocked (成功的机会还不到一半, 机会已超过一半, 约有一半, 把握还是三比一); the gap now takes an optional adverb and compound linking words (the reviewer's tested patch), with single optional spaces so it stays linear (the reviewer's \s* form was quadratic on long whitespace). The two-numeral tenths range covers 一两成 and skips a year in Chinese numerals (二〇二五成都车展). The docstring states that a sentence ending in a single-letter token is read with the next one (fail-closed).
…whatever the unit [RESEARCH-8]
_result_form reads a formula that keeps its data operands' unit as a
quantity whether or not that unit is a unit class, so a difference of
counts or bare numbers ('700%' from 107 - 100 units, '2,400%' from 37 - 13)
is no longer accepted as a ratio x 100. Failure cases and derived_quant_match
rows cover it; the flag-off pin is recomputed (base 30ab072 and dc89859 agree).
… x 100 [RESEARCH-8]
derived_numbers.is_quotient tells whether a formula's top operation divides
a term holding a data operand by another; _ResultForm.ratio replaces
quantity, so products, powers and calls of operands ('48,100%' from a*b,
'608%' from sqrt(a)) and any unit-keeping result state no percentage. The
flag-off pin is recomputed (bases 30ab072 and dc89859 agree with HEAD).
…ad single-digit results [RESEARCH-8] derived_numbers._dimension keeps the data operands' unit through + - abs min max only when every term is in that unit, the literal 0 aside, so 'a+100', '100-a' and 'min(a, 1000)' over GW or dollars state no figure in that unit (100-a over percentages is still a percentage). derived_quant_match reads the one number that may state a result (can_state), as a finding does, so a '7 %' row matches a 7% growth. The flag-off pin is recomputed (bases 30ab072 and dc89859 agree with HEAD).
A sum or difference of percentages is in percentage points and is stated in
points only ('26 percentage points', '26 pp', '26个百分点'; '26%' for
68% - 42% is result_mismatch); any other percentage result and a ratio's
x 100 reading reject the points form ('62 percentage points' for a relative
change, '162 pp' for a ratio). _NumberOccurrence records the points form
(_POINTS_NUMBER_RE). The flag-off pin is recomputed (bases 30ab072 and
dc89859 agree with HEAD).
Round 4 in full: a unit-keeping result is never a percentage whatever its
unit (medium); only a top-level quotient of data operands reads x 100; a
unitless literal offset drops the unit; derived_quant_match reads
single-digit percentage and unit rows; percentage points as above. Open
issue 5 (TIME-13 drop before derive) is left fail-closed as decided.
…trong-tier call feeding Part-2 and dated update triggers (published probabilities untouched) Reviewed in seven rounds (round 7 approved). The worded-probability wall is a documented lexicon rule (cross-sentence and detached-hedge forms out of scope, pinned by tests); over-length claims and triggers are dropped, never cut. Off by default (REPORT_COUNTER_CASE=false).
is_quotient requires both sides of the top division to be in the data
operands' unit (as keeps_unit reads them), so a quotient whose sides differ
in unit ('a*b/b' over GW, 'a/b/c') is no ratio and '3,700%' states no 37 GW.
New derived_numbers.gives_points: over percentage operands only sums and
differences (the literal 0 aside) or a complement taking them from 100 give
percentage points; _result_is_points uses it, so 'a+100' over 68% is no
'168 percentage points'. Merged feat/finharness-transplants 44a4561; the
flag-off pin is recomputed (bases 30ab072 and 44a4561 agree with HEAD).
Low 2 (explicit *100 on a unit-keeping formula) stays a recorded follow-up.
…ation key Gate wave27a: test_golden_tiering::test_gate_disabled_legacy_keys pins the pre-EVAL-8 bytes of the score-ledger report with GOLDEN_HEADLINE_GATE off. EVAL-5 (merged dc89859) adds n_unmatched_outcome to calibration_report as its spec requires (additive; legacy numbers unchanged), which that pin did not expect. The test now pops it, as it already does for EVAL-12's additive contamination block, and asserts it is 0 for the fixture.
…valuator, operand-on-page verification, derived supports Reviewed in five rounds (round 5 approved: the flag-off pin equals the merged base); its round-5 lows were fixed before the merge (d225ee4). Off by default (RESEARCH_DERIVED_FINDINGS=false).
… trust section DRF_ARCHITECTURE.md gains section 20 (where each transplanted mechanism lives, by pipeline stage, with its knob and default) and a key-file-map row; the README trust section gains '6. Time, settlement and evaluation integrity'. Every cited path and knob was checked against the merged branch, and every stated default against config.py.
…code A fact-check of every statement against the branch found: - hindcast admission lives in PipelineOrchestrator.start, not PipelineManager; - the TIME-8 gates live in source_dates.py and research_gateway.py; - some default-on knobs are not honesty fixes; - anchors carry price_time plus a basis (quoted_at and snapshot_as_of are on market rows); - page verification runs under RESEARCH_VERIFIED_FACTS. Opt-in features (counter-case, market-blend arithmetic, backbone check, settlement fold for the report) are now marked off by default. Rejected tool calls are free only up to the per-section cap. The scorecard and cost card are written by every run. Follow-ups now run through FU-12.
…sues Assembled from the program's essence chapters, roadmap, implementation status and gate records. It covers the essence of FinanceHarness, StockAgent, TradingAgents and five papers. It records how the 81 planned packages and 12 follow-ups landed: knobs, gates, integration commits and licensing. It also lists the open issues left after implementation. Mechanically validated: 1,066 code anchors and 514 paper anchors. Also adds the session's agent-progress entry.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 58818ec1f1
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| if self.pit is None: | ||
| return self.records | ||
| admissible = _memoized_sid_check(self._pit_admissible) | ||
| return {kid: pit_wall_record(record, admissible, per_claim=True)[0] for kid, record in self.records.items()} |
There was a problem hiding this comment.
Drop mixed co-citations from hindcast fallbacks
When a gated hindcast uses a deterministic fallback and a finding ends in a mixed cluster such as [S1][S2], where only S1 is available by the cutoff, per_claim=True strips S2 and publishes the entire finding under S1. If part of that finding was supported only by the post-cutoff S2, the fallback leaks future information into the hindcast and can invalidate its report and evaluation results; use the fail-closed whole-line wall applied to the writer digest instead.
Useful? React with 👍 / 👎.
| if not isinstance(stamp, str) or not stamp.strip(): | ||
| return None | ||
| try: | ||
| return datetime.fromisoformat(stamp.strip()).astimezone(timezone.utc) |
There was a problem hiding this comment.
Preserve the timezone of report completion timestamps
When a report generated in one timezone is backfilled on a machine in another timezone, completed_at is a naive local timestamp, so astimezone() interprets it in the replay machine's timezone rather than the generation machine's timezone. This shifts the cutoff used to stamp window-ended markets and can make the same archived report produce different bytes or labels across environments; persist an aware UTC timestamp and handle legacy naive values without assigning the replay host's zone.
Useful? React with 👍 / 👎.
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Unmatched outcomes corrupt calibration, the golden question changes its threshold, and scoring and provenance helpers have correctness gaps.
Review effort: Balanced
Findings: 3
Open (3)
What changed in this PR
This PR adds point-in-time research integrity, safer simulation/reporting, sealed evaluation infrastructure, and platform hardening across DRF.
Changes:
- Adds hindcast, provenance, typed evidence, honest parsing, and market-integrity controls.
- Expands evaluation, scoring, ledger, telemetry, and cost-accounting capabilities.
- Hardens identifiers, provider handling, configuration, testing, and deployment.
| File | Description |
|---|---|
setup.sh |
Deploys new bridge modules. |
scripts/doctor.sh |
Audits honesty pins and provider requests. |
README.md |
Documents trust guarantees. |
NOTICE |
Adds third-party attribution. |
frontend/tests/marketStatus.test.mjs |
Tests incomplete market states. |
frontend/src/utils/marketStatus.js |
Classifies market lookup failures. |
frontend/src/components/research/DossierViewer.vue |
Displays incomplete-market messaging. |
drf2/skills/custom/forecast-report/SKILL.md |
Reframes probability calibration guidance. |
drf2/skills/custom/deep-research/SKILL.md |
Distinguishes failed and empty searches. |
drf2/driver/state.py |
Hardens pipeline identifiers. |
deerflow_bridge/skills/forecast-visuals/scripts/render.py |
Filters future-dated actuals. |
deerflow_bridge/research_budget.py |
Adds quota circuit handling. |
deerflow_bridge/market_tools.py |
Adds delayed HTTP retries. |
backend/tests/test_worldstate.py |
Tests non-finite vote handling. |
backend/tests/test_utils_security.py |
Makes URL tests hermetic. |
backend/tests/test_utils_dates.py |
Tests date precision periods. |
backend/tests/test_sim_market_priors.py |
Tests expired-market filtering. |
backend/tests/test_sim_event_provenance.py |
Tests event provenance helpers. |
backend/tests/test_seed_ensemble.py |
Pins ensemble seed derivation. |
backend/tests/test_research_spend_flush.py |
Tests cached-token accounting. |
backend/tests/test_research_gateway.py |
Tests shell detection and ledger rollback. |
backend/tests/test_research_evidence_quality.py |
Tests future-date classification. |
backend/tests/test_research_engine_v3.py |
Extends v3 question-spec fixtures. |
backend/tests/test_research_engine_v3_text.py |
Updates engine test doubles. |
backend/tests/test_research_engine_v3_round3_synth.py |
Updates synthesis test context. |
backend/tests/test_research_engine_v3_round2.py |
Tests subprocess and model telemetry. |
backend/tests/test_report_market_evidence.py |
Tests typed market absence. |
backend/tests/test_report_lint_absence.py |
Tests absence-marker leakage removal. |
backend/tests/test_quantity_scoring.py |
Tests quantity scoring rules. |
backend/tests/test_prediction_markets.py |
Pins market clocks. |
backend/tests/test_point_in_time.py |
Tests strict hindcast dates. |
backend/tests/test_parallel_research_merge.py |
Tests market-status merging. |
backend/tests/test_market_influence_boundary.py |
Stabilizes market anchoring tests. |
backend/tests/test_llm_recovery.py |
Tests token-cap escalation. |
backend/tests/test_forecast_diff.py |
Tests unscoreable probabilities. |
backend/tests/test_foglamp_containment.py |
Tests pinned numeric policy. |
backend/tests/test_env_drift_pins.py |
Tests honesty-critical pins. |
backend/tests/test_ensemble_backtest.py |
Tests unmatched outcomes. |
backend/tests/test_deerflow_bridge_sync_guard.py |
Verifies bridge deployment parity. |
backend/tests/test_camel_context_delivery.py |
Tests CAMEL prompt delivery. |
backend/tests/test_bridge_overhaul_v3.py |
Pins market time in bridge tests. |
backend/tests/test_bridge_market_tools.py |
Tests retry and outage semantics. |
backend/tests/test_binary_source_rule.py |
Tests forecast provenance prompts. |
backend/tests/test_audit_fixes_infra.py |
Tests recovery and resume messaging. |
backend/tests/test_actor_context_runtime.py |
Removes test-time network dependencies. |
backend/tests/fixtures/question_spec_golden.json |
Adds question-spec golden data. |
backend/tests/eval/rubric.md |
Tightens grounding criteria. |
backend/scripts/run_twitter_simulation.py |
Uses guarded environment loading. |
backend/scripts/run_reddit_simulation.py |
Uses guarded environment loading. |
backend/scripts/preflight.py |
Centralizes provider request overrides. |
backend/scripts/model_comparison.py |
Preserves null probabilities. |
backend/scripts/cost_card.py |
Adds offline cost-card rebuilding. |
backend/scripts/batch_runs.py |
Preserves fork policies and lineage. |
backend/run.py |
Separates startup validation from audits. |
backend/pyproject.toml |
Enables strict offline test markers. |
backend/app/utils/sim_timeline.py |
Documents question-spec horizons. |
backend/app/utils/security.py |
Adds safe identifiers and containment. |
backend/app/utils/provider_overrides.py |
Centralizes provider-specific requests. |
backend/app/utils/point_in_time.py |
Adds strict temporal parsing. |
backend/app/utils/oasis_llm.py |
Applies resolved-provider overrides. |
backend/app/utils/llm_recovery.py |
Adds token escalation policy. |
backend/app/utils/env_loading.py |
Prevents test-time .env loading. |
backend/app/utils/dates.py |
Adds precision-aware date periods. |
backend/app/utils/ctxpool.py |
Propagates context into workers. |
backend/app/utils/canonical_json.py |
Centralizes canonical hashing. |
backend/app/utils/atomic.py |
Adds secure writes and strict JSON. |
backend/app/services/backtest.py |
Tracks unmatched outcomes in calibration. |
backend/app/services/zep_graph_memory_updater.py |
Honors pinned feedback policy. |
backend/app/services/zep_entity_resolver.py |
Rejects ambiguous aliases. |
backend/app/services/worldstate.py |
Rejects non-finite decisions. |
backend/app/services/simulation_manager.py |
Contains simulation paths. |
backend/app/services/simulation_config_generator.py |
Adds horizon and market gates. |
backend/app/services/sim_event_provenance.py |
Labels injected simulation events. |
backend/app/services/report_visualizer.py |
Adds typed and validity-aware charts. |
backend/app/services/quantity_scoring.py |
Adds quantitative forecast scoring. |
backend/app/services/oasis_profile_generator.py |
Propagates context and filters markets. |
backend/app/services/graphiti_client/runtime.py |
Contains Kuzu graph paths. |
backend/app/services/graph_pruner.py |
Protects all actor aliases. |
backend/app/services/graph_builder.py |
Contains layout paths. |
backend/app/services/exec_brief.py |
Reports research degradation. |
backend/app/services/ensemble.py |
Excludes unreadable forecast runs. |
backend/app/services/actor_role_prompt.py |
Hardens actor identity matching. |
backend/app/models/project.py |
Contains project paths. |
backend/app/mcp/sim_server.py |
Validates simulation IDs. |
backend/app/mcp/kg_server.py |
Validates graph IDs. |
backend/app/api/simulation.py |
Hardens simulation file access. |
backend/app/api/settings.py |
Reports persistence and reuses overrides. |
backend/app/api/research.py |
Adds strict admission and options. |
backend/app/__init__.py |
Adds host and identifier gates. |
agent-progress.txt |
Records implementation progress. |
.gitignore |
Excludes secret temporary files. |
.github/workflows/ci.yml |
Expands offline CI coverage. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| if sc.get("outcome_matched_a_scenario") is False: | ||
| n_unmatched += 1 |
| def covered80(q10: float, q90: float, y: float) -> bool: | ||
| """Whether y lies in [q10, q90] (bounds inclusive).""" | ||
| return _finite(q10, "q10") <= _finite(y, "y") <= _finite(q90, "q90") |
| "operational_question": "Will installed global data-centre capacity exceed 250 GW on 31 December 2027?", | ||
| "outcome_definition": "Global installed data-centre IT capacity is at least 250 GW (≥ 250 GW) on 31 December 2027.", |

What this is
This PR brings into DeepResearchForecast (DRF) what three codebases and five research papers do well. The codebases are FinanceHarness 0.1.0, StockAgent and TradingAgents 0.5.1. The papers are F²Agent, the FinanceHarness paper, Nexus, TIEM and the StockAgent paper. Every idea was checked against DRF's code before it was built.
The full write-up is
docs/research/2026-09-29-finharness-essence.md. It covers the essence of each source and the roadmap. It also has the implementation record (wave by wave, with knobs and gate results), the open issues and the appendices. Start with its synthesis chapter.How it was built
What changed, by area
as_ofpolicy.needs_reviewand is never coerced.FakeLLMClientparity.Review notes
Configknob documented in.env.example: 145 new entries, plus 2 former ghost knobs now declared. At its default a knob leaves output byte-identical, unless its.env.exampleentry says the default is on. Most default-on knobs are fail-closed honesty fixes; a few are provenance or shadow diagnostics.RESEARCH_DERIVED_FINDINGS,RESEARCH_V3_FORECAST_INPUTSandREPORT_COUNTER_CASE. Each needs a live check before its default flips.FakeLLMClient.NOTICE, with the licence inLICENSES/Apache-2.0.txt. It covers the FRED/EDGAR readers indeerflow_bridge/data_tools.pyand the EDGAR test fixtures.DRF_ARCHITECTURE.md§20 maps every mechanism to its module, and the README gains trust section 6. A final fact-check checked these docs against the code. It confirmed 30 errors, and all are corrected.deerflow_bridge/linear_research.py,backend/app/services/report_agent.pyandpipeline_orchestrator.py. About 86k of the ~148k added lines are tests.mainat c20ec60 (the currentmain).Testing
python -m pytest -q -p no:cacheprovider: 12,410 passed, 0 failed, 53 skipped, 11 xfailed;node --test: 79/79.Follow-ups (not in this PR)
The research document lists these under "Open issues after implementation". The main ones: