diff --git a/.env.example b/.env.example index 9ffc13e..8c05e41 100644 --- a/.env.example +++ b/.env.example @@ -166,7 +166,46 @@ LLM_PROVIDER=claude-cli # RESEARCH_LINEAR_GLM_EFFORT_AGENT=low # 仅 GLM:reasoning_effort(low | high | max)——采集代理步。GLM 恒开思考且默认 max,输出/时延代价极高,故显式压档 # RESEARCH_LINEAR_GLM_EFFORT_JSON=low # 仅 GLM:reasoning_effort——结构化 JSON 调用 # RESEARCH_LINEAR_GLM_EFFORT_WRITE=high # 仅 GLM:reasoning_effort——分节写作(质量优先) +# RESEARCH_VERIFIED_FACTS=true # v3: check each quantitative.json row's number against the fetched page of its cited source; adds verification (verified | unverified | snippet_only | none; snippet_only = source never fetched or its stored page unavailable, snippet not checked; absent = unchecked: no checkable number on a fetched page) and a verified bool, never changes values; zero model calls. REPORT-7: the same knob publishes evidence windows (the page sentence stating a verified figure next to >= 2 of its metric words, <= 360 chars, <= 2 per figure, <= 8 per source with every figure's best window first; windows over the cap are counted in windows_dropped) as sources.json supports, keeps quant source_ref, adds evidence_window / future_dated, and writes handoff verified_facts.json (drf.verified_facts/v1, SHA-manifested). false = every research artifact byte-identical. The orchestrator forwards it to the v3 child. +# RESEARCH_QUANT_TYPING=false # v3: type quantitative rows reported/projected (epistemic_class, date_precision, target-date repair when as_of_date holds a forecast target), bucket future-dated rows as future_dated instead of fresh in meta.quant_freshness, and add a date-semantics rule to the facts prompt. Off = byte-identical; needs a live A/B before the default flips. Forwarded to the v3 child. +# RESEARCH_QUANT_RECONCILE=true # TIME-4 quantitative sanity checks (every engine): quantitative rows on the same metric and unit (v3: also the same period end and length, geography and reported/projected class, but not the series name, so two entities' readings of one generic metric can still reconcile) that disagree by > 10% become contested.json claims (origin quant_reconcile; v3 adds at most 10, probable unit-scale errors first, actors.json keeps the extracted claims), a ~1000x gap is also a probable unit-scale error in meta.quant_unit_warnings, and claimed actuals dated after the research as-of (v3: also those whose as_of_date the bridge cannot read, or for a period ending after it) or with > 150% growth are listed in meta.quant_implausible (v3: rows typed reported or future_dated_reported, else value_type actual or absent; each meta list <= 20, meta.quant_sanity_truncated keeps the totals). Read-only: quantitative.json never changes. The legacy engine always ran them (its full run lists quant_implausible regardless of this knob); this restores them in v3 and the extract-only salvage. Parsed like the bridge (blank = true, else 1/true/yes/on); false = v3 and extract-only artifacts and meta byte-identical to before. Forwarded to every research child. FU-9: the report agent reads it too: when its contested-claims block (at most 15 claims) would cut quant_reconcile rows, up to 3 slots (more when the plain cut already shows more) go to them, probable unit-scale errors first, and a note counts those still cut; nothing changes when nothing is cut, and false = the plain first-15 cut +# QUANT_TYPED_RENDERING=false # RESEARCH-5: render RESEARCH_QUANT_TYPING's stamps on the report side: a persona's projected/unknown quantitative fact gets "(expectation by {source}, target {period})" (", as of {date}" when the row states only the source's date) and chart labels read epistemic_class first (projected -> forecast marker, reported -> actual). Off = byte-identical role prompts and charts; on changes role-prompt SHAs (PREPARE recomputes them). +# RESEARCH_FORECASTER_ATTRIBUTION=false # RESEARCH-6 v3: the facts prompt asks estimate/forecast/target rows for their forecaster (also written to analyst, so charts split by forecaster, not publisher) with a forecaster-free metric, plus the low/high range and n_forecasters the report states; a bound is kept only when its numbers are the report's, a count only next to a count noun ("40 economists", "n = 40", "40位经济学家"), a value written as a range becomes low/high (range_kind stated_range), actual rows never carry these keys; meta.forecaster_attribution counts kept and dropped fields. Off = byte-identical facts prompt, quantitative.json and meta; on adds ~5-10% extraction output and changes which quant rows match an actor in PREPARE context packs (a row that matched still matches and keeps its place at the 32-row pack cap; forecaster-only matches fill spare slots only). Forwarded to the v3 child. +# RESEARCH_EVIDENCE_HEADERS=false # RESEARCH-10 v3: each KIQ block of the writers' evidence digest opens with an engine count of its evidence (sourced findings by VERIFIED/REPORTED/UNVERIFIED tag, cited sources fetched vs snippet-only, distinct domains; counted before lines are dropped for length, '; k lines omitted for length' when they were) and a sufficiency label (insufficient: deterministic fallback notes, < 3 sourced findings or no fetched source; thin: < 2 VERIFIED, < 2 fetched or < 2 domains; else adequate; DRF-original thresholds, validate on stored kiq/*.json before the default flips); a legend opens the digest, the gap review's coverage matrix gains fetched/domains/sufficiency, the section rules gain one thin-evidence line, meta.kiqs gains evidence (per-KIQ profiles) and sufficiency (counts per label), and a degradation event fires when at least half of the researched KIQs are insufficient. Zero model calls. Off = digest, coverage matrix, section task and meta byte-identical. Forwarded to the v3 child. +# RESEARCH_TRUNCATION_FAIRNESS=false # RESEARCH-10 v3 fair deterministic truncation: the plan's 6,000-char scout digest is shared max-min fairly between the scout queries (each keeps at least an equal share), cut only between search results with '(k results omitted for length)', instead of a head-cut that silently lost the last queries; an over-cap digest block drops, among lines of equal priority, the one sharing the fewest terms with the KIQ question first (not simply the last). Zero model calls. Off = plan prompt and digest byte-identical. Forwarded to the v3 child. +# RESEARCH_AS_OF_PIN=true # TIME-1 honesty fix: v3 actors.json as_of_date is always the plan's as-of (UTC date fixed at plan time); a different value from the actor-extraction model is recorded in meta.json as_of_model_disagreement and never adopted (it anchors graph valid_at and the sim calendar). Fails closed: only 0/false/no/off disable it. false = the model's YYYY-MM-DD value is adopted as before. Forwarded to the v3 child. +# HINDCAST_ENABLED=false # TIME-7 hindcast admission: accept as_of (canonical YYYY-MM-DD, not in the future) on /api/research/run, /api/v1/run and PipelineOrchestrator.start; requires RESEARCH_ENGINE=v3. An admitted run pins hindcast_policy_v1 and an evaluation-run pin (eval_run_id hindcast_YYYYMMDD: evaluation ledger, no production calibration), runs v3 research with RESEARCH_AS_OF (brief dated as_of plus a point-in-time rule, fetched pages labelled LIVE PAGE) and prediction markets withheld, and anchors the graph at the pinned date (later-dated sources recorded in hindcast_violations, never adopted; hindcast_source_dates counts dated/undated sources, since v3 dates them only with RESEARCH_SOURCE_DATES). A research timeout that leaves no actors.json dated to the pin fails the stage (resumable) instead of running the legacy extract-only salvage. Characterization-only. false = any request carrying as_of is rejected (400 / ValueError), never run live. +# PIT_GATES=true # TIME-8 point-in-time evidence gates of a hindcast's v3 research, pinned at admission into hindcast_policy_v1.pit (a later config change never alters an admitted run) and effective only inside a pinned hindcast (live runs unaffected). The parent passes the pin to the v3 child as RESEARCH_PIT_GATES/SAME_DAY/UNDATED/PROVIDER_BOUNDS/OVERFETCH and forces RESEARCH_SOURCE_DATES=true (ambient RESEARCH_PIT_* values never reach a child). A source is available on the latest day consistent with its latest publication/update date (a month/quarter/year counts as its end); late search rows are dropped before they get an [S] id (all late = NO_IN_WINDOW_RESULTS, "not evidence that nothing happened"), URL-dated-late fetches are refused without budget, late pages are withheld (FETCH_WITHHELD) before storage, and each verdict is recorded as the ledger pit_status and counted in meta.tools.pit. TIME-9: the report then cites only sources admissible as of the as-of (References, sources.json and the writers' evidence digest are filtered), the research writes point_in_time.json (the gates' counts plus an independent date re-check of the cited sources) and its verdict is stamped into the pin before the report: forecast.json hindcast.integrity date_verified / date_verified_with_unverifiable / leak_suspected, run.json as_of_enforcement.retrieval_clamped. Only 0/false/no/off disable it +# PIT_SAME_DAY_POLICY=exclude # TIME-8: exclude (strict: a source available on the as-of day itself is late) | include (admitted, labelled same-day). Other values = exclude +# PIT_UNDATED_POLICY=drop # TIME-8: a fetched page with no readable date: drop (withheld as FETCH_WITHHELD(pit_undated); can starve undated data pages) | flag (stored with pit_status unverifiable, labelled undated). Undated search rows are always shown, labelled undated. Other values = drop +# PIT_PROVIDER_DATE_BOUNDS=true # TIME-8: gated searches ask the provider for a date bound at the as-of (Firecrawl tbs cdr:1,cd_max:M/D/YYYY; Serper/Tavily/DDG cannot and are counted unbounded; a bounded request Firecrawl rejects with HTTP 400/422 is retried once unbounded and the bound stays off for that process). Gated search cache entries are keyed apart from live ones (and bounded from unbounded) either way, and the row gate always runs. Only 0/false/no/off disable it +# PIT_SEARCH_OVERFETCH=1 # TIME-8: rows requested per gated search = 5 x this (clamped 1..4) so late rows can be dropped without starving the 5 render slots; Firecrawl still caps billed rows at RESEARCH_FIRECRAWL_SEARCH_LIMIT # RESEARCH_LINEAR_MODE= # 旧 linear v2 旋钮。v3 忽略 salvage 语义:绝不采纳非本引擎产出的既有 research_report.md;设为 salvage 仅等价于「续跑 v3 自身已持久化的阶段」(身份 = 问题/深度/模型/语言/引擎版本 匹配时 v3 本就总会续跑),设置时只记一条日志 +# RESEARCH_FETCH_SHELL_DETECTION=true # RESEARCH-1 honesty check: reader shells ("Markdown Content: undefined"), "page unavailable" pages, bot walls and short paywall teasers are failed fetches — never cached, they trigger provider failover, the v3 tool layer answers FETCH_FAILED() and never marks them fetched, and a shell stored before the check (resumed work dir) no longer counts as fetched and is published as cited (fetch_status=shell:). Fails closed: only 0/false/no/off disable it (blank, 1/yes/on or a typo keep it on). false = previous shell handling byte for byte (the direct-fallback PDF parser fix is unconditional) +# RESEARCH_FETCH_CALL_TIMEOUT_S=150 # RESEARCH-1 hard wall-clock bound (seconds) of one v3 web_fetch; on expiry the fetch returns FETCH_FAILED(fetch_call_deadline_exceeded) (not retried) without waiting for leftover executor threads. Successful fetches are unchanged; 0 = the previous asyncio.run path +# RESEARCH_EVIDENCE_QUOTES=off # RESEARCH-7 verbatim evidence spans for v3 findings: off | audit | enforce. Not off: the KIQ task asks each finding to end with EVIDENCE: "", the source ledger keeps every distinct snippet of a row, and each quote is located deterministically (zero model calls; exact, normalized or ordered elided segments) in the cited sources' stored pages / search text; KIQ facts get evidence, evidence_status (verified | failed | absent), evidence_near_miss, claimed_tag and a REPORTED-number audit (number_check, telemetry only), meta.json gets evidence and an [evidence] progress line per KIQ. audit changes no tag; enforce demotes facts whose quotes are not on what the agent was shown (UNVERIFIED, evidence_not_on_page) and VERIFIED facts whose numbers lie outside the quoted passages (REPORTED, numbers_outside_evidence), and flags a run where under half of the claimed-VERIFIED findings have a located quote. Off = byte-identical; blank/unknown = off; enforce only after an owner-approved audit run (located share >= 0.8, QA passing). Forwarded to the v3 child +# RESEARCH_EVIDENCE_SUPPORTS=false # RESEARCH-7: with RESEARCH_EVIDENCE_QUOTES not off, up to 3 located verbatim quotes per source (<= 280 chars, page quotes first) become its sources.json supports ahead of the REPORT-7 evidence windows. Changes the report's semantic-citation / quote-grounding inputs, so measure the gate delta in an audit run first. false = supports as before. Forwarded to the v3 child +# RESEARCH_DERIVED_FINDINGS=false # RESEARCH-8 v3 declarative derived findings: the KIQ task asks a finding that states a figure the agent calculated to end with (DERIVED: ; a= [S], b= [S]) and the notes postprocessor recomputes it with zero model calls (hardened Decimal evaluator: only the literals 0/1/100/1000, + - * / **, abs/min/max/sqrt/ln/exp/log10; every operand, and both years of a years(Y1,Y2) period, from one fetched page the finding cites and on that page at its full value; a stated number equal to the signed result at its display precision as written, sign and scale word included, a percentage result only as a percentage, a unit class only that of every operand through a formula that keeps it (sums, differences, literal scaling); every other number of the finding on that page too): pass = DERIVED (never VERIFIED), fail = UNVERIFIED with derivation_error. The digest shows the calculation, the section rules say how to state it, sources.json gains a separate derived_supports field (the report reads it without its calculation), unverified quant rows matching a DERIVED result gain derived_from, meta gains kiqs.derived and derived. Off = prompts, facts, digest, sources.json and meta byte-identical. Forwarded to the v3 child +# RESEARCH_SOURCE_TAXONOMY=false # RESEARCH-2 typed source outcomes: search budget denials no longer count as fetch failures; a Firecrawl 401/402 on search latches the v3 run (SEARCH_NOT_CONFIGURED, zero further backend calls) and on fetch disables Firecrawl for the process and opens the shared provider circuit; outages never enter the gateway run cache or the fetch negative cache; DDG "No results found" becomes SEARCH_EMPTY_UNCONFIRMED; fetch failures read FETCH_UNAVAILABLE (service) or FETCH_FAILED (page); writes meta.source_health. Off = byte-identical; needs a live comparison run before the default flips. Forwarded to every research child +# RESEARCH_ABSENCE_DISCIPLINE=false # RESEARCH-3 v3 absence discipline: an empty web search answers NO_RESULTS plus "Results are relevance-ranked and undated: an empty search is not evidence that something did not happen.", the KIQ task and the section rules each gain one line (state that something did not happen or was not reported only when a cited source says so; otherwise an open question / "the sources reviewed do not establish it"), and findings stating an absence are counted by tag in KIQ records (absence_cues) and meta.json (absence_findings: total, VERIFIED, REPORTED, UNVERIFIED) with no tag change. Off = tool text, prompts, records and meta byte-identical; needs a live comparison run before the default flips. Forwarded to the v3 child +# RESEARCH_QUESTION_SPEC=false # RESEARCH-11 v3 question spec: one post-scout JSON call (reusing the plan call's cached prefix) pins the operational question, outcome definition, resolution source (name/url/kind), horizon (label/date/basis; date kept only in (as_of, as_of+30y]) and reference class, and discloses at most 3 defaults it chose instead of asking (most important slot first). Written to handoff question_spec.json (drf.question_spec/v1, spec_sha256, SHA-manifested) before gathering, appended to the run brief, fed to the plan call (scenarios partition the spec's outcome; one KIQ establishes the reference-class base rate) and mirrored into actors.json (question_spec, horizon_date, forecast_inputs.base_rates). A failed call never fails the run: status unavailable + a degradation event. Off = brief, plan, actors.json and every prompt byte-identical. Forwarded to the v3 child +# QUESTION_SPEC_DOWNSTREAM=true # RESEARCH-12 backend consumers of a valid actors.json question_spec (schema drf.question_spec/v1, status ok/partial, spec_sha256 recomputed; anything else is ignored): the simulation calendar horizon below every deterministic prompt date (skips the LLM fallback; horizon_source question_spec) and the world-state seed's horizon fallback, a spec block first in the spine prompt's research inputs (only when the spec's deadline is the run's horizon: an explicit prompt date or an out-of-window day keeps it out, and forecast.json records horizon_applied=false), an operational-definitions/assumptions subsection in the report's resolution section and forecast.json question_spec. No spec (needs RESEARCH_QUESTION_SPEC) = byte-identical; false = shadow mode (persisted, unused) +# RESEARCH_V3_CITATION_STATS=true # RESEARCH-9 v3: the QA phase writes qa.json citation_stats (v3-citation-stats/1: markers before QA and published, orphan markers renumbering dropped (occurrences, distinct, sample), stale citation groups, cited fetched vs snippet sources and the snippet marker share, unused fetched pages, writer bibliographies and scaffold echo lines detected, prose numbers no VERIFIED/REPORTED finding, cited page or cited snippet traces), mirrored into meta.research_qa and meta.research_quality. Detection only: research_report.md and sources.json byte-identical; false = no key. Forwarded to the v3 child +# RESEARCH_JSON_STRICT_NUMBERS=true # INFRA-4 v3 research gateway: parse_json_object rejects NaN / Infinity / -Infinity and overflowing floats (1e999), and a dict nested inside such an object is not taken for the reply, so the JSON retry re-asks ("NaN or Infinity is not a JSON number") instead of handing a NaN to plan/facts/handoff artifacts; control characters stay tolerated; finite replies parse byte-identically; false = previous permissive decoder. Forwarded to the v3 child +# RESEARCH_V3_FORECAST_INPUTS=false # RESEARCH-11 v3: the facts extraction also asks for forecast drivers and dated leading indicators (precision-preserving dates, never padded) that fill actors.json forecast_inputs.drivers / .indicators (v3 wrote them empty). Needs RESEARCH_FORECAST_INPUTS=true too. Off = facts prompt and actors.json byte-identical; needs one live A/B before the default flips. Forwarded to the v3 child +# RESEARCH_SOURCE_DATES=false # TIME-2 v3 source publication dates: provider metadata (Firecrawl scrape metadata and search row dates, Exa published_date, direct-fetch JSON-LD//