Skip to content
TianjingLiu9Public

About

Auditable cross-domain enterprise-risk modeling with version-frozen structural portability, evidence-constrained stress testing, and reproducible experiments.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

RiskGraph

RiskGraph is an evidence-constrained study of whether a version-frozen enterprise-risk interface remains semantically usable across heterogeneous domains. It evaluates structural portability, not a calibrated universal predictor. NVIDIA and Global Payments are the design domains; CrowdStrike operational resilience is a third-domain structural test; outputs are mechanism stress-test results with explicit claim boundaries.

Working paper / preprint (v3.1.2 correction candidate). This repository is the public research companion to the externally shareable manuscript. It is not a peer-reviewed, submitted, or accepted paper. Version 3.1.2 corrects the paper authorship and coding-provenance record; it does not change the registered results.


Paper

RiskGraph v3: Version-Frozen Structural Portability and Evidence-Constrained Enterprise-Risk Stress Testing

Olivia Liu (also publishes as Tianjing Liu)

Sole-authored Working Paper / Preprint v3.1.2 — August 2026

paper/RiskGraph_v3.1.2_External_Sharing_Preprint.pdf

The v3.1.2 PDF is the canonical local correction candidate. paper/main_v3.tex is its source; paper/main_v3.pdf is the same file under the LaTeX build name. The earlier v3.1 and v3.1.1 PDFs are retained only as historical provenance and must not be shared as current manuscripts. The public GitHub tag and Release remain v3.1.1 until this correction is explicitly authorized, published, and remotely verified.


Key idea

  • NVIDIA (AI infrastructure) and Global Payments (digital payments) are design domains. Their graphs, probabilities, and interventions are registered synthetic stress specifications, not calibrated company models.
  • RiskGraph-Core-v2 is hash-frozen before any CrowdStrike mapping. The frozen-v2 score against 55 CrowdStrike concepts is the only held-out structural evaluation.
  • CrowdStrike exposes structural representation failures in the frozen core: most material concepts do not map cleanly, and some must remain local.
  • The v3 schema extensions are post-audit remediation on the same 55 concepts. They are not a second held-out test and do not validate v3 on a new domain.
  • Public filings and official incident documents constrain scale and interpretation. They do not calibrate event probabilities or identify causal effects.
  • Reported numbers are stress-test contrasts under declared seeds, draws, and evidence classes. They are not forecasts, causal treatment effects, or investment advice.

Main results

Registered computational outputs. Interpretation follows the manuscript.

Result Value How to read it
Frozen-v2 CrowdStrike mapping 8 / 30 / 13 / 4 clean / partial / failed / local of 55 Stage 1: version-frozen held-out structural audit
Post-audit v3 CrowdStrike mapping 26 / 25 / 0 / 4 of the same 55 Stage 2: representation fit after the failure audit, not a new held-out test
NVIDIA / Global Payments mappings all 18 registered targets preserved under v3 Backward-compatibility check, not new validation
NVIDIA matched-marginal ES95 0.5543 → 0.5744 (independent → dependency) Designed synthetic stress
Global Payments matched-marginal ES95 0.2236 → 0.2344 Designed synthetic stress
Global Payments rollback/redundancy NLRCR −0.1333 Retained negative mitigation result (break-even at 0)
CrowdStrike conditional incident ELR 0.0100, ES95 0.0527 Uncalibrated; conditioned on a defective content release
CrowdStrike mechanism policy with lag ELR 0.0083, ES95 0.0511, NLRCR 0.5989 Non-causal policy stress contrast
CrowdStrike manual-probability diagnostic mean loss within Monte Carlo noise; ES95 worsens by 0.0013 Tail shift is robust; NLRCR sign is not

A fourth prospectively unseen domain (Stage 3) has not been tested. All 55 concepts were coded by the sole author under the released rubric. No independent second coder or adjudication process existed, so no inter-rater reliability statistic is available; independent re-coding remains future validation.

The repository now also contains a separate, pre-coding Stage 3 extension under research_extensions/riskgraph_semantic_stage3/. It publishes an executable OWL/SHACL/PROV-O semantic profile, the frozen NTSB source-selection record, and a source-only 38-unit inventory. It contains no RiskGraph mappings, human labels, coder attestations, IRR estimate, or Stage 3 result; those gates remain explicitly blocked.


Repository structure

RiskGraph/
├── README.md
├── CITATION.cff
├── LICENSE
├── .gitignore
├── requirements.txt
├── requirements-lock.txt
├── pytest.ini
├── paper/
│   ├── RiskGraph_v3.1.2_External_Sharing_Preprint.pdf # correction candidate
│   ├── main_v3.tex / main_v3.pdf           # source + same PDF
│   ├── generated/v3_results.tex            # registered numerical macros
│   └── figures/v3/                         # ten figures (PDF/SVG/PNG)
├── src/                                    # ontology, adapters, inference, evaluation
├── configs/                                # frozen v2 core, v3 schema, domain adapters
├── data/                                   # public-safe derived ledgers (see data/README.md)
├── experiments/                            # registered result CSVs
├── research_extensions/
│   └── riskgraph_semantic_stage3/           # formal profile + pre-coding Stage 3 package
├── tests/                                  # 79-test suite
├── tools/                                  # release-inventory helpers used by tests
└── docs/
    ├── REPRODUCIBILITY.md
    ├── AUTHORSHIP_AND_SHARING.md
    └── EVIDENCE_AND_DATA.md

paper/main_v2.tex is a historical LaTeX record of the frozen-v2 panel. It is not the public paper.


Reproduction

From a fresh clone:

python3 -m venv .venv
source .venv/bin/activate
python -m pip install -r requirements.txt
pytest -q

pytest.ini adds the repository root to pythonpath so from src... imports resolve without a manual PYTHONPATH for the test suite.

Expected: 79 passed. The added clean-tree regression verifies that regenerating figures does not alter the shipped SVG bytes. For the environment known to reproduce the frozen numerical exports bit-for-bit, install requirements-lock.txt instead (Python 3.12, numpy 2.5.2, pandas 3.0.5, matplotlib 3.11.1, pytest 9.1.1).

The Stage 3 extension has its own pinned environment and 17-test suite. From the repository root:

python3 -m venv research_extensions/riskgraph_semantic_stage3/.venv
research_extensions/riskgraph_semantic_stage3/.venv/bin/pip install -r research_extensions/riskgraph_semantic_stage3/requirements.txt
cd research_extensions/riskgraph_semantic_stage3
PYTHONPATH=src .venv/bin/pytest -q
.venv/bin/python scripts/validate_semantics.py --out artifacts/semantic_validation.json
.venv/bin/python scripts/run_competency_questions.py --out artifacts/competency_question_results.json

The official NTSB PDF is not redistributed. The extension ships its URL, frozen SHA-256, source locators, unit hashes, and deterministic inventory. Rebuilding the inventory from the original source is an optional source-document check; see the extension README.

For a local, professor-shareable package (without committing, pushing, tagging, or creating a GitHub Release), run:

python tools/build_v3_1_2_external_sharing_candidate.py

The builder tests a fresh candidate copy, verifies its manifest again after tests, and writes the PDF, deterministic ZIP, checksums, and receipt under output/. It records sole authorship, historical-source permission, and the absence of independent-coder IRR and prospective Stage 3.

Registered experiment CSVs are already in experiments/. To regenerate them from the shipped ledgers and configs:

PYTHONPATH=. python -m src.evaluation.reconstruct
PYTHONPATH=. python -m src.evaluation.source_audit
PYTHONPATH=. python -m src.evaluation.cross_domain
PYTHONPATH=. python -m src.evaluation.interventions
PYTHONPATH=. python -m src.evaluation.ablations
PYTHONPATH=. python -m src.evaluation.uncertainty
PYTHONPATH=. python -m src.evaluation.sensitivity
PYTHONPATH=. python -m src.evaluation.export_empirical_validation
PYTHONPATH=. python -m src.evaluation.crowdstrike_source_quality
PYTHONPATH=. python -m src.evaluation.crowdstrike_experiment
PYTHONPATH=. python -m src.evaluation.v3_registry
PYTHONPATH=. python -m src.adapters.v3_schema
PYTHONPATH=. python -m src.evaluation.mapping_reliability
PYTHONPATH=. python -m src.evaluation.mc_robustness
PYTHONPATH=. python -m src.evaluation.export_v3_latex
PYTHONPATH=. python -m src.visualization.make_v3_figures

Production simulations use seed 5900 and 200,000 draws. python -m src.adapters.v3_schema must run before export_v3_latex because it regenerates experiments/cross_domain/portability.csv. mc_robustness must run before export_v3_latex because the LaTeX macros import that audit.

The manuscript can be recompiled with a TeX distribution that provides latexmk and BibTeX:

cd paper
latexmk -pdf -interaction=nonstopmode -halt-on-error main_v3.tex

See docs/REPRODUCIBILITY.md.


Evidence and data boundary

  • Original third-party and course-report PDFs are not redistributed.
  • Public derived ledgers in data/ contain the coded facts, hashes, official URLs, and extracted numbers required for registered computational reproduction.
  • Citations in data/crowdstrike/source_manifest.csv and the bibliography point to primary public sources (SEC filings, CrowdStrike and Microsoft incident reports).
  • Source-document reproduction (re-reading the original Capstone or course report) is distinct from computational reproduction (re-running the shipped code on the shipped ledgers).

Details: data/README.md and docs/EVIDENCE_AND_DATA.md.


Scientific scope

RiskGraph v3.1.2 is not:

  • a calibrated forecast of company loss, incident frequency, or enterprise value;
  • a causal-control or ROI estimate (no do-calculus identification);
  • investment, insurance, or operational advice;
  • a formal or universal enterprise-risk ontology;
  • a demonstration that v3 generalizes to a fourth domain.

The economic diagnostic is the net loss-reduction-to-cost ratio (NLRCR) = (ELR_base − ELR_policy) / E[cost], with control cost already inside the policy-arm loss. Break-even is 0. It is not a conventional benefit–cost ratio.


Author and historical source boundary

RiskGraph is an independently developed, sole-authored working paper by Olivia Liu, who also publishes as Tianjing Liu. She alone formulated the present research question, designed and implemented the framework and evaluation protocol, performed the source reconstruction and coding, ran and validated the experiments, analyzed the results, prepared the figures, and wrote and revised the manuscript and release materials. Correspondence: tl3393@columbia.edu.

The study builds on case materials from an earlier Columbia University ERM Capstone completed by a seven-person student team. The other Capstone team members contributed to that earlier course project and granted permission to reuse applicable team-authored material; they did not contribute to the present RiskGraph research, code, experiments, analysis, or manuscript and are not authors of this working paper. Private permission records are not distributed with this repository. See docs/AUTHORSHIP_AND_SHARING.md.


Citation

See CITATION.cff.

@unpublished{liu2026riskgraph,
  author = {Liu, Olivia},
  title  = {RiskGraph v3: Version-Frozen Structural Portability and
            Evidence-Constrained Enterprise-Risk Stress Testing},
  year   = {2026},
  month  = {8},
  note   = {Working paper / preprint},
  url    = {https://github.com/TianjingLiu9/RiskGraph}
}

Sharing and reuse boundary

This repository is not released under an open-source or Creative Commons license.

The author has approved public release and external sharing for scholarly review. Permission to reuse applicable historical Capstone material is recorded privately. Those permissions do not grant permission to republish, redistribute, modify, sublicense, commercialize, or train models on the code, manuscript, figures, or derived ledgers. Private, local execution solely to evaluate the stated reproducibility claims is permitted; redistribution is not. Third-party and course source documents are not included and are not licensed here.

See LICENSE. Any broader reuse requires separate written permission from the rights holders.

About

Auditable cross-domain enterprise-risk modeling with version-frozen structural portability, evidence-constrained stress testing, and reproducible experiments.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages