This repository contains reusable code for the conformal calibration layer described in:
Gene Ontology DAG-Aware Conformal FDR Control for Multi-Label Protein Function Prediction: A Maize Application
The repository is intentionally focused on the calibration method. It takes black-box protein--GO score matrices and calibration labels as input, then returns calibrated score thresholds, prediction sets, and protein-level FDP summaries under exact or GO-DAG-aware losses.
Included:
- Learn--Then--Test conformal calibration for protein-level FDP control.
- Exact multi-label FDP loss.
- GO-DAG-aware near-1 and near-2 FDP losses from a processed GO edge list.
- A small toy protein--GO example for local testing.
- A command-line script for running calibration on user-provided score and label matrices.
Not included:
- Raw maize protein sequences.
- ESM-2 embedding extraction.
- Neural-network score-model training.
- Full maize benchmark score matrices or trained checkpoints.
The intended use is:
black-box protein--GO scores + calibration labels + optional GO-DAG edges
-> conformal calibration layer
-> calibrated threshold, prediction sets, FDP summaries
Tested with Python 3.10--3.12.
python -m pip install -r requirements.txtOptional editable install:
python -m pip install -e .Run:
python scripts/run_toy_example.pyExpected runtime: less than 1 minute.
The script prints exact, near-1, and near-2 calibration/evaluation summaries using the small example files in data/example/.
Input CSV format:
- first column:
protein_id - remaining columns: GO terms, for example
GO:0004672 - score files: numeric scores/probabilities in
[0, 1] - label files: binary labels
0/1
Example command:
python scripts/calibrate_scores.py \
--cal-scores data/example/cal_scores.csv \
--cal-labels data/example/cal_labels.csv \
--test-scores data/example/test_scores.csv \
--test-labels data/example/test_labels.csv \
--go-edges data/example/go_edges.csv \
--mode near1 \
--alpha 0.10 \
--delta 0.10 \
--outdir outputs/example_near1For exact multi-label FDP without GO-DAG expansion:
python scripts/calibrate_scores.py \
--cal-scores data/example/cal_scores.csv \
--cal-labels data/example/cal_labels.csv \
--test-scores data/example/test_scores.csv \
--test-labels data/example/test_labels.csv \
--mode exact \
--alpha 0.10 \
--delta 0.10 \
--outdir outputs/example_exactcalibrate_scores.py writes:
summary.json: selected threshold, calibration risk, observed FDP if test labels are supplied, and yield summaries.selected_predictions.csv: selected protein--GO pairs on the test set.threshold_grid.csv: calibration risk and p-value over the threshold grid.
The code implements bounded-loss Learn--Then--Test calibration using the supplied calibration set. As in conformal risk control, the validity of the risk-control statement depends on exchangeability/representativeness between calibration proteins and future target proteins, and on evaluating the same loss used for calibration.
The GO-DAG near-1 and near-2 modes use undirected shortest-path distance on the supplied edge list, matching the GO-DAG-aware loss described in the paper.
If you use this code, please cite the associated STAI-X 2026 paper:
@inproceedings{wang2026conformalgo,
title={Gene Ontology DAG-Aware Conformal FDR Control for Multi-Label Protein Function Prediction: A Maize Application},
author={Wang, Min and Wang, Chong and Liu, Peng},
booktitle={STAI-X 2026},
year={2026}
}