Target: Replicate the core methodology from "Practical Black-Box Attacks against Machine Learning" (Papernot et al., 2017).
This repository has moved beyond a single reproduction checkpoint and now contains:
- a full-test-set FGSM transfer baseline
- substitute-training ablations (
vanilla,pss,pss_rs) - query-budget sweeps
- step-wise mechanism logging
- attack-family comparison (
FGSM,BIM,PGD, and a placeholderJSMA-likeimplementation) - baseline-vs-wide strong-oracle cross-comparison
- multi-seed robustness validation for the strong baseline oracle
The most up-to-date quantitative findings are stored in:
current_findings_2026-08-07.mdtechnical_report_strong_oracle.mdresults_index.mdresults/figures/
If you are opening this repository for the first time, read these in order:
current_findings_2026-08-07.mdtechnical_report_strong_oracle.mdresults_index.mdfigures/
What each file is for:
current_findings_2026-08-07.md- newest concise research summary
- includes baseline multi-seed robustness, cross-oracle mechanism dynamics, and cross-oracle attack-family findings
technical_report_strong_oracle.md- formal write-up of the strong-oracle story
- now includes compact cross-oracle tables and figure references
results_index.md- marks which
results/*.jsonfiles are primary, supporting, or legacy - prevents accidental mixing of new
crossoracle_*results with older default-named files
- marks which
figures/- contains the main visual evidence, including the new cross-oracle comparison plots
Current primary result files:
results/multiseed_validation_baseline_oracle_seed0-4.jsonresults/surrogate_training_seed0_steps8_init10_pss_mechanism_crossoracle_baseline.jsonresults/surrogate_training_seed20_steps8_init10_pss_mechanism_crossoracle_wide.jsonresults/attack_family_comparison_seed0_cross_oracle_baseline_pss.jsonresults/attack_family_comparison_seed20_cross_oracle_wide_pss.json
Current primary figure files:
figures/cross_oracle_transfer_asr_vs_substitute_step.pngfigures/cross_oracle_agreement_vs_substitute_step.pngfigures/cross_oracle_attack_family_transfer.png
pssagreement advantage is stable across seeds0..4under the strong baseline oracle.psstransfer advantage atepsilon = 0.3is stable across seeds0..4under the strong baseline oracle.- Early transfer emergence appears under both validated strong-oracle conditions.
- The wide strong oracle begins with slightly stronger early transfer, but ends with weaker medium- and high-epsilon transfer than the baseline strong oracle.
- Under matched full-budget
psssurrogates,FGSMis the most stable transfer family across the two strong-oracle conditions. - Simple gradient-alignment proxies remain too weak to serve as a standalone explanation for transfer differences.
-
Oracle (Target Model): A black-box CNN model trained locally on the full MNIST training set. The current checked-in
oracle.ptandartifacts/oracle_seed0.ptevaluate to about 98.87% test accuracy and should be treated as the current strong local oracle baseline. -
Surrogate Model Training (Jacobian-based Dataset Augmentation):
- Initial small training set: 100 samples from MNIST.
- Query the Oracle for initial labels.
- Iteratively augment the dataset (8 steps):
- Compute the Jacobian of the surrogate w.r.t. inputs.
- Shift inputs in the direction of the gradient (
$\lambda=0.1$ ), withvanilla, periodic step size (pss), and reservoir-sampling (pss_rs) variants available. - Query Oracle for labels of new points.
-
Adversarial Example Generation:
- Untargeted FGSM baseline in
attack.py. - Additional family comparison in
attack_family_comparison.pyoverFGSM,BIM,PGD, and a current placeholderJSMA-likeimplementation. -
Important: Perturbations are applied in the
$[0, 1]$ pixel space before normalization.
- Untargeted FGSM baseline in
-
Transferability Evaluation:
- Evaluate on the full 10,000-image MNIST test set, filtered to samples jointly classified correctly by both oracle and surrogate before attack.
This repository now supports a two-layer YAML system:
- a scientific master config in
configs/ - executable experiment orchestration configs in
experiments/
The experiment YAMLs reference the scientific master config and define runnable step sequences.
Current scientific master config:
configs/mnist_blackbox_transfer.yamlexport_config_summary.pycan turn it into an advisor-facing Markdown summary
This repository now supports reproducible experiment orchestration through YAML configs in experiments/.
Runner:
run_experiment_yaml.py
Current configs:
experiments/wide_oracle_minimal_cross_oracle_baseline.yamlexperiments/wide_oracle_pss_attack_family.yamlexperiments/wide_oracle_multiseed_validation.yamlexperiments/wide_oracle_budget_sweep.yamlexperiments/wide_oracle_reservoir_cap_sweep.yamlexperiments/wide_oracle_mechanism.yamlexperiments/baseline_oracle_multiseed_validation.yamlexperiments/paired_cross_oracle_mechanism.yamlexperiments/paired_cross_oracle_attack_family.yaml
Schema features:
schema_version: 1- grouped
defaultswithcommon,by_script, and optionalby_step config_reflinking each experiment config to a scientific master config- step-level
outputs - step-level
overwrite - run artifacts saved under
runs/<experiment>/<timestamp>/
Examples:
python3 run_experiment_yaml.py experiments/wide_oracle_minimal_cross_oracle_baseline.yaml --dry-run
python3 run_experiment_yaml.py experiments/wide_oracle_pss_attack_family.yaml
python3 run_experiment_yaml.py experiments/paired_cross_oracle_mechanism.yaml
python3 run_experiment_yaml.py experiments/paired_cross_oracle_attack_family.yamlCurrent direct scientific-config consumers:
attack.pyattack_family_comparison.pyrun_multiseed_validation.pysurrogate.pycheck_agreement.pyrun_budget_and_mechanism.pyoracle.py
These currently read the following master-config fields directly when --config is passed:
dataset.rootattack.primary_baseline.epsilonsevaluation.batch_size_defaultevaluation.agreement_batch_sizesurrogate_model.training_defaults.stepssurrogate_model.training_defaults.initial_size_per_classsurrogate_model.training_defaults.epochs_per_stepvictim_model.training_defaults.epochsvictim_model.training_defaults.batch_sizevictim_model.training_defaults.learning_rate
Python dependencies for YAML-driven experiments are listed in:
requirements.txtrequirements-lock.txtenvironment.lock.yaml
Example scientific-config summary export:
python3 export_config_summary.py configs/mnist_blackbox_transfer.yaml --output config_summary.mdgit clone <your-repo-url>
cd papernot-replication
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txtIf you want to match the currently recorded local runtime as closely as possible, consult:
requirements-lock.txtenvironment.lock.yaml
Then ensure MNIST raw files are available under:
./data/MNIST/raw
Run a minimal reproducibility sanity check:
python3 verify_reproducibility.pyYou can inspect the scientific master config with:
python3 export_config_summary.py configs/mnist_blackbox_transfer.yaml --output config_summary.mdAnd validate a runnable experiment without executing it via:
python3 run_experiment_yaml.py experiments/wide_oracle_mechanism.yaml --dry-runSource:
results/fgsm_transfer_eval_seed0.jsonresults/budget_mechanism_cap_sweep_seed0.json
| Epsilon ( |
Surrogate ASR (White-box FGSM) | Oracle ASR (Transfer FGSM) |
|---|---|---|
| 0.2 | 91.29% | 2.75% |
| 0.3 | 97.07% | 24.54% |
| 0.4 | 97.82% | 60.68% |
Source:
results/attack_family_comparison_seed0.json
At the current local oracle-surrogate pair:
FGSMis the strongest observed transfer baselinePGDandBIMare stronger white-box attacks on the surrogate but transfer substantially worse thanFGSM- the current
JSMA-likeimplementation is a placeholder and is not a faithful scientific baseline
- The current strong local oracle is much more realistic than the earlier weak local oracle used in prior runs, but it is still not the original remote MetaMind oracle from Papernot et al. Results should therefore be interpreted as local-replication results, not direct quantitative confirmation of the original paper's strongest reported numbers.
- Current results are based primarily on
seed=0unless otherwise noted. - Query-budget and reservoir-cap experiments are now separated conceptually:
query_budgetcontrols total oracle queries, whilereservoir_capcontrols retained substitute dataset size. - Mechanism results should be read carefully: current evidence supports several empirical findings, but gradient-level mechanism claims remain hypotheses until directly measured.
Moving from the earlier weak local oracle (~80.18%) to the current strong local oracle (98.87%) changed several conclusions:
- transfer rates dropped substantially across the board
FGSMbecame the strongest observed transfer attack familyPGDandBIMremained extremely strong white-box surrogate attacks, but no longer produced the strongest transfer
This ranking reversal is one of the most important current findings in the repository.
- Problem: ASR was initially <4%.
-
Solution:
- Re-trained Oracle on full MNIST. The previous 100-sample Oracle had non-transferable noisy boundaries.
- Moved FGSM attack to unnormalized
$[0, 1]$ space. Previous epsilon was effectively 3x smaller than intended due to being applied in normalized space. - Corrected Jacobian augmentation to follow Oracle labels for gradient direction.
current_findings_2026-08-07.md: most up-to-date short research summarytechnical_report_strong_oracle.md: formal report with cross-oracle updatesresults_index.md: primary vs supporting vs legacy result mapartifacts/oracle_seed0.json: current verified oracle accuracy recordartifacts/oracle_wide_seed10.pt: second strong-oracle condition used for cross-oracle comparisonexperiments/: YAML experiment configsconfigs/: scientific master configurationrun_experiment_yaml.py: minimal YAML experiment runnersurrogate.py: substitute training variants, query-budget control, and mechanism loggingrun_budget_and_mechanism.py: query-budget sweep, mechanism run, and reservoir-cap sweepattack_family_comparison.py: attack-family comparison pipelineplot_results.py: current figure generation
Recommended reading path for a new collaborator:
current_findings_2026-08-07.mdtechnical_report_strong_oracle.mdresults_index.mdconfigs/mnist_blackbox_transfer.yamlexperiments/surrogate.py,attack_family_comparison.py,run_experiment_yaml.py
oracle.py: Defines and trains the target model.surrogate.py: Implements the Jacobian augmentation training loop.attack.py: Generates and evaluates the FGSM baseline.attack_family_comparison.py: Compares FGSM/BIM/PGD and current JSMA-like placeholder.check_agreement.py: Utility to measure model consistency.