Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Papernot et al. (2017) Replication

Target: Replicate the core methodology from "Practical Black-Box Attacks against Machine Learning" (Papernot et al., 2017).

Current Status Note

This repository has moved beyond a single reproduction checkpoint and now contains:

  • a full-test-set FGSM transfer baseline
  • substitute-training ablations (vanilla, pss, pss_rs)
  • query-budget sweeps
  • step-wise mechanism logging
  • attack-family comparison (FGSM, BIM, PGD, and a placeholder JSMA-like implementation)
  • baseline-vs-wide strong-oracle cross-comparison
  • multi-seed robustness validation for the strong baseline oracle

The most up-to-date quantitative findings are stored in:

  • current_findings_2026-08-07.md
  • technical_report_strong_oracle.md
  • results_index.md
  • results/
  • figures/

Current Evidence Map

If you are opening this repository for the first time, read these in order:

  1. current_findings_2026-08-07.md
  2. technical_report_strong_oracle.md
  3. results_index.md
  4. figures/

What each file is for:

  • current_findings_2026-08-07.md
    • newest concise research summary
    • includes baseline multi-seed robustness, cross-oracle mechanism dynamics, and cross-oracle attack-family findings
  • technical_report_strong_oracle.md
    • formal write-up of the strong-oracle story
    • now includes compact cross-oracle tables and figure references
  • results_index.md
    • marks which results/*.json files are primary, supporting, or legacy
    • prevents accidental mixing of new crossoracle_* results with older default-named files
  • figures/
    • contains the main visual evidence, including the new cross-oracle comparison plots

Current primary result files:

  • results/multiseed_validation_baseline_oracle_seed0-4.json
  • results/surrogate_training_seed0_steps8_init10_pss_mechanism_crossoracle_baseline.json
  • results/surrogate_training_seed20_steps8_init10_pss_mechanism_crossoracle_wide.json
  • results/attack_family_comparison_seed0_cross_oracle_baseline_pss.json
  • results/attack_family_comparison_seed20_cross_oracle_wide_pss.json

Current primary figure files:

  • figures/cross_oracle_transfer_asr_vs_substitute_step.png
  • figures/cross_oracle_agreement_vs_substitute_step.png
  • figures/cross_oracle_attack_family_transfer.png

Current Headline Findings

  • pss agreement advantage is stable across seeds 0..4 under the strong baseline oracle.
  • pss transfer advantage at epsilon = 0.3 is stable across seeds 0..4 under the strong baseline oracle.
  • Early transfer emergence appears under both validated strong-oracle conditions.
  • The wide strong oracle begins with slightly stronger early transfer, but ends with weaker medium- and high-epsilon transfer than the baseline strong oracle.
  • Under matched full-budget pss surrogates, FGSM is the most stable transfer family across the two strong-oracle conditions.
  • Simple gradient-alignment proxies remain too weak to serve as a standalone explanation for transfer differences.

Key Components

  1. Oracle (Target Model): A black-box CNN model trained locally on the full MNIST training set. The current checked-in oracle.pt and artifacts/oracle_seed0.pt evaluate to about 98.87% test accuracy and should be treated as the current strong local oracle baseline.
  2. Surrogate Model Training (Jacobian-based Dataset Augmentation):
    • Initial small training set: 100 samples from MNIST.
    • Query the Oracle for initial labels.
    • Iteratively augment the dataset (8 steps):
      • Compute the Jacobian of the surrogate w.r.t. inputs.
      • Shift inputs in the direction of the gradient ($\lambda=0.1$), with vanilla, periodic step size (pss), and reservoir-sampling (pss_rs) variants available.
      • Query Oracle for labels of new points.
  3. Adversarial Example Generation:
    • Untargeted FGSM baseline in attack.py.
    • Additional family comparison in attack_family_comparison.py over FGSM, BIM, PGD, and a current placeholder JSMA-like implementation.
    • Important: Perturbations are applied in the $[0, 1]$ pixel space before normalization.
  4. Transferability Evaluation:
    • Evaluate on the full 10,000-image MNIST test set, filtered to samples jointly classified correctly by both oracle and surrogate before attack.

YAML Experiment Configs

This repository now supports a two-layer YAML system:

  1. a scientific master config in configs/
  2. executable experiment orchestration configs in experiments/

The experiment YAMLs reference the scientific master config and define runnable step sequences.

Current scientific master config:

  • configs/mnist_blackbox_transfer.yaml
  • export_config_summary.py can turn it into an advisor-facing Markdown summary

This repository now supports reproducible experiment orchestration through YAML configs in experiments/.

Runner:

  • run_experiment_yaml.py

Current configs:

  • experiments/wide_oracle_minimal_cross_oracle_baseline.yaml
  • experiments/wide_oracle_pss_attack_family.yaml
  • experiments/wide_oracle_multiseed_validation.yaml
  • experiments/wide_oracle_budget_sweep.yaml
  • experiments/wide_oracle_reservoir_cap_sweep.yaml
  • experiments/wide_oracle_mechanism.yaml
  • experiments/baseline_oracle_multiseed_validation.yaml
  • experiments/paired_cross_oracle_mechanism.yaml
  • experiments/paired_cross_oracle_attack_family.yaml

Schema features:

  • schema_version: 1
  • grouped defaults with common, by_script, and optional by_step
  • config_ref linking each experiment config to a scientific master config
  • step-level outputs
  • step-level overwrite
  • run artifacts saved under runs/<experiment>/<timestamp>/

Examples:

python3 run_experiment_yaml.py experiments/wide_oracle_minimal_cross_oracle_baseline.yaml --dry-run
python3 run_experiment_yaml.py experiments/wide_oracle_pss_attack_family.yaml
python3 run_experiment_yaml.py experiments/paired_cross_oracle_mechanism.yaml
python3 run_experiment_yaml.py experiments/paired_cross_oracle_attack_family.yaml

Current direct scientific-config consumers:

  • attack.py
  • attack_family_comparison.py
  • run_multiseed_validation.py
  • surrogate.py
  • check_agreement.py
  • run_budget_and_mechanism.py
  • oracle.py

These currently read the following master-config fields directly when --config is passed:

  • dataset.root
  • attack.primary_baseline.epsilons
  • evaluation.batch_size_default
  • evaluation.agreement_batch_size
  • surrogate_model.training_defaults.steps
  • surrogate_model.training_defaults.initial_size_per_class
  • surrogate_model.training_defaults.epochs_per_step
  • victim_model.training_defaults.epochs
  • victim_model.training_defaults.batch_size
  • victim_model.training_defaults.learning_rate

Python dependencies for YAML-driven experiments are listed in:

  • requirements.txt
  • requirements-lock.txt
  • environment.lock.yaml

Example scientific-config summary export:

python3 export_config_summary.py configs/mnist_blackbox_transfer.yaml --output config_summary.md

Minimal Clone-to-Reproduce Workflow

git clone <your-repo-url>
cd papernot-replication
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

If you want to match the currently recorded local runtime as closely as possible, consult:

  • requirements-lock.txt
  • environment.lock.yaml

Then ensure MNIST raw files are available under:

  • ./data/MNIST/raw

Run a minimal reproducibility sanity check:

python3 verify_reproducibility.py

You can inspect the scientific master config with:

python3 export_config_summary.py configs/mnist_blackbox_transfer.yaml --output config_summary.md

And validate a runnable experiment without executing it via:

python3 run_experiment_yaml.py experiments/wide_oracle_mechanism.yaml --dry-run

Current Baselines (Seed 0)

Full-budget FGSM baseline on vanilla surrogate

Source:

  • results/fgsm_transfer_eval_seed0.json
  • results/budget_mechanism_cap_sweep_seed0.json
Epsilon ($\epsilon$) Surrogate ASR (White-box FGSM) Oracle ASR (Transfer FGSM)
0.2 91.29% 2.75%
0.3 97.07% 24.54%
0.4 97.82% 60.68%

Attack-family comparison on current strongest final-transfer baseline

Source:

  • results/attack_family_comparison_seed0.json

At the current local oracle-surrogate pair:

  • FGSM is the strongest observed transfer baseline
  • PGD and BIM are stronger white-box attacks on the surrogate but transfer substantially worse than FGSM
  • the current JSMA-like implementation is a placeholder and is not a faithful scientific baseline

Scientific Caveats

  • The current strong local oracle is much more realistic than the earlier weak local oracle used in prior runs, but it is still not the original remote MetaMind oracle from Papernot et al. Results should therefore be interpreted as local-replication results, not direct quantitative confirmation of the original paper's strongest reported numbers.
  • Current results are based primarily on seed=0 unless otherwise noted.
  • Query-budget and reservoir-cap experiments are now separated conceptually: query_budget controls total oracle queries, while reservoir_cap controls retained substitute dataset size.
  • Mechanism results should be read carefully: current evidence supports several empirical findings, but gradient-level mechanism claims remain hypotheses until directly measured.

What Changed After Strengthening the Oracle

Moving from the earlier weak local oracle (~80.18%) to the current strong local oracle (98.87%) changed several conclusions:

  • transfer rates dropped substantially across the board
  • FGSM became the strongest observed transfer attack family
  • PGD and BIM remained extremely strong white-box surrogate attacks, but no longer produced the strongest transfer

This ranking reversal is one of the most important current findings in the repository.

Debugging Notes (Fixes implemented)

  • Problem: ASR was initially <4%.
  • Solution:
    1. Re-trained Oracle on full MNIST. The previous 100-sample Oracle had non-transferable noisy boundaries.
    2. Moved FGSM attack to unnormalized $[0, 1]$ space. Previous epsilon was effectively 3x smaller than intended due to being applied in normalized space.
    3. Corrected Jacobian augmentation to follow Oracle labels for gradient direction.

Recommended Entry Points

  • current_findings_2026-08-07.md: most up-to-date short research summary
  • technical_report_strong_oracle.md: formal report with cross-oracle updates
  • results_index.md: primary vs supporting vs legacy result map
  • artifacts/oracle_seed0.json: current verified oracle accuracy record
  • artifacts/oracle_wide_seed10.pt: second strong-oracle condition used for cross-oracle comparison
  • experiments/: YAML experiment configs
  • configs/: scientific master configuration
  • run_experiment_yaml.py: minimal YAML experiment runner
  • surrogate.py: substitute training variants, query-budget control, and mechanism logging
  • run_budget_and_mechanism.py: query-budget sweep, mechanism run, and reservoir-cap sweep
  • attack_family_comparison.py: attack-family comparison pipeline
  • plot_results.py: current figure generation

Recommended reading path for a new collaborator:

  1. current_findings_2026-08-07.md
  2. technical_report_strong_oracle.md
  3. results_index.md
  4. configs/mnist_blackbox_transfer.yaml
  5. experiments/
  6. surrogate.py, attack_family_comparison.py, run_experiment_yaml.py

Structure

  • oracle.py: Defines and trains the target model.
  • surrogate.py: Implements the Jacobian augmentation training loop.
  • attack.py: Generates and evaluates the FGSM baseline.
  • attack_family_comparison.py: Compares FGSM/BIM/PGD and current JSMA-like placeholder.
  • check_agreement.py: Utility to measure model consistency.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages