A Python pipeline for extracting light curves from variable systems observed by BlackGEM, detecting photometric variability via dip-flagging, and applying unsupervised machine learning (K-Means clustering) to help identify candidate white dwarf eclipsing binary systems.
lightcurves/*.fits → bg_main.py → bg_lc_flagging.py → ml_features.csv → bg_ml.ipynb
(input FITS) (orchestrate) (flag dips) (feature table) (cluster)
| Stage | Script | Purpose |
|---|---|---|
| 1 | bg_main.py |
Reads all *_LC.fits files, normalises flux per filter, calls dip-flagging, writes ml_features.csv |
| 2 | bg_lc_flagging.py |
Iterative sigma-clip dip detector; computes variability metrics per filter |
| 3 | bg_ml.ipynb |
Loads features, scales, tunes K-Means (elbow + silhouette), produces cluster plots |
bg_query.py(BigQuery interface) is present but its call inbg_main.pyis currently disabled pending cloud auth setup.
See requirements.txt for the full list of direct dependencies.
Install with:
python -m venv ENV_NAME
source ENV_NAME/bin/activate
pip install -r requirements.txtpython bg_main.py --config example_config.iniThen open and run bg_ml.ipynb once ml_features.csv has been generated.
Copy example_config.ini and fill in the paths for your environment:
[pipeline]
data_root = /path/to/data/root
project_id = your-gcp-project-id
variable_table_loc = dataset.variable_tabledata_root must contain a lightcurves/ subdirectory with *_LC.fits files. The pipeline creates an analysis/ subdirectory and writes ml_features.csv and logs/ there.
| Module | Description |
|---|---|
bg_main.py |
Top-level runner; reads config, loops over LC files, calls dip-flagging, writes ml_features.csv |
bg_query.py |
BigQuery interface for the BlackGEM detections and images tables on Google Cloud |
bg_lc_flagging.py |
Iterative sigma-clip dip detector; returns variability metrics (score, n_dips, depth, SNR, etc.) |
bg_logger.py |
Logging setup (rotating file handler + console) |
bg_ml.ipynb |
K-Means clustering on lc_flagging features; elbow, silhouette, PCA, and scatter plots |
| File | Description |
|---|---|
ml_features.csv |
ML feature table — one row per target per filter |
analysis/{target}/{target}_{filter}_flags_flagged.png |
Flagged light curve plot with dips highlighted |
logs/bg_pipeline.log |
Full pipeline run log |
outputs/cluster_centers.csv |
K-Means cluster centre values in original feature space |
outputs/elbow_method.png |
Inertia vs K (elbow method) |
outputs/silhouette_scores.png |
Silhouette score vs K |
outputs/cluster_scatter.png |
2-D cluster scatter (first two features) |
outputs/pca_clusters.png |
PCA-compressed 2-D cluster plot |
outputs/scaled_hist.png |
Per-feature histograms of scaled values |
| Feature | Description |
|---|---|
score |
Overall variability score (matched-filter SNR × run-length bonus) |
n_dips |
Number of detected dip events |
sigma |
Robust MAD-based flux scatter |
survival_fraction |
Fraction of points surviving sigma-clipping |
n_high_cut |
Number of points cut above the upper threshold |
p2p_over_mad |
Point-to-point scatter normalised by MAD |
duty_cycle |
Fraction of time span spent in dips |
spacing_frac_scatter |
Fractional scatter of inter-dip spacings (low → periodic) |
consistent_spacings |
1 if dip spacings are consistent (periodic), else 0 |
best_depth_flux |
Flux depth of the highest-SNR dip |
best_depth_frac |
Fractional depth of the highest-SNR dip |
best_depth_sigma |
Depth in units of sigma for the highest-SNR dip |
best_mf_snr |
Matched-filter SNR of the best dip |
best_duration |
Duration (days) of the best dip |
best_n_points |
Number of in-dip data points for the best dip |