Skip to content
pranay172Public

About

My Solution for Construction Cost Prediction competition on Solafune

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

🏗️ CCP · Construction Cost Prediction

Economic context, satellite imagery, and a surprisingly strong tabular baseline.

Rank Award Competitors Submissions

The challenge · Final leaderboard · Run the code · Dataset & metric

The idea

How much does it cost to build a square meter in a given region and quarter?

This was my solution to Solafune's Construction Cost Prediction challenge: estimate construction costs in Japan and the Philippines from economic indicators, geographic context, and quarterly satellite composites. The target was construction_cost_per_m2_usd, evaluated with RMSLE—lower is better.

I tried two routes: let AutoGluon work with the tabular data, then build a more involved pipeline with Prithvi satellite embeddings, image statistics, and country-specific tree ensembles. The simpler route finished ahead.

🥈 Six submissions, 21st place

I finished 21st out of 373 competitors, earning a silver badge. The final leaderboard lists six submissions under pranay212 and a best private RMSLE of 0.18540059576309284.

Approach Private RMSLE What went into it
AutoGluon tabular 0.185 Competition CSV features, bagging, and stacking
Prithvi + engineered-feature ensemble 0.192 Satellite embeddings, spectral/nightlight statistics, geographic features, and boosted trees

The approach-level scores above are rounded results recorded in my original project notes; the exact best score and placement are visible on the private leaderboard.

My main takeaway: keep a strong tabular baseline in the comparison. Adding satellite features made this a much richer experiment, but it did not improve the private score in these runs. That is a result for these implementations, not a verdict on the value of satellite imagery.

Two routes to a prediction

01 · Let the table do the work

solution_autogluon.py trains an AutoGluon TabularPredictor on the competition CSVs, without reading the images. It uses the extreme preset, five bagging folds, one stacking level, and a one-hour training budget.

The target is transformed with log1p, the model optimizes RMSE in that space, and predictions return to USD/m² with expm1. This was my best-scoring approach.

02 · Add a view from space

The second pipeline combines three kinds of information:

Feature family Implementation
Satellite embeddings Prithvi EO 2.0 300M temporal-location model; six Sentinel-2 bands, time/location coordinates, 1,024-dimensional pooled embeddings, then PCA to 64 components
Image statistics NDVI, NDBI, NDWI, band/index summaries, valid-pixel coverage, and VIIRS nightlight statistics
Tabular & geographic context Economic and hazard interactions, infrastructure access, regional target aggregates, temporal features, and spatial target-neighborhood statistics

extract_features.py creates the image features. train_model.py trains separate model sets for Japan and the Philippines using XGBoost, LightGBM, and CatBoost. generate_submission.py averages fold predictions and blends the model families in the original cost units.

Open the experiment notebook: validation scores & feature importance

These are the out-of-fold (OOF) results preserved from my original notes:

Model OOF RMSLE Final blend weight
XGBoost 0.2032 0.00
LightGBM 0.2021 0.43
CatBoost 0.2017 0.57
Weighted ensemble 0.2013 1.00

Validation uses five shuffled folds within each country. It does not hold out entire locations or future time periods. PCA is fitted before the folds, and blend weights are selected using the same OOF predictions reported here, so these scores are development estimates rather than an independent final test.

Saved XGBoost feature-importance plot

This saved plot comes from the last country's last XGBoost fold. It highlights the prominence of geographic target aggregates in that model; it is not an importance ranking for the full ensemble.

Why I stopped submitting

During the competition, I learned through the forum that the target could be derived from government statistics linked in the official overview: Japan's e-Stat construction statistics and the Philippine Statistics Authority's building-permit statistics.

That changed what I wanted to get out of the challenge. I was interested in developing a prediction model, so I chose to stop competing at that point. This repository preserves the modeling-based submissions I made before stepping away; it does not implement target reconstruction from those sources.

Explore the repo

CCP/
├── solution_autogluon.py       # Best private-score approach: tabular AutoML
├── extract_features.py        # Prithvi embeddings + image statistics
├── train_model.py             # Country-specific boosted-tree ensembles
├── generate_submission.py     # Load ensemble artifacts and predict
├── feature_importance.png     # Saved diagnostic from an XGBoost fold
├── CCP_Competition_Details.md  # Challenge, data, sources, and evaluation
└── docs/USAGE.md               # Setup, commands, artifacts, and caveats

Start with the tabular route: place the competition CSVs under data/ as described in the usage guide, then run from the repository root:

python -m pip install autogluon.tabular
python solution_autogluon.py

For the satellite pipeline, follow the full setup and extraction steps. Competition data, pretrained weights, trained models, and submission files are not bundled. This is an archive of competition experiments; package versions and the original training environment were not pinned.

About

My Solution for Construction Cost Prediction competition on Solafune

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages