I am interested in problems where the answer is not sitting in one clean dataset. The projects here use financial cash flows, fund metrics, environmental sensors, weather records, and time series to build evidence from incomplete information.
There are three main threads.
I built this to compare funds when cash-flow timing, fees, strategy mix, and terminal value all differ. It calculates XIRR and investment multiples, runs deterministic and Monte Carlo scenarios, tests when the preferred fund changes, and exports an Excel audit.
The demo uses synthetic or anonymized inputs. It is a decision model, not a return forecast.
This is the publishable quantitative core of a larger fund-finder prototype. It applies eligibility rules first, then ranks funds against peers using rolling-return and Sharpe percentile scores. It also keeps a separate record of exclusions so the ranking is easier to inspect.
The public repo uses synthetic fund observations and has automated tests for thresholds, ties, weighting, and peer-group logic.
To me, alternative-data work means using several imperfect sources together and checking whether they support the same conclusion.
I compared golf-course evapotranspiration with forest, urban, residential, and desert-adjacent sites, then joined the series to NOAA temperature and precipitation. The goal was to see whether ET could provide a useful water-demand signal across New York and Arizona.
The important limit is that ET is not metered irrigation. The project is useful because it combines sources and comparison groups, not because one proxy answers the whole question.
Raw station data can look precise while still being wrong. This R workflow checks physical bounds, sensor tilt, battery state, missingness, and sampling intervals before using the observations.
This project joins river-gauge readings to site-specific thresholds, finds the first action and flood-stage crossings, and measures major-stage exceedance.
A visualization-focused R exercise using long-run emissions, temperature anomalies, and space-launch records. These are separate descriptive questions; I do not treat space activity as an explanation for climate trends.
My strongest R coursework project. It covers transformations, multivariate regression, residual diagnostics, AIC-based selection, seasonal decomposition, autoregressive models, ARIMA forecasting, and prediction intervals.
An early R exercise in vectorized calculations, unit conversion, data frames, indexing, and subsetting.
A clearly labeled course archive covering audio processing, text analysis, CSV grids, image manipulation, and web scraping. I kept it separate because it shows where I started, not where my quantitative work is now.
The next step is a finance-focused alternative-data project with a real validation design: a signal defined before testing, a clean train/validation split, stability checks through time, and a short write-up connecting the result to an actual decision.