Exploratory Data Analysis (EDA) Portfolio Project Repository.
A comprehensive collection of data analysis projects showcasing advanced exploratory techniques, statistical modeling, and data visualization. This repository demonstrates practical applications of EDA for uncovering hidden patterns, identifying anomalies, testing hypotheses, and deriving actionable insights to support business strategy.
- Data Cleaning & Preprocessing: Handling missing values, outlier detection, data type casting, and structural error correction.
- Descriptive Statistics: Summary metrics including central tendency, dispersion, skewness, kurtosis, and distribution shapes.
- Univariate & Multivariate Analysis: In-depth feature profiling alongside correlation matrices, pair plots, and cross-tabulations.
- Advanced Visualizations: Dynamic plotting using libraries like Seaborn, Matplotlib, and Plotly for interactive distributions, heatmaps, and trend lines.
- Feature Engineering: Deriving new variables, encoding categorical fields, feature scaling, and dimensionality reduction (PCA).
- Customer & product Segmentation: Identified distinct behavioral cohorts leading to increase in marketing efficiency.
- Anomaly Detection: Flagged critical outliers and data entry errors that skewed historical baseline performance metrics.
- Predictive Indicators: Uncovered strong linear and non-linear correlations that served as primary features for downstream Machine Learning models.
- Programming: Python
- Libraries: Pandas, NumPy, Scipy, Seaborn, Matplotlib, Plotly
- Environments: Jupyter Notebooks, V.S. code
- Each project is contained within its own clearly named folder.
- Inside each folder, you will find the raw dataset, a clean Jupyter Notebook (
.ipynb), and any necessary supporting files. - To replicate any analysis, simply download the repository and run the notebooks using your preferred Python environment.