Coventry University (7150CEM)
Author: Rexley Olaoluwa Adio
Supervisor: Babak Jamshidi | Ethics Reviewer: Dr. Omid Chatrabgoun
This repository contains the implementation of a high-performance, robust, and explainable Network Intrusion Detection System (NIDS) designed for commercial IP-based surveillance systems and Internet of Things (IoT) networks.
The project employs a hybrid machine learning approach utilizing Random Forest and XGBoost to classify raw network packets into 6 distinct traffic classes (1 benign, 5 malicious). It incorporates Explainable AI (XAI) using SHAP (Shapley Additive Explanations) to provide model transparency and explain the key features driving the intrusion detection alerts.
graph TD
A[Raw PCAP Files] -->|1. Scapy Feature Extraction| B[Tabular CSV Dataset]
B -->|2. Data Cleaning & Normalization| C[StandardScaler & LabelEncoder]
C -->|3. Feature Selection| D[SelectKBest ANOVA F-Value]
D -->|4. Class Imbalance Resolution| E[SMOTE Over-Sampling]
E -->|5. Model Development| F[Random Forest / XGBoost]
F -->|6. Explainable AI| G[SHAP TreeExplainer]
-
Stage 1: Feature Extraction (Scapy)
- Parses raw
.pcappacket capture files. - Groups packets into time windows of 100 packets.
- Extracts 25 temporal, statistical, and signature features including TCP flags (SYN, ACK, RST, FIN, URG), average packet sizes, inter-arrival times, protocol ratios, and basic payload signature keywords (SQLi, XSS, shell commands).
- Parses raw
-
Stage 2: Preprocessing & Scaling
- Imputes missing/NaN values using the Mean Strategy.
- Encodes the multi-class labels into numeric values.
- Standardizes features using
StandardScalerto bring all values to the same scale.
-
Stage 3: Feature Selection (SelectKBest)
- Implements univariate ANOVA F-value (
f_classif) to reduce features from 25 down to the top 12 most statistically significant predictors. - Keeps features in their original physical form to facilitate explainability (instead of projecting them to abstract principal components like PCA).
- Implements univariate ANOVA F-value (
-
Stage 4: Handling Class Imbalance (SMOTE)
- Resolves class imbalance using SMOTE (Synthetic Minority Over-sampling Technique) on the training set, increasing representation from 5,790 observations to 13,500 balanced training samples.
-
Stage 5: Model Development & Evaluation
- Trains Random Forest (with class weight balancing) and XGBoost classifiers.
- Performs side-by-side comparative analysis of overall accuracy, precision, recall, confusion matrices, ROC-AUC, and inference latency per sample.
-
Stage 6: Explainable AI (SHAP)
- Integrates SHAP (Shapley Additive Explanations) to explain both global model behavior and individual packet classification decisions.
- Visualizes feature importance through summary bar plots, beeswarm plots, and decision flow plots.
Both models achieved a robust overall accuracy of 77% on the multi-class classification test set, with excellent performance on standard, brute force, and replay attack classes:
| Class Label | RF Precision | RF Recall | RF F1-Score | XGB Precision | XGB Recall | XGB F1-Score | Support |
|---|---|---|---|---|---|---|---|
| bruteforce | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 96 |
| exploits | 0.84 | 0.66 | 0.74 | 0.84 | 0.65 | 0.73 | 563 |
| malware | 0.80 | 0.73 | 0.76 | 0.80 | 0.72 | 0.76 | 385 |
| normal | 1.00 | 0.96 | 0.98 | 1.00 | 0.96 | 0.98 | 67 |
| replay | 0.99 | 0.94 | 0.96 | 0.98 | 0.95 | 0.96 | 140 |
| scanning | 0.49 | 0.88 | 0.63 | 0.47 | 0.88 | 0.61 | 197 |
| Overall Accuracy | - | - | 0.77 | - | - | 0.77 | 1448 |
- Random Forest: 18 microseconds per packet sample.
- XGBoost: 9 microseconds per packet sample. (XGBoost is 2x faster, making it highly suitable for real-time edge environments).
├── attackScenarios/ # Raw PCAP files of attack classes (e.g. Zeus, BlackEnergy)
├── Certificate-187226.pdf # Ethical Approval Certificate (Low Risk)
├── Checklist-187226.pdf # Ethics Review Application & Summary
├── Comments-187226.pdf # Institutional Ethics Review Comments
├── network-anomaly-dectection.ipynb # Primary Jupyter Notebook with Python Pipeline
├── rex-dissertation-edited copy.docx# Masters Thesis Dissertation Document
└── README.md # Project Navigation and Overview Documentation
Ensure you have Python 3.9+ installed. Clone the repository and navigate to the project folder:
git clone <your-repository-url>
cd "Rexley Adio Final Project"Install the required network analysis and machine learning dependencies:
pip install scapy pandas numpy scikit-learn xgboost imbalanced-learn shap matplotlib seabornOpen the Jupyter notebook locally:
jupyter notebook "network-anomaly-dectection (7).ipynb"-
Note: If running locally, make sure to update the Kaggle paths in Stage 1 & 2 to point to your local directories.
- Change
pcap_directory = '/kaggle/input/...'$\rightarrow$ pcap_directory = './attackScenarios/' - Change
output_csv = '/kaggle/working/...'$\rightarrow$ output_csv = './windowed_network_features.csv'
- Change
This project was submitted in partial fulfillment of the requirements for the degree of Master of Science in Data Science at Coventry University (Academic Year 2025/26). Special thanks to Babak Jamshidi (Supervisor) and the IEEE Dataport for providing the primary attack scenario pcap datasets.