This project implements a machine learning model to detect fraudulent credit card transactions using logistic regression. The code leverages Python's pandas for data manipulation, numpy for numerical operations, and scikit-learn for building and evaluating the machine learning model.
The dataset used is creditcard.csv, which contains transaction details and a target variable indicating whether a transaction is legitimate or fraudulent.
-
Data Loading and Exploration:
- The dataset is loaded using
pandasand initial exploration is performed usinghead()to preview the data. - The distribution of the target variable (
Class) is examined usingvalue_counts()to understand the imbalance between legitimate and fraudulent transactions.
- The dataset is loaded using
-
Data Balancing:
- To address class imbalance, a random sample of legitimate transactions is taken to match the number of fraudulent transactions, ensuring a balanced dataset for training.
-
Feature and Target Separation:
- Features (
x) and target (y) variables are separated. The target variable is the "Class" column, indicating transaction legitimacy.
- Features (
-
Data Splitting:
- The dataset is split into training and testing sets using
train_test_splitwith stratified sampling to maintain the distribution of the target variable.
- The dataset is split into training and testing sets using
-
Model Training:
- A
LogisticRegressionmodel is instantiated and trained on the training data (x_train,y_train) with a maximum iteration of 1000 to ensure convergence.
- A
-
Model Evaluation:
- The accuracy of the model is 94% on both the training and testing datasets, providing insights into the model's performance.
- The code outputs the accuracy of the model on both the training and testing datasets, demonstrating its effectiveness in detecting fraudulent transactions.
This project serves as a practical example of using logistic regression for binary classification tasks in financial fraud detection.