Skip to content

Repository files navigation

Banking-Fraud-Analytics

About This Project

India processes billions of banking transactions every year across UPI, NEFT, RTGS, IMPS, and card networks. As digital payments have grown, so have fraud risks, and the data quality issues that make it harder for regulators and analysts to get a clear picture of what is actually happening on the ground.

This project is my attempt to build a complete, end-to-end analytics pipeline on a real-world Indian banking transactions dataset, focused specifically on detecting and understanding fraud patterns. I did this independently to strengthen my data analytics skills and to create portfolio work that is relevant to banking and financial regulation contexts.

Starting from a raw, messy dataset, I cleaned and validated the data, ran exploratory analysis, wrote SQL queries for deeper fraud insights, automated an MIS reporting pipeline in Python, and built a Power BI dashboard that a non-technical stakeholder can actually read and use.

Problem Statement

Indian banking data is rarely clean, rarely simple, and rarely tells you what you need to know without some digging. Transaction volumes vary wildly across channels, states, and merchant categories. Fraud patterns are buried in the noise. Reporting is often manual and inconsistent.

The goal of this project was to cut through that and answer a few core questions:

  • Where is fraud risk concentrated, across which states, channels, and transaction types, and what patterns are associated with it?
  • How have digital payment channels grown over time, and which ones carry disproportionate fraud risk relative to their volume?
  • Can the reporting process be automated so analysts spend less time formatting and more time analysing?

What I Did

Data Cleaning (Python / pandas): Took the raw dataset, fixed missing values, removed duplicates, corrected data types, standardised inconsistent entries, and documented every single cleaning decision.

Exploratory Data Analysis (Python): Analysed transaction and fraud trends across payment channels, states, and merchant categories. Visualised fraud distribution, class imbalance, and correlation between transaction amount and fraud.

SQL Analysis (MySQL): Wrote queries using CTEs and window functions to rank states and channels by fraud rate, cross-tabulate fraud against KYC status, and surface the highest-risk segments for investigation.

MIS Automation (Python, MySQL, openpyxl, reportlab): Built a script that connects directly to the MySQL database and generates formatted Excel and PDF MIS reports on demand or on a schedule, removing the need for anyone to manually run queries or refresh reports.

Power BI Dashboard (DAX): Designed an interactive, multi-page dashboard with DAX measures covering fraud risk indicators, state and channel-level performance, and KYC-based risk signals. Built for a non-technical, senior stakeholder audience.

Tools Used

Python, pandas, numpy, matplotlib, seaborn, MySQL, Power BI, DAX, Excel, openpyxl

About

End-to-End Banking Fraud Analytics using Excel, Python, SQL and Power BI

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages