Title: Students Performance in Exams
Download Link: Kaggle Dataset
Dataset Description:
- Contains student scores across subjects, demographic info (gender, parental education), and other factors (lunch type, test preparation).
- Useful for analyzing performance trends, disparities, and correlations.
You're part of an educational data team aiming to assess student performance, learning trends, and institutional effectiveness across subjects and demographics.
The goal is to derive insights that can guide curriculum changes, targeted support, and performance optimization.
Use skills in data handling, SQL/statistics, and visual storytelling to answer key questions.
Student Performance & Institutional Effectiveness-Insights/
ββ data/ # Raw & processed datasets
ββ notebooks/ # Jupyter notebooks for EDA & analysis
ββ scripts/ # Python scripts for preprocessing & analysis
ββ README.md
- πΉ Are there significant score differences between genders across subjects?
- πΉ Does parental education level correlate with student performance?
- πΉ Which subject has the highest average score overall? Which is the most variable?
- πΉ How does lunch type or test preparation affect student performance?
- πΉ Is there evidence of bias or disparity in performance across demographic groups?
- Rename confusing column headers for clarity.
- Check and fix missing or inconsistent values.
- Convert categorical/text features into analyzable formats.
- Libraries Used: Pandas, Seaborn, Matplotlib
- Create 6β8 well-chosen visuals to compare, contrast, and interpret trends.
- Graph types may include count plots, histograms, box plots, and pair plots.
- Highlight comparisons or disparities among student groups.
- Comment on correlations or associations between variables.
- Optionally apply statistical tests (T-test, Chi-square) to validate insights.
- πΉ Use SQL for summarizing or filtering raw data.
- πΉ Create a correlation matrix and interpret relationships.
- πΉ Design an interactive dashboard to visualize insights.
This analysis, based on a dataset of 1,000 student exam records, provides initial findings on academic performance.
-
Lowest Performance in Math: Students show the lowest average score in Math (mean
$\approx 66.09$ ), while performing best in Reading (mean$\approx 69.17$ ). - Complete and Clean Dataset: The source data is highly reliable, consisting of 1,000 entries with no missing values across all 8 features (5 categorical, 3 numerical).
- Near-Normal Score Distribution: Scores in all subjects closely follow a normal distribution, with a slight negative skew, indicating the majority of scores cluster towards the higher end of the scale.
-
Highest Variability in Writing: The Writing score (Standard Deviation
$\approx 15.20$ ) shows the greatest spread in performance among students, closely followed by Math ($\approx 15.16$ ). - Presence of Extreme Low Scores: The data includes records of students with significantly low scores, such as a Math score of 0 and a Writing score of 10.
- Symmetrical Distribution Confirmed: The mean and median scores are closely aligned across all three subjects, confirming the overall symmetry of the score distributions.
- Team Size: 3 members
- Define clear division: data cleaning, analysis, visualizations
- Programming Languages: Python, SQL
- Libraries / Tools: Pandas, NumPy, Matplotlib, Seaborn, Jupyter Notebook
- Techniques: EDA, Statistical Analysis, Data Cleaning, Data Visualization
- π¨βπ» Suraj Mate β Sql Insights
- π¨βπ» Vaali nandhan β Exploratory Data Analysis (EDA)
- π¨βπ» Yash β Data Cleaning and preprocessing
- π Location: India
π Thank You for Visiting My Profile! π
π‘ I love building projects, exploring data, and learning new technologies!
π Keep Learning. Keep Growing. Keep Exploring! π