Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Spoken Language Identification

A classical-ML audio pipeline that identifies the spoken language of short clips from hand-engineered acoustic features — trained and evaluated across 4 languages with both supervised classifiers and unsupervised clustering.

Python librosa License: MIT


Overview

Given a ~1-minute audio clip, predict its language. Instead of an end-to-end deep model, this project uses an interpretable feature-engineering approach: each clip is reduced to a fixed-length acoustic descriptor with librosa, then fed to standard classifiers and clustering algorithms.

Pipeline

  1. Cleaning: resample to 16 kHz mono, silence trimming, peak normalization, 60 s cap.
  2. Feature extraction (53-dim/clip): MFCC (20×mean/std), spectral centroid/bandwidth/rolloff, chroma, ZCR, RMS, and tempo.
  3. Classification: KNN, SVM (RBF), MLP — StandardScaler + grid search with 3-fold CV (macro-F1), 80/20 stratified split.
  4. Clustering: K-Means (k chosen by silhouette) + Agglomerative, scored with silhouette and purity; embeddings visualized via PCA and t-SNE.

Dataset

712 clips across 4 languages (class-collected; recursively discovered under per-language folders):

Language Clips
Spanish 180
Korean 180
Italian 180
German 172
Total 712

The full audio corpus is a course-collected dataset and is not redistributed here. The extracted feature matrix and all result artifacts are included under artifacts/ so the modeling notebooks are fully reproducible without the raw audio.


Results

Classification (held-out 20% test, 143 clips): KNN, SVM-RBF, and MLP each reach 100% accuracy / macro-F1 = 1.00 on the test split.

The perfect score reflects how cleanly the engineered features separate these four languages on this small, speaker-consistent dataset; on a larger, speaker-disjoint corpus the task would be harder. Confusion matrices are in artifacts/.

Clustering (unsupervised):

Algorithm k Silhouette Purity
K-Means 10 0.30 0.87
Agglomerative 4 0.21 0.52

PCA and t-SNE scatter plots (artifacts/pca2_scatter.png, artifacts/tsne_scatter.png) show the language clusters in 2-D.


Repository Contents

.
├── notebooks/
│   ├── Data_Cleaning_and_Feature_Extraction.ipynb   # audio → features.parquet
│   ├── Classification.ipynb                          # KNN / SVM / MLP + grid search
│   ├── Clustering.ipynb                              # K-Means + Agglomerative
│   └── Evaluation.ipynb                              # metrics & comparison
├── artifacts/                                        # features, summaries, plots, confusion matrices
├── requirements.txt
└── README.md

Getting Started

pip install -r requirements.txt

Open the notebooks in order (feature extraction → classification → clustering → evaluation). The modeling notebooks read artifacts/features.parquet, so you can reproduce all results without the raw audio.


Tech Stack

Python · librosa · scikit-learn · pandas / numpy · matplotlib

License

MIT — see LICENSE.

Contact

Mahdiar Harandi — harandimahdiar@gmail.com GitHub · LinkedIn

About

Spoken language identification from librosa audio features (KNN/SVM/MLP + clustering).

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages