A classical-ML audio pipeline that identifies the spoken language of short clips from hand-engineered acoustic features — trained and evaluated across 4 languages with both supervised classifiers and unsupervised clustering.
Given a ~1-minute audio clip, predict its language. Instead of an end-to-end deep model, this project uses an interpretable feature-engineering approach: each clip is reduced to a fixed-length acoustic descriptor with librosa, then fed to standard classifiers and clustering algorithms.
Pipeline
- Cleaning: resample to 16 kHz mono, silence trimming, peak normalization, 60 s cap.
- Feature extraction (53-dim/clip): MFCC (20×mean/std), spectral centroid/bandwidth/rolloff, chroma, ZCR, RMS, and tempo.
- Classification: KNN, SVM (RBF), MLP —
StandardScaler+ grid search with 3-fold CV (macro-F1), 80/20 stratified split. - Clustering: K-Means (k chosen by silhouette) + Agglomerative, scored with silhouette and purity; embeddings visualized via PCA and t-SNE.
712 clips across 4 languages (class-collected; recursively discovered under per-language folders):
| Language | Clips |
|---|---|
| Spanish | 180 |
| Korean | 180 |
| Italian | 180 |
| German | 172 |
| Total | 712 |
The full audio corpus is a course-collected dataset and is not redistributed here. The extracted feature matrix and all result artifacts are included under
artifacts/so the modeling notebooks are fully reproducible without the raw audio.
Classification (held-out 20% test, 143 clips): KNN, SVM-RBF, and MLP each reach 100% accuracy / macro-F1 = 1.00 on the test split.
The perfect score reflects how cleanly the engineered features separate these four languages on this small, speaker-consistent dataset; on a larger, speaker-disjoint corpus the task would be harder. Confusion matrices are in
artifacts/.
Clustering (unsupervised):
| Algorithm | k | Silhouette | Purity |
|---|---|---|---|
| K-Means | 10 | 0.30 | 0.87 |
| Agglomerative | 4 | 0.21 | 0.52 |
PCA and t-SNE scatter plots (artifacts/pca2_scatter.png,
artifacts/tsne_scatter.png) show the language clusters in 2-D.
.
├── notebooks/
│ ├── Data_Cleaning_and_Feature_Extraction.ipynb # audio → features.parquet
│ ├── Classification.ipynb # KNN / SVM / MLP + grid search
│ ├── Clustering.ipynb # K-Means + Agglomerative
│ └── Evaluation.ipynb # metrics & comparison
├── artifacts/ # features, summaries, plots, confusion matrices
├── requirements.txt
└── README.md
pip install -r requirements.txtOpen the notebooks in order (feature extraction → classification → clustering →
evaluation). The modeling notebooks read artifacts/features.parquet, so you can
reproduce all results without the raw audio.
Python · librosa · scikit-learn · pandas / numpy · matplotlib
MIT — see LICENSE.
Mahdiar Harandi — harandimahdiar@gmail.com GitHub · LinkedIn