π Project Overview The goal is to classify sentiment in Hindi user reviews. The workflow includes:
Loading and preprocessing the dataset.
Tokenizing Hindi text for transformer compatibility.
Fine-tuning a pre-trained model (e.g., bert-base-multilingual-cased or Indic-specific models).
Evaluating performance with metrics such as accuracy and F1-score.
π§ Model & Dataset Model: You can use any transformer-based model like mBERT, IndicBERT, or mT5 depending on your use case.
Dataset: The IndicSentiment dataset's Hindi split is used.
π Getting Started Installation Install required libraries (as shown in the notebook):
bash Copy code pip install datasets transformers Running the Notebook Open the enhanced notebook and follow the step-by-step instructions:
bash Copy code jupyter notebook Hindi_Sentiment_Analysis_Enhanced.ipynb π Results and Insights The notebook includes cells for tracking model training, metrics logging, and evaluation insights. Update those cells after running your experiments.
π Repository Structure bash Copy code π Hindi-Sentiment-Analysis/ β βββ Hindi_Sentiment_Analysis_Enhanced.ipynb # Annotated notebook with explanations βββ README.md # Project overview and setup instructions π€ Contributing Feel free to fork this repo and contribute! PRs and suggestions are welcome.
π License This project is open-sourced under the MIT License.