This project implements a Persian Information Retrieval (IR) system for searching and ranking documents from the Hamshahri Persian News Corpus.
The goal of this project is to build a complete classical search engine pipeline, including Persian text preprocessing, indexing, TF-IDF based document representation, query processing, document ranking, and retrieval evaluation.
The system is implemented from scratch using Python and focuses on understanding the fundamental concepts behind information retrieval systems.
- Persian text normalization and preprocessing
- Tokenization and cleaning of Persian documents
- Stopword removal using a Persian stopword list
- Lemmatization and text normalization
- Vocabulary construction
- Inverted index generation
- TF-IDF document representation
- Query processing and ranking
- Document retrieval using cosine similarity
- Retrieval evaluation using standard IR metrics
- Query refinement and Persian spelling correction (optional module)
The project pipeline consists of the following stages:
Documents
|
v
Persian Text Preprocessing
|
v
Tokenization + Normalization + Stopword Removal
|
v
Vocabulary Creation
|
v
Inverted Index
|
v
TF-IDF Representation
|
v
Query Processing
|
v
Document Ranking
|
v
Evaluation
Persian documents are processed before indexing.
The preprocessing stage includes:
- Character normalization
- Persian text cleaning
- Removing unnecessary symbols
- Tokenization
- Stopword removal
- Lemmatization
An inverted index is created to efficiently map terms to the documents containing them.
Example:
term --> document IDs
This structure allows fast document retrieval during query processing.
Documents are represented using TF-IDF weighting.
The system calculates:
- Term Frequency (TF)
- Document Frequency (DF)
- Inverse Document Frequency (IDF)
TF-IDF vectors are then used for measuring document-query similarity.
The retrieval module ranks documents based on cosine similarity between:
- Query vector
- Document vectors
The system returns the most relevant documents with their similarity scores.
The retrieval system is evaluated using standard Information Retrieval metrics:
- Precision@K
- Recall@K
- F1 Score
- Average Precision (AP)
- Mean Average Precision (MAP)
The evaluation module compares retrieved documents with relevance judgments.
The optional notebook extends the main IR system with additional features:
- Query refinement
- Persian spelling correction
- Edit distance based correction
- Interactive query interface
This module is separated from the core retrieval pipeline.
Persian-Information-Retrieval-System
|
├── README.md
|
├── notebooks
│ ├── Persian_IR_System_Main.ipynb
│ └── Persian_IR_System_Optional.ipynb
|
├── resources
│ └── persian_stopwords.txt
|
├── data
│ └── README.md
|
└── assets
This project uses the Hamshahri Persian News Corpus.
Due to dataset size and distribution considerations, the original corpus is not included in this repository.
To run the notebooks, download the dataset separately and place it according to the instructions provided in the data folder.
Expected structure:
data/
|
└── HamshahriCorpus/
- Python
- Jupyter Notebook
- Pandas
- NumPy
- Hazm
- Natural Language Processing (NLP)
- Information Retrieval
- TF-IDF
- Cosine Similarity
This project demonstrates the implementation of a complete classical Persian search engine and covers important IR concepts:
- Text processing
- Indexing
- Ranking models
- Similarity measurement
- Retrieval evaluation
Computer Engineering - Artificial Intelligence
Project: Persian Information Retrieval System