Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Persian Information Retrieval System

Overview

This project implements a Persian Information Retrieval (IR) system for searching and ranking documents from the Hamshahri Persian News Corpus.

The goal of this project is to build a complete classical search engine pipeline, including Persian text preprocessing, indexing, TF-IDF based document representation, query processing, document ranking, and retrieval evaluation.

The system is implemented from scratch using Python and focuses on understanding the fundamental concepts behind information retrieval systems.


Project Features

  • Persian text normalization and preprocessing
  • Tokenization and cleaning of Persian documents
  • Stopword removal using a Persian stopword list
  • Lemmatization and text normalization
  • Vocabulary construction
  • Inverted index generation
  • TF-IDF document representation
  • Query processing and ranking
  • Document retrieval using cosine similarity
  • Retrieval evaluation using standard IR metrics
  • Query refinement and Persian spelling correction (optional module)

System Architecture

The project pipeline consists of the following stages:

Documents
    |
    v
Persian Text Preprocessing
    |
    v
Tokenization + Normalization + Stopword Removal
    |
    v
Vocabulary Creation
    |
    v
Inverted Index
    |
    v
TF-IDF Representation
    |
    v
Query Processing
    |
    v
Document Ranking
    |
    v
Evaluation

Main Information Retrieval Pipeline

1. Text Preprocessing

Persian documents are processed before indexing.

The preprocessing stage includes:

  • Character normalization
  • Persian text cleaning
  • Removing unnecessary symbols
  • Tokenization
  • Stopword removal
  • Lemmatization

2. Inverted Index

An inverted index is created to efficiently map terms to the documents containing them.

Example:

term  -->  document IDs

This structure allows fast document retrieval during query processing.


3. TF-IDF Weighting

Documents are represented using TF-IDF weighting.

The system calculates:

  • Term Frequency (TF)
  • Document Frequency (DF)
  • Inverse Document Frequency (IDF)

TF-IDF vectors are then used for measuring document-query similarity.


4. Retrieval and Ranking

The retrieval module ranks documents based on cosine similarity between:

  • Query vector
  • Document vectors

The system returns the most relevant documents with their similarity scores.


Evaluation

The retrieval system is evaluated using standard Information Retrieval metrics:

  • Precision@K
  • Recall@K
  • F1 Score
  • Average Precision (AP)
  • Mean Average Precision (MAP)

The evaluation module compares retrieved documents with relevance judgments.


Optional Module

The optional notebook extends the main IR system with additional features:

  • Query refinement
  • Persian spelling correction
  • Edit distance based correction
  • Interactive query interface

This module is separated from the core retrieval pipeline.


Repository Structure

Persian-Information-Retrieval-System
|
├── README.md
|
├── notebooks
│   ├── Persian_IR_System_Main.ipynb
│   └── Persian_IR_System_Optional.ipynb
|
├── resources
│   └── persian_stopwords.txt
|
├── data
│   └── README.md
|
└── assets

Dataset

This project uses the Hamshahri Persian News Corpus.

Due to dataset size and distribution considerations, the original corpus is not included in this repository.

To run the notebooks, download the dataset separately and place it according to the instructions provided in the data folder.

Expected structure:

data/
|
└── HamshahriCorpus/

Technologies

  • Python
  • Jupyter Notebook
  • Pandas
  • NumPy
  • Hazm
  • Natural Language Processing (NLP)
  • Information Retrieval
  • TF-IDF
  • Cosine Similarity

Project Objectives

This project demonstrates the implementation of a complete classical Persian search engine and covers important IR concepts:

  • Text processing
  • Indexing
  • Ranking models
  • Similarity measurement
  • Retrieval evaluation

Author

Computer Engineering - Artificial Intelligence

Project: Persian Information Retrieval System

About

A Persian Information Retrieval system implementing text preprocessing, inverted index, TF-IDF ranking, and document retrieval using the Hamshahri news corpus.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages