https://mqsygvcsmjbgregw7vpb44.streamlit.app/
AI-powered multilingual sentiment analysis and demographic analytics for Hajj & Umrah service improvement
Every year, more than 30 million pilgrims participate in Hajj and Umrah, generating millions of feedback comments in over 27 languages.
Traditional manual analysis of this feedback is:
- slow
- expensive
- inconsistent
- impossible at scale
Without automated analysis, authorities struggle to identify recurring service issues, understand demographic trends, and prioritize improvements.
This project demonstrates how Artificial Intelligence and Natural Language Processing (NLP) can transform multilingual feedback into actionable insights for evidence-based decision-making.
This project is an end-to-end AI application that combines multilingual text processing, sentiment analysis, and interactive analytics into a single Streamlit platform.
The system enables decision-makers to:
- Analyze multilingual pilgrim feedback
- Automatically translate comments into English
- Classify sentiment using transformer-based NLP models
- Explore demographic and cross-demographic trends
- Export processed results for further reporting
The application demonstrates how modern AI techniques can support public-sector decision making at scale.
This application leverages AI to analyze and interpret sentiments (comments) from pilgrims participating in Hajj and Umrah. It provides valuable demographic and cross-demographic insights—focusing on age, gender, and nationality—to help stakeholders better understand the experiences and backgrounds of pilgrims.
The Sentiment Analysis module classifies comments from pilgrims as either positive or negative, enabling service providers to assess satisfaction levels and identify recurring issues. This insight is crucial for enhancing service delivery and addressing pilgrims’ concerns.
Given the volume and linguistic diversity of over 30 million comments across 27+ languages, this tool offers a scalable, potential use case, and systematic approach for Saudi authorities and Hajj/Umrah service providers to make data-driven decisions.
The platform helps organizations:
- Monitor service quality using real-time feedback
- Identify recurring complaints across languages
- Understand demographic differences in satisfaction
- Improve operational planning
- Support data-driven policy decisions
Raw Comments|Multilingual Feedback
↓
Language Detection
↓
Translation to English
↓
Text Cleaning
↓
Transformer-based Sentiment Classification
↓
Interactive Analytics Dashboard
↓
Exportable Results
• Provides interactive visualizations of pilgrim demographics (age, gender, nationality).
• Filters available for gender, nationality, and age groups.
• Enables cross-demographic analysis.
• Analyzes textual comments and classifies them as positive or negative.
• Outputs include:
o Original comment
o Translated comment (in English)
o Sentiment label
o Confidence score
• Output is exportable/downloadable.
• The platform automatically:
o detects multilingual comments
o translates text into English
o predicts sentiment
o calculates confidence scores
o allows exportation of results
• Contains detailed guidance on system requirements, usage, inputs, and architecture.
• Programming Language: Python 3.10
• Libraries/Frameworks:
o pandas, numpy, matplotlib, seaborn, plotly (for data processing & visualization)
o nltk, gensim, huggingface transformers, deep-translator (for NLP and translation)
o torch (for model inference)
o streamlit (for interactive UI)
• Others: Git, Bash
• Supported Data Sources:
o Raw data
o File upload (CSV, Excel, others)
o API or URL
o Nationality, Gender, Age, Comments
Nationality Gender Age Comments Egypt Male 45 The experience was wonderful!
o Raw text
o File upload
o API or URL input
pilgrim-sentiment-analysis/
data/
docs/
images/
notebooks/
app.py
requirements.txt
README.md
• Interactive charts and tables showing pilgrim distribution across age, gender, and nationality
• Cross-tab analysis to explore patterns (e.g., satisfaction by gender & nationality)
• Table containing:
o Original comment
o English translation
o Sentiment label (Positive/Negative)
o Confidence score
- data shown is anonymized
- project intended for demonstration
- no confidential pilgrim information included
• Implement real-time data ingestion to replace the current manual data upload mechanism.
• Expand sentiment classification to include neutral or mixed categories.
• Integrate language-specific models to improve accuracy for underrepresented languages.
- Real-time streaming sentiment analysis
- Multi-class sentiment prediction
- Arabic-specific transformer models
- Topic modelling for issue discovery
- Interactive executive reporting
- Docker deployment
- REST API
- Python
- Natural Language Processing
- Transformer Models
- Deep Learning
- Machine Learning
- Data Visualization
- Streamlit
- Plotly
- Translation APIs
- Interactive Dashboards
- Business Intelligence
- AI Application Development
This project strengthened my skills in developing end-to-end AI applications by integrating multilingual natural language processing (NLP), machine translation, transformer-based sentiment analysis, and interactive dashboard development. It enhanced my ability to build scalable data pipelines, process multilingual text, visualize demographic insights, and translate complex analytical outputs into actionable information for decision-makers. Through this project, I also gained valuable experience in deploying AI-powered applications with Streamlit and applying data science to solve real-world challenges in large-scale public service environments.
This documentation provides insights about the application. It highlights valuable information about the application, purpose and audience, user manual, modules, libraries, and technology used, limitations of the app, and future improvements. It also details where to find codes for debugging and improvement.
The application is designed to provide analysis of the sentiments, demographics and cross-demographic analysis of pilgrims visiting the Umrah and Hajj. The models work by classifying comments either positive or negative in their respective key service areas. In reality, Hajj and Umrah experience more than 30 million visitors who leave comments and feedback in their native languages; more than 27 languages across the globe. Analyzing the comments and feedback is a real challenge considering the characteristics of the big data: volume, Velocity, Variety, Veracity, and Value. Accordingly, this AI-driven application remains valuable to authorities helping them to leverage on the comments towards addressing and sorting recurring concerns systematically. The application is significant in understanding the demographic nature of visitors significantly important for preparation, planning, management, and effecting hosting of Umrah and Hajj pilgrims.
The user of the application being administrators and authorities in Hajj and Umrah, can use the application to either analyze sentiments, demographics, or cross-demographic characteristics of visitors. The following sections highlights how to use the features for analysis and text classification.
This feature can be accessed from the homepage (https://sentimentalanalysispilgrim-mwaaysgfzdssubzst7dba2.streamlit.app/). Once on the home page, a user should scroll down to the Cross-Demographic and Demographic Analysis and Sentimental and Text Classification analysis (See the figure below) section on the bottom left of the home page.
Double click on the Cross-Demographic and Demographic Analysis button to navigate to demographic and cross demographic analysis section.
In this section, the user can get comprehensive and interactive insights into the demographics of Hajj and Umrah pilgrims. The key characteristics analyzed here in include:
• Age Distribution: Interactive visualizations illustrating the range and concentration of pilgrims’ ages.
• Statistical Overview: Key metrics including minimum, maximum, mean, quartiles, and mode of pilgrim ages.
• Nationality & Gender Breakdown: Detailed analysis of visitor nationalities segmented by gender.
• Cross-Demographic Insights: Integrated visualizations combining age, gender, and nationality to highlight deeper demographic trends.
To leverage on this feature, scroll down on the page to a section where can input data either as a file, url, API, or textual data. The section can be found on the middle left part of the page. A user should come across the following section:
Load Data Appropriately by Select the right Data Source
Upload File
Enter API URL
Paste Raw Text
Select your appropriate source and load data. A user should ensure that the data loaded has the following at least the following columns:
a. 'الجنسية Nationality',
b. 'الجنس Gender',
c. 'العمر Age'
Note: Without these columns, the analysis will throw an error.
Once data is loaded, the system analyzes the data and displays the visuals of the analysis.
Once data is loaded, a user has the ability to filter data accordingly.
Users should expect to see:
a. A line graphs showing descriptive and dispersion statistical distribution of age characteristics
b. An interactive plotly linear age distribution
c. An interactive Comparative bar graph of nationality by gender
d. An interactive bubble plot of mean ages of nationality and gender
e. An interactive histogram visualizing demographics by nationality, gender, and age
There are no limits to the data a user can input since the system is designed to work with Big data
Once done and need to access the sentimental and text classification feature, double click on the tab Back Home found at the left bottom of the dashboard board.
Once on the main page, https://sentimentalanalysispilgrim-mwaaysgfzdssubzst7dba2.streamlit.app/, scroll down to the bottom left of the page. Click on the button sentimental and text classification.
This feature accepts data inputs as files or text comments. Future improvements will include url and API to allow for real-time sentimental and text classification analysis. For sentiment and text classification analysis, a user can upload a file, type, or copy and paste comments appropriately. Once loaded the data is analyzed and output displayed as dataframe which can be downloaded for further analysis or documentation.
The expected output is:
Once done, a user can input more data or navigate to the main page using the Back Home tab at the left top of the page.
PilgrimageAI is built using a modern and modular tech stack, enabling real-time data visualization, multilingual NLP processing, and a rich interactive user experience. Here's a breakdown of the technologies used:
Backend & App Framework
• Python: Core programming language for logic, data processing, and NLP.
• Streamlit: Rapid web app framework used to build interactive dashboards and forms.
• Streamlit Autorefresh: Enables real-time dashboard updates by auto-refreshing at set intervals.
• Pandas: Efficient data handling and analysis of tabular feedback data.
• Plotly Express: Used for building interactive visualizations (e.g., line charts, histograms, scatter plots).
• Seaborn & Matplotlib: Statistical plotting libraries used for advanced visual analytics like age distribution with statistical markers.
• Transformers (Hugging Face): For sentiment analysis using pretrained models like distilbert-base-uncased-finetuned-sst-2-english.
• GoogleTranslator (Deep Translator): Translates multilingual user feedback into English before processing.
• Custom Text Classification Pipeline: Uses keyword token matching to categorize feedback into themes (e.g., transport, accommodation, staff behavior).
• pdfplumber: Extracts text from PDF feedback files.
• Requests: Fetches data from external APIs.
• StringIO: Converts text and raw CSV input into readable data frames.
• Custom HTML & CSS: Embedded styles add rich visual design (e.g., background images, overlays, styled buttons).
• Responsive Layouts: Uses st.columns, container sizing, and adaptive rendering to ensure the app works well across devices.
• Streamlit Watchdog Disabled: To avoid inotify limit errors on Linux systems, file system watchers are turned off with:
os.environ["STREAMLIT_DISABLE_WATCHDOG_WARNINGS"] = "true"
os.environ["STREAMLIT_WATCHDOG_MODE"] = "none"
OpenAI and Chatgpt: to improve generated codes and debugging the app end-to-end
For Full code structure and debugging visit the notebook and app.py on github public repo