Automated analysis of Telegram groups
"New technologies for investigations: a software to support communication analysis" Bachelor's thesis – University of Turin
The growing volume of unstructured data from instant messaging platforms poses a significant challenge for investigative analysis, generating information overload that slows down and complicates investigations. BigBrother was created in response to this challenge: a software prototype that orchestrates a complete pipeline to transform chaotic Telegram conversations into structured, actionable intelligence. Using advanced Natural Language Processing (NLP) and Large Language Models (LLM) techniques, the system automates data acquisition, sentiment analysis, and topic modelling. Finally, it presents the results in an interactive dashboard designed for investigative exploration. The goal is to provide analysts with a support tool to quickly identify patterns, key players, and critical themes, drastically reducing manual analysis time.
Try the demo at bigbrother.streamlit.app.
Click here to see a preview of the interface
The dashboard provides an overview with key metrics: total messages, number of users, and general emotional polarisation of the chat. User analysis is organised into practical tabs showing messages, average sentiment, and emoji usage for each participant.

Aggregate visualisations show the evolution of sentiment over time and the most common words used throughout the chat, with an interactive slider to adjust the number of words displayed.

By selecting a user from the sidebar, you can explore their specific activity in detail through three different sections: most used words, personal sentiment trends, and favourite emojis.
| Most Frequently Used Words by the User | Personal Sentiment Trend | Most Used Emojis |
|---|---|---|
![]() |
![]() |
![]() |
A search function allows you to analyse the use of specific keywords, showing who uses them, how often over time and the exact context of the messages, with the searched words highlighted in red.

Messages are grouped into thematic clusters to identify the main topics. The dashboard displays the distribution of topics, the average sentiment of each cluster, and an interactive map to explore the semantic proximity of messages.

You can select a specific topic from a drop-down menu to analyse its messages, most active users and popularity over time in detail.

-
Language and Core Libraries:
- Python 3.9.10: Main programming language.
- Pandas: For manipulation and analysis of tabular data.
-
Data Acquisition and Translation:
- Telethon: For interacting with Telegram API, acting as a user client to access the complete chat history.
- DeepL API: For automatic translation of English texts, chosen for its high quality, which is necessary to maximise the performance of NLP models.
-
Natural Language Processing (NLP):
- Hugging Face Transformers: For sentiment analysis, using the pre-trained model
tabularisai/multilingual-sentiment-analysis. - NLTK: For removing stopwords during the analysis of the most common words.
- Sentence-Transformers: For generating high-quality text embeddings that capture the semantic meaning of messages.
- Hugging Face Transformers: For sentiment analysis, using the pre-trained model
-
Topic Modelling and Clustering:
- BERTopic: Advanced framework for semantic clustering of messages and identification of latent topics.
- UMAP: Algorithm for dimensional reduction of embeddings, used for 2D visualisation of clusters.
-
Large Language Models (LLM):
- Ollama: For running local language models (e.g. Mistral) for the automatic generation of descriptive and interpretable labels for identified clusters.
-
Data Visualisation:
- Streamlit: For quickly creating an interactive web dashboard.
- Plotly: For generating interactive and dynamic graphs within the dashboard.
main.py: System entry point. Orchestrates the sequential execution of scraping, translation, analysis and clustering, then launches the dashboard.scraper.py: Module that connects to Telegram, prompts the user to select a chat and downloads messages in CSV format..traduttore.py: Module that uses DeepL's API to translate messages into English.Analysis.py: Module that contains logic for statistical analysis (message frequency, common words, emoji usage) and sentiment analysis. Saves results to CSV files.clustering.py: Module that performs topic modelling with BERTopic and, optionally, queries a local LLM via Ollama to label the identified clusters.dashboard.py: Module that defines the web user interface with Streamlit, displaying the complete analysis results interactively.config.json: Configuration file for storing API keys (Telegram, DeepL) and system settings (e.g. Ollama activation and model selection).clustering_prompt.txt: Prompt template used to instruct the LLM on how to generate labels for clusters.style.css: Style sheet for customising the appearance of the Streamlit dashboard.requirements.txt: List of Python dependencies required to run the project./data: Main directory where subfolders are created for each analysis session, containing all raw data and processed results.
- Scraping:
main.pystartsscraper.py. The user logs into Telegram, selects a chat and the number of messages to download. The data is saved in a[chat_name].csvfile inside a new folder in/data. - Translation:
traduttore.pyreads the CSV, translates the messages into English using the DeepL API, and saves the results in the file[chat_name]_translated.csv. - Analysis:
Analysis.pyprocesses the translated file, performing statistical and sentiment analysis. The results are saved in multiple CSV files (e.g.sentiment.csv,top_words.csv) in the subfolders/analysis. - Clustering:
clustering.pyuses thesentiment.csvfile and groups messages into semantic clusters using BERTopic. It saves the results incluster.csvand, if Ollama is active, generates labels incluster_label.csv. - Path Aggregation: At the end of the pipeline,
main.pycollects all the paths of the generated files and saves them in a single configuration file[chat_name].json. - Visualisation: Finally,
main.pylaunches the dashboard (dashboard.py) by passing the path to the[chat_name].jsonfile. The dashboard dynamically loads all the data and presents it to the user.
-
Clone the repository:
git clone https://github.com/botta0oss/BigBrother.git cd BigBrother -
Install dependencies:
pip install -r requirements.txt
-
Configure API and settings: Create the
config.jsonfile in the root directory of your project and enter your keys and settings.(To obtain
api_idandapi_hash, register at my.telegram.org. Forauth_key, sign up for an API plan at DeepL.com).Use this template for your
config.json:{ "api_id": "YOUR_TELEGRAM_API_ID", "api_hash": "YOUR_TELEGRAM_API_HASH", "phone": "+391234567890", "ollama": true, "modello": "mistral", "auth_key": "YOUR_AUTH_KEY_DEEPL" }phone: Your telephone number associated with your Telegram account, in international format.ollama: Set totrueto enable automatic cluster labelling,falseto disable it.modello: The name of the model you downloaded with Ollama (es.mistral,llama3).auth_key: Ensure you include the suffix:fxif you are using a free DeepL API account.
-
(Optional) Configure Ollama: To use the automatic cluster labelling feature, ensure that you have installed and started Ollama and that you have downloaded the template specified in
config.json.# Example of how to download the Mistral model ollama pull mistral -
Start up the complete system:
python main.py
The terminal will guide you through the authentication and scraping process. Once the entire analysis pipeline is complete, the Streamlit dashboard will automatically open in your browser.
- Demonstrate the feasibility of an end-to-end system to automate the entire analysis flow: from the acquisition of raw data from Telegram to its transformation into structured and viewable information.
- Apply and validate advanced NLP techniques to extract qualitative insights, in particular through Sentiment Analysis to map emotional polarity and Topic Modelling to identify the main topics of discussion without supervision.
- Introduce an innovative approach to the interpretability of results, using a local Large Language Model (LLM) to automatically generate semantic labels for message clusters, solving a common problem in topic modelling.
- Design a visualisation dashboard that is not just a simple static report, but an interactive exploration tool that guides the analyst in understanding the relational and temporal dynamics of the conversation.
- Validate the prototype's usefulness in a realistic operational context, demonstrating how it can support and accelerate an analyst's work in identifying relevant evidence and information within large volumes of text.
Francesco Bottacin
Bachelor's degree in Strategic and Security Sciences
University of Turin
15/09/2025



