A notebook-based data science project for predicting short-term market direction from Russian financial news using transformer-based sentiment classification.
This project fine-tunes a compact RuBERT model on labeled news texts and classifies each article into one of three classes:
0— Increase1— Stable2— Decrease
The predictions are then aggregated by date and combined with market data to build a simple trend indicator and backtesting workflow.
- clean and normalize financial news text
- train a Russian-language NLP classifier
- evaluate model quality with standard classification metrics
- aggregate news predictions into a daily indicator
- compare signals with Brent market data
The training workflow in eval_model.ipynb uses a labeled Excel dataset:
data/text_corpus_news.xlsx— training corpus referenced in the notebook- data/brent_march_2026.xls — market data used for analysis
Note: the Excel corpus is referenced in the notebook but is not included in the current repository snapshot.
The training notebook main.ipynb performs the following steps:
- loads the annotated dataset
- builds a single target label from
Increase,Stable, andDegrease - removes noisy text fragments and formatting artifacts
- splits the data into train / validation / test subsets
- tokenizes text with
cointegrated/rubert-tiny - fine-tunes a sequence classification model
- saves trained weights to models/model_rubert-tiny1
The custom dataset wrapper is CustomDataset, and evaluation is computed in compute_metrics.
- Base model:
cointegrated/rubert-tiny - Framework: PyTorch + Hugging Face Transformers
- Task: multi-class text classification
- Classes: increase / stable / decrease
- Training setup: mixed precision enabled (
fp16=True)
Install dependencies from requirements.txt:
pip install -r requirements.txtFor notebook execution and Excel loading, the following may also be needed:
pip install jupyter openpyxlOpen eval_model.ipynb and run the cells in order.
Open main.ipynb to test the end-to-end workflow, including:
import torch
from transformers import BertTokenizerFast, BertForSequenceClassification
tokenizer = BertTokenizerFast.from_pretrained("cointegrated/rubert-tiny")
model = BertForSequenceClassification.from_pretrained(
"cointegrated/rubert-tiny",
num_labels=3
)
model.load_state_dict(torch.load("models/model_rubert-tiny1", map_location="cpu"))
model.eval()The project tracks the following metrics during validation:
- Accuracy
- F1-score
- Precision
- Recall
See the metric calculation in compute_metrics and the reported outputs in eval_model.ipynb.
An example of calculating the indicator relative to Brent for March 2026:

- notebook-centric workflow
- training corpus is not stored in the repository
- compact model chosen for limited resources rather than maximum accuracy
- current pipeline is research-oriented and should be validated further before production use
- move preprocessing and inference into reusable Python modules
- add experiment tracking and configuration files
- improve train/validation/test splitting strategy
- add CLI or API for inference
- document final benchmark results in a separate report
Distributed under the MIT License.