Skip to content

About

A deep reinforcement learning system for portfolio optimization that combines historical price data with market sentiment (via FinBERT) to make risk-adjusted trading decisions. Uses PPO/DQN agents with a risk-penalized reward function, backtested against baseline strategies across varying market conditions.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

DRL Portfolio Robo-Advisor

A reinforcement-learning driven portfolio allocation system for Indian equities (Nifty 100), combining a PPO (Proximal Policy Optimization) trading agent with a hybrid sentiment veto system that lets a human operator review and override risky trades before they're executed.

Repo: Optimize-Trading-DRL


Overview

The system works in two stages:

  1. RL Policy Inference — A PPO agent (trained with Stable-Baselines3) analyzes recent price action, moving averages, and other technical signals across ~80+ Nifty 100 stocks, then outputs a target portfolio allocation.
  2. Hybrid Sentiment Guardrail — Before any trade is finalized, each proposed stock is checked against live news sentiment (via FinBERT, scored on headlines pulled from Google News RSS). Stocks with strongly negative sentiment are flagged for human review instead of being auto-executed or auto-blocked — the final call is made by the user, from the dashboard.

This human-in-the-loop veto step is the core idea: the RL agent proposes, live sentiment flags risk, and the user disposes.


Architecture

┌─────────────────────┐      ┌──────────────────────┐      ┌────────────────────┐
│   React + TS         │      │   Node.js / Express    │      │   FastAPI (Python)   │
│   Dashboard (Vite)   │ ───▶ │   API + MongoDB        │ ───▶ │   RL + Sentiment       │
│                       │ ◀── │   (persists portfolios) │ ◀── │   Engine               │
└─────────────────────┘      └──────────────────────┘      └────────────────────┘
  • FastAPI service — loads the trained PPO model, pulls live market data (yfinance), runs policy inference, scores news sentiment with FinBERT, and returns a structured portfolio result.
  • Node/Express API — triggers the FastAPI engine, persists each run to MongoDB, and serves the result to the frontend.
  • React dashboard — visualizes the allocation, performance vs. the Nifty 100 benchmark, live financial news sentiment, and the interactive veto review panel.

Key Features

  • PPO-based allocation — target weights generated by a trained RL policy, not a static rules-based strategy.
  • Live sentiment scoring — FinBERT-scored headlines (Google News RSS) per stock, refreshed on every run.
  • Interactive sentiment veto — flagged (high-risk) stocks are not auto-decided. The dashboard shows each one with its target weight and sentiment score; the user chooses Approve (invest) or Keep cash (reroute to cash reserve) per stock, applied instantly.
  • Performance vs. benchmark — Robo-Advisor performance plotted against Nifty 100 buy-and-hold, with Sharpe ratio and historical drawdown.
  • Live financial news feed — sentiment-tagged headlines with links back to source articles.
  • Cash reserve tracking — capital held back (either too small an allocation to act on, or vetoed by the user) is tracked separately from invested capital, shown as a live % + ₹ amount.

Tech Stack

Layer Technology
RL Engine Python, Stable-Baselines3 (PPO), Gymnasium-style custom env
Sentiment HuggingFace transformers (FinBERT — ProsusAI/finbert)
Market Data yfinance
News Google News RSS via feedparser
AI Service API FastAPI, Uvicorn
Backend API Node.js, Express, MongoDB (Mongoose)
Frontend React, TypeScript, Vite, Tailwind CSS, Recharts

Project Structure

Risk-Adjusted_Trading_Portfolio/
├── main.py                      # FastAPI entrypoint
├── phase7_hybrid_veto.py        # RL inference + hybrid sentiment veto pipeline
├── portfolio_env.py             # Custom trading environment
├── Nifity50service.py           # Nifty 100 benchmark performance endpoint
├── trained_models/              # Trained PPO model + VecNormalize stats
│   ├── ppo_portfolio_agent.zip
│   ├── vecnormalize_portfolio_stats.pkl
│   └── training_tickers.json
├── server/                      # Node/Express backend
│   ├── controllers/
│   │   └── portfolioController.js
│   └── models/
│       └── Portfolio.js
└── client/                      # React + TypeScript frontend
    └── src/
        ├── pages/
        │   └── Dashboard.tsx
        ├── components/
        │   ├── PortfolioAllocation.tsx
        │   ├── PerformanceAnalytics.tsx
        │   └── Newsfeed.tsx
        └── services/
            └── portfolioApi.ts

Getting Started

Prerequisites

  • Python 3.10+ with a virtual environment (e.g. conda or venv)
  • Node.js 18+
  • MongoDB (local or Atlas)

1. Python AI Engine

# from the project root
pip install -r requirements.txt

python main.py
# → FastAPI running on http://0.0.0.0:8000

2. Node/Express Backend

cd server
npm install
npm run dev
# → API running on http://localhost:5000

3. React Frontend

cd client
npm install
npm run dev
# → Dashboard running on http://localhost:5173

How the Veto System Works

For every stock the RL agent selects:

  1. Recent headlines are fetched and scored with FinBERT.
  2. If net sentiment falls below a threshold (SENTIMENT_THRESHOLD = -1.0), the stock is flagged, not auto-vetoed.
  3. Flagged stocks appear in the dashboard's "Veto interventions" panel, each with its target weight and sentiment score.
  4. The user clicks Approve (stock is invested) or Keep cash (its weight is rerouted to the cash reserve) — applied immediately, no separate confirmation step.
  5. The Current Portfolio Allocation table and Cash Reserve indicator update live to reflect the decision.

Disclaimer

This project is for educational and research purposes. It does not constitute financial advice, and the RL agent's outputs should not be used for actual trading decisions without independent due diligence.


License

MIT — see LICENSE for details.

About

A deep reinforcement learning system for portfolio optimization that combines historical price data with market sentiment (via FinBERT) to make risk-adjusted trading decisions. Uses PPO/DQN agents with a risk-penalized reward function, backtested against baseline strategies across varying market conditions.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages