This project recreates the “Sim-to-Real Reinforcement Learning for a Rotary Double-Inverted Pendulum” controller using the Truncated Quantile Critics (TQC) algorithm. The repo contains:
train_rdip_tqc.py: main training loop (PyTorch); saves TensorBoard logs and per-episode rollouts.rdip_env.py: rotary double-inverted pendulum simulator derived directly from the paper.tqc.py: TQC implementation (actor, critics, replay buffer).animate_latest_episode.py: visualizes saved episodes as 2D animations.
The instructions below walk through installation, running training, monitoring progress, animating episodes, and using the exported actor (rdip_tqc_actor.pt).
# Clone and enter workspace
git clone <repo-url>
cd rdip
# Create an isolated virtual environment
python3 -m venv .venv
source .venv/bin/activate
# Upgrade packaging tools (recommended)
python -m pip install --upgrade pip setuptools wheel
# Install runtime dependencies
python -m pip install torch numpy matplotlib tensorboard pypdf
# (Optional) install uv for faster dependency management
python -m pip install uvNote: If you are running on macOS and encounter pip / permission errors, prefer using the venv steps above and avoid pip install --user. On CUDA machines, verify that your PyTorch install matches the desired CUDA version.
Activate the virtual environment and launch the trainer:
source .venv/bin/activate
python train_rdip_tqc.pyKey defaults:
total_steps=5_000_000(≈5,000 episodes, 10 s each at 10 ms control intervals).- Episode transitions cycle through equilibrium modes EP0–EP3 to match the paper.
- TensorBoard logs write to
runs/TQC_<timestamp>_seed<seed>/. - Each completed episode is archived as
episode_XXXXX.npzfor later visualization. - The latest policy exports to
rdip_tqc_actor.pt(overwritten on each run).
Override any training argument by calling train(total_steps=..., seed=...) within a short script or REPL.
During training the loop saves the most recent rollout to runs/<run_dir>/latest_episode.npz and archives every episode as episode_00001.npz, episode_00002.npz, ….
Use the animation helper to replay a specific episode:
source .venv/bin/activate
# Point MPLCONFIGDIR to a writable cache if Matplotlib complains about ~/.matplotlib
MPLCONFIGDIR=/tmp/matplotlib python animate_latest_episode.py \
--run runs/TQC_20251012-231308_seed42 \
--episode 5Flags:
--run <path>: run directory (defaults to newest folder inruns/).--episode N: zero-based or integer episode index; the script automatically converts toepisode_0000N.npz.--path <file.npz>: specify a file directly.--loop: keep watching thelatest_episode.npzfile and re-display as new episodes complete.
The viewer shows a side-view stick figure plus angle traces over time.
TensorBoard logs are stored per training run under runs/TQC_<timestamp>_seed<seed>/.
Launch TensorBoard in another terminal:
source .venv/bin/activate
tensorboard --logdir runsNavigate to the printed URL (typically http://localhost:6006). You can track:
train/q_loss,train/pi_loss: critic and policy losses.train/alpha,train/entropy: auto-temperature statistics.episode/return,episode/ema_return: raw/EMA episode returns.episode/mode: equilibrium mode index per episode.
At the end of each training run the script saves the actor network as a TorchScript file:
source .venv/bin/activate
python train_rdip_tqc.py # produces rdip_tqc_actor.ptTo load and run the policy:
import torch
from rdip_env import RDIPEnv
actor = torch.jit.load("rdip_tqc_actor.pt")
actor.eval()
env = RDIPEnv(seed=0)
obs = env.reset(ep_mode=0) # 10-D observation: sin/cos angles, velocities, EP context
with torch.no_grad():
action = actor(torch.tensor(obs).unsqueeze(0))[0] # returns 1-D torque commandYou can step the simulation manually or integrate the actor into a control stack. For deployment, ensure the observation ordering matches RDIPEnv._obs() and feed the resulting angular acceleration into hardware or a higher fidelity simulator.
If you wish to keep multiple checkpoints, copy or rename rdip_tqc_actor.pt after each run.