Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
212 changes: 196 additions & 16 deletions _posts/2026-08-04-sail-revkl-post.md
Original file line number Diff line number Diff line change
@@ -1,32 +1,212 @@
---
title: "SAIL-RevKL: Stable and Provable Self-Improving Online LLM Alignment"
title: "SAIL-RevKL: Stable and Provable Self-Improving Online LLM Alignment (UAI 2026)"
image: images/logo.png
author: Xudong Wu
tags: technique
tags: publication llm-alignment optimization theory
---

[**Paper**](https://arxiv.org/abs/2606.31524) · [**MuJoCo Code**](https://github.com/xudongwu-0/SAIL_mujoco) · [**LLM Code**](https://github.com/xudongwu-0/SAIL_LLM)
Our paper, **"On the Convergence of Self-Improving Online LLM Alignment,"** appeared at the **42nd Conference on Uncertainty in Artificial Intelligence (UAI 2026)**. It addresses a basic theoretical gap behind self-improving alignment.

We are excited to share **SAIL-RevKL**, our work on making self-improving online alignment more stable and theoretically grounded. The paper, *On the Convergence of Self-Improving Online LLM Alignment*, has been accepted at **UAI 2026**.
## Research Question

## Why Online Alignment?
**Can we establish a global convergence guarantee for online RLHF formulated as a bilevel optimization problem?**

Most language-model alignment methods learn from a fixed preference dataset. While this approach is simple and effective, the dataset may not fully represent the responses produced by the model as it continues to improve.
[**Proceedings**](https://proceedings.mlr.press/v337/wu26c.html) · [**arXiv**](https://arxiv.org/abs/2606.31524) · [**Poster**](/images/sail-revkl-poster.png) · [**MuJoCo Code**](https://github.com/xudongwu-0/SAIL_mujoco) · [** Code**](https://github.com/xudongwu-0/SAIL_LLM)

Online alignment creates a continuous feedback loop: the current model generates new responses, receives fresh preference feedback, and then learns from that feedback. [SAIL](https://arxiv.org/abs/2406.15567) is an efficient framework for this setting, but an important question remained open: **will its learning process converge reliably?**
{% include figure.html image="images/sail-revkl-poster.png" link="images/sail-revkl-poster.png" width="100%" caption="Our UAI 2026 poster. Click the image to open it at full resolution." %}

## Our Approach: SAIL-RevKL
## TL;DR

We find that the original SAIL objective can have an unfavorable optimization landscape, which may make training unstable outside a local region. To address this issue, we introduce **SAIL-RevKL**, a simple extension that adds a reverse KL regularization term.
Our answer is conditional but positive. We first show that the original **Self-Improving Alignment (SAIL)** objective is guaranteed to have favorable curvature only near its initialization. We then introduce **SAIL-RevKL**, which adds

Intuitively, the reverse KL term acts as a guardrail. It keeps the updated policy anchored to a stable reference policy while still allowing the model to learn from new preference feedback. This small change makes the learning problem easier to optimize and allows us to establish a global convergence guarantee within a bounded parameter space.
$$
\gamma\,\mathbb{E}\_x\!\left[
D\_{\mathrm{KL}}\!\left(
\pi\_{\mathrm{ref}}(\cdot\mid x)\,\Vert\,\pi\_\theta(\cdot\mid x)
\right)
\right]
$$

## Main Highlights
as a penalty. For a fixed reference and a log-linear policy, the Hessian of this reverse KL is the Fisher information matrix. With enough regularization, that positive curvature dominates the adverse curvature of SAIL throughout the full bounded parameter region. This yields global strong concavity and a Polyak-Lojasiewicz (PL) condition for the regularized surrogate, together with near-linear dependence on the inverse target accuracy.

- We explain why vanilla SAIL only has favorable convergence properties locally.
- We introduce reverse KL regularization as a simple way to stabilize SAIL.
- We establish a global convergence guarantee with near-linear sample complexity under standard assumptions.
- We show improved training stability and performance on both continuous-control benchmarks and LLM alignment tasks.
## Why Bilevel-Formulated Online RLHF Is Hard

We evaluate SAIL-RevKL on Door Open, Walker Walk, Walker Stand, and Cheetah Run, as well as LLM alignment experiments using PKU-SafeRLHF and UltraFeedback. Across these settings, SAIL-RevKL consistently improves upon vanilla SAIL, showing that the reverse KL term is useful both theoretically and in practice.
Offline preference optimization trains on a fixed dataset. Online RLHF instead couples two learning problems: a reward model is fitted from preferences over responses generated by the current policy, and the policy is updated against that learned reward. This interaction is naturally formulated as a bilevel optimization problem because the solution of one level determines the objective and data distribution of the other.

SAIL uses the reward-policy equivalence to reduce this computationally expensive bilevel problem to an efficient single-level surrogate. If $\mathcal{D}\_\theta$ denotes the policy-dependent response and preference distribution, that surrogate has the form

$$
J(\theta)
=
\mathbb{E}\_{(x,y\_w,y\_l)\sim\mathcal{D}\_\theta}
\!\left[
\log\sigma\!\left(
\beta\log\frac{\pi\_\theta(y\_w\mid x)}{\pi\_{\mathrm{SFT}}(y\_w\mid x)}
-
\beta\log\frac{\pi\_\theta(y\_l\mid x)}{\pi\_{\mathrm{SFT}}(y\_l\mid x)}
\right)
\right].
$$

This dependence matters. Standard performance-difference arguments assume that the reward or evaluation measure is independent of the policy appearing inside the advantage. That assumption no longer holds here, so convergence cannot be inherited directly from ordinary policy-gradient analysis. We instead analyze the geometry of $J$ itself.

The negative Hessian of $J$ decomposes into four terms. Some are stabilizing, while interaction terms can become unfavorable as $\theta$ moves away from the initial SFT parameters $\theta\_0$. The proof must therefore control the objective's curvature throughout the region the algorithm may visit, not only at its starting point.

{% include alert.html
type="info"
content="**Hint: why the research question is hard.** In this bilevel formulation of online RLHF, the reward-learning level depends on data from the current policy, while the policy-update level depends on the learned reward. The two levels therefore evolve together, so a fixed-data convergence argument does not directly apply."
%}

## Assumptions and Scope

The global theorem is explicit about its scope. It applies to the following setting:

| Assumption | Mathematical condition | Role in the analysis |
|---|---|---|
| **A1: Log-linear policy** | $\pi\_\theta(a\mid s)\propto\exp(\theta^\top\psi(s,a))$, with $\max\_{s,a}\Vert\psi(s,a)\Vert\_2\le 1$ and $\Vert\theta-\theta\_0\Vert\_2\le B\_\theta$ | Makes the objective's curvature analyzable on a bounded parameter set. |
| **A2: Informative features** | $F\_\rho(\theta)\succeq\mu\_F I$ uniformly over the feasible set | Ensures the Fisher curvature is positive in every parameter direction. |
| **A3: Smooth objective** | $J\_\gamma$ is $L$-smooth | Controls the change of the gradient and permits a stable stepsize. |
| **A4: Stochastic gradients** | Each sample gradient is unbiased and has variance at most $\sigma\_g^2$ | Makes the mini-batch noise decay as $\sigma\_g^2/B\_s$. |

Two construction choices are also essential. The reference policy $\pi\_{\mathrm{ref}}$ is fixed during the analyzed update, and projected gradient ascent keeps every iterate inside $\Theta$, the radius-$B\_\theta$ Euclidean ball centered at $\theta\_0$.

**Scope of the result.** The word "global" below means global **within this full bounded feasible set**. The LoRA experiments test whether the method remains useful beyond these assumptions; the last-layer-only experiments are the closest empirical match to A1.

## Core Theorem: Global Convergence of SAIL-RevKL

The original SAIL surrogate is guaranteed to have favorable geometry only near its initialization. We add a reverse-KL penalty and optimize

$$
J\_\gamma(\theta)
=
J(\theta)
-
\gamma\,\mathbb{E}\_x\!\left[
D\_{\mathrm{KL}}\!\left(
\pi\_{\mathrm{ref}}(\cdot\mid x)\,\Vert\,\pi\_\theta(\cdot\mid x)
\right)
\right].
$$

For a fixed reference and a log-linear policy, the reverse-KL Hessian is the Fisher information matrix, so

$$
-\nabla^2J\_\gamma(\theta)
=
-\nabla^2J(\theta)+\gamma F(\theta).
$$

### Theorem (Global convergence on the feasible set)

Let $\Theta$ be the radius-$B\_\theta$ Euclidean ball centered at $\theta\_0$. Suppose A1-A4 hold, $\pi\_{\mathrm{ref}}$ remains fixed during optimization, and projected updates keep every iterate in $\Theta$. Choose $\gamma$ sufficiently large to satisfy the explicit curvature threshold in Theorem 2 of the paper, so that the Fisher term dominates the adverse curvature of the original SAIL objective throughout $\Theta$.

Then $J\_\gamma$ is $\mu$-strongly concave on all of $\Theta$ and satisfies the PL inequality

$$
\Vert\nabla J\_\gamma(\theta)\Vert\_2^2
\ge
2\mu\left(J\_\gamma(\theta\_\gamma^{\ast})-J\_\gamma(\theta)\right),
\qquad \theta\in\Theta,
$$

where $\theta\_\gamma^{\ast}=\arg\max\_{\theta\in\Theta}J\_\gamma(\theta)$ is the unique maximizer. Moreover, for projected mini-batch stochastic gradient ascent with a valid constant stepsize, there exist constants $q\in(0,1)$ and $C>0$ such that

$$
\mathbb{E}\!\left[
J\_\gamma(\theta\_\gamma^{\ast})-J\_\gamma(\theta\_T)
\right]
\le
q^T\left(J\_\gamma(\theta\_\gamma^{\ast})-J\_\gamma(\theta\_0)\right)
+
\frac{C\sigma\_g^2}{B\_s}.
$$

Thus the optimization error decreases geometrically until it reaches the mini-batch noise floor. Taking $B\_s=\widetilde{\mathcal{O}}(\varepsilon^{-1})$ and $T=\widetilde{\mathcal{O}}(\log\varepsilon^{-1})$ gives an $\varepsilon$ expected function-value gap with total sample complexity

$$
\widetilde{\mathcal{O}}\!\left(
\varepsilon^{-1}\log\varepsilon^{-1}
\right).
$$

{% include alert.html
type="info"
content="**Hint: what the theorem answers.** For the SAIL reduction of bilevel-formulated online RLHF, reverse KL turns a local convergence certificate into a global, non-asymptotic guarantee over the entire bounded feasible set."
%}

The theorem is proved in a deliberately controlled log-linear setting. We next test whether the same regularized objective is useful when training real language models.

## LLM Alignment Results

The main experiments use LoRA fine-tuning and GPT-4 as an LLM judge. PKU-SafeRLHF evaluates the balance between helpfulness and harmlessness, while UltraFeedback focuses on general instruction following. Pairwise win rate is the primary comparison below; tie rate and mean GPT score difference are reported separately because they capture different aspects of the judge outcomes.

**PKU-SafeRLHF, Qwen 0.5B**

| Method | Pairwise win rate | Tie rate | Mean GPT score difference |
|---|---:|---:|---:|
| DPO | 26.0% | 14.0% | -2.209 |
| SAIL | 29.0% | 0.0% | -2.004 |
| **SAIL-RevKL** | **43.0%** | 6.0% | **-0.885** |

**UltraFeedback**

| Backbone | Method | Pairwise win rate | Tie rate | Mean GPT score difference |
|---|---|---:|---:|---:|
| Qwen 0.5B | DPO | 0.0% | 7.0% | -2.580 |
| Qwen 0.5B | SAIL | 1.0% | 8.0% | -2.110 |
| Qwen 0.5B | **SAIL-RevKL** | **3.0%** | **18.0%** | **-1.300** |
| Phi-3 3.8B | DPO | 11.0% | 15.0% | -1.160 |
| Phi-3 3.8B | SAIL | 15.0% | **60.0%** | **-0.150** |
| Phi-3 3.8B | **SAIL-RevKL** | **30.0%** | 40.0% | -0.300 |
| LLaMA-3 8B | DPO | 22.0% | 24.0% | -0.930 |
| LLaMA-3 8B | SAIL | 23.0% | **29.0%** | -0.700 |
| LLaMA-3 8B | **SAIL-RevKL** | **34.0%** | **29.0%** | **-0.230** |

The headline comparison is the pairwise win rate against vanilla SAIL: **43% vs. 29%** on Qwen with PKU-SafeRLHF, **30% vs. 15%** on Phi-3 with UltraFeedback, and **34% vs. 23%** on LLaMA-3 with UltraFeedback.

### A theory-aligned test

LoRA changes nonlinear features and is not a direct instantiation of the log-linear policy in A1. We therefore also freeze each backbone and train only its final linear layer on UltraFeedback U10. This isolates the regime that most closely matches the theory.

| Backbone | Method | $\gamma$ | Win rate | Lose rate | Tie rate |
|---|---|---:|---:|---:|---:|
| Qwen 0.5B | DPO | - | 1.0% | 91.0% | 8.0% |
| Qwen 0.5B | SAIL | - | 1.0% | 93.0% | 6.0% |
| Qwen 0.5B | SAIL-RevKL | $10^{-3}$ | 3.0% | 91.0% | 6.0% |
| Qwen 0.5B | **SAIL-RevKL** | $10^{-2}$ | **4.0%** | 86.0% | 10.0% |
| Qwen 0.5B | SAIL-RevKL | $10^{-1}$ | 3.0% | **83.0%** | **14.0%** |
| Phi-3 3.8B | DPO | - | 13.0% | **37.0%** | **50.0%** |
| Phi-3 3.8B | SAIL | - | 16.0% | 40.0% | 44.0% |
| Phi-3 3.8B | SAIL-RevKL | $10^{-3}$ | 16.0% | 40.0% | 44.0% |
| Phi-3 3.8B | SAIL-RevKL | $10^{-2}$ | 21.0% | 43.0% | 36.0% |
| Phi-3 3.8B | **SAIL-RevKL** | $10^{-1}$ | **23.0%** | 43.0% | 34.0% |
| LLaMA-3 8B | DPO | - | 0.0% | 93.0% | 7.0% |
| LLaMA-3 8B | SAIL | - | 5.0% | 69.0% | 26.0% |
| LLaMA-3 8B | SAIL-RevKL | $10^{-3}$ | 7.0% | 67.0% | 26.0% |
| LLaMA-3 8B | SAIL-RevKL | $10^{-2}$ | 7.0% | **58.0%** | **35.0%** |
| LLaMA-3 8B | **SAIL-RevKL** | $10^{-1}$ | **13.0%** | 59.0% | 28.0% |

SAIL-RevKL improves the best win rate over vanilla SAIL on all three backbones. The preferred coefficient is scale-dependent: $10^{-2}$ for Qwen 0.5B and $10^{-1}$ for Phi-3 3.8B and LLaMA-3 8B. This is consistent with the theory's curvature-bias trade-off and shows why $\gamma$ should be tuned rather than treated as a universal constant.

{% include alert.html
type="info"
content="**Hint: how to read the evidence.** The LoRA results show that SAIL-RevKL improves pairwise win rates beyond the policy class covered by the theorem, but they do not prove global convergence for LoRA. Frozen-last-layer training is the closer empirical check of the log-linear regime."
%}

## Citation

```bibtex
@inproceedings{wu2026convergence,
title = {On the Convergence of Self-Improving Online {LLM} Alignment},
author = {Wu, Xudong and Liu, Pangpang and Aggarwal, Vaneet and Chen, Jiayu},
booktitle = {Proceedings of the 42nd Conference on Uncertainty in Artificial Intelligence},
pages = {7433--7467},
year = {2026},
volume = {337},
series = {Proceedings of Machine Learning Research}
}
```

## Acknowledgements

This work was supported in part by the Seed Fund for PI Research, Basic Research, from The University of Hong Kong under Project Code 2502251784.
Binary file added images/sail-revkl-poster.png
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading