Finetuning with Sampling: SFT Learns Better Than You Think
SFT with more on-policy data
SFT causes catastrophic forgetting and generalization issues due to being off-policy. But RL struggles with capabilities that the base model didn’t have to begin with (e.g. 0/8 rollouts solve it)
Theoretical Motivation
Their method’s idea is to try to transform SFT data to be more on-distribution:
Suppose that $x \in \mathcal{C}$ is the set of all rollouts where the answer is correct, and $p(x)$ is the base model’s distribution over rollouts. The closest distribution (in KL) $q$ to $p$ which has support only on $C$ is \(p_{\mathcal{C}}(x) \propto p(x) \cdot \mathbf{1}\{ x \in \mathcal{C} \}\). We’d like to get our SFT data to look like it was drawn from $p_{\mathcal{C}}(x)$.
Of course, you can get this distribution by just sampling from $p$ and rejecting until you get $x \in \mathcal{C}$, but what if it’s very low probability? We can use Metropolis Hastings instead
For metropolis hastings, you need your target stationary distribution up to normalization constant, and a proposal distribution $\kappa(x \mid x’)$ for the transition probability between states.

- this is the acceptance probability that you multiply by when moving from $x’$ to $x$
- recall Metropolis-hastings makes it so that the probability flow from $x \to x’$ is and $x’ \to x$ are both the same, and equal to the min under the proposal distribution
- you can get rid of the \(\mathbf{1}\{ x \in \mathcal{C} \}\) if you make sure that your transitions are only on $\mathcal{C}$.
Empirics and Results
In practice, the method looks like
- Take off a suffix of the current generation, use the base model prompted with the correct answer as the proposer (this is an approximation to actual conditioning on $x \in \mathcal{C}$)
- Accept the new sample if it has higher likelihood under the base model than the current completion (approximation to the metropolis-hastings acceptance condition)
Compare to on-policy self-distillation: https://siyan-zhao.github.io/blog/2026/opsd/: they also use the model prompted with the correct answer as a teacher model, but instead of changing the training data (which is still generated by the model being trained), they just edit the loss to be a distillation KL loss
Qwen2.5-7B, Olmo-3-7B, Chemistry, MMLU, AMC

- Seems like it does result in more on policy SFT data

- Less forgetting, higher accuracy.
- they have Table 1 showing that it beats GRPO too