Finetuning with Sampling: SFT Learns Better Than You Think

SFT with more on-policy data

SFT causes catastrophic forgetting and generalization issues due to being off-policy. But RL struggles with capabilities that the base model didn’t have to begin with (e.g. 0/8 rollouts solve it)

Theoretical Motivation

Their method’s idea is to try to transform SFT data to be more on-distribution:

Suppose that $x \in \mathcal{C}$ is the set of all rollouts where the answer is correct, and $p(x)$ is the base model’s distribution over rollouts. The closest distribution (in KL) $q$ to $p$ which has support only on $C$ is \(p_{\mathcal{C}}(x) \propto p(x) \cdot \mathbf{1}\{ x \in \mathcal{C} \}\). We’d like to get our SFT data to look like it was drawn from $p_{\mathcal{C}}(x)$.

Of course, you can get this distribution by just sampling from $p$ and rejecting until you get $x \in \mathcal{C}$, but what if it’s very low probability? We can use Metropolis Hastings instead

For metropolis hastings, you need your target stationary distribution up to normalization constant, and a proposal distribution $\kappa(x \mid x’)$ for the transition probability between states.

  • this is the acceptance probability that you multiply by when moving from $x’$ to $x$
    • recall Metropolis-hastings makes it so that the probability flow from $x \to x’$ is and $x’ \to x$ are both the same, and equal to the min under the proposal distribution
  • you can get rid of the \(\mathbf{1}\{ x \in \mathcal{C} \}\) if you make sure that your transitions are only on $\mathcal{C}$.

Empirics and Results

In practice, the method looks like

  1. Take off a suffix of the current generation, use the base model prompted with the correct answer as the proposer (this is an approximation to actual conditioning on $x \in \mathcal{C}$)
  2. Accept the new sample if it has higher likelihood under the base model than the current completion (approximation to the metropolis-hastings acceptance condition)

Compare to on-policy self-distillation: https://siyan-zhao.github.io/blog/2026/opsd/: they also use the model prompted with the correct answer as a teacher model, but instead of changing the training data (which is still generated by the model being trained), they just edit the loss to be a distillation KL loss

Qwen2.5-7B, Olmo-3-7B, Chemistry, MMLU, AMC

  • Seems like it does result in more on policy SFT data

  • Less forgetting, higher accuracy.
  • they have Table 1 showing that it beats GRPO too