School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs
reward hacking SFT
They basically just created a SFT dataset of reward-hackable easy prompts, and then show that the resulting models exhibit emergent misalignment
I think it might be useful to get a sense of the type of tasks that I could use, maybe, which are susceptible to reward hacking but not that hard? But other than that, I don’t think SFT on this kind of data set is the way I want to get reward hacking in my model organism
