Towards RL for Superhuman Text: Unslopping AI
Reinforcement Learning from eXpert-Aligned Rubrics
Trying to make AI write better using RL on rubrics which themselves are evolved
Training process:
- meta-optimize a rubric by having an LLM score the human and model generated writings according to the rubric, and while model scores higher than human, show the LLM the most wrong assessment pairs, and tell it to revise the rubric.
- RL the writer LLM against this rubric
- Generate completions from the writer LLM, and repeat
It’s like RL with adversarial rubrics
Data
Need many samples of expert human writing.
Papers: 561 papers + 90 validation from S2ORC, hold out abstract, intro, related works, or conclusion one at a time, have the model complete it given the rest of the paper.
Detail: they tell the model how long to write roughly
Books: Continue from a scene. Filtered for
- quality: in a Pulitzer or Nobel Prize winning book
- memorization: they do some checks to make sure that the model hasn’t memorized the book
Wikipedia: Write the body section of a Wikipedia article from only the title and section heading.
- also filtering for quality, memorization
Evaluations
Final rubric score: They try to measure validation performance by using a different judge (GPT-5.6) to grade by the rubrics, and the average of the 3 rubrics used during each of the 3 steps. Obviously, though, they’re still going to be overreporting a little bit because they trained on a similar metric.
Human evaluation check Standard
Eval for the rubric: Better human writing is scored higher on the rubrics than worse human writing
they took low, medium, and high quality papers

Preliminary Experiments
Optimizing the rubrics is necessary: The LLM judges by default’s preferred model responses (left), and initial rubrics that the LLM comes up with actually prefer the model responses (middle)

Optimizing the rubrics works The rubrics caused the judges to now prefer the human responses over the model ones

The prompts themselves also seem reasonable according to the authors
- sectional ownership (i.e., not a miniature of the whole paper), disciplined selection, economy, and precise on-scope detail rather than breadth of coverage
Experiments and Results
Main experiment
3 iterations of RL-XAR on papers, 1 for stories and wikipedia
Qwen3.5-27B as the writer being RL’d, Qwen3.8-2.4T-A95B as the judge during RL, Kimi K2.6 or Muse Spark 1.1 for the meta optimizer
Papers: beats Opus 5.5


- left: baseline abstract
- right: RL-XAR abstract, it’s noticeably better
Story writing does as well, Wikipedia writing kind mid (they say it’s the hardest / their training judge wasn’t capable enough)

Ablations
Judge needs to be smart enough
Weaker judges can’t tell the difference between humans and the model, even given the rubric

Rubric meta-optimizer needs to be smart enough
Opus and Kimi both produced meta prompts at some point which had a positive gap, Muse didn’t

Misc
Iterative RL does help
- can even do the rubric updating fully online, future work
Rubrics generalize across judges
Without length constraints, judges prefer longer responses