LLMs have stagnated on long-form non-fiction writing, and the author of a new post-training textbook, 'Reinforcement Learning from Human Feedback,' has the receipts. Nathan Lambert used current frontier models throughout the writing process and found a consistent ceiling: models can check individual sentences, catch typos, and unblock writer's block at the section level, but they fall apart across a full chapter. GPT 4.5 and Kimi K2 remain benchmarks for writing quality despite being old releases, while models have gone from zero to superhuman at coding and math in the same window.
The failure mode is specific and worth understanding. Models increase entropy in long-form writing. They compound errors as they stack additions, they over-reach stylistically where precision is needed, and they cannot organize and compress knowledge into coherent argument. Lambert draws a direct line from this to AI science ambitions: organizing knowledge into insight is a prerequisite for solving open-ended problems, not a downstream skill. On the same day Anthropic published progress on the Riemann Hypothesis, Lambert is arguing that coverage across scientific breadth is far thinner than assumed.
The original piece is worth reading in full for two reasons. First, Lambert distinguishes exactly where each model excels: GPT found deep typos across a 200-300 page PDF, Claude provided structural editorial judgment. Second, he connects the writing failure to the absence of RLVR-style training signals for prose, the same mechanism that made math and code tractable. If that analogy holds, the path to fixing this is known but unbuilt. Read 'Interconnects' to see where he thinks the ceiling actually is.
[READ ORIGINAL →]