Vol. I · No. 1 · MMXXVI
NextChapterBench A Benchmark of Literary Continuation
13 models · 10 items · 3 judges

Can a model write the next chapter?

Each item gives a model one chapter of a public-domain novel and a deliberately divergent brief for the next; the continuation cannot be retrieved from training data. Three LLM judges grade each chapter on compliance and craft, and the headline score is their product. The full recipe is on the Method page.

§ I

The Standings

Models ranked by headline score at their best reasoning-effort setting: compliance × craft, averaged over all item–judge verdicts.

Run ‘first’ · 10 items · 3 judges · each model name links to its detail page.
Rank Model Headline Compliance Craft Words Cost
1 Fable 5.1 Anthropic · xhigh reasoning effort · via Claude CLIother settingshigh0.77low0.76medium0.75 0.820.990.822,588$15.15
· Fable 5.1 high Anthropic · high reasoning effort · via Claude CLI 0.770.970.802,660$4.67
Fable 5.1 low Anthropic · low reasoning effort · via Claude CLI · 9 of 10 items 0.760.970.782,238$2.30
· Fable 5.1 medium Anthropic · medium reasoning effort · via Claude CLI 0.750.970.782,552$2.78
2 Fable 5 Anthropic · xhigh reasoning effort · via Claude CLIother settingshigh0.80medium0.75 0.800.990.802,194$8.83
· Fable 5 high Anthropic · high reasoning effort · via Claude CLI 0.800.980.812,157$5.13
Fable 5 medium Anthropic · medium reasoning effort · via Claude CLI · 9 of 10 items 0.750.980.771,958$2.44
3 Claude Opus 5 Anthropic · high reasoning effort · via Claude CLIother settingsmedium0.77 0.780.970.802,430$2.02
· Claude Opus 5 medium Anthropic · medium reasoning effort · via Claude CLI 0.770.960.802,433$1.34
4 GPT‑5.6 Sol OpenAI · medium reasoning effort · via Codex CLI 0.740.980.752,235$1.42
5 GPT‑5.5 OpenAI · xhigh reasoning effort · via Codex CLI 0.700.980.722,098$2.28
6 Muse Spark 1.3 Meta · xhigh reasoning effort · via OpenCode CLI 0.600.940.642,656$0.59
7 GPT‑5.6 Terra OpenAI · medium reasoning effort · via Codex CLI 0.590.850.632,164$0.74
8 Gemini 3.7 Flash Google · medium reasoning effort · via Antigravity CLIother settingshigh0.56 0.560.930.603,216$0.36
· Gemini 3.7 Flash high Google · high reasoning effort · via Antigravity CLI 0.560.940.593,371$0.92
9 Gemini 3.8 Flash Google · medium reasoning effort · via Antigravity CLIother settingshigh0.55 0.560.920.613,596$0.57
· Gemini 3.8 Flash high Google · high reasoning effort · via Antigravity CLI 0.550.950.573,092$1.67
10 Muse Spark 1.2 Meta · via OpenCode CLI 0.550.910.602,978$0.42
11 GPT‑5.6 Luna OpenAI · medium reasoning effort · via Codex CLI 0.520.920.562,583$0.08
12 Gemini 3.1 Pro Google · high reasoning effort · via Antigravity CLI 0.460.930.502,792$2.26
13 Hy3 Tencent · via B.AI API 0.450.920.481,662$0.06

  Generation cost, estimated: the model’s input and output tokens for this run priced at current API rates, with cached input charged at the input price. Judge-call costs are excluded. Models without a configured price show an em dash. Laurels are awarded once three or more models have been evaluated. Two scores fewer than about 0.03 apart are within run-to-run noise and should be read as a tie; the measurement is on the Method page.


§ II

Headline Score vs. Cost

Every evaluated model’s headline score against the estimated generation cost of its run; a line joins one model’s effort settings. Tick a model in the key to show or hide it.

Headline score vs. estimated generation cost · run ‘first’, 10 items

Anthropic OpenAI Google Tencent Meta
med Fable 5.1 xhigh Fable 5 xhigh med Claude Opus 5 GPT‑5.6 Sol GPT‑5.5 Muse Spark 1.3 GPT‑5.6 Terra med Gemini 3.7 Flash med Gemini 3.8 Flash Muse Spark 1.2 GPT‑5.6 Luna Gemini 3.1 Pro Hy3
x: estimated generation cost of the 10-item run, USD, at current API list prices; judge-call costs excluded · y: headline score, 0.40–0.85 · a line joins one model’s reasoning-effort settings, low to xhigh; darker shades of a vendor’s hue are newer generations.

§ III

From the Judges’ Verdicts

Each verdict names the chapter’s single worst flaw, with evidence. Quoted verbatim from run ‘first’.

The brief's requirements are executed rather than dramatised — Bruff's letter is a bare transcript of the plot instructions (\"deliver it to old Betteredge to lock in the plate-room overnight\") and Lady Verinder's prohibition arrives unprompted before a word of greeting is exchanged, so the chapter's two central events have no motivation beyond obedience to an unexplained command.

Fable 5, on Hy3’s The Moonstone (item 011)

Lodging Mary in the West Wing — the wing the source establishes as Mr Craven's forbidden private domain — contradicts the story's central geography, and the surrounding prose compounds it by inflating Burnett's plain style with adjective-heavy padding (\"vast, impenetrable void,\" \"immense, rushing sea of cold air\") that runs the chapter half again over its target length.

Fable 5, on Gemini 3.7 Flash (medium)’s The Secret Garden (item 013)

The chapter turns its figures into mouthpieces for its theme—most damagingly the silent lady in black, who becomes a gothic oracle intoning \"One must not listen too closely to its sighing,\" and Robert, whose \"waiting for the sea to call her by some new name\" announces Edna's awakening instead of letting it surface, flattening Chopin's obliquity into statement.

Fable 5, on Gemini 3.1 Pro’s The Awakening (item 016)

§ IV

Read the Chapters

Every chapter each model wrote for the approved briefs, printed in full.


§ V

Further Reading