ChronicleBenchThe Chronicle Result

Archived Chronicle result · ChronicleBench v1.0

01 The finding

Frontier models write excellent novellas. Chronicle writes complete novels.

Every system on this page was asked for the same thing: a complete 85,000-word novel from a fixed brief, scored with the same pinned instrument at fixed positions through the manuscript. Frontier models produce wonderful short fiction and then stop, stall, or decay. Chronicle turns one of those models into a complete-novel system — and improves its quality at long range, most of all at the ending.

2.4×
The natural length of its own underlying model. Raw Sonnet 4.6 delivers ~30K words and stops; Chronicle delivers 70–76K
+5.4
Full-length quality over the same model pushed to the same horizon (71.0 vs 65.6 ChronicleBench Score)
+12.9
Ending quality vs sustained raw Sonnet (66.9 vs 54.0) — the position where raw models collapse hardest

ChronicleBench is built by the Chronicle team — read that disclosure as a reason to check our work, not to take our word. Contract frozen and committed before any Chronicle run · one documented amendment, direction favored competitors · no cherry-picks, no selective reruns · every manuscript, score and negative result published. Full board & methodology · raw data.

02 The architecture premium

Same model. Same briefs. Chronicle added.

The cleanest experiment in the benchmark: Claude Sonnet 4.6 writing raw under a sustained-continuation harness, versus the identical model inside Chronicle's generation architecture. Same underlying model, briefs and scoring instrument — the generation architecture is the treatment. If Chronicle were "just a wrapper," these columns would match.

DimensionRaw Sonnet 4.6ChronicleChronicle effect
Natural delivered length~30,000 words70–76K words≈2.4× longer
Complete narrative endingat novella scale onlyyes, at novel scale6/6 books ended properly
Opening quality71.9278.15+6.23
Midpoint quality (50K)68.869.4+0.6
Ending quality54.066.9+12.9
Full-length ChronicleBench Score65.671.0+5.4
Generation cost per book~$2.41~1.6–2 h wall-clock
Raw Sonnet 4.6, sustained Chronicle (same model inside)
Opening
71.9
Opening
78.2
50K
68.8
50K
69.4
Ending
54.0
Ending
66.9

Position-scored quality on the pinned instrument. The raw model's curve is the familiar long-form failure: strong start, mid-book slide, collapse at the ending. Chronicle's architecture protects quality precisely where raw generation loses it.

Chronicle is not "more tokens." It changes what arrives at the end of the book — the part readers judge a novel by.

† Read the +5.4 with its caveats. Chronicle's full-length score currently rests on its two eligible books (both brief N1), and and all v1.0 entries — Chronicle and models alike — were scored over four windows under Amendment 1 (v1.1 moved to five), so the full-length figure is not directly comparable across benchmark versions. The positional deltas (opening, 50K, ending) are window-for-window comparisons and stand on their own. The preregistered v1.1 cohort — corrected 80–95K eligibility gate, uniform five-window scoring — has since completed (144/144 runs); Chronicle's v1.1 system entry is deferred because the current production engine naturally delivers 70–76K-word books, below that floor.

03 The complete-novel task

The complete-novel task: current evidence.

Everything tested against the 85K commission so far, exactly as the contract computes it — one SYSTEM entrant and four raw models under the standard continuation harness. They are not the same kind of entrant; the Full-Novel Systems track proper starts when real system competitors join Chronicle. GPT-5.6 Sol leads at 72.5 with Chronicle at 71.0 — a preregistered statistical tie (Δ1.5 against a 3.0 tie band). Chronicle's asterisk is self-inflicted and disclosed: a production length-calibration defect left only 2 of its 6 books above the eligibility floor. The v1.1 cohort has since completed (144/144 runs); Chronicle's v1.1 system entry is deferred because the current production engine naturally delivers 70–76K-word books, below the cohort's 80K eligibility floor. The same-model architecture comparison on this page remains the published controlled result.

#System / modelChronicleBench ScoreEligible runsStatus
T1GPT-5.6 Sol MODEL+HARNESS72.56/6completes; decays ~15 points over the book
T1Chronicle SYSTEM71.02/6statistical tie with Sol; length defect fixed, rerun preregistered
3Claude Sonnet 4.6 MODEL+HARNESS65.66/6completes at the lowest eligible quality
GPT-5.6 TerraDNF0/4no ending within the standard budget
Claude Fable 5DNF0/2best prose in the field; refuses novel length (stops at 23–35K)

No entrant yet combines top-tier prose, reliable 85K delivery and protected ending quality. Chronicle comes closest to the complete product — its length calibration still has to prove itself in v1.1. The v1.1 cohort has since completed; this table remains archived as the v1.0 result.

04 Frontier models on their own terms

Wonderful short fiction is not a novel.

Raw models deserve a race they can actually win — and lose informatively. Their story on this corpus:

Best short-horizon prose: Claude Fable 5

Quality 82–86 on what it chooses to write — the highest in the field — with endings scored 79–86. But it stalls at 23–35K words, under half the commission, on every run. Brilliance that declines the distance.

Best sustained quality: GPT-5.6 Sol

The only raw entrant that completes every 85K run — at the cost of a ~15-point slide from opening to ending. Endurance with erosion.

Completes at the floor: Claude Sonnet 4.6

Finishes the distance under the harness, at the lowest eligible quality — and its ending window (54.0) is the collapse Chronicle's architecture exists to prevent.

Did not finish: GPT-5.6 Terra

No ending inside the standard call budget on any run. The DNF is itself the finding: sustained completion is a capability, not a given.

A model that writes an 86-quality novella did not lose a tiebreak to a novel system — it entered a different event. That's why systems and models are ranked apart.

05 What happens next

ChronicleBench v1.1 — preregistered before results exist.

  • Rerun Chronicle on the current production engine (2 briefs × 3 books, explicit 85K commission) after the validated length-calibration fix — no cherry-picking; all six count.
  • Fix the contract: eligibility gate moves to 80–95K so every eligible entry supports all five scoring windows; all entries rescored uniformly.
  • Cross-lab robustness rescore: a non-Anthropic judge rescoring of opening/ending windows, published as a robustness annex — answering the same-family objection before anyone raises it.
  • Real system competitors join the systems track. Until then we claim the architecture premium over raw Sonnet — not "#1 full-novel system."
  • This result stays archived as the prior benchmark version. History doesn't get rewritten here.
  • Every future model addition follows the same frozen protocol — pinned slug, pinned provider, both arms, full replicates, published manifests: the standing methodology.