Archived Chronicle result · ChronicleBench v1.0
01 The finding
Every system on this page was asked for the same thing: a complete 85,000-word novel from a fixed brief, scored with the same pinned instrument at fixed positions through the manuscript. Frontier models produce wonderful short fiction and then stop, stall, or decay. Chronicle turns one of those models into a complete-novel system — and improves its quality at long range, most of all at the ending.
ChronicleBench is built by the Chronicle team — read that disclosure as a reason to check our work, not to take our word. Contract frozen and committed before any Chronicle run · one documented amendment, direction favored competitors · no cherry-picks, no selective reruns · every manuscript, score and negative result published. Full board & methodology · raw data.
02 The architecture premium
The cleanest experiment in the benchmark: Claude Sonnet 4.6 writing raw under a sustained-continuation harness, versus the identical model inside Chronicle's generation architecture. Same underlying model, briefs and scoring instrument — the generation architecture is the treatment. If Chronicle were "just a wrapper," these columns would match.
| Dimension | Raw Sonnet 4.6 | Chronicle | Chronicle effect |
|---|---|---|---|
| Natural delivered length | ~30,000 words | 70–76K words | ≈2.4× longer |
| Complete narrative ending | at novella scale only | yes, at novel scale | 6/6 books ended properly |
| Opening quality | 71.92 | 78.15 | +6.23 |
| Midpoint quality (50K) | 68.8 | 69.4 | +0.6 |
| Ending quality | 54.0 | 66.9 | +12.9 |
| Full-length ChronicleBench Score | 65.6 | 71.0† | +5.4 |
| Generation cost per book | — | ~$2.41 | ~1.6–2 h wall-clock |
Position-scored quality on the pinned instrument. The raw model's curve is the familiar long-form failure: strong start, mid-book slide, collapse at the ending. Chronicle's architecture protects quality precisely where raw generation loses it.
Chronicle is not "more tokens." It changes what arrives at the end of the book — the part readers judge a novel by.
03 The complete-novel task
Everything tested against the 85K commission so far, exactly as the contract computes it — one SYSTEM entrant and four raw models under the standard continuation harness. They are not the same kind of entrant; the Full-Novel Systems track proper starts when real system competitors join Chronicle. GPT-5.6 Sol leads at 72.5 with Chronicle at 71.0 — a preregistered statistical tie (Δ1.5 against a 3.0 tie band). Chronicle's asterisk is self-inflicted and disclosed: a production length-calibration defect left only 2 of its 6 books above the eligibility floor. The v1.1 cohort has since completed (144/144 runs); Chronicle's v1.1 system entry is deferred because the current production engine naturally delivers 70–76K-word books, below the cohort's 80K eligibility floor. The same-model architecture comparison on this page remains the published controlled result.
| # | System / model | ChronicleBench Score | Eligible runs | Status |
|---|---|---|---|---|
| T1 | GPT-5.6 Sol MODEL+HARNESS | 72.5 | 6/6 | completes; decays ~15 points over the book |
| T1 | Chronicle SYSTEM | 71.0 | 2/6 | statistical tie with Sol; length defect fixed, rerun preregistered |
| 3 | Claude Sonnet 4.6 MODEL+HARNESS | 65.6 | 6/6 | completes at the lowest eligible quality |
| — | GPT-5.6 Terra | DNF | 0/4 | no ending within the standard budget |
| — | Claude Fable 5 | DNF | 0/2 | best prose in the field; refuses novel length (stops at 23–35K) |
No entrant yet combines top-tier prose, reliable 85K delivery and protected ending quality. Chronicle comes closest to the complete product — its length calibration still has to prove itself in v1.1. The v1.1 cohort has since completed; this table remains archived as the v1.0 result.
04 Frontier models on their own terms
Raw models deserve a race they can actually win — and lose informatively. Their story on this corpus:
Quality 82–86 on what it chooses to write — the highest in the field — with endings scored 79–86. But it stalls at 23–35K words, under half the commission, on every run. Brilliance that declines the distance.
The only raw entrant that completes every 85K run — at the cost of a ~15-point slide from opening to ending. Endurance with erosion.
Finishes the distance under the harness, at the lowest eligible quality — and its ending window (54.0) is the collapse Chronicle's architecture exists to prevent.
No ending inside the standard call budget on any run. The DNF is itself the finding: sustained completion is a capability, not a given.
A model that writes an 86-quality novella did not lose a tiebreak to a novel system — it entered a different event. That's why systems and models are ranked apart.
05 What happens next