Rename note (2026-09-01): the benchmark was developed and frozen under the working name SagaBench and is now published as ChronicleBench; the metric formerly labeled SagaScore is the ChronicleBench Score. Branding only — no protocol content, threshold, instrument or data changed with the rename, and the original name remains in archived artifacts and data files (sagascore fields). The old domain redirects here.
2026-08-21. One contemporaneous cohort, all models via OpenRouter, uniform harness, no reuse of v1.0 generations. v1.0 is archived, never rewritten. This document is the canonical methodology reference and is mirrored publicly at chroniclebench.vercel.app/protocol-v1.1.html; every future model addition follows §6 verbatim.
Amendment A2 (2026-08-21, mid-generation, owner-directed, operational only): budget guards lowered from $900 total / $80 per-model to $250 total / $60 per-model at 63/168 runs complete ($41.05 spent; full-cohort projection ~$120–180). Guards are abort ceilings, not methodology — no generation parameter, roster, brief or scoring change rides along. Applied by restarting the per-model runners from their manifests (in-flight partial calls discarded and re-run clean, as with any mechanical restart).
Amendment A6 (2026-08-21, late final stretch): N1 sustained runs proved costlier than every projection (more stalls → more continuation calls); at 132/144 ($317.50) the remaining need (~$60–70) exceeded both the A5 guards and the account balance (~$56). Guards moved to $390/$110 — above the account ceiling — so that any stop records as an Insufficient credits infrastructure failure (resumable under §2) rather than a guard abort. The account balance, controlled solely by the owner, is the true binding budget from here; the owner was notified before departure. If the balance exhausts short of 144, the missing runs auto-resume on any future top-up via the standing supervisor.
Amendment A5 (2026-08-21, owner-approved final stretch): guards $320→$360 total, $80→$95 per-model at 113/144 roster runs ($249). Opus 5 and Sol Pro each projected ~$78–82 — a per-model trip within the last replicates would have left a top model partial. Final guard change; a box-side supervisor relaunches any guard-aborted session under the recalibrated guards, with the abort visible in the manifest.
Amendment A4 (2026-08-21, owner-directed mid-generation): (1) Roster reduced 14→12: Moonshot Kimi K3 (0/12 complete) and Z.AI GLM 5.3 (2/12) withdrawn by owner budget decision; their partial outputs publish as diagnostics only, never on any board; either may rejoin later via the §6 full-protocol procedure. (2) At 109/168 runs ($178.90) the OpenRouter account exhausted its credits; all affected runs failed with recorded Insufficient credits errors and were re-run after an owner top-up under §2's mechanical infrastructure-failure rule (transport failure, identical parameters). (3) Budget guards recalibrated $250→$320 total, $60→$80 per-model, owner-approved with the ~$265 expected total stated — the A2 guards sat below the corrected projection and a guard abort near completion would have forced a partial cohort, which §6.2 forbids on headline boards. No generation parameters changed.
Amendment A3 (2026-08-21, BEFORE any reference-corpus scoring): the first seeded Gutenberg selection surfaced candidates violating the preregistered eligibility rules in ways the automated filter could not see from gutendex metadata (publication language ≠ original language; collected-works volumes; multi-volume fragments; authorless anthologies). Enforcement was mechanized — translator-field + non-anglophone-author exclusion, title patterns for volumes/collections/abridgements, empty-author exclusion — and the selection re-run under the unchanged seed; superseded manifests are preserved. Additionally, one narrow content exclusion: works whose primary subject is the promotion of racial violence are excluded from the public reference corpus, recorded individually (applied once: Dixon's *The Clansman*). All exclusions are title-level and pre-scoring; no book was ever removed after being scored.
Amendment A1 (2026-08-21, BEFORE any generation): roster extended to 14 (GPT-5.6 Sol added alongside Sol Pro — both OpenAI configurations were requested); free-run arm upgraded to full 2 briefs × 3 replicates (parity with the sustained arm); one infrastructure-only smoke call per model precedes the scored cohort and never enters the benchmark; budget guards raised for the larger design ($900 total / $80 per model); §7 adds the preregistered Human Reference Corpus expansion. Nothing had been generated when A1 was committed.
| Lab | Entry | Pinned slug | Pinned provider | Note |
|---|---|---|---|---|
| OpenAI | GPT-5.6 Sol | openai/gpt-5.6-sol | openai | v1.0 leader, rerun cleanly |
| OpenAI | GPT-5.6 Sol Pro | openai/gpt-5.6-sol-pro | openai | best available OpenAI configuration |
| Anthropic | Claude Opus 5 | anthropic/claude-opus-5 | anthropic | flagship |
| Anthropic | Claude Fable 5 | anthropic/claude-fable-5 | anthropic | creative-writing specialist; no temperature support |
| Anthropic | Claude Sonnet 4.6 | anthropic/claude-sonnet-4.6 | anthropic | required architecture control (Chronicle's underlying model) |
| Gemini 3.1 Pro | google/gemini-3.1-pro-preview | google-ai-studio | ||
| xAI | Grok 4.6 | x-ai/grok-4.6 | xai | |
| DeepSeek | DeepSeek V4 Pro 0813 | deepseek/deepseek-v4-pro-0813 | alibaba | documented substitution (smoke phase): first-party deepseek endpoint is blocked by the account's OpenRouter data-policy settings; same model slug, highest-capacity permitted host |
| Alibaba | Qwen 3.8 Max | qwen/qwen3.8-max | alibaba | sole provider |
| Meta | Llama 4 Maverick | meta-llama/llama-4-maverick | deepinfra | no first-party hosting exists; documented third-party host |
| Mistral | Mistral Medium 3.5 | mistralai/mistral-medium-3-5 | mistral | sole provider |
| Moonshot | Kimi K3 | moonshotai/kimi-k3 | deepinfra | no first-party hosting on OpenRouter; documented host |
| Z.AI | GLM 5.3 | z-ai/glm-5.3 | z-ai | sole provider |
| MiniMax | MiniMax M3 | minimax/minimax-m3 | deepinfra | no first-party hosting on OpenRouter; documented host |
Every recorded call captures: returned model, returned provider, request id, timestamp, request parameters and cost. Dropped from the headline roster: GPT-5.6 Terra (another OpenAI SKU, not another lab; its v1.0 result stays in the historical dashboard). Moving latest/alias slugs are prohibited.
finish=length; the model's own stop is the voluntary-length datum) and three sustained replicates (neutral CONTINUE after every call, including voluntary stops, until ≥85,000 words at a stop / stall / cap). Per-call word offsets are recorded so every voluntary stop survives into analysis. 14 models × 2 briefs × 6 runs = 168 scored runs.params_transformed); max_tokens = min(32,768, provider completion cap); SSE streaming; 40-min per-call ceiling; hard word cap 94,000 (below the 95K eligibility ceiling so runaway generation cannot convert a run into an over-length DNF).provider: { only: [pinned], allow_fallbacks: false }; the returned provider and model of every call are recorded in the per-run JSONL. A pinned-provider failure is a recorded failure, not a silent reroute.failed entry, never a quiet retry. Infrastructure failures (transport death before any content) may be rerun exactly once under the mechanical rule in §2; nothing else reruns.Every score on this site — every model window, every human reference novel, every Chronicle book — comes from one instrument: OBR (Objective Book Review), Chronicle's book evaluator, pinned at system-prompt hash obr-v2.1-2026-03-08. Judge model: Claude Sonnet 4.6, temperature 0, three independent runs per text, median-aggregated. The instrument was frozen before the cohort generated and is never modified mid-cohort (§6).
Three layers per evaluation:
sentence_cv), lexical diversity (type_token_ratio), AI-tell phrase density per 1,000 words (ai_tells_per_kw), repetition and structural metrics, plus a prose fingerprint. These numbers are handed to the judge as cross-checks it must stay consistent with, and feed the release-score gates.strong_hook, tension_spike, filler_scene, pacing_sag, strong_close, …), plus 3–5 top issues from a fixed diagnosis taxonomy. Tags outside the vocabulary are dropped in validation.The eight dimensions and their default_v2 baseline weights. Each text is scored under the profile mapped from its genre — the B2 sci-fi brief under scifi_v2, the N1 family drama under literary_v2, reference novels under their stratum topic’s profile — all re-weighting the same eight dimensions; the profile used is recorded in every score file (weight_profile):
| Dimension | Weight | What a 10 means |
|---|---|---|
| Momentum | 0.15 | sustained curiosity, no perceptible drag, no mid-book sag |
| Voice & craft | 0.15 | indistinguishable from a confident human author, first page to last |
| Emotional arc | 0.15 | genuine emotion shift across the book; stakes that matter; payoff lands |
| Psychological specificity | 0.13 | inner lives that could not belong to anyone else |
| Characters | 0.12 | identifiable by dialogue alone; clear desires, conflict, change |
| Coherence | 0.12 | zero logic breaks, timeline errors or dropped threads |
| Causal inevitability | 0.10 | every major turn surprising yet inevitable in retrospect; no deus ex machina |
| Promise delivery | 0.08 | the book delivers exactly what its premise and genre promised |
Composite: weighted mean of the eight dimensions × 10 → a 0–100 score, with deterministic hard caps applied afterwards (for example, a manuscript-level canonical contradiction of a central identity caps the composite regardless of how well the prose scores — caps and reasons are recorded in every score file). ChronicleBench Score = the mean of the five positional-window composites (opening / 25K / 50K / 75K / ending; 10,000-word windows cut on paragraph boundaries, ending window tail-kept), each window itself a median of three judge runs. A whole-book composite is additionally recorded as a diagnostic.
Integrity measures and disclosed limitations:
Three separate boards, never blended: (1) complete-novel SYSTEM entrants; (2) sustained model+standard-harness leaderboard; (3) short-horizon voluntary-generation leaderboard. The controlled Chronicle-vs-same-Sonnet exhibit stands independent of all three.
v1.0 (contract v1.0.1) remains published as the prior benchmark version with its own data file; nothing is rescored retroactively into it.
/protocol-v1.1.html) is updated in the same change that adds the model.Goal: grow the public-domain human reference from 14 to 50 novels; the existing 14 stay as legacy calibration anchors, never replaced or rescored away.
sagabench-v1.1-gutenberg, SHA-256 over sorted Gutenberg IDs) → commit the title list + manuscript SHA-256 hashes before scoring.Clarification (2026-08-31, post-scoring): the published v1.1 five-window human band is computed over the 36 newly selected books only. The 14 legacy calibration novels were scored under the v1.0 contract, whose ChronicleBench Score aggregates four windows rather than v1.1's five, so their scores are a different metric object; they remain published as v1.0 calibration anchors rather than being mixed into the v1.1 band — the "legacy calibration anchors, never replaced or rescored away" treatment preregistered above.
"#1 full-novel system" only after real system competitors are tested. "Best-performing complete-novel generator tested" only if Chronicle wins under the frozen contract and uncertainty treatment. Ties reported as ties. No superiority-to-classics claims from small or selected subsets — percentile language against the full 50-book reference. Chronicle losses publish exactly like any other loss. The live site (/, /pilot.html, /chronicle.html) is not restructured until the cohort completes and results are reviewed; v1.1 outputs build separately for review.