ChronicleBench v1.1 — Cohort Protocol (FROZEN before generation)

Rename note (2026-09-01): the benchmark was developed and frozen under the working name SagaBench and is now published as ChronicleBench; the metric formerly labeled SagaScore is the ChronicleBench Score. Branding only — no protocol content, threshold, instrument or data changed with the rename, and the original name remains in archived artifacts and data files (sagascore fields). The old domain redirects here.

2026-08-21. One contemporaneous cohort, all models via OpenRouter, uniform harness, no reuse of v1.0 generations. v1.0 is archived, never rewritten. This document is the canonical methodology reference and is mirrored publicly at chroniclebench.vercel.app/protocol-v1.1.html; every future model addition follows §6 verbatim.

Amendment A2 (2026-08-21, mid-generation, owner-directed, operational only): budget guards lowered from $900 total / $80 per-model to $250 total / $60 per-model at 63/168 runs complete ($41.05 spent; full-cohort projection ~$120–180). Guards are abort ceilings, not methodology — no generation parameter, roster, brief or scoring change rides along. Applied by restarting the per-model runners from their manifests (in-flight partial calls discarded and re-run clean, as with any mechanical restart).

Amendment A6 (2026-08-21, late final stretch): N1 sustained runs proved costlier than every projection (more stalls → more continuation calls); at 132/144 ($317.50) the remaining need (~$60–70) exceeded both the A5 guards and the account balance (~$56). Guards moved to $390/$110 — above the account ceiling — so that any stop records as an Insufficient credits infrastructure failure (resumable under §2) rather than a guard abort. The account balance, controlled solely by the owner, is the true binding budget from here; the owner was notified before departure. If the balance exhausts short of 144, the missing runs auto-resume on any future top-up via the standing supervisor.

Amendment A5 (2026-08-21, owner-approved final stretch): guards $320→$360 total, $80→$95 per-model at 113/144 roster runs ($249). Opus 5 and Sol Pro each projected ~$78–82 — a per-model trip within the last replicates would have left a top model partial. Final guard change; a box-side supervisor relaunches any guard-aborted session under the recalibrated guards, with the abort visible in the manifest.

Amendment A4 (2026-08-21, owner-directed mid-generation): (1) Roster reduced 14→12: Moonshot Kimi K3 (0/12 complete) and Z.AI GLM 5.3 (2/12) withdrawn by owner budget decision; their partial outputs publish as diagnostics only, never on any board; either may rejoin later via the §6 full-protocol procedure. (2) At 109/168 runs ($178.90) the OpenRouter account exhausted its credits; all affected runs failed with recorded Insufficient credits errors and were re-run after an owner top-up under §2's mechanical infrastructure-failure rule (transport failure, identical parameters). (3) Budget guards recalibrated $250→$320 total, $60→$80 per-model, owner-approved with the ~$265 expected total stated — the A2 guards sat below the corrected projection and a guard abort near completion would have forced a partial cohort, which §6.2 forbids on headline boards. No generation parameters changed.

Amendment A3 (2026-08-21, BEFORE any reference-corpus scoring): the first seeded Gutenberg selection surfaced candidates violating the preregistered eligibility rules in ways the automated filter could not see from gutendex metadata (publication language ≠ original language; collected-works volumes; multi-volume fragments; authorless anthologies). Enforcement was mechanized — translator-field + non-anglophone-author exclusion, title patterns for volumes/collections/abridgements, empty-author exclusion — and the selection re-run under the unchanged seed; superseded manifests are preserved. Additionally, one narrow content exclusion: works whose primary subject is the promotion of racial violence are excluded from the public reference corpus, recorded individually (applied once: Dixon's *The Clansman*). All exclusions are title-level and pre-scoring; no book was ever removed after being scored.

Amendment A1 (2026-08-21, BEFORE any generation): roster extended to 14 (GPT-5.6 Sol added alongside Sol Pro — both OpenAI configurations were requested); free-run arm upgraded to full 2 briefs × 3 replicates (parity with the sustained arm); one infrastructure-only smoke call per model precedes the scored cohort and never enters the benchmark; budget guards raised for the larger design ($900 total / $80 per model); §7 adds the preregistered Human Reference Corpus expansion. Nothing had been generated when A1 was committed.

1. Roster (pinned slugs + pinned providers, verified against the live OpenRouter catalogue 2026-08-21)

LabEntryPinned slugPinned providerNote
OpenAIGPT-5.6 Solopenai/gpt-5.6-solopenaiv1.0 leader, rerun cleanly
OpenAIGPT-5.6 Sol Proopenai/gpt-5.6-sol-proopenaibest available OpenAI configuration
AnthropicClaude Opus 5anthropic/claude-opus-5anthropicflagship
AnthropicClaude Fable 5anthropic/claude-fable-5anthropiccreative-writing specialist; no temperature support
AnthropicClaude Sonnet 4.6anthropic/claude-sonnet-4.6anthropicrequired architecture control (Chronicle's underlying model)
GoogleGemini 3.1 Progoogle/gemini-3.1-pro-previewgoogle-ai-studio
xAIGrok 4.6x-ai/grok-4.6xai
DeepSeekDeepSeek V4 Pro 0813deepseek/deepseek-v4-pro-0813alibabadocumented substitution (smoke phase): first-party deepseek endpoint is blocked by the account's OpenRouter data-policy settings; same model slug, highest-capacity permitted host
AlibabaQwen 3.8 Maxqwen/qwen3.8-maxalibabasole provider
MetaLlama 4 Maverickmeta-llama/llama-4-maverickdeepinfrano first-party hosting exists; documented third-party host
MistralMistral Medium 3.5mistralai/mistral-medium-3-5mistralsole provider
MoonshotKimi K3moonshotai/kimi-k3deepinfrano first-party hosting on OpenRouter; documented host
Z.AIGLM 5.3z-ai/glm-5.3z-aisole provider
MiniMaxMiniMax M3minimax/minimax-m3deepinfrano first-party hosting on OpenRouter; documented host

Every recorded call captures: returned model, returned provider, request id, timestamp, request parameters and cost. Dropped from the headline roster: GPT-5.6 Terra (another OpenAI SKU, not another lab; its v1.0 result stays in the historical dashboard). Moving latest/alias slugs are prohibited.

2. Generation protocol (identical for every model, byte-identical harness to wave-1)

3. Eligibility & scoring (v1.1 contract)

3a. The scoring instrument — how a book is actually judged (OBR)

Every score on this site — every model window, every human reference novel, every Chronicle book — comes from one instrument: OBR (Objective Book Review), Chronicle's book evaluator, pinned at system-prompt hash obr-v2.1-2026-03-08. Judge model: Claude Sonnet 4.6, temperature 0, three independent runs per text, median-aggregated. The instrument was frozen before the cohort generated and is never modified mid-cohort (§6).

Three layers per evaluation:

The eight dimensions and their default_v2 baseline weights. Each text is scored under the profile mapped from its genre — the B2 sci-fi brief under scifi_v2, the N1 family drama under literary_v2, reference novels under their stratum topic’s profile — all re-weighting the same eight dimensions; the profile used is recorded in every score file (weight_profile):

DimensionWeightWhat a 10 means
Momentum0.15sustained curiosity, no perceptible drag, no mid-book sag
Voice & craft0.15indistinguishable from a confident human author, first page to last
Emotional arc0.15genuine emotion shift across the book; stakes that matter; payoff lands
Psychological specificity0.13inner lives that could not belong to anyone else
Characters0.12identifiable by dialogue alone; clear desires, conflict, change
Coherence0.12zero logic breaks, timeline errors or dropped threads
Causal inevitability0.10every major turn surprising yet inevitable in retrospect; no deus ex machina
Promise delivery0.08the book delivers exactly what its premise and genre promised

Composite: weighted mean of the eight dimensions × 10 → a 0–100 score, with deterministic hard caps applied afterwards (for example, a manuscript-level canonical contradiction of a central identity caps the composite regardless of how well the prose scores — caps and reasons are recorded in every score file). ChronicleBench Score = the mean of the five positional-window composites (opening / 25K / 50K / 75K / ending; 10,000-word windows cut on paragraph boundaries, ending window tail-kept), each window itself a median of three judge runs. A whole-book composite is additionally recorded as a diagnostic.

Integrity measures and disclosed limitations:

4. What the public page may claim, by result

Three separate boards, never blended: (1) complete-novel SYSTEM entrants; (2) sustained model+standard-harness leaderboard; (3) short-horizon voluntary-generation leaderboard. The controlled Chronicle-vs-same-Sonnet exhibit stands independent of all three.

5. Archival

v1.0 (contract v1.0.1) remains published as the prior benchmark version with its own data file; nothing is rescored retroactively into it.

6. Standing procedure for adding ANY future model

7. Human Reference Corpus expansion (preregistered with A1)

Goal: grow the public-domain human reference from 14 to 50 novels; the existing 14 stay as legacy calibration anchors, never replaced or rescored away.

Clarification (2026-08-31, post-scoring): the published v1.1 five-window human band is computed over the 36 newly selected books only. The 14 legacy calibration novels were scored under the v1.0 contract, whose ChronicleBench Score aggregates four windows rather than v1.1's five, so their scores are a different metric object; they remain published as v1.0 calibration anchors rather than being mixed into the v1.1 band — the "legacy calibration anchors, never replaced or rescored away" treatment preregistered above.

8. Claims discipline (binding)

"#1 full-novel system" only after real system competitors are tested. "Best-performing complete-novel generator tested" only if Chronicle wins under the frozen contract and uncertainty treatment. Ties reported as ties. No superiority-to-classics claims from small or selected subsets — percentile language against the full 50-book reference. Chronicle losses publish exactly like any other loss. The live site (/, /pilot.html, /chronicle.html) is not restructured until the cohort completes and results are reviewed; v1.1 outputs build separately for review.