# C5 results: paraphrase compliance

> **Public version.** Run 2026-07-31 via [`c5-paraphrase-check.mjs`](c5-paraphrase-check.mjs), published in full. Deterministic, no model calls, no API cost.
> One episode, nine sections, three models = **27 measurements**.
> **Thresholds:** longest verbatim run greater than **12 words**, or 8-gram overlap greater than **10%**, fails a section. Either condition alone is sufficient.
> **Mandated brand phrases are stripped from both texts before comparison.** Three phrases: a required signature opening, a presenter self-identification, and a branded sign-off. Reproducing those is compliance with the outline, not copying. The phrases themselves are not published.
> Writers appear as Model 1, 2 and 3. The mapping to named models is not published, because the finding is not a model ranking.

---

## Results

Each cell is *longest verbatim run in words / proportion of the section's 8-grams also present in the example column*.

| Section | Model 1 | Model 2 | Model 3 |
| --- | --- | --- | --- |
| HOOK | 5 / 0.0% PASS | 7 / 0.0% PASS | 5 / 0.0% PASS |
| ACT 1 | 49 / 38.3% FAIL | 44 / 22.8% FAIL | 49 / 29.3% FAIL |
| ACT 2 | 104 / 64.4% FAIL | **21 / 25.7%** FAIL | 104 / 61.4% FAIL |
| CHECKPOINT | 23 / 43.2% FAIL | 21 / 46.7% FAIL | 23 / 43.2% FAIL |
| ACT 3 | 54 / 36.0% FAIL | 75 / 65.1% FAIL | 38 / 26.3% FAIL |
| STORY BEAT | 40 / 86.8% FAIL | 32 / 75.8% FAIL | 25 / 62.5% FAIL |
| ACT 4 | 70 / 50.5% FAIL | 66 / 41.4% FAIL | 58 / 44.1% FAIL |
| ACT 5 | 94 / 55.3% FAIL | **210 / 100.0%** FAIL | 94 / 55.3% FAIL |
| CALL TO ACTION | 6 / 0.0% PASS | 11 / 16.7% FAIL | 6 / 0.0% PASS |
| **Fail rate** | **7 of 9 (78%)** | **8 of 9 (89%)** | **7 of 9 (78%)** |

**22 of 27 sections fail.**

---

## 1. The pre-registered hypothesis was refuted

Written into the rubric before this data was seen:

> *"Two of the three writers hew close to the outline's example column while the third paraphrases further."*

**Wrong.** The model identified as the paraphraser, on the evidence of Act 1 and Act 2, fails **more** sections than the other two: 8 of 9 against 7 of 9. The inference generalised from two sections to nine, and nine sections disagreed.

This is reported first, and prominently, on purpose. **A refuted pre-registration is worth more than a confirmed one.** It demonstrates that the prediction preceded the data rather than being fitted to it, and it is the clearest available evidence that the evaluation was not steered toward a convenient story. An evaluation whose every hypothesis is confirmed should be read with suspicion, not admiration.

---

## 2. Section-level variance dominates model-level variance

The same model copies heavily in one section and paraphrases well in another.

- **Model 2's ACT 2** has the *lowest* verbatim run in the entire table: 21 words, against 104 for the other two. The original Act 2 reading was correct, for Act 2.
- **Model 2's ACT 5** is a 210-word section with a **210-word verbatim run and 100% 8-gram overlap.** The example column, reproduced in full, word for word.

The same model produced both. **"Which model follows the paraphrase instruction" is therefore the wrong question.** The better one is *"which sections get copied"*, and the pattern points at the examples rather than at the models. That reframing is the single most useful thing this run produced, and it is a direct consequence of measuring all nine sections instead of the two that prompted the hypothesis.

---

## 3. Only the opening passes reliably

All three models pass HOOK once the mandated signature phrases are stripped: runs of 5 to 7 words, 0% overlap. Given a one-line brief and three short example variants, all three wrote genuinely original copy.

**The call to action is the threshold-sensitivity case.** Models 1 and 3 pass cleanly (run 6, 0%). Model 2 fails on **overlap alone**: its run of 11 words is under the 12-word limit, but 16.7% overlap is over the 10% limit. A single marginal fail driven by one of two thresholds is a useful reminder that both numbers are author-chosen judgements, not derived values.

---

## 4. Confound: the example column is finished prose, not a sketch

**This is the most important caveat in the report, and it is why the headline is not "the models disobeyed."**

The outline's own Synthesis Decisions block shows the example column was assembled from prior complete drafts. It is polished, on-voice, correctly paced copy. Not a guideline, but a shippable draft.

Instructing a model not to copy a passage that is *already exactly what is wanted* is close to an impossible brief. **This is plausibly a pipeline design finding rather than a model compliance failure.**

### The decision this forces

1. **Accept it.** If polished examples are what the pipeline wants, reproduction is the expected outcome, and C5 should not be framed as a violation check at all. Reframe it as a *divergence measurement*: a number that says how far each writer moved from the reference, with no pass line attached.
2. **Make the examples skeletal.** Bullet fragments and tonal notes rather than finished prose, leaving the model something to actually write. Then C5 measures compliance meaningfully.

Option 2 makes C5 a real criterion. Option 1 keeps the current pipeline and drops the pretence that the instruction is being followed. **This decision is open at the time of publication.** It is recorded here rather than resolved quietly, because which option is chosen determines whether the 22 of 27 figure means anything at all.

---

## 5. Limitations

- **One episode, one topic, one outline author.** Directional. Not a rate that generalises.
- **Sections within one episode are not independent observations.** 27 measurements, far fewer than 27 independent trials.
- **One output per model per section.** No repeat runs, so sampling noise is not separated from model behaviour.
- **Both thresholds (12 words, 10%) are author-chosen** and were not tuned. The call to action result for Model 2 shows the outcome is threshold-sensitive at the margin.
- **N-gram overlap detects copying, not close structural paraphrase** that reuses shape while changing every word. A writer who follows the example sentence by sentence with entirely different vocabulary scores as compliant.
- **The measurement code is published; the input data is not.** A reader can audit exactly how the numbers were produced but cannot reproduce them, because the scripts and the outline stay private. That is a real limit on verifiability and it is stated rather than glossed.
- **Model identities are known to the author.** Irrelevant to this particular result, since the measurement is deterministic code with no judgement in it, but stated for completeness.

---

## 6. What this result actually supports

**Defensible:**

> *"Across one episode of nine sections, three frontier models reproduced the outline's example prose beyond a 12-word verbatim threshold in 22 of 27 sections. Copying tracked the section rather than the model, and one section was reproduced in full. The example column consists of finished prose, which the report treats as the likeliest cause."*

**Not defensible:** any statement ranking the models on instruction-following, any pass rate presented as generalising beyond this episode, and any claim that the models "ignored" the instruction without noting the confound.
