# Rubric v1: rendering quality in a multi-agent script pipeline

> **Public version.** Part of the evaluation artifact. See [`00-findings.md`](00-findings.md) for the results and [`README.md`](README.md) for what is published and what is held back.
>
> **Scope:** Act 1 sections. Not hooks, not later acts, not conclusions. Section type demonstrably changes what "adequate development" means, so the rubric does not claim to travel beyond the type it was calibrated on.

---

## 1. What is being evaluated

Three frontier models, spanning three different labs, each render the same **Act 1 section** from a shared outline produced in an earlier stage of the pipeline. The outline carries two relevant columns: a beat-by-beat column with **example content**, under an explicit instruction not to copy it word for word, and a structural column.

**The writers are not generating substance. They are rendering supplied substance in their own words.** The rubric measures rendering quality, not invention.

### What the evaluation may claim

> Given a fixed outline, which model produces the densest, best-developed Act 1 prose while following the paraphrase instruction.

### What it may not claim

> Which model "writes better scripts." The outline is held constant, so the comparison is of marginal variation on a fixed skeleton. That is a genuine strength, because most model comparisons are confounded by models choosing different structures, but it caps the claim and the cap has to be stated.

---

## 2. What "good" means

Written before any output was scored.

An Act 1 section is good when it **covers every beat it was assigned, develops each one rather than merely asserting it, adds nothing the outline did not ask for, and does so in its own words rather than the outline's.**

Act 1 is the setup act. It carries more obligation to develop than a fifteen-second hook, and less than a body section. It is expected to move quickly. It is not expected to be a list.

---

## 3. Gate

Checked before quality scoring. **Never contributes to the score.**

### G1: Section length

The section meets the length required for video pacing.

- **Grader:** code (word count).
- **Result:** pass or fail. A fail is reported separately and the section is still scored, so a length miss never masquerades as a quality result.

> **Why this is a gate and not a criterion.** Word count is the proxy that was already being gamed: writers hit the target by restating the same point in different words. Scoring length as quality would rebuild the exact incentive this rubric exists to remove, and would bake verbosity bias into the instrument. Length is a real production constraint, so it is checked. It is never rewarded.

---

## 4. Quality criteria

### C1: Outline coverage

**List the beats assigned to this section in the outline. For each, mark whether it appears in the section as a distinct claim, regardless of wording. Fail if fewer than N of the assigned beats appear.**

- **Grader:** model, human-validated.
- **Scale:** pass or fail, plus the raw count for analysis.
- **`N` is unset.** It is calibrated on the pilot sample before any scoring run. See §7.

> **Granularity comes from the outline, but the outline has to be read at the right level.** The beat list is nested: headed groups containing bullets. **The unit is the individual bullet, never the heading.** One act in the pilot episode has 4 headed groups containing 15 bullets, and C1 reads completely differently at 4 than at 15.
>
> *Corrected after the pilot.* An earlier draft of this rubric claimed the outline settled granularity by itself, so graders could not disagree. That was wrong. Nested structure leaves the unit ambiguous unless the rubric names the level.

**Declared limitation:** because the outline supplies the beats, C1 measures *rendering fidelity*, not the model's capacity to generate substance. It would not transfer to a free-form writing task.

---

### C2: Development

*Revised after the pilot. The original version failed the best section in the sample. See "Why C2 was rewritten" below.*

**Assess development at the level of the beat GROUP, not the individual bullet.** A group is the headed block in the outline.

**A group is DEVELOPED if the section supplies at least two distinct elaborating moves across that group's material, where a move is: a mechanism, a consequence, an example, a contrast, or a qualification. Moves count whether they appear in prose or in enumerated form. Fail if more than one third of the section's groups are undeveloped.**

- **Grader:** model.
- **Scale:** pass or fail, plus the per-group move counts and what each move was.
- **Primary quality signal.** With substance supplied, depth of treatment is the model's actual job.

> **Restating is not elaborating.** Repeating a beat in new words supplies none of the five moves, so padding cannot pass, and length gains a section nothing.

> **Form is not depth.** A numbered list item that states a mechanism *is* a mechanism. The criterion counts what the text supplies, never how it is laid out.

**Declared limitations:**

- Beatable by **decorative elaboration**: text that technically adds a consequence or example while carrying no real content.
- The five moves are a **closed list chosen by one author**. A section could develop a group in a way none of them capture.
- **Both thresholds (two moves per group, one third of groups) are judgements, not derived values.** Neither has been empirically tuned.
- Grading at group level means **a section can pass C2 while leaving individual bullets undeveloped**, provided the group as a whole is treated with depth. That is deliberate, but it does make C2 coarser than C1.

#### Why C2 was rewritten

The original criterion counted **elaborating sentences per bullet** and failed any section where more than a third of bullets had fewer than two.

In the pilot, one section rendered a four-item group as four one-line numbered items, mirroring the outline's own numbered layout. Under the original C2 that is four undeveloped beats and an automatic fail. **But a numbered set of tactics is good script writing, not padding.**

The result was that the original C2 **failed the section with perfect outline coverage, no scope creep and full paraphrase compliance, while passing the two sections that copied the example column near-verbatim and overran into the next act's material.** The rubric was ranking prose form, not quality.

**The fix uses the outline's own structure.** The beat list is nested, so each level is read for what it is good at: **bullets give C1 an unambiguous unit for coverage, groups give C2 the right unit for depth.** A group rendered as a list is developed if its items supply the moves. A group rendered as prose is judged the same way. Layout stops mattering, which is the point.

> **Consequence still to be checked:** the three fixtures in [`04-test-cases.md`](04-test-cases.md) were built against the original C2. **T3 (breadth-padded) has to be re-verified against this revision.** It should still fail, because twelve one-line assertions with no mechanism, consequence, example, contrast or qualification leave every group undeveloped. That expectation is a prediction, not a result, until it is actually run.

---

### C3: Factual support

**Removed from the text rubric.** Checked at the research stage against the production research document, not against the script.

> **Why.** Act 1 sections make hard factual claims (dates, percentages, named people) and carry no citations, because scripts do not carry citations. Applied to script text, C3 fails every section from every model, which makes it useless rather than strict. The claim's correctness is a property of the research that fed the outline, and that is where it has to be checked.

---

### C4: Additions

**Identify any material in the section that does not correspond to a beat in the outline. Classify each as either a transition (connective material serving assigned beats) or scope creep (new claims, teases, or arguments the outline did not ask for). Fail if the section contains any scope creep.**

- **Grader:** model finds candidates, human classifies.
- **Scale:** pass or fail, plus a list of additions with their classification.

> **Why this exists.** C1 counts what should be there. C4 catches what should not. Without it, a model that adds three unrequested claims looks identical to one that followed the brief, and in the labelled sample one model's extra material was mostly appended closing teases.

**Declared limitation:** the transition versus scope-creep line is a judgement call and is expected to be the lowest-agreement criterion in the set. It is kept because the distinction matters to production, and its disagreement rate is reported rather than hidden.

---

### C5: Paraphrase compliance

**Compute the longest common contiguous word sequence and the proportion of overlapping 8-grams between the section and the outline's example column for the corresponding beats. Fail if any single verbatim run exceeds 12 words, or if more than 10% of the section's 8-grams appear in the example column.**

- **Grader:** code. No judgement involved.
- **Scale:** pass or fail, plus both raw measures.

> **Why this exists.** The pipeline already instructs writers not to copy the example column word for word. **Nobody was measuring compliance.** It is the only criterion in the set that is fully mechanical, and therefore free to run on every sample.

**Hypothesis, stated in advance and refuted by the run.** The prediction was: two of the three writers hew close to the example column while the third paraphrases further.

**Result: wrong.** The model identified as the paraphraser fails *more* sections than the other two, 8 of 9 against 7 of 9. **22 of 27 sections fail C5**, and copying tracks the *section* rather than the model. Full results in [`05-c5-results.md`](05-c5-results.md).

> **Pre-registering the hypothesis was deliberate, and its refutation is the more valuable outcome.** A prediction that survives is evidence. A prediction that is published and then reported as wrong is proof the prediction preceded the data rather than being fitted to it.

**Confound found by the run, which may invalidate C5 as a compliance check:** the example column is finished, on-voice, shippable prose assembled from prior drafts, not a sketch. Instructing a model not to copy a passage that is already exactly what is wanted is close to an impossible brief. See [`05-c5-results.md`](05-c5-results.md) §4.

**Declared limitations:** both thresholds (12 words, 10%) are author-chosen and not derived. N-gram overlap detects copying, not close paraphrase that reuses structure while changing every word.

---

### C6: On brief

**The section addresses the topic and audience the outline specifies, and does not drift into a different framing. Fail if the section's framing could not be predicted from the outline.**

- **Grader:** model.
- **Scale:** pass or fail.

> Expected to pass almost always, since the outline is prescriptive. It is retained as a **guard rather than a discriminator**: a criterion that catches a rare catastrophic failure earns its place even when it almost never fires.

---

## 5. Grader allocation

| Criterion | Grader | Cost | Runs on |
| --- | --- | --- | --- |
| G1 length | Code | free | every sample |
| C5 paraphrase | Code | free | every sample |
| C1 coverage | Model, human-validated | per call | every sample |
| C2 development | Model | per call | every sample |
| C4 additions | Model finds, human classifies | per call plus human time | every sample |
| C6 on brief | Model | per call | every sample |

**Principle: the cheapest grader that can do the job.** Code is free and deterministic, so it runs on everything. Model graders cost money and vary between calls, so they are spent where judgement is genuinely required. Human time is scarcest and is spent validating the judge and reading disagreements, never on bulk grading.

---

## 6. Run protocol

**One episode. All sections. Three models. 27 real outputs, at zero additional generation cost.**

Generating new episodes costs real API spend, so the run uses data that already existed. The episode already had nine sections (a hook, five acts, a checkpoint, a story beat and a call to action), each written by all three models. That is a genuine dataset that was sitting unused.

### Two tiers, because the criteria have different requirements

**Tier 1: C5 paraphrase compliance, across all 27 sections.**
C5 compares each section to **its own** example column, so it is **section-type agnostic**: it works on the hook and the call to action exactly as well as on Act 2. It is a code check, so it needs no judge, no blind protocol, no human labeller and no agreement statistics.

This is the tier that could be finished and published quickly, and it carries the pre-registered hypothesis. **It is the tier reported in this artifact.**

**Tier 2: C1, C2, C4 and C6, on the Act sections only.**
These need judgement, so they need the judge and a second human labeller. Restricted to sections of a single type at a time, since C2's thresholds are type-specific. **Tier 2 has not been run and nothing about it is reported here.**

### What this costs the claim, stated plainly

- **One episode means one topic, one research base, one outline author.** Anything found here is directional and may not generalise. It cannot be reported as a rate.
- **Sections within one episode are not independent observations.** They share a topic, a research document and an outline style, so 27 sections is not 27 independent trials.
- **No repeat runs**, so sampling noise is not separated from model difference. A single output per model per section is partly luck.

> These are real limitations and they are the price of not spending weeks and an API budget. **A narrow finding published honestly beats a broad one that never ships**, but the narrowness has to be on the page, not left for a reader to work out.

**Reported per model:** result per criterion per section, with raw counts. Never a single composite score.

---

## 7. Calibration, before any scoring run

1. **Pilot-label 6 sections** against C1, C2, C4 and C6, blind.
2. **Set `N` in C1** at the threshold that separates sections judged adequate from sections judged thin. If no such threshold exists, C1 is not discriminative and has to be reworked rather than tuned.
3. **Record every ambiguity** encountered while pilot-labelling. Each one is a criterion that is not yet checkable enough for a second person.
4. **Only then** freeze the rubric and begin the run.

> The pilot exists so that the second human labeller is not spent on a rubric that still has holes in it.

---

## 8. Judge validation

Full procedure and the judge prompt: [`02-judge-protocol.md`](02-judge-protocol.md).

What has to be true before any judge output is reported:

- Labels written **blind**, before any judge output is seen.
- Judge drawn from **outside the three candidate families**, or the family-preference effect measured explicitly.
- Answer order **randomised**.
- Agreement reported as **percent agreement and Cohen's kappa**, not correlation, because the criteria are binary.
- Agreement compared against **human to human agreement**, not against 100%.
- The judge has to **separate the matched-length dense and padded fixtures** in [`04-test-cases.md`](04-test-cases.md). Failing that, it is not usable regardless of its aggregate score.
- If the judge prompt was iterated, the reported number comes from a set it was not tuned on, or the iteration is stated plainly.

---

## 9. Known limitations

Published in full. Volunteered limitations read as rigour. The same limitations found by a reader read as sloppiness.

**Design**

- All criteria authored by one person.
- All thresholds (`N`, one third, 12 words, 10%) are author-chosen judgements, not derived values.
- C4's transition versus scope-creep boundary is expected to be the weakest-agreement criterion in the set.
- C2 is beatable by decorative elaboration. C5 is beatable by close structural paraphrase.

**Scope**

- Act 1 sections only for the judged criteria. The rubric is not calibrated for hooks, body sections or conclusions, and section type demonstrably affects what "adequate development" means.
- Because the outline supplies the beats, results measure rendering fidelity and would not transfer to free-form writing.
- One content domain, one production pipeline, one language.

**Scale**

- One episode, nine sections, three models. Small enough that a single unusual section still carries weight.
- The three writer models are not fully independent in their output: two produced near-verbatim passages in the pilot sample, so they should not be treated as three separate observations of "how frontier models write."

**Measurement**

- The outline stage itself is ungraded, so outline quality is an uncontrolled confound beneath every writer comparison.
- Trajectory (retries, latency, cost per stage) was not logged for this run, so the comparison is on output quality alone.
- If no second human labeller is available for Tier 2, no inter-human baseline exists and judge agreement cannot be calibrated against the task's own ceiling. **That limitation has to be stated explicitly rather than omitted.**

**Prior-work context**

- The pipeline's original quality checker shared a model family with one of the writers, and applied no order randomisation. Comparative results it produced before this evaluation carry an unquantifiable bias. Both defects are corrected in this design.
