# Adversarial fixtures

> **Public version.** Three purpose-built test sections, published in full so a reader can see the criteria operating rather than taking the author's word for it.
>
> **These are test fixtures. They are not pipeline output and they are not findings.** They were written as test material, on a neutral topic, tied to no channel and no client.

---

## Why these exist

The real generated scripts stay private. That leaves a reader unable to see the criteria actually working, and "trust me, I applied it" is the weakest sentence in any evaluation. Fixtures solve that: they can be published in full.

There is a second reason, independent of privacy. **Real data gives you the distribution. Constructed cases probe specific failure modes on demand.** Each case below is engineered to trip exactly one criterion, which is how you find out whether a criterion works at all.

## Note on factual accuracy

Content is **illustrative and not fact-checked**. This is deliberate and does not weaken the fixtures, because **C1 and C2 measure structure, not truth**: claim distinctness and depth of treatment. Factual accuracy is C3's job and is checked separately against the production research document. Do not cite these as history.

## Design

Three sections on one brief, **matched to roughly 300 words each**, so that length cannot explain any score difference. Length matching is the point. It is what stops a verbosity-biased judge from passing the test by accident.

| Case | Engineered flaw | Expected C1 | Expected C2 |
| --- | --- | --- | --- |
| **T1 Dense** | none, this is the control | Pass | Pass |
| **T2 Repetition-padded** | restates the same points in new words | **Fail** | Pass, on the few real claims |
| **T3 Breadth-padded** | many claims, none developed | Pass | **Fail** |

T2 and T3 are the two padding modes, separated. A judge that catches one and not the other is only half working, and an aggregate agreement score would hide that.

**Brief used for all three:** *Act 1 opening, roughly 300 words, on how the shopping cart was adopted.*

> **Status note.** These fixtures were written against the original C2, which counted elaborating sentences per bullet. C2 has since been rewritten to grade at beat-group level. **T3 is expected to still fail C2 under the revision**, because twelve one-line assertions supply no mechanism, consequence, example, contrast or qualification, leaving every group undeveloped. That is a prediction, not a result, until it is run.

---

## T1: Dense (control, should pass both)

### ACT 1: Nobody Wanted It (~300 words)

Picture a grocery store in 1937. You are carrying a wire basket that cuts into your fingers, and it is half full, because half full is all you can lift. When it gets heavy, you stop shopping. You go and pay.

That last sentence was the whole problem, and one store owner was the only person who noticed it. Sylvan Goldman ran a chain in Oklahoma City, and he worked out that his customers were not leaving because they had everything they needed. They were leaving because their arms hurt. Every basket that got heavy was a sale that never happened.

So he built a cart: two wire baskets on a folding frame with wheels, borrowed from the folding chair in his office. By any reasonable measure it was a good idea. It solved a real problem that was costing him real money.

Nobody used it.

Men refused, because pushing a cart looked weak and they were not going to be seen doing it. Women refused too, because after years of pushing prams they had no appetite for pushing another one around a shop. Goldman had solved the engineering problem and walked straight into a social one, and the social one was harder.

His fix was to lie. He hired actors, men and women of different ages, and paid them to push carts around his stores looking perfectly content. He put a greeter at the door to hand a cart to every arriving customer, so refusing meant refusing a person rather than an object. Within a year the carts were never idle.

The shopping cart did not win because it was better. It won because Goldman manufactured the appearance that everyone else had already accepted it.

---

## T2: Repetition-padded (should fail C1)

### ACT 1: Nobody Wanted It (~300 words)

Picture a grocery store in 1937. You are carrying a wire basket, and it is heavy. It is genuinely heavy, the kind of heavy that you feel. Baskets in those days were heavy things to carry, and carrying them was hard work.

When the basket got heavy, shoppers stopped shopping. They would stop. The weight made them stop. Once a basket became too heavy to carry comfortably, the shopping trip was effectively over, because nobody wants to carry something uncomfortable. So they went and paid, and that was the end of it.

One store owner noticed this. Sylvan Goldman ran a chain in Oklahoma City, and he noticed what was happening in his stores. He saw the problem clearly. He understood it. He recognised that something was going wrong, and he was the one who spotted it when others had not.

So he built a cart. It was a cart, with wheels, designed for carrying shopping. The idea was that the cart would carry the things the basket had been carrying, so the customer would not have to.

It was a good idea. It was, in fact, a very good idea, and a sensible one, and it addressed something real. Good ideas like this one deserve to succeed.

Nobody used it. People did not use it. Adoption was essentially zero, and the cart sat there unused, because customers were not using it.

Goldman had a problem on his hands. The problem was that his solution was not being taken up, which is a problem when you have built something and nobody takes it up. He had built the thing. The thing was not being used. That was the situation he found himself in as 1937 went on.

---

## T3: Breadth-padded (should pass C1, fail C2)

### ACT 1: Nobody Wanted It (~300 words)

Picture a grocery store in 1937. Shoppers carried wire baskets. The baskets were small. They cut into your fingers. Most people could manage about fifteen pounds. Heavy baskets ended shopping trips early. Lost sales followed.

Sylvan Goldman ran a chain in Oklahoma City. He noticed the pattern. He had a background in retail. His stores were doing reasonably well. He wanted them to do better.

He built a cart. It used a folding frame. The design came from a folding chair. It held two wire baskets. It had four wheels. It could be stacked for storage. It cost very little to make. He patented it.

Nobody used it. Men thought it looked weak. Women associated it with prams. Older shoppers found it unfamiliar. Staff were unsure how to explain it. The carts sat by the entrance. Weeks passed.

Goldman hired actors. They pushed carts around the stores. They looked relaxed. They were different ages. Some were men. Some were women. He also hired a greeter. The greeter stood at the door. The greeter handed out carts.

It worked. Adoption climbed. Within a year the carts were in constant use. Stores had to buy more. Other chains copied the design. Goldman licensed it. He became wealthy. The cart got bigger over the following decades. Average basket size grew with it. Supermarkets redesigned aisles around cart width. The layout of retail changed permanently.

The lesson is about social proof. People follow other people. Objects need permission before they get adopted. Engineering was never the hard part. The hard part was persuasion. Goldman understood this before most people did. That insight was worth more than the patent.

---

## How to use these

1. **Hand-label all three against C1 and C2 first**, blind if possible, before showing them to any judge.
2. **Run each judge on all three.** A judge that passes T2 or T3 is broken for this purpose regardless of its aggregate agreement number on real sections.
3. **Report the results in full**, so a reader can see the criteria operating.
4. **Report per case, never averaged.** A judge that catches repetition padding but not breadth padding has a specific, nameable weakness, and averaging hides it.

## Known weaknesses of these fixtures

- **Written by one author in one sitting**, so they may share stylistic tells that a judge picks up on instead of the criteria. A judge could in principle learn "the padded ones sound flat" rather than applying C1 and C2.
- **Deliberately extreme.** Real padding is subtler than T2. These test whether a criterion works at all, not where its threshold sits. Threshold calibration needs real sections.
- **One section type only.** Section type appears to matter, so equivalent fixtures are needed for body sections and conclusions before the rubric can be applied across a whole script.
