Culture Source-backed 4-minute read

What a score does to the thing it scores

We run an AI gate that scores everything we publish. The nearest research is about rewards and interest, and reading it across to scores is my own leap.

The gate we run

Curio is a swipe feed of short knowledge cards. Every card is drafted by a model and then judged by one. So is every post on this blog, and so is the narration in our videos — three separate judges, one reading a card, one reading a post, one listening to a take.

The judge scores a draft out of 10. An 8 sends it to a human. To publish with nobody reading it first, a card needs a 9.2 and zero flagged issues. We set that bar high because nobody does read it first.

Most drafts miss on the first try. That is what a high bar is for. But a bar changes how you work toward it, and the closest research I can find is not about scores at all. It is about rewards.

Paying children to draw

In 1973 Lepper, Greene and Nisbett gave preschoolers felt-tip pens. Some were promised a certificate for drawing. Some got one as a surprise. Some got nothing.

Later the researchers watched how much the children chose to draw during free play. The ones who had expected a reward drew less than the others.

Their reading was that the reward gave the child a second reason to draw, and the second reason crowded out the first. That account is the writers' interpretation of the pattern, not something the drawing itself proved.

How big, and which way

Deci, Koestner and Ryan pooled 128 studies in 1999. They looked at two things: how much people kept doing the task once they no longer had to, and how interested people said they were.

Rewards for doing a task, for finishing it, or for doing it well all reduced the first. The rewarded groups came out −0.40, −0.36 and −0.28 below the comparison groups, measured in standard deviations, a standard deviation being the typical spread of scores in a group. Small to moderate gaps, not collapses.

One result runs the other way, and it matters here. Groups given positive feedback came out 0.33 above their comparison groups on that same measure of sticking with the task, and 0.31 above on how interested they said they were. So this is not a finding that any score sours the work. In these experiments, a tangible reward and a piece of praise did not do the same thing.

Campbell's argument in 1979 was narrower than the slogan it turned into. He was writing about indicators used to make social decisions: the more weight one carries, the more pressure builds to corrupt it, and the more it distorts the very process it was put there to watch.

What we actually do about our own bar

Here the evidence stops and my reading starts. None of this work involved a model marking a draft out of 10. Lepper handed out certificates. The studies Deci and colleagues pooled used many kinds of reward and feedback. Campbell was not running an experiment at all. The line from any of it to our gate is one I am drawing.

We re-record narration at about $0.39 a take, three at a time, until the middle score clears the bar. The same locked script has scored a 9 on one run and a 4 on the next. We noticed takes ending on a question score higher, so more of our scripts end on questions now.

None of that makes the work better in a way I could defend. It raises the odds of passing.

The shape is an old one. A standard sits where nobody reliably reaches it, and next to it grows a set of small things anyone can do: record the take again, end the script on a question.

Those are cheap and repeatable, and they are what you reach for when the standard itself is outside your control. Writing a resume out of the job posting's own words is the same move.

When the practice is superstition

Our judge is noisy. We measured it scoring one identical audio file 4, then 7, then 2. An internal audit made us write the word superstition into our own rule about chasing scores above 8, and the caveat is still sitting in it.

That is the real risk of a noisy verdict. Whatever you happened to do before a pass starts to look like the reason for it.

I am not against the gate. It is the only reason this blog can publish with nobody watching, and it stays strict.

What I want from any score, ours included, is three things. I want to know that a threshold exists and where it sits. I want to know what the judge is actually reading: ours works from a rubric covering the hook, the writing, whether a claim matches its source, and whether the piece says anything. And I want to know how much the verdict moves when the work does not move at all.

I have never once been told all three about a score someone else was applying to me.

Sources

  1. Lepper, Greene & Nisbett, Undermining children's intrinsic interest with extrinsic reward (Journal of Personality and Social Psychology, 1973)
  2. Deci, Koestner & Ryan, A meta-analytic review of experiments examining the effects of extrinsic rewards on intrinsic motivation (Psychological Bulletin, 1999)
  3. Campbell, Assessing the impact of planned social change (Evaluation and Program Planning, 1979)