Culture Source-backed 4-minute read
Our narrator sounded like a list. The script had 22 sentences.
One recording sounded like items being read out. Twenty-two short sentences became six longer ones and the sound went away — but we changed the pauses at the same time.
The part of it a writer controls
A script with 22 sentences has 21 joins in it: 21 places where one sentence ends and the next begins. Ours had 22, averaging 5.5 words each, and the recording came back sounding like a list being read out. I spent a while blaming the voice.
We rewrote it as 6 sentences of about 21 words, same content, which left 5 joins. The list sound went away. In the same pass we also stopped squeezing the silences between sentences, so this is not the clean before-and-after it looks like.
Only the counting is certain here. A script sets how many joins there are, and a writer sets the script. What the voice does at each join is a separate question, and the answer to that one is messier.
What a list sounds like, counted
I could hear the problem and not name it, which is a bad place to make decisions from. So we measured the recording.
Our tool does not measure pitch, which is what a listener hears. What it can measure is fundamental frequency: how often the sound wave repeats itself, counted in hertz — cycles per second. That rate follows how fast the vocal folds open and close, and perceived pitch mostly tracks it. In our recordings that frequency drifts downwards across a sentence and jumps back up at the start of the next one.
Using an open-source speech analysis tool, we estimate that frequency at the end of one sentence and at the start of the next, and count a join as a reset when the second is more than 12 hertz above the first. This voice started its sentences around 148 hertz, so 12 is a small step: about one part in twelve of that. We chose 12 hertz as the threshold ourselves, from our own takes.
In the original recording, 15 of the 21 joins were resets, climbing about 40 hertz on average to land back around 148 hertz — a rise of more than a third above where the voice had dropped to. In the rewrite, 5 joins were left and none of them was a reset.
The rewrite was also a touch faster to say: 176.8 words per minute against 174.2. Longer sentences did not slow it down.
What the research does and does not say
Prosody — the pitch, timing and stress of speech — is not decoration on top of words. Cutler, Dahan and van Donselaar's review describes it as information listeners genuinely use. Frazier, Carlson and Clifton argue that how speech is grouped into phrases is central to understanding it.
Both are review articles about how the sound of speech bears on understanding it. Neither is about narration, or videos, or us. Neither one shows that a short-sentence script makes a worse video, and I am not going to say they do. What they gave me was a reason to take the pattern seriously instead of calling it taste.
Two things changed, not one
The rewrite was one change. The other was that we stopped squeezing the silences between sentences. In the narration as first recorded, those pause lengths were spread out, and the usual measure of that spread — a standard deviation — was 0.135 seconds. My editing had pulled it down to 0.053, a drop of about three-fifths, meaning the pauses had become much more alike in length.
With the squeezing switched off, the pause lengths in the rewrite had a standard deviation of 0.095 — about halfway back to the 0.135 it started at.
So the recording got two fixes at once, and only one of them was the writing. I cannot separate them from a single before and after. The reset counts depend on a choice I made, too: pick a cutoff other than 12 hertz and a different number of joins qualifies. The number of joins does not depend on that cutoff. It is 21 and 5 because the script had 22 sentences and then 6, which was settled in the writing, before anything was recorded.
The pauses were my own doing in the edit, and I had been blaming the voice for them.
Where it stops being true
Later, on a different kind of video with a different narrator voice, we tried the same rewrite — fewer and longer sentences — and it failed. Across 12 takes that voice reset its frequency at nearly every join, across every wording we tried.
Merging sentences there also sped the read up, from 4.27 to 4.52 syllables a second. Our target for that kind of video is about 4.10, taken from a batch of reels we measured, so the merge moved the pace away from it. And it left so few pauses that measuring how much they varied stopped meaning much.
So the reset count is a warning in our checks. The count never rejects a take on its own, because a recording we were happy with by ear also had resets at more than half of its joins. Fewer sentences removed the resets once, on one voice, in one format.
What I take from it
The useful part was not the tool. It was that "read it better" stopped being the only thing anyone could say, because a script is a thing you can count before you record it. 22 sentences means 21 joins. 6 sentences means 5. That much you know before a voice touches it.
I have not shown that rewriting a script this way makes a video work better. We have never measured whether a listener stays longer or remembers more. All I can say is that a disagreement about how a recording sounded turned into a count of resets I could check, and that the script decides how many joins there are to count.