AI Source-backed 3-minute read
Jensen Huang, AI Judgment, and One Leaner Workflow
In one internal video-idea review, I knew which workflow made each idea. The leaner setup scored higher, raising a question Jensen Huang's words helped frame.
The short version
- I knew which workflow made each idea, and I rated the leaner setup higher.
- The workflows changed in several ways, including their prompts and amount of context.
- The review cannot show which difference led to my ratings.
The larger workflow did not win the review
I compared two AI workflows in one pre-release review. I gave the leaner workflow's video ideas higher ratings.
No audience saw the ideas. I knew which workflow made each idea, so the review was not blind. It does not show which ideas would have held viewers longer.
The larger workflow used Perplexity, an AI research tool, to gather material before drafting. Another model then reviewed that material before the draft.
The research covered how curiosity starts, how attention holds, and how long viewers keep watching. The last measure is retention.
The larger workflow received more material. I did not give its ideas higher ratings to match the added context.
The ratings favored specific story choices
The research output gave broad rules for how curiosity works. In this review, I preferred the detail that the leaner workflow chose for the opening line.
Research and selection are different jobs. Research gives a system more options. Selection chooses the fact for the opening and saves another fact for a later reveal.
A report can describe a prediction gap, the distance between what a viewer expects and what comes next. It can still fail to find the fact that creates that gap.
The same problem applies to a hook, the opening meant to make someone stay. General advice about hooks does not write the right first line for one story.
The larger pipeline supplied useful facts. In this review, I rated its choices lower.
The smaller workflow changed the question
I tried a lighter setup. I passed the source material to GPT-5.6 Sol, the OpenAI model used in this test. I also gave it the limits and checks learned from earlier projects.
Then I asked it to find the strongest story already in the source. I rated its openings cleaner, its reveals stronger, and its reasons to continue clearer.
Several things changed at once. They included the prompt, the instructions and source text sent to the model. The amount of background material changed too.
That makes cause hard to assign. The result does not prove that Sol always wins or that research hurts creative work.
The result gives Curio a reason to run a controlled test of the research-heavy workflow. The current review cannot isolate the effect of any one change.
Intelligence and judgment are different jobs
Jensen Huang used words that fit this review. In a January 2026 interview, he argued that raw technical intelligence was becoming a commodity. Here, commodity means that such ability was becoming widely available and less distinctive.
He valued judgment that could "see around corners." I take that metaphor to mean spotting important shifts and likely results before they are clear.
NVIDIA calls systems that make large amounts of AI output AI factories. The term is NVIDIA's framing. It does not show that research lacks value.
Judgment is the choice. It decides which facts change the plan and which option gets more time.
Research still matters when a fact is missing or a claim needs a check. Each research step should answer a named doubt. A named doubt is one specific question or point of uncertainty. A longer prompt is not useful merely because it is longer.
OpenAI's GPT-5.6 guide recommends testing leaner prompts on real work. A token is a small chunk of text a model reads or writes. Using fewer tokens matters only if the work still passes its checks.
Every layer needs a named test
One review did not make research obsolete. It raised a narrower question: does the current research layer improve Curio's idea ratings?
Both workflows gave me options. My largest rating gap concerned the story and opening each workflow chose.
An agent is a model system that can take steps or use tools. More agents can add more handoffs without improving the choice.
Before release, each layer should improve one check stated in advance. It might sharpen the opening, clarify the reveal, catch a false claim, or shorten the path to an approved idea.
After release, viewing data can test whether the chosen opening and reveal kept people watching. Until then, the honest result is small: one unblinded review of two workflows that differed in several ways favored the leaner setup.