Curio Evidence Atlas

The surprising result is only half the story.

Popular claims often grow larger than the study beneath them. This map keeps the interesting finding, the test, and the boundary in the same place.

Showing every reviewed evidence note.

CuriositySource-backed explainer

Waiting for an Answer Helped an Unrelated Face Stick

People remembered faces a little better when they appeared during a high-curiosity wait, even though nobody told them to learn the faces.

The interesting result

Faces shown during high-curiosity waits were recognized a little better later.

What was checked

Trivia questions created high- and low-curiosity waits. Unrelated faces appeared before the answers, followed by an unexpected face-recognition test.

What it does not prove

It does not show that curiosity guarantees memory, releases a measured burst of dopamine, or makes everything nearby memorable. The group-level advantage was modest.

Why it stays interesting

Brain scans found related activity in reward and memory regions without proving a cause.

Read the full explanation and sources

AISource-backed essay

One Case Study Found More Confidence, Not More Correct Answers

In Stuart Oskamp's small case study, everyone read the same sections in the same order. Confidence rose without a clear gain in correct answers.

The interesting result

Everyone read the same case sections in the same order; confidence rose without a clear gain in correct answers.

What was checked

Judges read one difficult case in cumulative sections, answered the same fixed-choice questions, and rated their confidence after each stage.

What it does not prove

It does not show that more information always causes overconfidence or that AI research is useless. The study used one case, a small mixed sample, and a fixed order.

Why it stays interesting

When technical answers become cheap, choosing which detail matters becomes the scarce part.

Read the full explanation and sources

Technology & SocietySource-backed explainer

A Facebook Like for Curly Fries Helped Predict a Reasoning Score

A Facebook Like for Curly Fries helped predict a reasoning score in old data. Instagram uses different signals; the bridge is what patterns can reveal.

The interesting result

A Facebook Like for Curly Fries helped predict a reasoning score as part of a wider pattern.

What was checked

Older Facebook studies modeled quiz scores, profile labels, and self-reported traits from Like patterns. Meta separately describes behavior signals used for Instagram ranking.

What it does not prove

It does not show that Instagram predicts personality, knows a person's real self, or uses those Facebook studies. Group prediction is not certainty about one person.

Why it stays interesting

Instagram's guides show another use of behavior patterns, but not the same trait-prediction system.

Read the full explanation and sources

Building CurioBuilding Curio · case note

I Changed Sentences and Pauses. The Voiceover Felt Less Like a List.

I preferred one Curio rewrite, but sentence structure and pause timing changed together. The comparison cannot credit either edit.

The interesting result

I judged one Curio rewrite less list-like.

What was checked

Curio compared one short-sentence voiceover with a joined-sentence rewrite and inspected pitch movement around sentence boundaries.

What it does not prove

It does not show that fewer sentences always improve narration or that pitch resets cause list-like delivery. Sentence structure and pause timing changed together.

Why it stays interesting

A boundary measure locates changes; controlled listener ratings would test whether they track perception.

Read the full explanation and sources

Building CurioBuilding Curio · case note

My Pinned-Card Check Stopped Before the Phone Screen

My server check sent one Curio card first. The phone then rebuilt and ranked the list, so the visible result still needed a screen test.

The interesting result

My check proved which card the server sent first, not which card the phone finally showed.

What was checked

A server test confirmed which card was returned first. The phone then rebuilt and ranked the list before drawing the visible screen.

What it does not prove

It did not prove which card a person finally saw, whether the position changed, or whether position affected reader behavior. A live screen test was still required.

Why it stays interesting

Ranking research shows that position can shape clicks, so each human claim needs a matching test.

Read the full explanation and sources

Technology & SocietySource-backed essay

A High-Stakes Score Can Change the Work It Measures

United States law schools changed admissions work and resource use around a public ranking. The case shows how a score can enter the process it measures.

The interesting result

Law schools reported changing admissions work and resource use around a public ranking.

What was checked

Researchers interviewed United States law-school staff about how admissions work and resources changed around a public ranking.

What it does not prove

It does not show that every metric causes gaming, that all scores are invalid, or that a higher score never reflects better work. The AI-writing link is an application, not a tested result.

Why it stays interesting

Curio can audit the edits made to pass a score and check clarity apart from the grade.

Read the full explanation and sources

LearningSource-backed explainer

More Study Won First; Trying to Remember Without the Passage Won Later

In two short-passage college studies, repeated study produced higher recall after minutes. Trying to remember without the passage led after waits of days.

The interesting result

Repeated study won after five minutes; repeated retrieval won after waits of days.

What was checked

Undergraduates learned short passages through repeated study or repeated free recall, then took a final recall test after waits ranging from minutes to a week.

What it does not prove

It does not show that any quiz improves learning or that retrieval transfers equally to every task. The foundational studies used short passages and a particular final test.

Why it stays interesting

Transfer gains were stronger when practice and final tasks asked people to produce answers in similar ways.

Read the full explanation and sources

AttentionSource-backed explainer

Closely Timed Signals and Switched Rules Slow Responses Differently

Signals arriving close together can delay a second response. In separate task-switching tests, changed rules also slowed responses against repeat trials.

The interesting result

Closely timed signals delayed the second response; in separate tests, switching rules was slower than repeating them.

What was checked

Dual-task studies changed the interval between two signals. Separate task-switching studies compared trials that repeated a rule with trials that changed it.

What it does not prove

It does not establish one fixed cost of multitasking or explain everyday phone behavior. A similar delay can come from different tasks and does not identify its cause by itself.

Why it stays interesting

Equal pauses can come from different test conditions and need different questions.

Read the full explanation and sources

MemoryResearch summary

One Fear Test Lost the Threat. Another Still Found It.

After propranolol and a brief fear reminder, a loud-noise blink response no longer clearly separated threat from safety. People's expected-shock ratings still did.

The interesting result

One group's loud-noise blink response no longer clearly separated its threat and safe cues.

What was checked

Volunteers learned threat and safe cues. After propranolol and a brief reminder, researchers measured both loud-noise startle and expected-shock ratings.

What it does not prove

It does not show that the memory was erased or that propranolol removes fear. Human reconsolidation is inferred indirectly and the startle result has a mixed replication record.

Why it stays interesting

A later study did not repeat the eye-blink result, leaving the broader claim unsettled.

Read the full explanation and sources

MemorySource-backed explainer

Unfinished Tasks Can Pull You Back Without Better Recall

A review found no clear recall edge for unfinished tasks, even though other studies found signs of faster access and a pull to resume.

The interesting result

Unfinished tasks were not clearly easier to recall across the reviewed studies.

What was checked

Separate intention studies measured access to planned-action words; a recent synthesis separately reviewed interrupted-task recall and the pull to resume.

What it does not prove

It does not show that every unfinished task stays active in working memory, that interruption improves recall, or that planning reliably removes anxiety.

Why it stays interesting

The urge to resume and the power to recall are different outcomes.

Read the full explanation and sources

MemorySource-backed explainer

September 11 Reports Changed While Memories Felt Certain

Later September 11 accounts conflicted with earlier reports, while people still described those memories as vivid and trustworthy.

The interesting result

Reports of September 11 memories changed while confidence in those reports stayed high.

What was checked

People recorded September 11 and ordinary-event memories soon after the attacks, then reported them again over later intervals.

What it does not prove

It does not show that vivid memories are necessarily false. The studies measured agreement with an earlier report, not objective truth, and did not isolate one cause of change.

Why it stays interesting

Confidence shows how remembering feels now, not whether every detail stayed the same.

Read the full explanation and sources

LearningResearch summary

People Reading on Screens Sometimes Misjudged What They Learned

Matched-text reviews found a small average paper edge. Some people reading on screens also estimated their own test results less accurately.

The interesting result

Matched-text reviews found a small average comprehension edge for paper.

What was checked

Meta-analyses compared matched text on paper and screens. Separate experiments compared predicted test scores with the scores readers actually earned.

What it does not prove

It does not show that screens make people less intelligent or that paper always wins. Results vary with the text, device, timing, and task.

Why it stays interesting

Comparing predicted and earned scores separates test performance from self-judgment.

Read the full explanation and sources

Building CurioExperiment plan

Curio Will Compare Two AI Workflow Bundles on Fixed Briefs

Curio will compare two AI workflow bundles on fixed briefs and review rules. The test asks whether splitting idea work from building helps a solo builder.

The question

Can splitting idea work from implementation improve complete work under fixed conditions?

How it will be tested

Curio proposes fixed briefs, source packets, tools, prompts, limits, and review rules, followed by a blind review of two reversed workflow bundles.

What it does not prove

No Curio result exists yet. The comparison will test complete workflows, not prove that either displayed model label is universally better at one job.

Why it stays interesting

Human creativity research separates making from judging, but it does not assign those jobs to AI models.

Read the full explanation and sources

How to read this map

Curio keeps the boundary visible.

Each entry comes from a complete article with its sources linked. “What was checked” describes the task or comparison. “What it does not prove” prevents a narrow result from turning into a universal rule.

This is an editorial map, not a medical or scientific review service. When evidence changes, the article and its place in the map should change with it.