What the Blind Test on AI-Written Fiction Actually Shows
1.54 versus 0.97, on a scale that runs from minus three to plus three. Both scores are positive, and nobody could reliably tell which story came from a machine.

1.54 versus 0.97, and the missing piece
1.54 versus 0.97. On its own, that pairing looks like the verdict that has been making the rounds for a week: readers supposedly prefer fiction written by artificial intelligence.
One detail is missing next to those two numbers. They sit on a scale running from minus three to plus three, seven notches, from "strongly disagree" to "strongly agree." Both scores are positive, and the gap between them is 0.57 points.
Readers liked both stories, in other words, and liked one a bit more than the other. That is roughly the distance between two restaurants rated 4.1 and 4.4 out of five. You can build a headline out of that, as long as you mention the food was good at both places.
What the study actually measured
The study itself is solid, and that needs saying up front. Sydney Sears and Deena Skolnick Weisberg, of Villanova University, published it on August 5 in Judgment and Decision Making, the journal published by Cambridge University Press. The data and analysis scripts are public on OSF.
Three short stories by human authors, paired with three ChatGPT 4.0 generated versions, all around a thousand words. A detail few of the writeups mention: the prompts handed the AI the same theme, point of view, and setting as the matching human story. The recipe was given away before anyone ordered the dish.
In the first study, 1,682 American adults recruited on Prolific each read one of the six stories. They were told its origin, correctly half the time, falsely the other half. They then rated perceived quality and absorption on the same minus-three-to-plus-three scale. Quality: 1.54 for AI, 0.97 for human. Absorption: 1.42 versus 1.00.
The result the headline buried
Two more studies follow, with 905 participants. This time, each person sees both versions of a story and has to pick which one was machine-written.
In Study 2, 167 out of 424 people guessed right, or 39.39%. A proportion test places that result significantly below 50%: a compass pointing south would be more useful. In Study 3, 51.97%, which the same test cannot distinguish from 50%. The abstract sums both up in one line: participants were "no better than chance."
The same text gets a higher rating when readers are told a human wrote it, regardless of where it actually came from. The two effects are independent, measured within the same experimental design: readers penalize the AI label, and they cannot detect it in the first place.
The abstract puts it this way: AI programs produce creative work judged "at least as good as" that of humans, but are not perceived as capable of it. At least as good, not better.
Who gave you the scale, and who did not
The six writeups we checked all correctly report the labeling paradox. What gets lost along the way is the unit of measurement.
Slate publishes all four averages and specifies that readers rated stories "on a scale running from -3 to 3."
ActuaLitté goes further, with the scale, the word count, the titles of all six stories, and a methodological note on the prompts. The Decoder also spells out the scale in full.
Futura publishes the same four averages and never names the scale. Its headline, to be fair, is accurate: it captures the paradox and claims no preference. The loss happens in the body of the piece, even as the headline holds up.
A number without its unit is like a street address missing the city name. It exists, it is spelled correctly, and it leads nowhere.
The press release oversells its own abstract
The press release that set off the wave was published on August 4 on EurekAlert, credited to Cambridge University Press. Its headline: "People prefer stories written by AI." The publisher of the paper, in other words, titled its own release more aggressively than the abstract it posted the next day.
The release, for its part, gives no effect size at all. No 1.54, no 0.97, no scale. It states that AI-generated texts scored higher, without ever saying by how much: a tenth of a point or three, there is no way to know from the release alone.
A number that shows where everyone actually read
One more detail worth flagging. The press release states that participants correctly identified the author 39.93% of the time. The paper itself prints 39.39%, and the count confirms it: 167 correct answers out of 424. Two digits swapped.
The conclusion does not hinge on it, both figures sit below chance either way. But a slip like that behaves like a copyist's error in a manuscript: it lets you trace who copied whom. earth.com repeats 39.93% verbatim. ActuaLitté prints 39.39%, a figure that only exists in the original paper.
Full disclosure: this piece nearly repeated the error too. The research file used to prepare it carried 39.93%, believing it was quoting the study, when it was actually quoting the press release. The discrepancy only surfaced when the PDF was reopened. The mechanism spares no one, including a series built around catching distorted numbers.
The scope, as the authors themselves define it
Six stories of roughly a thousand words, realistic fiction, an entirely American sample, versions generated between November 2023 and December 2024. The paper's limitations section opens with exactly that: the stories were "very brief," and longer works, "such as entire novels," might produce a different result, since the machine may be less capable of sustaining characters over greater length.
The researchers also note that their perceived-quality measure had not been validated in prior work. Six short stories is a tasting plate, not the full menu.
There is a precedent worth noting here: AI already beat the human average on creativity tests, without beating actual creatives.
What holds up is more interesting than the headline
The explanation the researchers offer is not literary merit, it is fluency. Readers liked these texts because they found them "easier to read or understand." And the way ChatGPT builds a narrative, by averaging millions of examples, may "smooth out the rough edges" of human writing, the way a face composed of many real faces can look more harmonious than any single one of them.
The generated versions state their theme instead of letting readers work it out. That is more comfortable to read, and it is exactly what a good short story refuses to do.
Weisberg drops one more detail into the release: people familiar with AI spot its writing more reliably, relying on tells like the em dash or constructions such as "it's not just X, it's Y." We covered that em dash back in April. What this study adds is that this kind of pattern recognition is learned, and not by reading a lot of novels: participants' literary expertise, on its own, made no difference.
Topics covered:
Frequently asked questions
Do readers actually prefer stories written by AI?
Can readers tell when a story was written by AI?
What happens to a story's rating once it is labeled written by AI?
Why did the AI-generated texts score slightly higher?
How far does this study's finding actually reach?
What is Declic Media's Verified Number of the Week series?

Alexandre Noto
Co-founder & Tech Expert
Alexandre has been in tech for over 20 years. Entrepreneur, software architect and AI enthusiast, he translates complex concepts into accessible explanations. At Declic Media, he is the technical voice that makes AI understandable for everyone.
All articles by Alexandre →