At 350 words, AI passes for a novelist
A blind test pitted GPT-5 against four award-winning novelists. Readers couldn't tell them apart. Here's what still gives the machine away.

Eight very short stories. Four written by award-winning novelists, four by GPT-5. Nearly a thousand readers were asked to guess which ones came out of a machine. The verdict: three correct calls out of eight. Coin-flipping does about as well.
One theme, eight texts, no names
The experiment comes from Mark Lawrence, a fantasy author and former research scientist, who ran it on his blog in August 2025. The setup is simple. One theme imposed on everyone, the devil, and 350 words to handle it. On the human side, heavyweights: Robin Hobb, Janny Wurts, Christian Cameron, and Lawrence himself. On the machine side, GPT-5, with the same brief.
Then you shuffle the texts, strip the names, and ask readers to decide, text by text, before rating the quality. Around 964 people voted on the first story. Enough that the result doesn't hang on a handful of opinions, and enough to turn a parlor game into something closer to a measurement.
The imposed theme matters more than it looks. Giving everyone the same prompt strips away the usual signature of a writer, the choice of subject, the personal obsession. What's left on the page is closer to pure craft, which is exactly the terrain where a model trained on millions of pages feels most at home.
Let's be clear up front: this is nowhere near a novel. It's nano-fiction, the shortest format there is. We'll come back to that, because it decides everything.
Readers ended up flipping a coin
They couldn't untangle the machine from the human. Across the eight texts, readers placed only three correctly. For the other five, they either guessed wrong or split evenly between the two options.
More awkward for the human team: the AI texts scored higher on average, and the top mark in the entire test went to a GPT-5 story. One of the invited authors played along on his own side. He got four of his five answers wrong, and even put two AI texts at the top of his personal ranking.
The result isn't a one-off. A New York Times quiz, run with 86,000 readers across five different registers, saw 54% of participants prefer the AI-written passage. On short formats the pattern holds: the average reader no longer sees the seam.
What still gives the machine away
The test isn't just a score. In the comments, voters explained how they decided, and one signal came up more than the rest: the metaphor that doesn't hold up. AI lines up images that sound good taken one at a time, but that no longer interlock when you read them together. One comparison summons another, and the logic gets lost along the way.
Picture the kind of chain where a blade whispers like a promise, then dances like a starving shadow two lines later. Each image works on its own. End to end, they tell you nothing, they just decorate. A human author would have kept one of those flourishes and cut the rest.
Novelist Rachel Neumeier, who has dissected several of these tests, makes it her number-one detector. She reads until the first wobbly metaphor, and she votes. In her view it's the fastest and most reliable way to spot a generated text.
The other tells are more familiar. The overuse of dashes, adjectives stacking up, vague endings that land without surprise, dialogue tags always poured from the same mold.
Taken alone, none of these signals proves anything. Stacked together, they start to smell like a template. What the machine lacks, Neumeier sums up, is wit, the precision of the right word, humor, the unexpected.
The real safeguard is length
There's a point Lawrence owns himself. At 350 words, AI is playing on its best turf. On a text that short, a model has almost nothing to hold over time.
No character to develop across two hundred pages, no plot to close thirty chapters later, no consistency to keep in memory from one end of the book to the other. The machine sprints well. The marathon, where you have to remember the first mile by the time you reach the thirtieth, is another matter.
Lawrence reckons that at the scale of a 20,000-word novella, the professionals would win every time. Neumeier agrees: at novel length, she says, AI fools no one. What the test measures fits in a single line. At 350 words, an ordinary reader can no longer tell the two apart. At book length, the gap reopens.
There's one last limit, an honest one. These authors make their living from novels, not nano-fiction. They were pulled out of their discipline to race on a format that isn't theirs. The result says as much about the chosen ground as about the machine's actual level.
The unsettling part isn't today's score. It's that this safeguard, length, shrinks with every generation of model. The day an AI can hold 20,000 words without losing the thread, the real question won't be who wrote it, but what that changes about the way we read.
Topics covered:
Frequently asked questions
What is Mark Lawrence's blind test?
Did readers spot the AI texts?
What still gives an AI-written text away?
Can AI write a full novel just as well?

Alexandre Noto
Co-founder & Tech Expert
Alexandre has been in tech for over 20 years. Entrepreneur, software architect and AI enthusiast, he translates complex concepts into accessible explanations. At Declic Media, he is the technical voice that makes AI understandable for everyone.
All articles by Alexandre →