Graphite Finds 13,000 Statistical Tells in AI-Generated Prose
A Graphite study found thousands of phrases that appear disproportionately in AI-generated prose, with Claude and GPT models showing distinct linguistic patterns.

Key takeaways · 3
- 01
Treat familiar stylistic clues cautiously because models can shed old habits while developing new, version-specific patterns.
- 02
Phrase-frequency analysis offers stronger evidence than flagging a single word, punctuation mark, or sentence construction.
- 03
When reviewing Opus 5.5 output, scrutinize repeated explanatory framing such as “this matters” and “why X matters.”
How Graphite Tested Prose
Graphite examined the writing habits of frontier models and identified 13,000 phrases that appeared at least twice as often in AI content as in human content, its threshold for a writing “tell.” [1] The researchers began with 10,000 articles published before ChatGPT’s release as a human-generated control group, then asked different models to rewrite the articles from summaries to reduce source bias. [1] Matching human and model samples allowed the researchers to compare the frequency of words and phrases alongside broader sentence-construction patterns. [1]
Opus Has Distinct Habits
Graphite chief AI officer Greg Druck told TechCrunch that Claude models are moving closer to the human word distribution over time, while GPT models are moving further away. [1] Claude Opus 5.5 used “dependable” 23 times more often than the human samples and favored “more than an X, it’s a Y,” even as it avoided the older “it’s not X, it’s Y” construction. [1] The model used “this matters” 116 times more often than human writers and “why X matters” 92 times more often. [1]
What it means
The findings make AI-writing detection look less like a checklist of permanent giveaways and more like statistical profiling of particular model versions. Claude Opus 5.5 can abandon a familiar contrast construction while retaining other favored vocabulary and explanatory framing. The Claude-GPT comparison also shows that progress toward human-like word distributions is not uniform across model families, so one detector or style heuristic may not transfer cleanly between them. What the sources don't address: whether these measured tells remain reliable across different prompts, subject areas, languages, and editing workflows.
Simple rules such as flagging em dashes or isolated buzzwords are unlikely to remain dependable as models change. Practitioners evaluating authorship or editing model output should examine recurring patterns, account for model versions, and avoid treating any single phrase as proof.
Why it matters
Put this to work — one session a day, built for your industry.
Create a free account for a daily session — eight questions and one real-work challenge, on the news that affects your role.
Start freeHow this developed
1 October 2026
Graphite Finds 13,000 Statistical Tells in AI-Generated Prose
1 October 2026
Event created from source cluster.