Person writing in a notebook with a pen, laptop open on the table, plant in the background, and a cup nearby.

AI Writing Still Leaves Fingerprints, New Study Says — Even as the Obvious Tells Fade

A new Graphite study says AI writing tells persist even as models drop obvious clues like em dashes and develop new signature phrases.

In short

Graphite says frontier AI models still produce detectable writing tells, even after labs reduced obvious markers like em dashes. The study found distinct phrase patterns across models, suggesting AI prose is becoming subtler rather than indistinguishable.

  • Graphite identified 13,000 phrases that appeared at least twice as often in AI text as in human writing.
  • Claude Opus 5.5 and OpenAI Astra each showed distinctive style habits, from “this matters” framing to corrective contrasts.
  • Frontier labs have reduced obvious tells like em-dash overuse, but new patterns keep emerging.
  • The study suggests AI prose is improving in readability while remaining statistically identifiable.

New research from the marketing analytics firm Graphite says frontier AI models are still leaving identifiable writing fingerprints, even after labs have trimmed some of the most obvious giveaways. The study, published this week, found that systems such as Anthropic’s Claude Opus 5.5 and OpenAI’s Astra still rely on distinctive phrases, sentence patterns and framing habits that appear far more often in AI-generated prose than in human writing.

The findings matter because they suggest the race to make AI text sound more natural is far from over. Developers may have reduced famous cues like the overuse of em dashes, but Graphite says other habits are emerging, making model authorship easier to detect in some cases and harder in others.

For businesses, publishers and security teams trying to identify synthetic content at scale, the study is a reminder that AI prose is evolving rather than disappearing into a perfect imitation of human language.

What Graphite found in frontier model writing

Graphite’s core argument is that AI-generated text still contains recurring “tells” — word choices and constructions that show up disproportionately often in model output compared with human writing. The company says it identified 13,000 phrases that appeared at least twice as frequently in AI text as in human samples.

The pattern was not limited to a few famous clichés. Instead, Graphite found a wide spread of model-specific habits that changed from one system to another. Some models leaned heavily on contrastive phrasing, while others preferred hedging language or explanatory set-ups that framed a topic as newly important.

That breadth is notable because early AI detectors and human spotters often focused on a few stereotyped giveaways — words like “delve,” repetitive transitions, or the dash-heavy style that once made AI text easier to flag. Graphite says many of those signals have weakened, but they have not gone away in a broader sense. They have simply been replaced by newer habits.

How did the researchers test for AI writing tells?

Graphite says it built its analysis around a comparison between pre-ChatGPT human writing and rewritten versions produced by AI models. The approach was designed to reduce bias from source material and isolate stylistic differences more cleanly.

The company used a control set of 10,000 articles published before ChatGPT’s launch, treating them as human-written baseline material. It then asked models to rewrite those articles from summaries. By matching the samples, researchers could compare vocabulary, phrasing and sentence structure across humans and machines under similar prompts.

This method matters because it avoids one of the biggest problems in AI-text analysis: source contamination. If an AI model is asked to rewrite existing copy, it may imitate the subject matter, structure or tone of the original. Using paired samples gives researchers a clearer view of what the model adds on top of the prompt.

Model Most notable tell Relative frequency Style pattern
Claude Opus 5.5 “dependable” 23x more than human samples Likes “this matters” and “more than X, it’s a Y” constructions
OpenAI Astra “corrective framing” More than 100x more common Uses “not simply X” and “rather than relying on X” style contrast
Gemini 3.1 Pro Em-dash use Nearly eliminated Has sharply reduced a long-known punctuation tell

Why does Claude Opus 5.5 keep saying “this matters”?

Claude Opus 5.5’s biggest giveaway, according to Graphite, is an apparently benign habit: it leans heavily on the phrase “this matters.” The study says that expression appears 116 times more often in Opus 5.5 output than in human writing, while the related phrase “why X matters” shows up 92 times more often.

Graphite also found that the model often reaches for the adjective “dependable,” which it used far more frequently than human writers in the comparison set. Another recurring pattern is a preference for a particular kind of upgrade construction: instead of saying “it’s not X, it’s Y,” Opus 5.5 often shifts into “more than X, it’s Y.”

That distinction may sound small, but style analysis often depends on small differences repeated at scale. When a model uses the same rhetorical move hundreds or thousands of times, the pattern becomes visible in aggregate even if any single sentence feels ordinary.

Graphite chief AI officer Greg Druck said the company’s analysis suggests Claude models are drifting closer to the human word distribution over time, while GPT-family models appear to be moving in the opposite direction.

His comment points to a larger issue in AI style tuning: making a model sound more human is not the same as making it less detectable. A model can reduce one famous tell and still develop a different signature.

What is OpenAI Astra’s biggest writing habit?

OpenAI’s Astra appears to have a very different stylistic fingerprint. Graphite says the model often reaches for what it calls “corrective framing,” a pattern in which a topic is defined by contrast — for example, describing something as “not simply X” or presenting an alternative “rather than relying on X.”

According to the study, those constructions appeared more than 100 times as often in Astra-generated text as in human samples. Astra also frequently used phrases such as “another dimension” and tended to hedge with formulations like “may provide” or “can provide” when discussing benefits.

That combination gives Astra’s writing a particular tone: cautious, explanatory and contrast-driven. It is the kind of prose that can sound polished and persuasive, but also slightly over-structured, as if every claim is being scaffolded for clarity.

Graphite’s finding matters because it suggests that AI models do not just produce a generic “AI voice.” They seem to develop individual habits that are consistent enough to identify, at least in controlled testing.

How much have the famous em-dash tells changed?

They have changed a lot, but Graphite says the broader problem remains. The most publicized AI writing tell in recent years — the overuse of em dashes — has been sharply reduced in many frontier models.

In Graphite’s sample, Opus 5.5 used em dashes 99% less often than the model version it was compared against. Astra reportedly used them 88% less than human samples, while Gemini 3.1 Pro had nearly eliminated the punctuation mark from its writing altogether.

That is an important win for model trainers who want text that feels less mechanical. But it may also be a warning sign for anyone relying on a short checklist of AI markers. Once one giveaway becomes common knowledge, model makers can tune it down. The result is not human writing — only a new version of machine writing.

Key telltale habits Graphite highlighted

  • Contrast-heavy phrasing: Models often define ideas through opposition or correction.
  • Importance framing: Some systems repeatedly explain why a topic “matters.”
  • Hedging language: Words such as “may” and “can” soften claims.
  • Model-specific vocabulary: Certain words, like “dependable,” show unusual frequency spikes.
  • Punctuation trimming: Em dashes are down, but other patterns remain.

Why do AI writing tells keep surviving?

Graphite argues that the persistence of these patterns is less surprising than it might seem. Large language models are trained to produce plausible, fluent text, not to mimic the full statistical messiness of human writing in every context. That means style tuning can remove visible quirks without fully flattening the model’s underlying preferences.

Druck said the company’s working assumption is that developers may have less control over these patterns than outsiders expect. With systems built from billions of parameters, labs can run only so many tests and only catch so much before release. Some quirks will inevitably slip through.

That explanation fits how large models are built and deployed. Improvements often come in layers: better instruction following, fewer obvious artifacts, cleaner punctuation, more natural transitions. But each layer can expose another edge case, and every new model generation may develop its own signature in the process.

In other words, the hunt for AI writing tells is not ending. It is becoming a moving target.

What the labs are saying versus what the data shows

The study arrives at a moment when major AI companies are openly marketing improved prose quality. Anthropic said in its Opus 5.5 announcement that the model communicates more naturally than earlier versions and that early users found the writing clearer and easier to follow.

OpenAI has made similar claims about newer GPT-6 versions, saying users should expect more clarity, less jargon and fewer awkward turns of phrase. Those statements reflect a real product priority: good writing is a competitive feature, not just a cosmetic improvement.

But Graphite’s data suggests that polish and detectability are not opposites. A model can become easier to read while still remaining statistically distinctive. That distinction is important for editors, fact-checkers, school administrators, trust and safety teams and anyone else trying to decide whether a piece of prose came from a person or a machine.

According to Druck, the labs may be underestimating how difficult it is to purge these patterns completely because the models are so large and so complex that even extensive testing cannot catch every habit before release.

That is an especially relevant warning in an era when AI text generation is being woven into customer support, marketing, search, coding and newsroom workflows. The closer model prose gets to human norms, the harder it becomes to rely on intuition alone.

How should readers and editors interpret AI tells now?

The answer is that no single tell is enough on its own. A lone em dash, a hedged phrase or a formal-sounding transition no longer proves anything. But clusters of repeated habits can still be useful indicators, especially when combined with metadata, context and workflow clues.

For editors, that means checking for tone consistency, unexpected rhetorical repetition, generic framing and over-explained transitions. For platform teams, it means detector tools should be updated regularly and should avoid overconfidence. And for readers, it is a reminder that polished prose is not always human prose.

The larger lesson is that AI writing is developing a style of its own, even as its makers try to make it disappear into the background. The machine voice may be getting subtler, but it is not becoming silent.

Timeline: how the AI writing debate evolved

Period Development Why it mattered
Early chatbot era Em dashes, “delve” and other formulaic phrases were common People could spot AI text with a quick glance
Model tuning phase Labs worked to remove obvious stylistic giveaways AI prose became more fluent and less obviously mechanical
Current frontier models New habits emerged, including contrastive framing and “this matters” language Detection became more subtle but still possible in aggregate

What happens next?

The likely next step is a continuing arms race. As model makers suppress one set of stylistic fingerprints, researchers, publishers and detection tools will look for the next set. The overall number of tells may stay roughly stable, but the specific phrases and structures will keep changing.

That dynamic helps explain why AI text analysis remains important even as models improve. The aim is not just to catch obvious machine writing. It is to understand how synthetic text behaves, how it differs from human prose and how those differences change over time.

Graphite’s study suggests the field is moving beyond simple clichés and toward more nuanced pattern recognition. That may make detection harder, but it also makes the debate more useful. Instead of asking whether AI can write like a person at all, the more relevant question may be: what kind of person does it most want to sound like?

For now, the answer appears to be that no model has fully escaped its own habits. The dashes may be gone. The fingerprints are not.

Bottom line: AI writing is becoming more polished, but it still leaves statistical traces — and those traces may be getting more sophisticated rather than disappearing.

Frequently asked questions

What did Graphite’s study find about AI writing tells?

Graphite found that frontier AI models still leave recognizable stylistic fingerprints in their writing. The firm identified 13,000 phrases that showed up at least twice as often in AI text as in human samples, suggesting that subtle tells remain even as obvious giveaways fade.

Which AI model had the most distinctive writing habit?

Claude Opus 5.5 stood out for repeatedly using phrases like “this matters” and for favoring words such as “dependable.” OpenAI’s Astra showed a different pattern, leaning into corrective framing such as “not simply X” and “rather than relying on X.”

Have em dashes stopped being an AI writing tell?

Yes, mostly. Graphite says many frontier models have dramatically reduced em-dash use, with some nearly eliminating it. But the study argues that removing one obvious tell does not erase the broader pattern of machine-specific writing habits.

Can AI-generated text still be detected reliably?

Yes, but not with a single shortcut. The study suggests detection works better when multiple patterns are considered together, including word choice, sentence construction and repetitive framing, rather than relying on one famous clue like punctuation.

Why does this study matter for publishers and editors?

It matters because AI text is becoming more polished and harder to spot by eye. Editors, platforms and fact-checkers may need more sophisticated methods to identify synthetic content as model-generated prose becomes less obviously robotic.

Share this 🚀