Why Your Most Clearly-Written Content Might Be the Least Visible to AI
Flying V Group and GEO Genius published a study this month measuring content across four large online publishers, comparing how frequently individual pages get cited in AI-generated answers against how those same pages rank in traditional Google search. The two showed little correlation. More striking: content that's easy for an AI model to summarize using what it already knows from training was consistently less likely to be cited than content offering information the model likely hadn't already learned.
If your content strategy was built around clarity and easy comprehension, the two things that made content perform well in Google search for over a decade, that same clarity may now be working against your visibility in AI-generated answers.
Why This Happens: Adaptive Retrieval
Modern AI systems don't search the web for every query. They answer directly from training data by default, and only reach for outside sources when their internal confidence in the answer falls below a certain threshold. Researchers call this adaptive retrieval. Ask an AI who wrote a novel published two centuries ago, and it answers instantly from training data because that fact has been settled and repeated everywhere for years. Ask it for a current statistic or a genuinely niche detail, and it searches, because its training data can't be trusted to be current or complete on that specific point.
This is the mechanism behind the study's finding. A page that clearly and accurately explains a well-established concept is answering a question the model can already answer itself. There's no reason for the model to search out and cite that page specifically. A page containing something genuinely outside the model's training, a specific data point, a real client pattern, an original argument, gives the model an actual reason to retrieve and cite it.
The Metric Behind the Study: Commodity Content Score
GEO Genius built a specific tool to measure this directly, called the Commodity Content Score. It analyzes a given page and returns a score reflecting how saturated, replicated, and well-trodden the topic already is across the web. The more a page simply restates what's already widely available elsewhere, the higher its commodity score, and the study found an inverse relationship between that score and AI citation frequency. Low commodity content, genuinely differentiated material, earned more visibility in AI-generated answers.
This Doesn't Mean Write Confusing Content
It's worth being precise here, because the finding is easy to overcorrect. Separate, well-corroborated research on what winning AI-cited content actually looks like (analysis of over a million ChatGPT citations by independent researcher Kevin Indig, among others) found that cited content also tends to use simple, direct language, states key information early in the piece rather than after a long introduction, and favors clear declarative statements over hedged or complex phrasing.
Put together, these two findings aren't in tension. Both are true simultaneously: content needs to say something the model doesn't already know, and it needs to say that thing simply and immediately, not buried in jargon or introduction. Confirm what's already established briefly if needed, then move quickly to the genuinely new information. Clarity earns the model's trust. Novelty earns the citation. You need both.
What This Means for How You Write Thought Leadership Content
Most B2B content, including a fair amount of what gets published as thought leadership, is written to be maximally clear and accessible: a plain-English explanation of a concept the audience may not fully understand yet. That instinct made sense when the audience reading it was a human unfamiliar with the topic. It matters less, and can actively hurt AI visibility, when the actual audience increasingly includes a model that already knows the concept cold and is deciding whether your specific page adds anything worth citing.
Before publishing, it's worth asking one direct question of any piece of content: does this contain something a well-informed AI system couldn't already generate on its own? A specific number from your own business. A genuinely contrarian take, argued with evidence. A real pattern observed across actual client work. A piece of content that passes this test has a structural reason to be cited. One that doesn't is competing against the model's own default answer, and losing.
Practical Ways to Build This In
Include specific, sourced statistics rather than general claims wherever possible, ideally numbers that aren't already the most commonly repeated figure on the topic. Share a genuine operational detail or pattern from real work, something a generic explainer wouldn't include because it isn't generic. Take an actual position on a genuinely contested question rather than a balanced, hedge-everything summary of both sides. And when citing a widely known fact as necessary context, do it briefly, then move immediately to what's actually new in the piece.
Frequently Asked Questions
Does this mean well-written, clear content no longer matters for AI visibility?
No. Clarity and simple language remain strongly associated with AI citation in separate research. The finding here is specifically that content restating common knowledge, however clearly written, is less likely to be cited than content offering genuinely new information, however it's written. The two findings work together rather than against each other.
How would I know if my own content scores as commodity content?
A rough gut check: could a knowledgeable person in your industry write the same core content from memory, without any research specific to your business? If yes, it's likely high on the commodity scale. If your content includes something only your business could credibly say, a real number, a real pattern, a genuine position, it's more likely to score as differentiated.
Is this only relevant for large publishers, or does it apply to smaller business blogs?
The mechanism applies at any scale. A small business blog post that repeats widely available generic advice faces the same structural disadvantage in AI citation as a large publisher's commodity content. The fix, adding something genuinely specific and original, is also available at any scale.
This builds directly on our coverage of what actually determines AI citation eligibility and why reputation is replacing reach as the thing that gets businesses recommended. If you're not sure whether your own content is differentiated enough to earn AI visibility, that's exactly the kind of gap a content development partnership is built to close.