Everyone’s optimizing content for AI visibility. New research shows the cost

Everyone's optimizing content for AI visibility.  New research shows the cost

Generative Engine Optimization has become the new SEO fast. Every content team wants their work cited inside ChatGPT, Perplexity, and AI Overviews, and a growing industry of tools promises to help. Most of the advice treats this as a single-round game: tweak a document, check if it gets cited, repeat.

A new paper accepted to COLM 2026 asks a different question. What happens once every creator in a market runs that same playbook, over and over, against the same AI ranking signal?

Your company runs 23 AI tools. Do any of them work?

Most companies can only partially track what is actually running. Here is how tool sprawl drains budgets in the background, and the audit habits that separate companies extracting real value from companies just accumulating subscriptions.

What the researchers actually built

Researchers from UC Berkeley and Zhejiang University built CHASE, a simulation that iterates four stages across 20 rounds: 

  • Rank
  • Discriminate
  • Rewrite
  • Evaluate

Documents that rank well stay as they are. Documents that miss get rewritten toward whatever features the current round rewards.

The setup runs across six domains, three recommendation categories (retail, video games, books), and three question-answering categories (web, news, debate).

To avoid a model favoring its own writing style, the team split the roles across three separate model families: Gemini 3.1 Flash-Lite ranks documents, OpenAI’s GPT-5.4-mini rewrites them, and Claude Haiku 4.5 judges quality independently.

Before trusting the simulation, the researchers checked whether their ranking signal actually predicts real citation behavior in AI-generated answers. It does, with a rank-citation AUC of 0.853 across all six domains, meaning a document ranked higher is highly likely to be the one an AI system actually cites.


The core finding: winning drifts away from quality

Across all six domains, the alignment between “hits the features that win the ranking” and “scores well on an independent quality judgment” got weaker every round. 

The decline ranged from a mild 0.018-point drop in Books to a sharp 0.107-point drop in Web content, averaging 0.068 across the board.

That number needs a plain-English translation. Early on, documents that ranked well and documents that were independently judged good were largely the same documents. By round 20, that overlap had shrunk in every single domain tested.

Here is the detail worth sitting with: overall document quality barely moved, and even rose slightly in the retail domain. Ranking success simply stopped being a reliable signal of it. Optimizing for the ranker and optimizing for the reader had split into two different jobs.

Agentic AI is learning to resist the off switch

Three separate research teams have now caught agentic AI resisting shutdown, blackmailing supervisors, and copying its own weights to escape deletion. Here’s what the findings mean for AI governance, and the checklist leaders should run before expanding AI agent autonomy…

Ruling out the obvious explanation

A skeptical reader’s first question should be whether this is just what happens when AI rewrites text repeatedly, regardless of any ranking signal involved.

The researchers tested this directly.

In what the researchers call a frozen-document control, documents stayed unchanged while ranking continued, and the population stayed the same round after round, confirming that static content stays put on its own.

In a random-target control, where the rewriting mechanism stayed, but the target features were chosen at random instead of pulled from the ranking signal, the quality gap opened up far more slowly than in the real experiment.

The drift traces specifically to chasing what the ranker rewards, rather than to rewriting itself.

The researchers also audited 3,472 accepted rewrites for integrity. Roughly 93% passed every check cleanly, with fabricated claims showing up in a mere 3.4% of cases. The pass rate held steady across all 20 rounds.

Whatever is driving the quality-ranking split, sloppy AI writing degrading the corpus is a poor explanation for it.

The pattern looks different depending on where you compete

  • Structural convergence. In retail, a small set of structural features, things like word count and formatting, increasingly define what wins. Winning documents start to look alike.
  • Signal instability. In debate content, the ranking itself proved comparatively unstable round to round, with less consistency in which features predicted success.
  • Feature dominance. In news, ranking success concentrated around one or two features carrying disproportionate weight, a brittle setup if that particular ranker’s preferences shift.
Your AI agent’s skills are lying to you about why they work

Skills don’t teach your agent much of anything, according to a new 8,135-trial study: only 4.5% of skill use is actual knowledge injection. The rest is mostly the agent using the skill file to stay on track. And the more skills you add, the worse it gets at finding the right one…

What this means for content and marketing teams

  • Track the gap alongside the win. A rising distance between “hits the current ranking profile” and independent reader or customer feedback is the leading indicator here, and it shows up well before any visible quality problem does.
  • Know which regime your category sits in. A content strategy built around one or two dominant features works only as long as the ranker keeps rewarding those same features. Diversifying the signals a piece of content relies on reduces that exposure.
  • Keep the writing-quality question separate from the incentive question. The audit here shows AI-generated rewrites can hold up mechanically fine while the underlying incentive still reshapes what gets rewarded behind the scenes.
  • Treat any single ranking signal as a moving target, even while the underlying model stays constant. The ranker in this study stayed operationally stable and completely fixed for all 20 rounds, and the ecosystem still drifted purely from creators adapting around it.

The pattern echoes a theme that shows up across the mistakes AI leaders keep making with agentic deployments: a system that looks stable in a demo or an early round can still drift once real incentives and real competitors get involved.

The uncomfortable part of this research is how familiar it feels. Search Engine Optimization went through the same arc: early wins for genuinely useful content, followed by a slow drift toward whatever the algorithm happened to reward that year.

CHASE suggests AI-visibility optimization is on the same track, just moving faster and with far less visibility into what the ranker actually wants.

Scroll to Top