Analytics

We Tracked AI Citations for 90 Days: 7 Findings

·2026-08-11·17 min read
Editorial illustration of AI citation instability. A large grid of circular markers on off-white paper records repeated checks of the same prompt: solid brand-red markers where the brand was named, hollow outlined markers where it was absent. The filled markers are scattered irregularly with no trend, so no row or column resembles any other. Below the grid a single hollow marker sits alone, labelled ONE CHECK, arguing that one reading cannot represent the field it was drawn from.

A client asked us a reasonable question in April: are we showing up in ChatGPT yet?

We checked. The answer was yes. We put it in the deck. Two days later, preparing for the call, someone re-ran the same prompt to grab a screenshot and the brand was gone. Not demoted, not described differently - absent. Same question, same account, same country, forty-eight hours apart.

That is an uncomfortable thing to discover the morning of a client call, and it exposed something worse than an embarrassing slide. Every AI visibility claim we had made until that point, ours and everyone else's, rested on a single check. One person, one prompt, one moment. We had been reporting the output of a non-deterministic system as though it were a rank position.

So we built a panel and ran it properly for a quarter. Sixty-two buyer questions, fourteen client brands across nine categories, three engines, three repeats per engine per week, thirteen consecutive weeks. Just over 7,200 recorded answers.

This article is the whole dataset and what it changed about how we work. It is not a list of tips for getting cited - we cover that in how to rank in AI Overviews, how to rank on ChatGPT and how to rank on Perplexity. This is about what the surface actually looks like when you stop taking one photograph of it and start filming.

The Short Answer

Over 13 weeks and just over 7,200 recorded AI answers across ChatGPT, Gemini and Perplexity, the strongest findings were these. Three repeats of the same prompt in the same week agreed with each other only 61% of the time. When a brand appeared at all, it appeared on all three engines only 18% of the time. 63% of every citation pointed at a domain the brand does not own. Restructuring an existing ranking page produced a first citation in a median 11 days, while new pages took 19 and schema work moved nothing inside the window. And brands whose citation rate rose sharply saw branded search rise a median 14% while sessions to the cited pages barely moved at all.

How We Ran It

The method matters more than usual here, because almost every disagreement about AI visibility data turns out on inspection to be a disagreement about sampling. So here is exactly what we did, in enough detail that you can argue with it.

The prompt set. We assembled 62 questions across the fourteen brands - between three and six per brand - drawn from real sales-call transcripts and search console query data rather than written from imagination. Four types, deliberately mixed: category questions ("what is the best X for Y"), comparison questions ("A vs B for Z"), problem questions ("how do I fix W"), and shortlist questions ("who are the leading providers of X in India"). The list was frozen on day one and never edited, which is the single most important discipline in the whole exercise. A panel you keep improving is a panel you can no longer compare across time.

The engines. ChatGPT, Google Gemini, and Perplexity. We excluded Google's AI Overviews from this dataset because Overview presence itself fluctuates by query and location in a way that would have confounded the brand-level analysis, and because we track that surface separately. We also excluded Claude and Copilot for capacity reasons, not because they do not matter.

The sampling rule. Every prompt, on every engine, three times per week, in a fresh signed-out or clean-profile session, checked from India and from the United States. Three repeats is not statistically generous. It was the most we could sustain for thirteen weeks without the collection quietly degrading, and it was enough to expose the variance problem, which was the point.

What we recorded. Four fields per answer, and nothing more: was the brand named, was any URL cited, which domains were cited, and was the brand's description factually accurate. We resisted every temptation to add sentiment scoring, position weighting or prominence scores. Richer schemas sound better in a methodology section and collapse in week five when the person collecting is tired.

ParameterValue
Window13 consecutive weeks
Brands14, across 9 categories
Prompts62, frozen on day one
EnginesChatGPT, Gemini, Perplexity
Repeats3 per prompt, per engine, per week
LocationsIndia and United States, English only
Total recorded answers~7,200
Fields per answerNamed, cited, domains, accuracy

What this design cannot tell you. It is observational. We recorded what happened while normal client work continued around it, which means the timing findings later in this article describe association and not proven causation. The panel is small. Personalisation could not be fully eliminated. And at least two engines shipped model or retrieval changes mid-window, so some of the movement is the platform moving, not the brands. We will come back to all of this at the end rather than burying it.

Finding 1: A Single Check Is Close to Worthless

This is the finding that changed our practice most, so it goes first.

Across every prompt-week in the panel, we asked a simple question of the three repeats: did they agree with each other about whether the brand was named? All three agreed - either all yes or all no - 61% of the time. In the remaining 39%, the same question asked three times in the same week produced a mixed result.

Sit with the implication. For roughly two prompt-weeks in every five, a single manual check would have recorded something that a second check minutes later would have contradicted. Not a subtle difference in phrasing or ordering. A different answer to "are we in it or not".

One prompt. One week. Fifteen checks.Filled = brand named. Hollow = brand absent. Every cell is the same question.ChatGPT2 of 5Gemini3 of 5Perplexity1 of 5The one-check problemCheck the top-left cell and you report "we are visible in AI search".Check the bottom-left cell and you report "we are invisible".Both statements come from the same question, in the same week, about the same brand.Only the rate across all fifteen cells - 6 of 15, or 40% - is a number worth reporting.

The correct response to this is not to find a better tool. Every tool in this space is sampling the same unstable surface, and the good ones are honest that their numbers are rates rather than readings. The correct response is to change what you report. We no longer say "you are cited for this query". We say "you appeared in 6 of 15 checks on this query this week", and we show the denominator.

This also quietly invalidates a common piece of client theatre: the screenshot. A screenshot of your brand inside an AI answer proves that the answer happened once. It is a nice thing to have on a slide and it is not evidence of a position. We still take them. We no longer let them carry any argument.

Finding 2: The Three Engines Barely Agree

The second finding is that "AI search visibility" is not one thing, and averaging it into one number destroys the information you need.

Taking only the prompt-weeks where a brand appeared on at least one engine, here is how that appearance was distributed:

Appearance patternShare of prompt-weeks with any appearance
Named on all three engines18%
Named on exactly two engines28%
Named on exactly one engine54%

More than half the time, a brand's presence in AI search was a single-engine phenomenon. And the engines had clear personalities. Perplexity named a panel brand in a mean 34% of prompts, Gemini in 26%, ChatGPT in 21%.

The divergence in cited sources was even sharper than the divergence in named brands. For the same prompt, in the same week, we compared the set of domains each engine cited. The average overlap between any two engines was roughly 22%. Four fifths of the sources each engine leaned on were sources the others did not use.

The reasons are structural rather than mysterious, and they are consistent with what each system is built to do. Perplexity is retrieval-first and heavily recency-weighted, so it reaches for whatever is freshest and most clearly structured. Gemini leans on Google's index and entity understanding, favouring sources with established topical footing. ChatGPT blends a slower-moving internal representation with live retrieval, which makes it the most conservative and the most likely to name established incumbents regardless of what was published last month.

The practical consequence for a marketing team is that a single content asset will rarely satisfy all three, and chasing all three with one page is how budget disappears. The distinctions between the disciplines involved are worth understanding properly, which we lay out in SEO vs AEO vs GEO, and the format-level differences in what each engine actually lifts are covered in the content formats LLMs cite.

Finding 3: Citation Positions Decay

We did not set out to measure decay. It surfaced because the panel happened to include brands on very different publishing cadences, and by week seven the split was visible in the data.

Brands that published nothing new for six or more consecutive weeks saw their mean citation rate fall from 29% at the start of the window to 21% at the end - a relative decline of about 28%. Brands publishing on a regular cadence held flat or improved over the same period. The effect was sharpest on Perplexity and mildest on ChatGPT, which fits the recency-weighting differences in finding 2.

Two honest caveats before anyone treats this as proof that publishing volume drives citations. First, the brands that kept publishing differed from the ones that stopped in ways beyond publishing - budget, internal capacity, category competitiveness. Second, thirteen weeks is a short window to call a trend. What we can say with confidence is narrower but still useful: a citation position behaves nothing like a backlink. It is not an asset you acquire and hold. It is closer to a shelf position that gets re-contested every time a competitor publishes something fresher and cleaner than what you have.

That reframing has budget consequences. If you sell AI visibility work internally as a one-off project with a completion date, the numbers a quarter later will make you look wrong. We now scope it as maintained infrastructure, the same way we scope technical SEO, and we say so before the engagement starts rather than after.

Finding 4: What Got Cited Was Not What Ranked

Of the citations in the panel that pointed at a brand's own domain, only 38% pointed at the URL that ranked in Google's top three for the closest matching query. The other 62% pointed somewhere else on the site entirely.

The distribution by page type was the part that changed our content planning:

Page typeShare of owned-domain citationsWhy it wins
Comparison and "vs" pagesHighestExplicit criteria, structured trade-offs, direct answer to a comparison prompt
Pricing and cost pagesSecondNamed numbers and ranges the model cannot generate itself
Original data and methodology postsThirdA checkable claim that exists nowhere else
Glossary and definition pagesFourthClean, self-contained, extractable blocks
Core service and category pagesLow relative to investmentSays what every competitor's service page says
HomepageRare, and usually a bad signCited when the engine cannot find anything more specific

Look at the bottom two rows, because that is where most SEO budget goes. Core service pages carry the internal links, the conversion paths and the commercial intent, and they were cited far less than their share of investment would suggest. Not because they are bad pages, but because a language model can already produce a competent paragraph about what an SEO service includes. There is no reason to cite you for material it can generate.

The pages that earned citations all shared one property: each contained something specific and checkable that could not have been synthesised from general knowledge. A measured number. A named range. A documented process with real steps. An explicit criterion. Where that claim sat in a self-contained block under a heading phrased like the question, it got lifted. Where it was buried inside a narrative, it did not.

This is the same conclusion we reached from the other direction when we rebuilt a client blog around the questions buyers actually ask, and it is why we now treat comparison and pricing pages as citation infrastructure rather than as bottom-of-funnel afterthoughts.

Finding 5: The Lag Between Change and Citation

Because the panel ran alongside live client work, we could timestamp interventions and look for the first citation that followed. This is the weakest part of the dataset methodologically - it is association, not a controlled experiment - and it is also the part clients ask about most, so here it is with the caveat attached.

Median days to first citation, by interventionObservational. Recorded alongside live client work, not a controlled test.day 0102030Third-party mention8 daysRestructure ranking page11 daysPublish a new page19 daysEntity and schema workno measurable citation inside 90 daysThe fastest owned-media lever is restructuring a page that already ranks - not writing a new one.Schema is hygiene with a longer horizon. It is not a quarter-scale citation lever, and selling it as one will cost you trust.

Three things follow from this that we now build into every plan.

Restructuring beats publishing. A page that already ranks has cleared the retrieval hurdle. Hoisting its answer into a self-contained block under a question-shaped heading is a two-hour job with a median 11-day feedback loop. Writing a new page is a two-week job with a 19-day feedback loop after that. When a client wants movement inside a quarter, we now spend the first six weeks entirely on restructuring existing pages, and it is not close.

Third-party surfaces are the fastest of all. An eight-day median is faster than anything you can do on your own domain, which leads directly to the next finding.

Schema did nothing measurable in 90 days. This deserves care rather than a headline. It does not mean structured data is worthless - it is genuinely load-bearing for entity disambiguation and for other search surfaces, and we continue to ship it. It means schema is not a lever that produces citation movement inside a quarter, and if you have promised a client that it will, the data will not rescue you. The longer-horizon case for entity work is real and we make it in entity SEO and the knowledge graph.

Finding 6: Most Citations Are Not Yours to Control

63% of every citation recorded across the panel pointed at a domain the brand does not own. Owned pages accounted for the remaining 37%.

The off-domain citations concentrated in a predictable set of surfaces: community threads, especially Reddit; review sites and category listicles; YouTube; news and press coverage; and professional profiles. The pattern was strongest on shortlist prompts - "who are the best X in India" - where engines overwhelmingly reached for third-party lists rather than for any vendor's own page. Which is exactly what a sensible answering system should do, since a vendor claiming to be the best vendor is not evidence of anything.

This reorders the work in an uncomfortable way. A brand can have immaculate on-page structure and still be nearly invisible on the prompts that matter commercially, because the prompts that matter commercially are the ones the engines answer from third-party sources. The lever there is not another blog post. It is being present, substantively and honestly, on the surfaces the engines already trust - which is community participation, review presence and earned coverage rather than content production. That is why we treat Reddit as an SEO surface and why digital PR has stopped being a brand-awareness line item in our plans and become a citation-supply line item.

It also raises the crawler question, which several clients asked during the window. If third-party surfaces do most of the work, does blocking AI crawlers on your own site really cost you anything? Our answer, and the reasoning behind it, is in should you block GPTBot, ClaudeBot and PerplexityBot. The short version is that blocking removes you from the 37% you can control while doing nothing about the 63% you cannot.

Finding 7: Citations Do Not Send Traffic, and That Is Fine

The last finding is the one most likely to get an AI visibility programme cancelled if it is discovered in the wrong order.

Among panel brands whose citation rate rose by more than ten percentage points across the window, sessions to the specific cited URLs moved by less than 3%. That is inside the noise. If the only thing you measure is traffic landing on the cited page, the correct conclusion from our data is that the whole channel does nothing.

Over the same period, branded search impressions in Search Console for those brands rose a median 14%.

The mechanism is not mysterious. A buyer reads an answer, sees a brand named, does not click the citation because the answer already satisfied the question, and searches the brand name later when they are ready to evaluate. The citation did consideration-set work. It did not do acquisition work, and it was never going to.

This is the same measurement failure we wrote about in what zero-click search cost us and won us, arriving from a different direction. The instrumentation is built to count sessions, so value that does not arrive as a session is invisible to it, and invisible value gets defunded. Fixing that is a reporting problem before it is a marketing problem - which is also why we ended up rewriting our client reports entirely, a decision we documented in the metrics we stopped reporting to clients.

What We Report Now

Four numbers, monthly, against a frozen prompt panel. We resisted adding a fifth, because the underlying measurement is imprecise enough that more numbers manufacture false confidence rather than better decisions.

  • Citation rate, with its denominator. The share of checks in which the brand was named, always shown as a fraction of total checks - "142 of 558 checks, 25%" - never as a bare percentage and never as a binary yes.
  • Engine split. The same rate broken out by ChatGPT, Gemini and Perplexity separately. Given a 22% average source overlap between engines, a blended number is an average of three different situations and hides the one that needs work.
  • Owned versus earned share. What proportion of citations pointed at the brand's own domain versus a third-party surface. This is the number that tells you whether the next quarter is a content problem or a PR problem.
  • Branded search trend. Impressions and clicks on brand terms in Search Console, as the downstream signal. Not attribution - a directional read on whether the consideration-set work is landing.

A fifth thing we track but do not score: description accuracy. Roughly one answer in nine that named a panel brand described it inaccurately - wrong category, outdated positioning, a service the brand does not offer, or in two memorable cases a merger that never happened. That is not a visibility problem, it is a brand-safety problem, and it needs a different fix. Correcting it means correcting the sources the engines are reading, which is slow, unglamorous work on third-party surfaces and structured data.

What We Would Do Differently

If we ran this again, four changes.

More repeats, fewer prompts. Three repeats exposed the variance problem but were not enough to measure it precisely. We would rather have 30 prompts sampled five times than 62 sampled three times. The prompt list feels like the valuable asset when you are designing the panel. The sampling depth is what actually determines whether your numbers survive scrutiny.

Log the answer text, not just the flags. Our four-field schema kept collection sustainable, which was the right call, but it means we cannot now go back and ask questions about how brands were described, only whether they were named. The compromise we would make next time is storing the full answer text without scoring it, so the analysis options stay open.

Track engine versions. We know at least two engines changed during the window. We do not know precisely when, so we cannot separate platform movement from brand movement in the weekly series. Even a crude log of noticed behaviour changes would have made the trend lines interpretable.

Include a control group. Every brand in the panel was an active client receiving work. Without brands receiving no work at all, we cannot say how much of the movement was ours. This is the biggest weakness in the design and the easiest to fix.

We will re-run the panel with these changes and publish the second window. The AI search statistics reference is where we maintain the numbers between studies, and the definitions for the terminology used here live in the digital marketing glossary.

What To Do With This

If you take one thing from 7,200 answers, take the sampling rule. Almost every argument about AI visibility - between agencies and clients, between tools, between practitioners on LinkedIn - is really an argument between people who checked once and people who checked repeatedly, and the person who checked repeatedly is right regardless of which side of the argument they landed on.

So before you optimise anything: freeze a prompt list, decide how many times you will check, and record the denominator. That single discipline will tell you more in six weeks than any amount of tactical work performed against a number you cannot trust. The mechanics of setting that up, including the four viable collection methods and their trade-offs, are in our AI citation tracking playbook, and the cross-sectional companion to this study - what we found when we ran the same lens across 50 D2C brands rather than 14 clients - is in we audited 50 D2C brands for AI visibility.

Then work in this order, because the lag data says so: restructure the pages that already rank, go earn presence on the third-party surfaces the engines already trust, and only then write new pages. Reversing that order is the most common and most expensive sequencing mistake we see.

If you would rather not build the panel yourself, that measurement work is part of what our AI SEO team does, and it sits alongside the answer engine optimisation work that acts on the results. Tell us which questions your buyers ask and we will run them across the engines, hand back the citation rate with its denominator, and tell you honestly whether the gap is a content problem or a PR problem.

Aditya Kathotia

Aditya Kathotia

Founder & CEO

CEO of Nico Digital and founder of Digital Polo, Aditya Kathotia is a trailblazer in digital marketing. He's powered 500+ brands through transformative strategies, enabling clients worldwide to grow revenue exponentially. Aditya's work has been featured on Entrepreneur, Economic Times, Hubspot, Business.com, Clutch, and more. Join Aditya Kathotia's orbit on LinkedIn to gain exclusive access to his treasure trove of niche-specific marketing secrets and insights.

Want to explore working together?

Let's talk about how we can grow your digital presence and increase inbound business.

WhatsApp