What Is an AI Citation, and How Is It Different From a Mention?
An AI citation happens when an answer engine names your domain or links to one of your pages as the source behind a specific claim in its response. A mention is different and looser: the engine says your brand name in the answer text without pointing to anything you own. DeepSmith's guide to no-budget AI citation tracking treats these as two separate observations that need their own columns in a tracking sheet, not one blended score.
Tracking that distinction is the measurement layer underneath AEO (answer engine optimization) and GEO (generative engine optimization), the disciplines concerned with getting cited and recommended inside AI-generated answers. That's distinct from LLMO (large language model optimization), which covers the content and technical work that earns those citations, and from traditional SEO, which still measures rankings and clicks in classic search. All four feed the same reporting problem: knowing what actually counts as a win.
Most teams report a single "AI visibility" percentage that hides which of several distinct outcomes actually happened. DeepSmith recommends logging at least five states per prompt run: a direct URL citation, a domain-only citation with no specific page named, a brand mention with no citation at all, full absence, and a run that produced no answer surface at all (a timeout, a refusal, or an empty response). That fifth state matters because it isn't a loss, it's missing data, and folding it into your denominator quietly deflates every rate you calculate later.

How Many Prompts Do You Need to Track AI Citations Across 5 Engines?
Ten to twenty prompts a month is the practical ceiling for tracking AI citations by hand, according to CitedSpy's guide to manual AI citation tracking; past that volume, the time spent running and logging checks tends to outweigh what the data tells you. That ceiling holds across the five engines worth checking for most brands: ChatGPT, Perplexity, Gemini, Copilot, and Claude, which CitedSpy names specifically because each draws on different retrieval and ranking logic, so the same prompt can produce a citation on one engine and nothing on another.
DeepSmith's starter method uses a fixed pack of 10 prompts, small enough to rerun consistently without burning a full afternoon. citability.dev reports that a 20-query baseline sweep across engines can be completed in under an hour, which puts a first full measurement pass within reach of a single sitting. Checking fewer than five engines leaves you blind to gaps a wider sweep would catch.
How Do You Calculate Your AI Citation Rate Without Inflating It?
Your AI citation rate is the percentage of valid prompt runs where the engine's answer includes at least one qualifying source from your own domain: (valid runs with a qualifying citation ÷ total valid runs) × 100. DeepSmith's method for measuring AI citation rate treats the denominator as the part of the calculation most likely to get abused, and it's the part that quietly wrecks every number downstream of it.
The rule that trips up most manual trackers is what counts as a valid run. A prompt that times out, gets refused, or returns no citations for anyone because the engine couldn't ground an answer is not a zero for your brand, it's missing data. DeepSmith is explicit that timeouts and no-answer surfaces should be excluded from the denominator entirely rather than logged as failed attempts. Include them and your citation rate looks worse than it is, since every dead run drags the average down even though nobody else got cited either.
Get the denominator wrong and the rate looks precise while it measures the wrong thing, which is why DeepSmith treats invalid-run exclusion as a rule, not a judgment call. Ten prompts run across five engines produce 50 checks; if six of those return no answer surface, your real denominator is 44, not 50.
How Many Times Should You Rerun a Prompt Before Trusting the Result?
Three to five reruns is the practical floor for trusting a result, according to Formative Digital's research on measuring AI search citations, because large language models are non-deterministic: the same prompt sent twice to the same engine can return different citations, different sources, or no citation at all.
Run a prompt once and get cited, and you might conclude you're winning that query when the engine actually cites you on only one of five attempts. Run it once and get skipped, and you might log a gap that a second attempt would have closed. Formative Digital's May 2026 scrape, built on repeated sampling across 1,732 citations, only holds up as evidence of real per-engine divergence because it wasn't built on one-off checks.
How Do You Track AI Citations Without Paying for a Tool?
You can build a usable AI citation baseline with a browser, a spreadsheet, and about two hours, according to DeepSmith's no-budget tracking method. The workflow rests on a locked prompt pack, frozen test conditions, consistent logging, and repeat runs, and it works because no major answer engine exposes citation data through a direct API, which makes prompt testing the primary inference layer available, per genalphai.com's review of AI citation measurement approaches.
- Build a fixed pack of 10 prompts, DeepSmith's recommended starting size, phrased the way real buyers ask, and don't edit the wording once you start.
- Freeze test conditions: same browser, consistent logged-in or logged-out state, same location settings, so results are comparable week over week.
- Run each prompt through your target engines and log four outcomes per run: direct URL citation, domain-only citation, mention without citation, or absence.
- Mark timeouts, refusals, or no-answer responses as invalid runs, not zeros, and exclude them from your rate calculation.
- Rerun the full pack on the next cycle rather than editing prompts mid-cycle, so any movement you see reflects the engine, not a change in your own method.
genalphai.com recommends layering this manual pass with two supporting signals once it's running: crawler-log analysis to confirm bots like GPTBot and OAI-SearchBot are reaching your pages, and referrer-header matching to catch AI-driven traffic your analytics might otherwise miss.
Why Do ChatGPT and Perplexity Cite Different Sources for the Same Query?
ChatGPT and Perplexity draw from different retrieval systems, so they frequently cite different sources for the identical query, and the gap is large enough to change your read on where you actually stand. Formative Digital's May 2026 scrape analyzed 1,732 citations across engines and found that 83.7% of cited sources did not overlap between engines, meaning an aggregate "we got cited" number can be almost entirely driven by one engine while you're invisible everywhere else.
That's why ai-advisors.ai's tracking framework insists on per-engine reporting instead of one combined score: a brand can show a rising blended citation rate while losing ground in Perplexity and gaining it only in Copilot, and a blended number will never surface that swap. Engine-specific behavior also isn't consistent between engines, so what works to show up in Perplexity doesn't automatically transfer to ChatGPT.
What Does It Mean When Your URL Is Cited but Your Brand Isn't Named?
A source-only result occurs when an answer engine cites your URL to support a factual claim but never names your brand in the visible answer text, a different failure than being absent and a different outcome than being recommended. MentionWell's own scan of mentionwell.com, completed September 12, 2026, shows exactly this pattern across several SGE-related prompts, where owned citations appeared without a single brand mention.
| Prompt | Owned citations | Brand mentions | Source-only runs (from scan) |
|---|---|---|---|
| pricing for sge optimization services | 1 | 0 | 1 |
| what is sge seo | 2 | 0 | 2 |
| best sge optimization tools in 2026 | 1 | 0 | 1 |
| which blog engine is best for ai search citations | 0 | 0 | 0 |
| best llmo tools for chatgpt citations | 0 | 0 | 0 |
The first three rows are source-only wins: the page got pulled in as supporting evidence, but no engine described it as a recommendation, per the scan's own per-prompt data. That's meaningfully different from the last two rows, where the scan recorded zero owned citations and zero brand mentions across all eight completed engines, a true absence rather than an under-credited citation. Logging both as "not visible" would hide that the fix for the first three is a comparison page built to earn the recommendation, while the fix for the last two starts with getting cited at all.
This is one site's own scan data, not a cross-brand benchmark, but the pattern generalizes: any brand can be cited without being named, and a measurement system that only asks whether it was mentioned will miss it.
When Should You Move From Manual Tracking to a Paid AI Citation Tool?
Manual tracking stops paying off once you exceed roughly 10-20 prompts a month, per CitedSpy's guide to manual AI citation tracking, or once your reporting needs more engines and a tighter cadence than a spreadsheet can carry without someone spending hours on it weekly.
The arithmetic explains why. A responsible manual setup runs a fixed prompt pack across five engines with three to five reruns per prompt, following Formative Digital's reliability threshold. Even a modest 10-prompt pack at three reruns across five engines works out to 150 individual checks per cycle. Rahil Jain's own weekly loop, described on LinkedIn, runs 20 buyer prompts across three engines and calls it a 30-minute exercise; scale that to five engines with proper reruns and the time cost climbs fast.
That time cost is what paid tools and agency retainers are priced against. Discovered Labs' pricing guide puts full-service AEO agency retainers at $10,000 to $20,000 a month, and Gigawatt Group's breakdown of GEO and AEO pricing puts enterprise-tier retainers above $10,000 a month once a brand needs continuous multi-engine coverage and reporting. That's the budget a spreadsheet is competing against: the manual method above gets comparable baseline data without the retainer, at least until prompt volume or engine count outgrows what one person can log by hand.
Once measurement turns up a specific gap, closing it is a separate job from tracking it. The scan above found three prompts where mentionwell.com earned an owned citation with zero brand mentions, exactly the kind of source-only result a comparison page is built to fix, per the scan's own intervention recommendation for that pattern. Mentionwell's pipeline is built to publish that comparison page, then rerun the scan to confirm the brand mention shows up alongside the citation. Before scaling prompt volume further, it's also worth confirming your site clears the technical bar for being crawled and cited at all, which the AI visibility audit checklist walks through step by step.