← All posts
Behind The Build

A Prompt Gives You an Answer. A System Gives You a Process.

A Prompt Gives You an Answer. A System Gives You a Process.

A prompt gives you an answer. A system gives you a process. Most “AI SEO” work right now is someone pasting a URL into ChatGPT and asking if it’s optimized. That’s a one-time opinion, not a repeatable process, and the two produce completely different kinds of proof. Here’s how that distinction played out concretely, building the actual tools this practice runs on.

The Gap I Ran Into First

The first time this came up was tracking whether client pages were actually getting cited by AI search, not just ranking for it. The obvious off-the-shelf option for that is licensed through my day job, not something available for Small Factory 5 client work. So instead of manually prompting each platform every time a client asked “are we showing up,” I built the measurement instead of asking the question.

What “Built, Not Prompted” Actually Means

A single prompt to ChatGPT asking “am I cited for X” is a sample size of one. AI answers vary run to run, and the same question can come back different two minutes apart. The Citability Index runs the same prompt three times per platform, per check, and averages the mention and citation rate across those three calls instead of trusting one pass. Citation counts more than mention in the score, weighted 60/40, because being the actual cited source means the page did the work, not just brand-name recognition. Every raw response gets saved before anything is scored, so if a client ever questions a number, the receipts exist.

None of this is about the specific weighting being sacred. It’s that the process is fixed and repeatable, so a score means the same thing this month as it did last month, instead of resetting every time someone asks a new question.

The Same Pattern Showed Up Twice More

Once one measurement replaced one prompt, the same gap kept showing up elsewhere.

Rank tracking is the obvious case. Asking an AI assistant “where do I rank” gets a guess pulled from stale training data, not a real position. Checking rank the shallow way, top 10 results only, misses almost everything: a page moving from position 45 to 22 is real, visible progress, but at that depth both read as “not found” until the page reaches page one. The actual tool checks the top 100 results every time, so movement shows up long before a page is anywhere near ranking. The same check also pulls whether Google’s AI Overview actually cites the client for that query: a direct yes or no, not an inference from watching traffic and guessing.

The second case was keyword search volume. A one-off prompt asking “how many people search for X” produces a plausible-sounding number with no way to check it. A real system checking real volume data caught something a prompt never would: a live, reproducible bug where the underlying data provider silently dropped real volume numbers for a keyword depending on which other keywords happened to share the same batch request. That’s not the kind of thing you notice by asking once and getting an answer that sounds fine. It’s the kind of thing you only catch by running the same check enough times, on enough keywords, to notice the pattern doesn’t hold. Every keyword that comes back empty now gets individually re-checked before it’s trusted, specifically because a one-time check would have quietly reported a wrong number as fact.

What This Means If You’re Evaluating Anyone’s AI-Search Work

This isn’t really a post about my own tools. It’s a diagnostic question worth asking about anyone doing AI-search work, including your own: if you asked the same “am I cited” or “where do I rank” question twice in a row, would you get the same answer? If not, what you have is a mood, not a measurement.

The five-lever framework this whole practice runs on treats Citability the same way: not a one-time check, but a lever you monitor, because a page’s citation status isn’t fixed the moment you happen to look at it.

Common Questions

What is a Citability Index / AI visibility score? A 0-100 score, per platform, measuring how often a brand gets mentioned or actually cited as a source in AI-generated answers to real buyer questions. Weighted 60/40 toward citation over mention, since being the cited source means the page itself did the work.

How is this different from traditional SEO rank tracking? Rank tracking measures position in a results list. This measures whether an AI system actually surfaces and cites the brand when answering a real question. Related, but not the same signal.

Which AI platforms does this actually check? ChatGPT, Perplexity, and Gemini, through each provider’s real developer API with grounding or web search turned on, not the consumer chat apps. Copilot isn’t in scope yet: there’s no clean API for it without a paid scraping service, so it stays out rather than get faked.

How often should citation tracking actually run? On a fixed, defined cadence, monthly by default, so the result is a trend line, not a snapshot. Checking once whenever you’re curious is exactly the prompt-versus-system gap this post is about.

Want This Run Against Your Own Pages?

That’s exactly what a page review checks. See how it works.


Written by David Cox, GEO consultant at Small Factory 5. More on how the toolkit behind this practice actually works.

Ranking isn't the same as getting cited.

If you're not sure whether AI search actually cites this page, that's exactly what a page review checks.

See how it works