A two-person Shopify store came close to paying an agency about $3,000 to audit how the brand shows up in ChatGPT and Gemini, then $400 a month to keep watching. The owner had noticed that tools now do roughly the same thing for a fraction of that, and asked what an agency does that a tool cannot. Eighty-one replies later, the useful ones fit in three sentences, and they are the spine of this article.
Below: the quote against the tools that answer engines actually name, with prices copied from each vendor’s page this week; what happened when we asked the same question to ChatGPT and Google AI Mode; and the do-it-yourself version step by step, including where the built-in SeoGeoAgent in Orkas does the tedious parts.
What the $3,000 actually buys
Strip the proposal down and an agency audit and a software audit have the same shape: a set of prompts, a list of engines, a table of where you were mentioned or cited, the competitors that appeared instead, and a fix list. What varies is who wrote the prompts, how many engines were asked, and whether anyone executes the fix list.
The thread gave its own answer on price. One agency owner replied that they do the full audit for £450 and £50 a month. Another reader said an agency only earns a premium when the audit turns into work you cannot do yourself: correcting product data, improving source coverage, writing the comparison content that is missing, and measuring whether any of it moved qualified traffic or sales. A third suggested picking twenty to thirty questions real customers would ask before buying, and testing them by hand first.
So the question to put to any quote is not “what will you report” but “what will you change”. A report you could have generated for $299, or read for free in a tool’s trial, is not worth $3,000. Six weeks of page rewrites and outreach might be.
The price table
Prices below are copied from each vendor’s own page on 16 September 2026. Vendors re-tier without much ceremony; check before you buy.
| Option | What you pay | What you get | Where it stops |
|---|---|---|---|
| Agency quote from the thread | $3,000 audit, then $400 a month | An audit of how the brand shows up in ChatGPT and Gemini, plus ongoing monitoring | Ends at recommendations unless execution is written into the scope |
| Agency counter-quote, same thread | £450 audit, then £50 a month | “The full audit”, as described by the agency owner who posted it | Same deliverable shape; the spread between the two quotes is about six to one |
| CiteScore | Free 20-prompt scan; $299 one-time audit; tracking from $69 a month (50 prompts) or $149 a month (100 prompts); $1,500 “Fix Sprint” | The $299 audit runs 100 prompts across ChatGPT, Gemini, Claude and Perplexity, maps citations, audits 25 pages and returns a 90-day plan in about 30 minutes. The Fix Sprint adds six weeks of page rewrites, articles and citation outreach | Self-serve; execution is the $1,500 tier |
| Otterly.AI | $29 a month for 15 prompts; $189 for 100; $489 for 400; free trial | Daily tracking on ChatGPT, Google AI Overviews, Perplexity and Copilot. Gemini and Google AI Mode are an add-on from $9 a month, Claude from $29 | Monitoring only; you write the prompts and the fixes |
| AppearAI (Shopify App Store) | Free plan; $19.90 one-time full report | An AI visibility score, 20 buyer questions, competitor snapshots and content tips per question, from inside the Shopify admin | Shopify only; launched June 2026 with a single review at the time of writing |
| Profound | Free trial of 10 prompts on ChatGPT, once; everything else is custom-priced | Up to nine answer engines tracked on the enterprise plan, with SSO and SOC 2 | No public price; a third-party comparison lists it from $399 a month |
| Yourself, in Orkas | The app is free; model usage is metered, or bring your own API key (pricing) | Your own prompt list, the engines your buyers actually use, and a file you can re-run next month. Our reference run: 24 questions on two engines, one afternoon including reading every answer | Your time, and a sample rather than a measurement — see the limits at the end |
What happened when we asked the question itself
We put the thread’s question, almost word for word, to ChatGPT and to Google’s AI Mode in September 2026 and recorded what each cited.
ChatGPT retrieved and cited four Shopify official pages and nothing else: no tool names, no agencies. AI Mode answered “do it yourself” with a three-step frame — a free scan, a manual prompt check, then off-site trust signals — and cited a Yotpo audit guide, a seven-step checklist from an outreach agency, the Reddit thread itself, and three tool pages: CiteScore’s audit page, AppearAI’s App Store listing, and a third-party pricing comparison that named Otterly.
Two things follow. The tools that got named are the ones with a public price and a page a comparison article can link to. And the answer engines themselves recommend the do-it-yourself route for a two-person store, which is the route described next.
Doing it yourself in five steps
This is the same check as our five-step product-recommendation check, priced. Where a step is something the built-in SeoGeoAgent in Orkas does, it says so; nothing here requires it.
1. Write the questions, unbranded
Twenty to forty questions phrased the way a shopper types them before buying: best <product type> for <use>, <product type> that <constraint>, alternatives to <the obvious brand>. Keep a handful of brand-name questions too, but only as a control: asked “what is <your brand>”, a model names you by construction, so a branded row cannot measure visibility. Every priced option above runs some version of this list; a $299 audit whose prompts miss how your buyers actually ask is a $299 answer to the wrong question.
In Orkas, the agent’s geo-probe skill generates a candidate list from a crawl of your site and rejects candidates that carry your brand, run under three or over twelve words, or use an ambiguous acronym an engine would answer for the wrong industry. It names the reason for every rejection.
2. Run them on the engines your buyers use
ChatGPT and Google AI Mode at minimum; Perplexity and Claude if your category shows up there. For each answer record four things: mentioned, cited with a link, which competitors appeared and in what order, and every URL cited. Use ChatGPT’s temporary chat so earlier questions do not leak into later answers, and AI Mode’s “try without personalisation”. Pause every ten or so queries — Google shows a verification page if you do not.
In Orkas you can let an agent drive the shared browser tab in front of you, or paste the answers in; either way they are fed back to the probe for scoring. The metered cost is the model tokens spent reading about fifty answers.
3. Score it, and label what kind of number it is
Share of voice over the unbranded rows only. Citation rate — an actual link to your domain — separately, because it is the reliable signal. Competitor share. And a bucket for ambiguous mentions, where the brand word appears without any context that proves it is you rather than a homonym. The geo-probe score operation returns exactly these, and marks the result Measured only when every answer came from a retrieval-capable run; a mention from a model’s memory is not a citation and is not reported as one.
4. Fix the source, not the score
Two kinds of fix. On your own pages, the usual work: stale facts, missing comparison content, structured data, whether the page is quotable in three sentences. SeoGeoAgent’s diagnose mode scores a page for GEO citability alongside technical SEO and returns an action plan; apply mode makes the edits in your local repository after you confirm each one, then re-crawls the edited source.
Off your pages, the bigger lever. In our own runs, non-brand answers cited Reddit threads, vendor blogs, App Store listings and comparison articles far more often than store homepages. The action list there is to correct stale facts wherever they physically live, to add the comparison page nobody in your category has written, and to answer the actual threads your buyers are reading. Why one passage gets quoted and another does not is in our note on getting cited by ChatGPT.
5. Baseline now, re-run in four weeks
Save run one. Monitor mode keeps a baseline with the health and GEO scores, so run two can be compared instead of re-argued. Do not re-measure inside a week of a change; model-visible metadata lags by weeks. And treat every run as a sample: repeat a question on three days before believing a change.
The whole loop — tracking which questions you appear in and turning gaps into pages — lives in the search and AI answer visibility use case.
When paying is the right answer
- You cannot write the missing pages yourselves. A six-week execution package, or the £450 agency, is buying writing and outreach, not a dashboard. That can be worth it.
- You need daily tracking of a hundred prompts across four engines. A subscription in the $69 to $189 range is cheaper than the hours.
- You are on Shopify and want a twenty-question read in ten minutes. $19.90 is a reasonable price for a first look, as long as you then write your own list.
- Do not pay for monthly monitoring before you have changed anything. There is nothing for the number to move on. Nor for a dashboard whose prompts you did not write.
What none of the options give you
- Traffic. Google Search Console’s generative-AI report has no API and, structurally, reports impressions only; no option in the table can compute an AI click-through rate.
- A competitor’s real share. Every option is sampling answers, not measuring a population.
- A guarantee that a fix propagates. Several of the cited sources are pages you do not own.