Since we started running AI optimization on our own site in June, the question I hear most from other founders and marketers is no longer whether this matters. It is what to buy. A tool market has exploded to fill that gap, most of it launched in the last eighteen months. This is a practitioner's map of the landscape, organized by category, with an opinion at the end about where the money is well spent and where it is theater. Prices are mid-2026 and move constantly, so treat them as signals, not quotes.
The category that matters: AI-visibility tracking
If AIO has a flagship tool category, this is it: software that tells you whether ChatGPT, Perplexity, Google's AI Overviews, Gemini, and Copilot are citing your brand, on which questions, and how accurately. The money agrees. Profound, the category leader, raised a $96 million Series C in February 2026 at a $1 billion valuation, led by Lightspeed, and says it serves more than 700 enterprises including Target, Walmart, and Figma. When a measurement tool becomes a unicorn in eighteen months, you are looking at a gold rush.
Here is the shortlist worth knowing, from enterprise down to budget, plus the one free option.
| Tool | Engines tracked | Entry price (mid-2026, verify live) | Best for |
|---|---|---|---|
| Profound | ChatGPT, Claude, Perplexity, Gemini, Google AIO/AI Mode | $99/mo Starter (ChatGPT only); $399/mo Growth is the real multi-engine entry | Enterprise, original research |
| Scrunch AI | ChatGPT, Perplexity, Google AIO, Copilot | ~$250/mo | Mid-market, compliance buyers |
| Peec AI | ChatGPT, Perplexity, Google AIO, others | ~€199/mo | Agencies, client reporting |
| Semrush AI Toolkit | ChatGPT, AI Overviews, AI Mode | $99/mo add-on | Existing Semrush users |
| Ahrefs Brand Radar | 6 engines incl. Copilot | from $199/engine (~$800+ all-in) | Existing Ahrefs users, data depth |
| Otterly.ai | Google AIO, ChatGPT, Perplexity, Copilot | from ~$29/mo | Solo or small budget |
| LLMrefs | all major engines | from $79/mo | Flat-rate prompt tracking |
| RankPrompt | ChatGPT, Perplexity, Google AI Mode, AI Overviews, Claude, Gemini, Grok | ~$89/mo (Pro) | Prompt-level tracking with an API you can automate |
| Bing Webmaster Tools (AI Performance) | Copilot / Bing only | free | Everyone, as a baseline |
At least twenty more exist (AthenaHQ, Nightwatch, Goodie, Bluefish, Knowatoa, and a new one seemingly every week). The category is crowded and consolidating at once. Do not agonize over the perfect pick, because of the problem in the next section.
The catch nobody prices in: these tools disagree with reality
Here is the opinion this whole piece is built around, and the part vendors gloss over. AI-visibility numbers are less reliable than the dashboards imply because what they measure is unstable.
Petra Labs ran 900 trials in early 2026, asking identical prompts across three ways of reaching ChatGPT: logged-in, logged-out, and the API. The same brand's visibility swung by up to 32 percentage points depending only on which surface was queried. One brand appeared in 15 to 18% of chat trials and zero API trials. Most tools sample one surface, usually the API, so a brand doing fine in the chat window can read as flatly invisible. Separate 2026 research found citations held for only about 33% of queries over a 28-day window, and that 87% of query-and-platform pairs fell below a basic stability threshold.
Our own weeks of measurement taught the same lesson from the other side of the glass. The purpose-built answer-engine modes we tried were enterprise-gated. Google's AI Overviews resisted measurement for weeks, serving bot walls and the occasional wrong-country result, then recovered. Perplexity, our most readable engine for most of that stretch, went dark behind a login wall in late July. The engines trade places without notice, so any reliability assumption about a specific engine ages in weeks. Any tool that hands you a single confident "share of voice" number across all engines is averaging over this mess.
Then, in late July, we put real money down and started a thirty-day head-to-head: two commercial trackers running against our own measurement pipeline. Day one was an education. One tool read our brand as a coworking-space company and proposed WeWork and Regus as our competitors. The other derived our positioning correctly but surfaced a subtler trap: engines will sometimes answer from training data without searching the live web at all, which turns "zero citations" into a number that means nothing. Our pipeline wasn't innocent either. It had quietly swapped the surface it measured, reporting classic search results where an AI answer had been the week before, with no error anywhere. And a per-query cost assumption we had budgeted on turned out to be wrong by 50x, caught only on the first live API call.
The discipline that separates a useful tool from an expensive screenshot: verify the instrument before you trust the number. Run one real query by hand before you model anything on a tool's output. Check which surface each number actually comes from. Treat a same-day repeat that flips from hit to miss as normal, because a citation is a distribution, not a screenshot. And if you can, run two cheap tools and make them argue; their disagreements catch setup defects that either one alone would have hidden, and where they agree, you know the finding is real.
Turn on the free tools first
Before you pay anyone, switch on the first-party tools, because the platforms have started reporting on themselves. Microsoft shipped an "AI Performance" report inside Bing Webmaster Tools in February 2026, the first native dashboard for AI citations, though it covers only Copilot and Bing. Google Search Console added generative-AI performance reports on June 3, 2026, showing impressions and pages for AI Overviews and AI Mode (not clicks), and only for some sites so far. Both are free, and neither needs a vendor. While you are in there, confirm your robots.txt lets the AI crawlers in, decide whether to allow or block Google-Extended, and wire up IndexNow so new pages get discovered fast. This plumbing is unglamorous, and it is where half of any real audit's value lives.
Optimization, content, and one thing that is theater
The rest of the stack helps you produce and structure content the engines can lift. Content optimizers like Surfer (around $89 per month), Clearscope (from $129), and MarketMuse (around $149) predate AI search and have repositioned around it. They are solid for writing answer-shaped, well-structured pages, which is one of the few levers that reliably moves citations. Schema markup is cheap and still worth doing for classic search and Bing, but the 2026 evidence on whether it moves AI citations is split: Ahrefs tracked 1,885 pages that added it and found no meaningful lift, while another vendor reported large gains. Do it for the old reasons, not the new ones.
Then there is llms.txt, the proposed file you place at your site root to feed AI models. Skip it. Ahrefs studied 137,000 domains in June 2026 and found that 97% of these files received zero requests, and that the AI retrieval bots the file exists to serve made up about 1% of the traffic it did get. GEO audit tools crawl it more than real AI assistants do. Google's John Mueller has called it a "temporary crutch" for coding tools, not a search feature. It is a cottage industry studying a file format before its audience exists.
If you would rather buy the outcome than the software, a services market has appeared too. The number of agencies selling AEO or GEO work has at least tripled since early 2025. Demand has outrun the supply of people who can deliver it, so vet capability, not the nav bar.
Current view, subject to change
If I were spending a marketer's budget today, I would turn on the two free native reports, pick one mid-priced tracker that reads the engines my buyers use (Scrunch, Peec, or Semrush's add-on if I already paid for the suite), and hold the enterprise platforms until I had proven the channel was worth it. That is what we are doing ourselves: a budget tracker and a prompt-level tracker with an API, running thirty days against our own pipeline, keep-or-cut decision at the end. I would treat every number as a range, sample it more than once, and spot-check by hand.
One more thing the head-to-head taught us that no dashboard will: the same tools that mismeasured us also diagnosed us. Both independently showed our visibility living in exactly one query frame, which told us the problem was authority, not indexing, and reordered our entire work plan. A tool you have verified is worth ten you have merely bought. What would change my mind: when one tool can reliably read every major engine across access surfaces and publish an independent accuracy audit, paying up for it becomes the obvious move. No one is there yet.
Final thoughts
This category is eighteen months old and priced like it is five years mature. The tools are improving fast, and the good ones earn their keep by making an invisible shift measurable. But remember what you are buying: an estimate of a moving target, not a truth. Measure often, trust the trend over the snapshot, and spend what you save on being the kind of source the engines want to cite in the first place.
Regards,
Charles Stack