A few weeks ago, a client asked me a question that stuck with me: "We rank first for our category on Google. Why did ChatGPT just recommend three of our competitors and not us?"
I didn't have a clean answer for them on the spot, and that bothered me. So I spent the next few weeks doing what I usually do when something bothers me: I ran the experiment myself, across a handful of client accounts, using nothing but a spreadsheet and a lot of patience. What I found is the reason I'm writing this.
Ranking well on Google and being visible in AI answers are no longer the same job. They used to overlap almost completely. Now they barely touch. And most brands still only have a system for the first one.
Two Different Games, Two Different Scoreboards
Traditional rank tracking answers one question: where does our URL land on the results page for a given keyword? It's tidy. Position 3. Position 7. You can chart it and hand it to your CFO without a footnote.
AI visibility doesn't work that way, because large language models aren't handing back a list of ten blue links. They're generating a fresh, synthesized answer every single time, built from whatever sources the model decided were trustworthy enough to pull from. Ask the same question twice in the same afternoon and you can get two different answers with two different sets of brands mentioned.
That's not a bug. It's just how these systems work. And it means the old mental model, "get to position one and stay there," doesn't transfer. What you're actually trying to understand is something closer to a batting average: across dozens of runs of similar questions, how often does your brand show up, and does it show up saying something true and positive?
Why This Isn't Optional Anymore
I get skepticism about this topic. Marketing has a long history of chasing shiny new metrics that turn out to be vanity numbers six months later. I've written before about why chasing raw traffic is already the wrong obsession in 2026, so I'm not about to tell you to chase a new number just because it's trendy.
But the buyer behavior data here is hard to wave away. This isn't a hypothetical shift happening someday. It's already happened for a meaningful share of your buyers.
That last stat is the one I keep coming back to when I talk to clients about budget. Most SEO budgets are still built almost entirely around owned-site optimization: better titles, better internal links, better page speed. All still worth doing. But if the majority of what an AI model cites when it talks about your category lives on someone else's domain, then a strategy that only touches your own site is optimizing for a shrinking slice of the pie.
What "Showing Up" Actually Means
There's a trap here worth naming early: getting mentioned and getting recommended are not the same outcome, and neither is the same as getting cited.
- Mentioned means the model says your brand's name somewhere in the answer.
- Cited means a page from your domain (or one that features you) is listed as a source the model pulled from.
- Recommended means the model actually tells the person to consider you, not just references that you exist.
You can be cited constantly and still lose the sale, because the model pulled a fact from your page but recommended a competitor in the actual answer. That gap between being sourced and being suggested is one of the more frustrating patterns brands run into once they start paying attention, and it's usually a sign that your content is technically accurate but not persuasive enough to earn the final nod.
Building a List of Prompts Worth Watching
You cannot track "everything people might ask an AI about your category." That's not a strategy, it's a hobby. Start with 20 to 30 prompts that map to actual buying moments, organized loosely into four buckets.
| Type | What it looks like | Why it matters |
|---|---|---|
| Evaluation | "Best [category] software for [specific use case]" | Shows whether you exist in the consideration set at all |
| Comparison | "[Your brand] vs. [competitor]" or "[Competitor] alternatives" | Reveals how you're framed head-to-head, including things you'd never say about yourself |
| Reputation | "Is [your brand] worth the price?" or "Is [your brand] reliable?" | Surfaces sentiment and outdated claims models may be repeating |
| Gap | Prompts where a specific competitor consistently owns the answer | Tells you exactly where to focus content and PR effort next |
A quick filter I use with clients: for every candidate prompt, ask "if someone typed this into an AI tool, would I actually want my brand in the answer, and would I be annoyed if it wasn't?" If the answer to both is yes, keep it. If you're only adding a prompt because it makes your numbers look better, like a branded query your own name is baked into, leave it out of your core tracking set. Save branded prompts for a separate cluster you only use for comparison and reputation checks.
A Simple Process You Can Run This Week
You don't need a platform subscription to start. You need a spreadsheet and about half an hour a week.
Run in a fresh session, every time
This one gets skipped constantly. If you're logged into your regular ChatGPT or Gemini account, the model has context on you and your past chats that a random buyer never will. Use a temporary or logged-out session so you're seeing what an actual prospect would see.
Test across more than one model
ChatGPT, Gemini, Perplexity, and Claude pull from different sources and weight them differently, so an answer that looks great in one can look thin in another. If you only have bandwidth for two, check which platforms your specific audience actually leans on before picking.
Don't panic over one bad week
A single week of weak visibility is closer to noise than signal. These systems are non-deterministic by design, and platforms push model updates that quietly shift source weighting without any announcement. Give it four consecutive weeks of the same pattern before you treat it as something worth fixing.
Turning the Data Into Actual Fixes
Once you've got a few weeks of data, three patterns are worth watching closely.
A steady decline in a specific cluster. If your visibility for "best [category] for enterprise teams" has been sliding for a month, something changed, either a competitor published new comparison content, or the piece you were relying on has gone stale. Refresh it with current data, or build the comparison page you don't have yet.
The same third-party sources keep showing up. If G2, a specific trade publication, or a particular YouTube channel keeps appearing across your evaluation prompts, that's the model telling you where it trusts information in your category. Check whether your listing there is current and complete before you spend a dollar anywhere else.
You're cited but not recommended. If a page of yours keeps showing up in the source list but a competitor gets the actual nod, that's usually a signal that your content answers the informational question but doesn't make the case for why someone should pick you. Fixing that is often more about adding proof (pricing clarity, specific outcomes, direct comparisons) than about adding more pages.
None of this happens overnight, and I'd be lying if I told you otherwise. Models take time to crawl and re-weight new content, sometimes weeks. The brands that get real value out of this aren't the ones checking their score daily. They're the ones tracking a stable set of prompts, reviewing trends monthly, and picking one deliberate fix at a time instead of trying to boil the ocean.
Where I'd Start If I Were You
If you do nothing else after reading this, do these three things:
- Build a list of 20 to 30 prompts split across evaluation, comparison, reputation, and gap questions.
- Run them in fresh sessions across at least two AI platforms and log what comes back.
- Check in monthly, not daily, and fix one cluster at a time.
The brands that will win the next few years of search aren't the ones with the flashiest AI visibility dashboard. They're the ones who noticed early that the game changed, and quietly went and did something about it while everyone else was still arguing about whether it mattered.
It does. My client's question is proof enough of that for me.

