AI Comparison Queries: What ChatGPT and Google Actually Cite
Everyone tells you to optimise for AI comparison queries. Almost nobody has checked what those answers are actually made of.
So we ran five B2B software matchups through ChatGPT, Claude and Google AI Mode on 8 and 9 September 2026, all on the free consumer tier of each, and logged every source that appeared.
Then five more “best alternative to” prompts through ChatGPT alone, for twenty runs in total.
Ten of the fifteen comparison answers cited nothing at all. No links, no sources panel, no footnotes. Claude cited zero sources across all five comparisons. ChatGPT cited on one of five. Only Google AI Mode showed its working with any consistency.
That finding reverses the usual advice. Before you can worry about whether your comparison page gets cited, you have to accept that in most of these answers, nothing gets cited. The model is answering from what it already believes about your product.
What we ran
Five pairs, chosen because both vendors publish their own comparison page: HubSpot vs Salesforce, Zapier vs Make, Notion vs Confluence, Klaviyo vs Mailchimp, Ahrefs vs Semrush.
The prompt was the bare pair name, nothing else. No “which is better for a 20-person team”, no pricing qualifier. The point was to see what each engine invents when you give it nothing to work with, because that is what a buyer types.
Three engines: ChatGPT, Claude and Google AI Mode. Fifteen runs. We logged every visible citation, the structure of each answer, and any specific figure quoted.
We also ran five “best alternative to [vendor]” prompts through ChatGPT alone, as a control on the alternatives format that the comparison pages pillar treats as a separate page type with separate intent.
If you want the background on why one prompt turns into a dozen retrievals before any of this happens, that mechanic is covered in how search behaviour changed and in the AI visibility glossary.

Everything here ran on free tiers
This matters enough to state before any of the findings.
Every run used the free consumer tier: free ChatGPT, Claude on the free plan running Sonnet, Google AI Mode signed in but unpaid. Paid tiers get different models, different tool access and often more aggressive web search. A paid seat may well have produced citations where ours produced none, and we have not tested that.
The capture was uneven too, because the products do not treat their own output the same way.
ChatGPT copied cleanly with links intact, which is why its citations appear here as exact URLs.
Claude on the free plan would not let us copy the comparison tables at all. Only the short verdict section came out as text, so the tables in our dataset are screenshots. It also ran all five prompts in one thread rather than five fresh ones, which may have shaped retrieval after the first answer.
Google AI Mode split citations between inline numbers and a side panel, with video cards and a “Show all” control hiding the rest. Our counts are visible citations, not total retrievals, so the real number is higher.
One caveat we would fix first: the ChatGPT account was personal and had history. In the Semrush answer the model referred to the kind of SEO and backlink work the account holder had been doing before making its pick. That is memory shaping an answer meant to be generic. Run logged out.
Small sample, one snapshot, one country, free tiers, personalised account, uneven capture. What follows is a method worth copying and a set of findings worth checking, not a benchmark to plan a quarter around.
Finding 1: most answers cited nothing
Claude produced five confident comparison tables. Ease of use, pricing model, learning curve, integrations, a verdict. Not one source behind any of it. We have written before on where each model is actually useful for content work, and this is a reminder that answering well and sourcing well are separate skills.
ChatGPT cited on HubSpot vs Salesforce and nowhere else. The other four answers, including detailed pricing commentary, arrived with no links at all. On Zapier vs Make it hedged that pricing structures had changed and advised comparing current plans, which is a model declining to guess rather than a model checking.
Google AI Mode was the outlier, showing sources on four of five.
The implication is uncomfortable if your AI visibility plan is a page. When there is no retrieval step, no page wins, because no page was consulted. What decides the answer is whatever the model absorbed about your category during training, which is the aggregate of years of reviews, forum threads, documentation, video and third-party coverage. Our wider citation study found the same split by question type, and it is the reason showing up in ChatGPT is rarely a one-page fix.
A comparison page still earns its place. It just cannot be the whole plan.
Finding 2: when they cite, vendor pages win
This is the part that cuts against the received wisdom.
The standard line, backed by Rampiq’s finding that around 85% of citations for broad B2B category queries come from third parties, is that vendor content loses to review sites. On broad queries, it does.
On two-brand queries, it did not. Every one of the four AI Mode answers that cited anything included at least one vendor-owned source. Salesforce’s own comparison page for HubSpot vs Salesforce. Make’s and Zapier’s own vs posts for Zapier vs Make. Semrush’s own vs page and Ahrefs’ compare page for Ahrefs vs Semrush. Klaviyo’s own YouTube channel for Klaviyo vs Mailchimp.
ChatGPT went further. All five URLs in its single cited answer were vendor-owned, three of them HubSpot properties including HubSpot’s own comparison page. HubSpot effectively wrote the answer to a query about its biggest competitor.
The five alternatives runs behaved the same way. Where ChatGPT cited at all it cited vendor pricing pages, Ahrefs, Semrush and SE Ranking, rather than roundups. A useful check on our own alternatives posts: the format earns traffic and trust, but on these prompts it was not what got cited. More on those runs below.
So the narrow query really is the winnable one. That is the argument the comparison pages pillar is built on, and this is the first evidence we have gathered for it directly.

Finding 3: video is doing work you are not
YouTube showed up in three of the four AI Mode answers that cited anything, and in Zapier vs Make it was the single most common source type: three separate videos in the numbered citations, three more cards in the sources panel, from creators rather than the vendors.
Klaviyo vs Mailchimp pulled an agency video and Klaviyo’s own channel. HubSpot vs Salesforce pulled a consumer review channel.
There is no B2B SaaS marketing team on earth whose comparison strategy includes filming a fifteen-minute head-to-head walkthrough. There is an entire cottage industry of independent creators who do exactly that, and Google is reading their work back to your buyers.
This is not a suggestion to start a YouTube channel on the strength of one test. It is a flag that the sources shaping these answers sit outside the surfaces most teams measure. We made a similar argument about the channels B2B teams ignore, and this run puts video at the top of that list for comparison intent specifically.
One Reddit thread also surfaced, in Ahrefs vs Semrush, from an r/SEO discussion posted more than two years before the query. Worth reading next to what the G2 and Capterra data shows about which third parties actually carry weight.
Finding 4: all three engines produced the same shape
Run the three answers side by side and the structure is close to identical.
A one-line positioning split at the top. Then a feature table. Then a section per real dimension. Then an explicit “choose X if” and “choose Y if” block. Then, in AI Mode’s case, three follow-up questions.
Nobody prompted for that format. Three different systems from three different companies converged on it, which tells you it is the shape the underlying question wants.
That is a gift, and almost nobody uses it. The answer format is the page format, which is why the pillar’s page structure puts the table and the who-should-pick-what split where it does. If three independent systems all decide a comparison needs those blocks, a page missing any of them is missing something the machine expects to find.
The follow-up questions are the more useful half. AI Mode asked, unprompted, about team size, existing stack, expected volume, technical skill and whether the buyer had dedicated admin resources. Those are the fan-out sub-questions made visible. They are also, almost word for word, the qualifying questions a good sales rep asks on a first call.
Your comparison page should answer them before the call happens.

Finding 5: they quote prices, and not from your pricing page
Every engine put numbers in front of the buyer.
AI Mode gave Zapier’s entry tier at roughly twenty dollars a month for 750 tasks against Make at nine to twelve dollars for 10,000 credits, called Make roughly three times cheaper at entry, and put a 15,000-contact list at around $350 a month on Klaviyo against $230 on Mailchimp. ChatGPT listed four Salesforce tiers by name and price.
Some of that came from vendor pricing pages, including HubSpot’s and Salesforce’s. Some came from third-party comparison posts of uncertain vintage: a Knack guide with 2025 in its title, a Major post, an email software comparison site, a CRM switching blog and a Forbes comparison. None of those authors will update their numbers when you change your pricing.
Two things follow. Your pricing page needs to be crawlable, current and readable as text, because it is the only version of your pricing you control. And a stale third-party post is not a nuisance, it is a live misquote of your prices being delivered to buyers with an authoritative tone and no visible date.
The same logic applies to the competitor rows on your own comparison page, which is why the pillar argues for a quarterly audit and a verified date on the table.
Finding 6: the alternatives prompt returns the same duopoly
The five extra runs were meant to be a control. They turned into the sharpest result we got.
For each pair we asked ChatGPT for the best alternative to one half of it. Every single time, the top recommendation was the other half. Salesforce to HubSpot. Zapier to Make. Semrush to Ahrefs. Confluence to Notion. Mailchimp to Klaviyo. Five for five.

So the vs query and the alternatives query resolve to the same rivalry. If a buyer asks either way, they land on the same two names, which is a strong argument for treating your single biggest competitor as a page rather than one row in a roundup.
The shortlists themselves ran six to seven names deep. Salesforce alternatives brought Dynamics, Zoho, Pipedrive, Freshsales, Close and Monday. Zapier alternatives brought n8n, Pipedream, Power Automate, IFTTT and Workato. Being absent from that list is being absent from the shortlist, at the exact moment it forms.
The citation pattern held too. Four of the five alternatives answers cited nothing. The one that did, Semrush, cited three vendor pricing pages and no roundups.
One more wrinkle worth knowing. ChatGPT’s headline pick and its B2B SaaS pick sometimes differed. It named Klaviyo the best Mailchimp alternative overall, then ranked ActiveCampaign first once it applied a B2B lens it had introduced itself. The “best” in these answers is not stable, and it moves on context the model supplies rather than context the buyer gave.
What we would change on a comparison page tomorrow
Match the answer shape. Positioning line, table, dimension sections, choose-if split. All three engines produced it, so build to it.
Answer the follow-ups on the page. Team size, existing stack, volume, admin resources, technical skill. Those five appeared unprompted across the AI Mode runs. Each is an H2.
Publish the vs page anyway. Vendor pages were cited in every AI Mode answer that cited anything. The narrow query is winnable in a way the category query is not.
Make the pricing machine-readable. Text, not an image. Current. Dated. The same rule that keeps homepage content legible to a crawler applies here.
Build the page for your single biggest rival first. The alternatives prompt returned the same competitor as the vs prompt every time. One rivalry, two queries, one page worth building properly.
Look outside your own site. Video, forums and third-party comparison posts carried a large share of what got cited. You cannot publish your way to those, but you can be the source they use.
Stop treating rank as the metric. Ten of fifteen answers had no sources to rank among. Track what the answer says about you, not just whether you are in it. Our AI visibility checklist covers the wider audit, and HubSpot’s own traffic collapse is the cautionary tale for anyone still counting sessions. If the page is getting seen and still not converting, that is a different problem.

Run it yourself
Five pairs from your own category, three engines, fresh session each, bare prompt with nothing appended.
Log four things per run: which sources appeared, whether any vendor page appeared, the order of topics in the answer, and any figure quoted about price. That fourth column is the one that catches misquotes before a prospect does.
Twenty minutes a run. Repeat monthly and judge on the trend, not on any single answer. If the gaps you find turn into new sections, our brief generator turns them into something a writer can work from.
Add Perplexity if you have the time. It behaves differently again, and getting cited there is close to a separate discipline. If you are wondering whether an llms.txt file helps with any of this, the short answer is not yet.
One warning. Do not ask the model which sub-questions it generated. It will produce a plausible list on request, and that list is not a record of anything. Infer the fan-out from the shape of the answer instead, which is observable and defensible.
The raw data
Twenty runs, 8 and 9 September 2026, free consumer tiers. Fifteen comparison prompts across three engines, plus five “best alternative to” prompts on ChatGPT alone. Counts below are visible citations: numbered sources in the answer body plus cards in the sources panel. Google AI Mode hides the rest behind a “Show all” control, so its real totals are higher.
Comparison prompts
| Prompt | ChatGPT | Claude | Google AI Mode |
|---|---|---|---|
| HubSpot vs Salesforce | 5, all vendor-owned | 0 | 5 |
| Zapier vs Make | 0 | 0 | 10 |
| Notion vs Confluence | 0 | 0 | 0 |
| Klaviyo vs Mailchimp | 0 | 0 | 5 |
| Ahrefs vs Semrush | 0 | 0 | 4 |
Runs with at least one visible citation: 5 of 15. Runs with none: 10 of 15.
Alternatives prompts, ChatGPT only
| Prompt | Top pick | Visible citations |
|---|---|---|
| best alternative to Salesforce | HubSpot | 0 |
| best alternative to Zapier | Make | 0 |
| best alternative to Semrush | Ahrefs | 3, all vendor pricing pages |
| best alternative to Confluence | Notion | 0 |
| best alternative to Mailchimp | Klaviyo | 0 |
In all five the top pick was the other half of the original comparison pair.
Source types
Thirty visible citations across all twenty runs, after deduplicating cases where a source appeared both as a numbered citation and a panel card.
| Type | Count |
|---|---|
| Vendor-owned pages | 16 |
| YouTube videos | 6 |
| Third-party comparison posts | 6 |
| Mainstream press | 1 |
| Forum | 1 |
Every source we could capture a URL for
HubSpot vs Salesforce, ChatGPT: HubSpot’s comparison page, HubSpot pricing, HubSpot CRM, Salesforce sales pricing, Salesforce CRM.
HubSpot vs Salesforce, AI Mode: CRM Switch, Salesforce Trailhead, Salesforce’s own comparison page, a Forbes comparison, a Consumer Reviews video.
Zapier vs Make, AI Mode: Make’s vs post, Zapier’s vs post, Zapier’s app directory, Knack, Major, and three creator videos: one, two, three.
Klaviyo vs Mailchimp, AI Mode: Email Software Insights, an agency video, a Zapier comparison post, Klaviyo’s own YouTube channel, an Exposure Ninja video.
Ahrefs vs Semrush, AI Mode: Semrush’s own vs page, an r/SEO thread from April 2024, Ahrefs’ own compare page, an SE Ranking comparison post.
Alternatives control, ChatGPT: Ahrefs pricing, Semrush pricing, SE Ranking pricing.
Notion vs Confluence returned no citations on any engine.
Sources named without a link appeared as panel cards with no copyable URL
FAQ
Does this mean AI visibility work is pointless for comparison queries? No. It means the lever is different. In answers with no citations, what matters is the model’s existing impression of your product, which is built from the wider web over time rather than from a page you shipped last week.
Why did Claude cite nothing? It answered from model knowledge rather than searching. Our five prompts also shared one thread, which is a limitation of this run rather than a finding about Claude.
Is a comparison page still worth building? Yes. Vendor pages appeared in every Google AI Mode answer that cited anything, including both vendors’ own pages in one case. The narrow two-brand query is where vendor content competes best.
Did the alternatives prompts behave differently? Not in citation terms, four of five cited nothing. But in every case the top alternative ChatGPT named was the competitor from our original pair, so both query types funnel buyers to the same two vendors.
How big a sample is this? Small. Twenty runs, one snapshot, two days. Treat it as a method to copy rather than a benchmark to plan against.
Would longer prompts change the results? Almost certainly. Adding context does the engine’s fan-out work for it. We used bare prompts because that is what buyers type.
The buyer comparing you to a competitor is getting an answer either way. Two thirds of the time nobody is checking a source before it is written.
That is the job we do. See how we write comparison pages and how we approach AI visibility.
Think B2B is boring?
That's coz you are not subscribed to ours newsletter. We share a 100 bad days (of B2B marketing) made a 100 good stories. A 100 good stories make YOU interesting at parties.
Almost there. Check your inbox and click the confirmation link.
