Insights · AI & Search
AI & SearchMethodology

The GEO audit stack, in dependency order.

Most GEO advice starts at the content calendar, which is the fourth question. On the discovery queries most teams are buying GEO to win, the first question is whether your brand is in the answer pool at all, because when it isn't, nothing you publish on your own site changes that particular outcome. Here are the four checks, where the ordering holds, and where it breaks.

On this page

Two of the best studies published on AI search this year point at different layers of the same problem, and both are right. One tested 100 B2B products and found agents could answer a pricing question from the vendor's own site only 79% of the time. The other ran 270 queries and found one brand, Microsoft Entra ID, taking the top slot in 71% of Okta alternatives answers.

Those are different problems with different fixes and very different timelines, and the order you check them in decides whether your next six months of work does anything. Most GEO advice skips both and starts at the content calendar, which is the fourth question.

Two tests before you read the rest of this. Open ChatGPT and ask for alternatives to your closest competitor, ten times, in fresh chats, and count how often your own brand comes back. Then ask it to find your pricing, and note whether it quotes your site or a directory. The first number matters more, and most of this post is an argument about why.

01The two studies, and what each measures

Kevin Indig and David Kaufman gave agents three buyer tasks across 100 B2B products. Find pricing and features, find integrations, find security and compliance, five runs each, with no starting links supplied. In their July 2026 study, that 79% first-party answer rate on pricing sits against 93% for integrations and 92% for security. Pricing is the outlier, and it fails at the exact moment a buyer is comparing.

Separately, Ayomide Joseph ran 270 query variations across ChatGPT, Gemini and Perplexity, asking each for alternatives to three B2B SaaS products. In his experiment, published the same week and available in full here, each category consolidated around four or five brands. Okta's pool was Microsoft Entra ID, JumpCloud, OneLogin, Ping Identity and Auth0. Kustomer's was Zendesk, Gorgias, Freshdesk and Intercom, with Zendesk taking first position in 55% of those queries.

Read together, those two studies answer different questions about different kinds of query, and that distinction does more work than either study claims for it.

Joseph asked unbranded discovery questions: alternatives to a named competitor, where the engine assembles a consideration set from scratch and your brand has to earn its way in. Indig asked branded retrieval questions: find pricing for this specific product, where the name is already supplied and the only question is whether an agent can lift the fact off your site.

So the dependency is real, and it has a boundary worth stating plainly. On a discovery query, pool membership gates everything under it, because a flawless pricing page changes nothing when the question never surfaces your name. On a branded query, extractability matters whether or not you sit in any pool, because the buyer already arrived with your name from a rep, a referral, a review site or an ad. Indig's 100 products were never screened for pool membership, and pricing still failed 21% of the time.

Indig's own recommendation is to fix opacity and machine-readability first, and for the queries he measured, that's right. My argument is narrower than a disagreement with him. Most teams are buying GEO to win the queries Joseph measured, and pricing that work against Indig's ordering is how a quarter disappears into schema markup while the discovery problem sits untouched.

On a discovery query, find out whether you're in the answer before you optimize it.

02One. Are you in the answer at all?

This is the uncomfortable check, so it goes first.

Joseph's thesis, and he labels it a thesis rather than a finding, is that AI search behaves less like a ranking system and more like an eligibility filter. Once a category settles on its four or five brands, challengers don't argue their way in with on-page work. The pool was decided by years of accumulated reviews, listicles, analyst coverage and comparison content, and the engines are retrieving that consensus rather than forming a new one.

One retrieval detail underneath this matters more than it looks. On those 270 alternatives queries, ChatGPT searched the live web only 44% of the time, answering the rest from training data with no fresh retrieval step, against 74% for Gemini and 100% for Perplexity. That 44% belongs to this query type and shouldn't be read as ChatGPT's general retrieval rate, but for the query type in question it is the whole ballgame. When more than half the answers never touch the current web, publishing something new this quarter cannot reach them. You aren't editing a page the engine will re-read. You're waiting on a model that already formed its opinion of your category.

So a bad score here reads as a timeline, not a task list. Joseph puts the compounding work at 18 to 24 months.

03Two. Can the engine quote you?

This is where Indig says to start, and on the branded queries he measured, he's right. It is also the layer most GEO checklists open with regardless of query type, which is the part I'd argue with.

Indig's team found three ways a site fails an agent. Opacity, where the price isn't published. Machine-readability, where the price exists but sits inside JavaScript, a calculator widget, a screenshot or a PDF. And access friction, where the fetch fails outright.

Access errors showed up in just 7% of runs, so they're rare. But in pricing runs where they happened, third-party fallback jumped to 77%, against 17% without them. A rare failure with a severe consequence is exactly the kind that hides from a dashboard.

Publishing a price doesn't close this on its own. Even when the vendor showed a real numeric price on the page, agents still cited at least one third-party source in 18% of runs. The number was visible to a human reader and still not clean enough for a machine to lift and attribute.

The fixes are unglamorous and fast: publish real prices as server-rendered text, keep one canonical pricing URL, explain usage-based pricing in prose rather than only in a calculator, and check that robots.txt actually admits GPTBot, ClaudeBot and PerplexityBot rather than admitting Googlebot and quietly blocking the rest.

One caveat on the numbers here. Indig's post reports that schema.org Product and Offer markup moved a page from 73 to 93 on his readiness score, and reports separately that a B2B SaaS pricing page scored 73 and then 92.5 within the same hour, with that second jump attributed entirely to how the page was fetched. He never says those are the same page, so this isn't a contradiction so much as two similar-looking numbers with different causes sitting close together. Either way I'd treat schema markup as sound practice with an unproven effect size rather than a lever with a number on it. Worth noting the study is partly paywalled, though both of those readings sit above the wall.

04Three. Which sources decide question one?

Question one is not a fixed score. It's the output of whichever third-party sources the engines lean on for your category, which makes those sources the actual work.

They differ by engine, which is why a single-engine check misleads. In Joseph's data, about a third of ChatGPT's citations came from listicles, Gemini pulled more than half of its citations from vendor sites, and Perplexity mixed vendor pages with analyst content and product documentation at a citation density three to four times ChatGPT's. Pull the cited domains from your own question-one runs and you have a target list that took twenty minutes to build.

Indig's study shows where the fallback actually lands. Across 580 third-party pricing citations, 52% were editorial pages (blogs, comparison guides, explainers) and 46% were directories (G2, Capterra, Vendr and similar). Those two categories are the shadow pricing page most B2B vendors don't know they have, and only one of them is reachable by anything you write.

The obvious exploit is manufacturing that corroboration, and the market found it early. Gaetano DiNardi has documented GEO vendors buying brand mentions on spam sites and selling it as citation growth.

A different failure mode sits next to it, and it deserves separating rather than blurring together. Lily Ray's analysis of more than 220 sites using AI content-scaling platforms found strategies that work until they don't: sites publishing templated pages on their own domains at volume, then losing Google organic traffic when the algorithm caught up. Her patterns include comparison pages, self-promotional listicles and off-topic content published at scale. That is Google search traffic rather than AI citations, so it is not the same measurement. It is worth sitting with anyway, because those are the formats this discipline currently recommends, and the same page can win a citation and lose a ranking.

The signal these engines use is third-party corroboration from sources with real authority, which is the same signal Google has always used. The channel changed. The mechanism didn't.

05Four. What your own content is for

This is where most GEO advice lives, and it isn't wrong. It's fourth.

The standard playbook is a bottom-funnel keyword and prompt matrix, comparison and alternatives pages built to be extracted, and a refresh cadence that keeps them current. It works, and it compounds, on the condition that the three layers under it are functioning. A comparison page that gets cited assumes the engine can fetch it, and assumes your brand is in the consideration set the question produces.

The reason this layer keeps getting sold first is that it's the only one that produces a deliverable. An editorial calendar looks like work. A robots.txt correction takes an afternoon, and an honest answer to question one is sometimes just a number you don't like with an eighteen-month timeline attached. Neither of those fills a statement of work.

The volume here will look small for a while, and that's worth pricing in rather than panicking about. When Ahrefs published its own traffic data in June 2025, AI search visitors were 0.5% of visits and drove 12.1% of signups, a 23x conversion rate against organic search. That's one company reporting on itself, a year before the two studies above, so treat it as a directional reading rather than a benchmark. The shape is the useful part: very little traffic, converting far above its weight.

06Run the audit

FOUR CHECKS, ONE AFTERNOON RUN IN ORDER · A FAILURE UPSTREAM CHANGES WHAT DOWNSTREAM MEANS
1. Pool presence: "[category] alternatives" ×10, three engines~45 min
2. Extractability: "find pricing for [product]", note the cited domain~15 min
3. Source map: pull every domain cited in runs 1 and 2~30 min
4. Content gap: existing pages against bottom-funnel prompts~90 min
Then: robots.txt admits GPTBot, ClaudeBot, PerplexityBot?~5 min

Checks one and two together cost about an hour, so the ordering was never about which one to run. Run both. The order governs what you fund afterward, which is where it actually costs money to get wrong.

Read them against each other. A brand that shows up in eight of ten pool runs and still gets quoted from Capterra has a question-two problem worth fixing this month. A brand at zero has a question-one problem, and on discovery queries the markup work will not move it, though it still matters for every buyer who arrives already knowing the name.

Ten runs per engine is enough to tell whether you're in the pool or nowhere near it. It is not enough to measure a change. Ask the same engine the same question twice and the named set moves, so two readings can disagree without anything in the market having moved at all. Tracking movement across a quarter needs a fixed prompt set, a fixed number of repetitions, and the discipline to change nothing else between runs. I'm running exactly that against a real category now, and I'll publish the method with the numbers attached rather than ahead of them.

WHAT THIS DOESN'T SETTLE: READ BEFORE YOU QUOTE IT

This post is a synthesis of two studies plus an ordering argument, and the ordering argument is mine, not theirs. Neither study set out to test which layer to fix first, so the dependency is an inference from what each one measured, and it's the part most likely to be wrong. The sharpest objection is the one I've tried to answer in the body rather than bury here: the two studies used different query types, Joseph testing unbranded discovery and Indig testing branded retrieval, and my ordering only holds for the first. Where a company already has branded demand arriving from referrals or a sales team, question two is the cheaper and faster win and the order should flip. Indig himself recommends starting there. Both studies are single reads of fast-moving systems, 100 products in one and three categories in the other, both from July 2026, in a market where retrieval behavior changes between model releases. The Ahrefs figure is one company reporting on itself a year earlier. And my own prompt-tracking study is still running: the numbers are not in yet, and when they land they will either support this ordering or complicate it. I'll say plainly which.

08 Who wrote this
[ FOUNDER
PHOTO ]
The practice, in one person

Marvin Viachica

I run Searchline Partners, a senior search and demand generation practice for B2B companies. I diagnose search and AI-visibility problems and then build the systems that fix them (measurement engines, content architectures, reporting), installed in the client's own stack. The person who writes the study is the person who runs it and the person you'd hire.

09 Get the next one

One email when the next study is done.

No drip sequence, no gated PDF. When a post with real numbers and a stated opinion ships, I send it once. That's the whole newsletter.

No spam · No sharing your address · ~1–2 emails a month, max
10 Related