FREE Website Analysis & SEO Audit
Instant website audit with SEO insights to help you rank higher and perform better—100% free.
Related Articles
The click math changed. Pew Research Center data published in July 2025 found that when Google displays an AI summary, the share of users who click a traditional search result drops from 15 percent to 8 percent, and only 1 percent click a link inside the summary itself.
For a low-consideration purchase, that shift mostly costs traffic.
For a service with a long sales cycle, it moves the entire shortlist stage into a place you cannot see. Buyers ask an assistant who does this kind of work, get three names, and start their research already narrowed.
Being one of those three names is a different job from ranking. Here is what actually influences it, and what has changed in the last few weeks.
First, check whether machines can reach your pages
This step gets skipped, and in 2026 it stops being safe to skip.
On July 1, 2025, Cloudflare became the first major infrastructure provider to block AI crawlers by default for newly onboarded domains, alongside a Pay Per Crawl beta that let site owners charge for access.
Because Cloudflare sits in front of a large share of the web, a lot of sites had their AI access changed by a setting nobody on the marketing side chose.
It moved again this month. On July 1, 2026, Cloudflare split AI crawlers into three categories — search, agent, and training — and announced that from September 15, 2026, training and agent crawlers will be blocked by default on pages displaying ads for new domains, while search crawlers stay allowed.
Crawlers that blend purposes get judged on all of their behavior. Pay Per Crawl is being replaced by a model that pays out when content appears in an answer rather than when a bot fetches a page.
Two details inside that change matter more than the headline.
Crawlers serving more than one purpose, including Googlebot, Applebot, and BingBot, fall under the most restrictive rule that applies to them, so a site that blocks training can lose search crawling as a side effect.
And existing customers can opt out before September 15 through their zone security settings, which preserves the current treatment.
The volume behind all of this is the reason the defaults keep moving: Cloudflare reported AI-related crawlers accounting for 52 percent of crawler requests on its network in June 2026, up from 22 percent in spring 2025 (Cloudflare, July 2026).
Two practical actions follow. Open your CDN dashboard and confirm which crawler categories are allowed on your own site and on client sites, because defaults have shifted twice in twelve months. Then read your robots.txt with fresh eyes.
What the robots.txt tokens actually control
Most of the confusion sits here, because the tokens do different jobs and the names look alike. Allowing one does not imply allowing another.
User-agent token | Operator | What allowing it affects |
|---|---|---|
| GPTBot | OpenAI | Content used to train models |
| OAI-SearchBot | OpenAI | Eligibility to appear in ChatGPT’s search results and citations |
| ChatGPT-User | OpenAI | Live fetches triggered when a user asks about your page |
| ClaudeBot | Anthropic | Content used to train models |
| Claude-SearchBot / Claude-User | Anthropic | Search indexing and live, user-triggered fetches |
| PerplexityBot / Perplexity-User | Perplexity | Search indexing and live, user-triggered fetches |
| Googlebot | Google Search indexing, which is what AI Overviews and AI Mode draw on | |
| Google-Extended | Gemini and Vertex AI use of your content. Separate from Search | |
| Applebot / Applebot-Extended | Apple | Siri and Spotlight indexing / Apple model training |
| CCBot | Common Crawl | Inclusion in an open dataset many models train on |
Vendors add and rename tokens regularly. Confirm the current list in each provider’s own documentation before you edit the file, and check server logs to see which agents are actually hitting the site.
One misconception worth clearing up while you are there. Google-Extended controls whether your content is used for Gemini and Vertex AI, and it is separate from Search indexing.
Blocking it does not remove you from Google’s AI summaries, which draw on the search index. Marketers have blocked the wrong token in both directions.
The reverse error is more expensive. Blocking Googlebot to keep a page out of AI Overviews removes the page from Google Search altogether.
Write a definition a model can lift
Answer engines assemble responses from passages, not from whole pages. The passage most likely to get quoted about your service is the one that defines it in a self-contained way.
Position on the page matters too. An analysis of citations across AI Mode, ChatGPT and Perplexity found roughly 44 percent of quoted passages came from the first third of a page (Search Engine Land, 2026). The opening is not a warm-up.
Most service pages fail this at the first sentence, because they open with a benefit claim addressed to a reader who already knows the context.
A model has no context. It needs a sentence that survives being copied out of the page with nothing above it.
A page selling IoT consulting works better when it opens by naming the service, the buyer, and the output: what the engagement covers, who commissions it, and what document or decision the client ends up holding.
Three sentences, no first-person, no adjectives doing the work of nouns.
What that looks like written out
Before. “We help ambitious operators unlock growth through tailored technology strategy.” A model cannot tell what is being sold, to whom, or what the buyer receives.
After, “IoT consulting is a paid engagement in which an external team assesses a company’s connected-device estate and specifies what to build, buy or replace.
It is usually commissioned by a VP of Engineering or an operations director at a manufacturer, utility or logistics business.
The engagement ends with an architecture document, a vendor shortlist and a costed implementation roadmap.”
The same shape works for a trade. “Water damage restoration is emergency work that extracts standing water, dries the structure to a measured moisture standard, and documents the loss for an insurance claim.
It is commissioned by homeowners and property managers, usually within hours of a burst pipe or a storm.
The job ends with a dried, moisture-tested property and a report an adjuster will accept.”
Nothing there is clever. All of it is liftable.
Test it directly. Copy your opening paragraph, paste it somewhere with no surrounding text, and read it as a stranger.
If it does not answer what this is and who it is for, rewrite it before touching anything else.
Answer the sub-questions on the same page
Buyers ask assistants about cost, timeline, scope, and risk. Pages that answer those questions in labeled sections get quoted for those specific queries.
Give each one a heading that matches how people ask it, then answer in the first two sentences underneath. Ranges beat silence.
A page saying a discovery engagement typically runs 4 to 6 weeks and produces an architecture document and a cost roadmap is quotable. A page saying every project is unique is not.
| Buyer question | Heading to use | What a quotable answer contains |
|---|---|---|
| What does it cost? | How much does [service] cost? | A range, the unit it is charged in, and the two or three variables that move it |
| How long does it take? | How long does [service] take? | A typical duration, the phases inside it, and what makes a project run long |
| What do I get? | What’s included in a [service] engagement | Deliverables named as nouns, plus what sits outside scope |
| Is it right for us? | When [service] is not the right fit | The situations where you decline the work, stated plainly |
A range is not a price. A range with its variables named beside it lets a buyer self-qualify, gives an assistant something it can attribute to you, and filters out the calls you did not want anyway.
The instinct on complex services is to withhold numbers until a sales call. That instinct now costs visibility at the exact stage where the shortlist forms.
Numbers need a date and a source
Models quote figures more readily than opinions, and they quote figures with attribution more readily than bare claims.
Two habits make a page citable. Attach the source and the year to every number in the text itself, not in a footnote. And date the page visibly, since freshness affects whether a model treats the figure as current.
The format is unglamorous: the figure, the organization that produced it, and the year, inside the sentence. A link is good practice for a human reader, but only the text reliably survives extraction into an answer.
Content built this way holds up better as a citation target.
A roundup of IoT trends that names the research firm and the year behind each figure gives an answer engine something it can attribute, while a page asserting rapid growth with no source gives it nothing to stand on.
The same applies to your own data.
Anonymized results from your own engagements are the one category of number no competitor can copy, and they are the reason a model has a reason to name you rather than a research firm.
Put a review date on evergreen pages and keep it. A figure that has aged two report cycles quietly undermines every other claim on the page.
The query class that decides shortlists
Assistants get asked to recommend, compare, and shortlist. Those prompts pull heavily from pages structured as comparisons: category overviews, provider lists, and versus pages.
For a service business, this creates two plays.
Publish honest comparison content in your own category, including where your approach fits poorly, because one-sided comparisons read as promotional to both humans and models.
And get accurately described in third-party comparisons, which means checking what directories and roundups currently say about you.
The evidence for that second play is stronger than most teams expect.
Analysis of AI brand mentions found that only about 13 percent originate on a brand’s own domain, making a brand roughly 6.5 times more likely to be named because of a third-party page than because of its own (AirOps, 2025).
In professional services categories, neutral third-party lists out-cite self-published lists by roughly 81 percent to 19 percent (Search Engine Land, 2026).
And in a study of 100 grounded local service queries, directory and ranking sites were cited in 78 percent of answers (Acromatico, 2026).
A workable order of operations
- Run your five highest-value buyer prompts and write down every third-party page the assistant cites.
- Check how you appear on each one: company name, category, service description, service area, and whether you appear at all.
- Correct what is wrong, claim what is unclaimed, and pitch inclusion where you are absent.
- Repeat quarterly, because those pages get rewritten and the rankings inside them move.
That second part is unglamorous list maintenance, and it moves the needle more than most on-page work.
Say your service name the same way everywhere
Models build associations between entities and topics from repeated, consistent phrasing across many sources.
If your site says strategic technology advisory, your LinkedIn says digital transformation consulting, and directory listings say IT consulting, you have diluted three descriptions into no clear association.
Pick the phrase buyers actually use, then use it identically on your site, your profiles, your directory entries, and in any bylined article your team publishes.
Consistency does not mean writing the same sentence forever. It means the noun phrase for what you sell does not change.
Write one description in a shared document, then align, in roughly the order these get cited: your main service page, your Google Business Profile category and description, your LinkedIn company page, the directories your category actually uses, the review platforms your buyers check, and every speaker or podcast bio your team submits.
Include a plainly worded page stating who you are, where you operate, what you do, and who you serve. It reads like boilerplate to a human and functions as a source of truth for a machine.
Measuring this without fooling yourself
Referral traffic from assistants will look tiny, and Pew’s 1 percent figure explains why. Judging the work by referral sessions alone will tell you it failed.
Small is not the same as worthless. Across studies published through 2026, visitors arriving from assistants convert at several times the rate of standard organic traffic.
Semrush put the cross-industry figure at 4.4 times (Semrush, 2026), and Adobe Analytics found AI-referred visitors to US retail sites converting 42 percent better than non-AI traffic in March 2026, a channel that had converted 38 percent worse a year earlier (Adobe Analytics, April 2026).
The volume is thin. The intent is not.
Expect your analytics to understate it further. A large share of assistant-referred sessions arrive with no referrer header and land in Direct, so build a custom channel group in GA4 that catches the known hostnames before you judge the numbers.
Four measurements that work better:
- Run a fixed set of 15 to 20 buyer prompts monthly across the assistants your clients use, and record which brands get named.
- Track brand mentions in AI answers over time rather than clicks from them.
- Add a self-reported attribution field to your forms, since buyers will tell you an assistant recommended you when asked directly.
- Watch branded search volume, which rises when assistants start naming you.
| Metric | How to capture it | Cadence | What good looks like |
|---|---|---|---|
| Prompt citation rate | Fixed set of 15–20 buyer prompts run across the assistants your buyers use | Monthly | Named in a rising share of prompts. A first appearance on a category prompt is a real result |
| Brand mentions in answers | Record every brand named per prompt, including yours | Monthly | You appear alongside the incumbents, not only when asked about by name |
| Self-reported attribution | An open “how did you hear about us” field on every form | Continuous | A steady trickle of replies naming an assistant |
| Branded search volume | Search Console impressions on brand terms | Monthly | Branded impressions rising with no campaign behind them |
| AI referral sessions | Custom GA4 channel group for known assistant hostnames | Monthly | Low volume, conversion rate well above non-branded organic |
Keep the prompt list stable so month-over-month comparison means something. Changing the prompts every cycle produces noise.
Three things not worth your time
Filling the page with FAQ schema for questions nobody asks. Structured data helps machines parse what is already there, and it does not create authority the content lacks.
Analyses of pages that added JSON-LD have not found a reliable lift in AI citations on that basis alone (Ahrefs, 2026).
Treating llms.txt as a solution. The proposal exists and costs nothing to add, but major providers have not confirmed they use it. Add it if you like, then get back to the content.
Writing a keyword-stuffed definition paragraph. Models are unusually good at ignoring text that reads like it was written for a machine, which is the reverse of how keyword stuffing worked in 2012.
A checklist you can run this week
- Confirm your CDN and robots.txt allow the crawler categories you actually want, given the September 15, 2026 default change.
- Rewrite the opening paragraph of your main service page so it stands alone.
- Add labeled sections answering cost, timeline, scope, and fit, with real ranges.
- Date every statistic and name its source in the sentence.
- Make your service description identical across site, profiles, and directories.
- Pull the third-party pages assistants cite for your top five buyer prompts, and correct how you appear on each.
- Add a self-reported attribution field to every form on the site.
- Set a fixed prompt list and record which brands get named this month.
Frequently asked questions
Is this different from SEO, or the same work with a new name?
The foundations overlap, since both need crawlable pages and content that answers a real question. The differences are real: passage-level self-containment, explicit sourcing, and consistent entity description matter more, while link volume matters less than it does for rankings.
Should we block AI crawlers to protect our content?
That decision depends on whether you sell content or sell services. A publisher losing traffic to summaries has a case for blocking. A service business whose buyers now shortlist through assistants is choosing not to appear at the moment of consideration.
Which pages should we fix first?
Start with the single page that would answer your highest-value buyer prompt, which is usually the main service page rather than the homepage. Rewrite its opening paragraph so it stands alone, add the cost, timeline, scope and fit sections, and date it. One page done properly beats a shallow pass across fifty, because the thing that gets quoted is a passage, not a domain. Depth per page also beats publishing volume here: twelve pages that each answer a real buyer question cleanly will out-cite two hundred that circle the topic.
Our pricing genuinely varies. Do we have to publish numbers?
You have to publish a range and the variables that move it, not a price list. “Discovery engagements run four to six weeks and $15,000 to $30,000, driven by the number of sites and systems in scope” is quotable, and it still leaves the actual quote to the sales call. “Every project is unique” is true, and it gives an assistant nothing to name you for.
Do we need schema markup for this?
Schema helps machines parse content that already exists; it does not manufacture authority. Organization, Service and LocalBusiness markup are worth having because they resolve identity questions cleanly. FAQPage markup is worth adding when the questions are ones buyers genuinely ask. Treat all of it as hygiene rather than strategy, since adding JSON-LD on its own has not been shown to lift AI citations (Ahrefs, 2026).
Our AI referral traffic is almost nothing in GA4. Is the channel real?
Probably, and the report is probably understating it. Many assistant-referred sessions arrive without a referrer header and get bucketed as Direct, so build a custom channel group for the known assistant hostnames before drawing conclusions. Then judge the channel by conversion rate rather than sessions: cross-industry data puts AI-referred conversion at roughly 4.4 times standard organic (Semrush, 2026).
An assistant describes our service incorrectly, or names competitors that are not really comparable. What do we do?
Treat it as a sourcing problem rather than a bug to report. Find the pages the assistant cites for that prompt and check what each one says about you. Most bad descriptions trace back to a stale directory entry, an old press release, or a third-party list that summarized a version of your site you have since rewritten. Correct the source, then re-run the prompt a few weeks later. No support ticket fixes this faster than fixing the record.
Does this apply to local service businesses, or only to long B2B sales cycles?
It applies to both, but the leverage sits in different places. Local queries lean heavily on directories, review platforms and map data; one analysis of grounded local service queries found directory and ranking sites cited in 78 percent of answers (Acromatico, 2026). For a local business, the Google Business Profile, the review corpus and listing accuracy do more work than website copy in the first ninety days. The page-level advice still holds, it just is not where the first gains come from.
Will the September 15, 2026 Cloudflare change affect a site we already run?
The new defaults apply to new domains onboarding to Cloudflare, and existing customers are being notified with an opt-out available in zone security settings before that date. The part worth checking either way is the multi-purpose crawler rule, because from that date a restrictive training policy can also catch crawlers that perform search, including Googlebot, Applebot and BingBot. Open the dashboard and set the three categories deliberately rather than inheriting whichever default arrives.
How long before this shows up in the pipeline?
Expect brand mentions in answers to move before anything reaches your CRM, and expect the CRM signal to arrive through self-reported attribution rather than a referral source. On a long sales cycle, assume two to three quarters before the connection is visible.
One test before you start
Take the three questions your best-fit buyers ask on a first call. Put each one to two different assistants and write down which companies get named. That list is your current competitive set at the shortlist stage, and it is usually different from the one in your quarterly ranking report.




