Jason Burns / jasonburns.co.uk
Available - taking new work Contact

Updated

Updated What changed on
  • Added the overlap measurement showing that ranking on Google explains very little about ChatGPT citations.
  • Added the freshness evidence from 17 million citations.
Diagnostic ChatGPT-specific

Why is my site not showing in ChatGPT? Five silent blockers.

The 30-second answer

Five silent blockers cover almost every case I diagnose. In order of frequency: GPTBot or OAI-SearchBot blocked (usually by accident), weak Bing index coverage, content too buried for a clean quote, no named-author or schema credibility, off-site authority gap. The first two account for half the cases I see and both take 30 seconds to verify with the free bot tester.

Pipeline ChatGPT actually uses: OpenAI GPTBot crawls Bing Bing indexes ChatGPT ChatGPT cites
ChatGPT visibility diagnostic - typical UK B2B site Each cell is one diagnostic check. Red = blocked; amber = present but weak; green = sorted. RED GPTBot in robots.txt Target: allowed Actual Blocked accidentally RED OAI-SearchBot access Target: allowed Actual Often missed in older robots.txt AMBER Bing Webmaster Tools Target: verified Actual Set up but no sitemap AMBER Bing index coverage Target: 90%+ Actual Patchy on key URLs RED Article + FAQ schema Target: on key pages Actual Not present AMBER Named author + Person schema Target: set up Actual Author byline only, no schema GREEN Content depth on priority queries Target: 1500+ words Actual Solid AMBER Off-site authority (links, mentions) Target: present Actual Branded only, no third-party Fix red cells first (blocked = invisible). Amber second (present but weak). Green stays maintained.
Traffic-light diagnostic grid. Most UK B2B sites score 1-2 green, 3-4 amber, 2-3 red. The reds are the quick wins.

What the citation studies show

12%

of links cited by ChatGPT, Gemini and Copilot appear in Google's top 10 for the same prompt. Perplexity is the outlier at closer to one in three.

Only 12% of AI Cited URLs Rank in Google's Top 10 for the Same Prompt, Ahrefs. Published 11 August 2025. Checked 19 August 2026.

How it was measured: 15,000 prompts. This is 2025 data and has not been re-run, so treat it as directional rather than current.

25.7%

fresher is the content AI assistants cite, compared with the content ranking in organic results.

New Study: AI Assistants Prefer to Cite "Fresher" Content, Ahrefs. Published 28 July 2025, revised 27 April 2026. Checked 19 August 2026.

How it was measured: 16.975 million cited URLs across ChatGPT, Perplexity, Gemini, Copilot, AI Overviews and organic Google. Average age of AI-cited pages was 1,064 days against 1,432 days for organic results. Published July 2025 and last revised April 2026, so it predates the 2026 shift Ahrefs measured in its own citation study.

1,432 days

is the average age of Google's top three AI Overview citations, which is the same as its organic results. The preference for fresh content does not hold for AI Overviews.

New Study: AI Assistants Prefer to Cite "Fresher" Content, Ahrefs. Published 28 July 2025, revised 27 April 2026. Checked 19 August 2026.

How it was measured: Same 17 million citation dataset. This is the counterweight to the headline freshness figure and is usually left out when that figure is quoted.

Before the five blockers: what ChatGPT actually does

ChatGPT answers from two surfaces. One is the training corpus, frozen at the model's cutoff; the other is live browsing through a Bing-backed retrieval layer. CXL's 2026 analysis put the live-search trigger rate at 34.5% of prompts, with the other two thirds answered from training data alone. The same study found 88.46% of ChatGPT citations come from pages already in a standard search index. If you are not in Bing's index, you are out of the running on the live surface, full stop.

The implication: there are two failure modes, not one. Either the model never saw you at training time (slow fix, six to eighteen months) or the live retrieval layer cannot reach or rank you (fast fix, days to weeks). The five blockers below are ordered by which one is most likely to be the actual problem on a typical UK B2B site.

The five silent blockers, ordered by frequency

1GPTBot or OAI-SearchBot blocked by accident

The single most common cause and the easiest to miss. Cloudflare's Super Bot Fight Mode defaults to blocking unknown AI crawlers on new sites. WordPress security plugins (Wordfence, Sucuri) add bot rules quietly after updates. Old robots.txt directives left behind by a privacy-focused agency. A staging robots.txt copied to production. None of these surface in any dashboard you check day-to-day.

Fix: test it. The free AI bot tester on this site checks all three OpenAI crawlers (GPTBot, OAI-SearchBot, ChatGPT-User) plus Google-Extended, ClaudeBot and PerplexityBot in one pass. Red rows are unblocked work; do those first.

2Bing index coverage is weak

ChatGPT's live browsing routes through Bing, not Google. A site that ranks page one in Google but has half its URLs missing from Bing is invisible on live queries regardless. Common causes: never verified in Bing Webmaster Tools, no Bing sitemap submission, the Bing crawler being deprioritised by a host that throttles unknown bots, or a JavaScript-heavy stack where Bing renders less reliably than Googlebot.

Fix sequence: verify the property in Bing Webmaster Tools, submit the XML sitemap, fix any crawl errors flagged there, plug IndexNow in if you publish more than once a fortnight. Then check coverage URL by URL on the priority pages. Effects on the live ChatGPT surface show within days once Bing crawls.

3The answer is buried, not extractable

The page exists, it is crawlable, it ranks. But the answer to the question someone might ask is buried three paragraphs into brand voice. ChatGPT cannot lift a clean quote so it picks a cleaner source. Kyenna's 2026 piece puts it well: a 2,000-word post that ranks on Google can still get zero AI citations because the answers are not in clearly labelled, extractable blocks.

Fix: lead each section with a 20-30 word direct answer in plain prose, then expand. Turn at least one H2 per priority page into a real question ("What does X mean for Y?" not "Our approach to X"). Add 5-8 FAQs at the bottom of pages where the topic naturally invites questions, and mark them up with FAQPage schema so the question/answer pairing is machine-readable.

4No named-author credibility signal

Anonymous brand voice. No author byline, no real bio, no credentials. When ranking and content depth look similar across competing pages, ChatGPT skews toward sources with a named expert behind them. Schema on its own is not the lever - Ahrefs' May 2026 controlled test found JSON-LD presence is hygiene, not a ranking input. The lever is the named expert; the schema just makes the signal readable.

Fix: name the author at the top of every priority page, write a credible bio that links to LinkedIn and any professional profile, add Person + Article schema referenced by @id from a single shared Person node in your homepage schema. If the author is also the business owner, link the Organization schema's founder to the same @id. Then run the page through the AI Schema Validator to check the author, publisher and sameAs chain.

5Off-site authority gap

Wikipedia entry absent, Wikidata entry absent, Knowledge Graph entity weak, few citations from credible publications. These count less for ChatGPT than for Perplexity (which is explicitly retrieval-heavy) but still tip the tiebreak when two pages look otherwise equivalent. The slowest fix on the list; do it after the four faster ones, not before.

Fix sequence: a verifiable Wikidata entity (start there, it is the lowest bar), a handful of editorially earned mentions on trade press your readers actually recognise, and a clean sameAs property in your Person/Organization schema pointing at every profile that proves you are who you say you are.

The freshness lever almost no one uses

From Typescape's 2026 citation study: pages updated within the last 30 days earned 3.2x more ChatGPT citations than stale pages on the same topic. Set yourself a quarterly review cycle on the ten pages that matter and the lift compounds.

If your evergreen post has not been touched in six months and a competitor updated theirs this quarter, the model now has a reason to prefer theirs. A genuine refresh - new figures, a new example, a clarification of something that has changed in the field - is enough. Editing one paragraph and pushing the dateModified is not enough; the model retrieval layer compares actual content, not just timestamps.

How to check whether ChatGPT is citing you today

Two checks worth doing now before you fix anything:

  • Brand check. Open ChatGPT (logged-out, web search on), ask "what does [your company] do" and "who is [your name] in [your field]". If the answer is wrong, vague or attributed to a competitor, the model does not have a clean entity record. That is a training-data and authority problem.
  • Category check. Pick three queries a buyer would type ("best X consultant in [city]", "how to fix Y", "is Z worth the money"). Ask ChatGPT with browsing enabled. Note which domains it cites. If the same three competitors keep appearing and you do not, those are the pages to study for content gaps.
  • Server log check. Grep for GPTBot, OAI-SearchBot and ChatGPT-User in your access logs. No hits means OpenAI's crawler is not reaching you - usually the Cloudflare or robots.txt issue in blocker 1. Hits but no citations means the crawler can read you but is choosing a different source - blockers 3 to 5.

What is not the problem (despite what people will tell you)

  • Buying ChatGPT Plus. Does nothing for visibility in other users' answers. ChatGPT's retrieval is the same for everyone.
  • Stuffing keywords for ChatGPT. The model does not weight keyword density the way 2015 Google did. Clear prose that actually answers the question beats keyword-stuffed prose every time.
  • Submitting to OpenAI directly. There is no submission form. The crawler discovers via the Bing index, links and crawl seed lists.
  • Adding llms.txt and waiting. OpenAI has not confirmed it as a retrieval signal. Publish one as hygiene if you want, but expect zero direct lift on ChatGPT specifically.
  • More content volume. Three sharp 1,500-word pages on questions your category asks beat thirty thin pages on broad terms. Citation rate scales with extractability and authority, not URL count.

The order to do the work in

  1. Run the bot test. If anything is red, fix the access layer first. Half the cases I see end here.
  2. Verify Bing Webmaster Tools, submit the sitemap, fix flagged crawl errors. Wait a week, re-check coverage.
  3. Pick the ten pages that matter (homepage, primary service, top blog posts). Rewrite each so the first 50 words answer the page's core question directly. Add 5-8 FAQs with schema where the topic supports them.
  4. Name the author on every priority page. Add Person schema referenced from the homepage. Link sameAs profiles that prove the person exists outside your domain.
  5. Refresh those ten pages on a quarterly cycle. Real updates, not just dateModified bumps.
  6. Off-site: Wikidata entry, three to five editorially earned mentions on trade publications. Slowest payoff, do last.

Primary sources

// questions I get

More on the diagnostic.

Can I check if ChatGPT can crawl my site?

Yes. The free AI Bot Robots.txt Tester checks GPTBot, OAI-SearchBot and ChatGPT-User in 30 seconds. Half the cases I see end at this step, with a Cloudflare default or a security plugin block the owner did not know they had.

My site ranks page one in Google but not in ChatGPT. Why?

ChatGPT browsing uses Bing, not Google. Weak Bing index coverage means ChatGPT cannot find you on live queries regardless of Google ranking. Verify the site in Bing Webmaster Tools, submit the Bing sitemap, fix flagged crawl errors. CXL's 2026 analysis found 88.46% of ChatGPT citations come from pages already in a standard search index.

How often does ChatGPT actually search the web?

About a third of the time. CXL's 2026 study put it at 34.5% of prompts triggering live web search; the other two thirds answer from training data. If your topic is covered in training, ChatGPT will not search for your page. The fix is publishing on questions the model did not have a good answer for at training time, then ensuring Bing has indexed your version.

Does paying for ChatGPT Plus help my visibility?

No. Subscriptions do nothing for your visibility in answers other users receive. ChatGPT visibility is determined by retrieval, citation and synthesis behaviour, not by account status.

How quickly can I expect to be picked up by ChatGPT?

Once Bing indexes the page and GPTBot can crawl, live browsing citations can appear within days. Training-data citations follow model release cycles and take months. The live surface is where the quick wins are.

Should I add a llms.txt to help ChatGPT?

OpenAI has not confirmed llms.txt as a retrieval signal for ChatGPT. Publishing one costs nothing and may help adjacent engines. Expecting it to lift ChatGPT visibility specifically would be optimistic. Do the five fixes above first.

Does updating a page help it get cited?

Typescape's 2026 data found pages updated within 30 days earned 3.2x more ChatGPT citations than stale pages on the same topic. Real updates - new figures, new examples - not just dateModified bumps.

Jason Burns, independent UK SEO, GEO and AI consultant
Written by

Jason Burns

Independent UK SEO, GEO and AI consultant. 17 years in search. Portfolio includes 3M, BlackRock, Unilever and E.ON. Owner of SEO Moves Ltd since 2014.

More about Jason →
// next step

Want me to run the full diagnostic?

Send three queries where ChatGPT should be mentioning you. I run them live, log what shows up and which of the five blockers applies. Free check, no follow-up sequence.

Send the queries