glimanaDocs Open Glimana →

Issue guide · AI access

robots.txt blocks AI search bots

SeverityHigh
CategoryAI access
ScopeWhole site
EffortSmall
Verified byRe-crawl after you mark the task done
Rule keyai_search_bot_blocked

Short answer

Your robots.txt contains a Disallow rule that applies to one or more AI search or user bots: OAI-SearchBot (ChatGPT search), Claude-SearchBot, PerplexityBot, DuckAssistBot, or user-triggered fetchers such as ChatGPT-User and Claude-User. These bots read your pages to cite them in answers; if they can't read a page, they can't show it as a source. Remove the Disallow rule for these bots. If you only want to keep your content out of AI training, block the training bots (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot) instead and set your AI visibility preference to "Appear in answers, but block training".

Why it matters

This issue affects whether your site can be used as a source in AI assistants' answers. A growing share of research-type queries now ends in a chat answer with a handful of cited sources; a site whose search bots are blocked is invisible there, however well it ranks in Google. Many sites blocked these bots by copying a "block all AI" robots.txt snippet in 2023, before search and user bots were separate from training bots.

How Glimana detects it

Glimana parses your robots.txt and evaluates the rules for each bot's user agent exactly as the bot would, group by group (a User-agent: * block applies when no specific group matches). The rule fires when any search or user bot is disallowed from / or from the sampled pages, and the finding is confirmed, because robots.txt is a deterministic signal. The AI access page lists each bot with its verdict; the rule's severity follows your AI visibility preference (high when you want to appear in answers, information when you don't care, not reported when you chose to block all AI bots).

How to fix it

  1. Find the affected pages in Glimana. Open AI access in the panel. The bot rows marked blocked show the verdict (robots.txt = confirmed, live request = indicator); Show samples lists the URLs that were tested and the status each one returned.
  2. Apply the fix on your platform.

    Open robots.txt and look for the user-agent groups. Three shapes cause this:

    # 1. A blanket block that also catches search bots
    User-agent: *
    Disallow: /
    
    # 2. The search bot named explicitly
    User-agent: OAI-SearchBot
    Disallow: /
    
    # 3. A "block AI" list that mixes training and search bots
    User-agent: GPTBot
    User-agent: OAI-SearchBot
    User-agent: ChatGPT-User
    Disallow: /
    

    Keep the training bots in their own group if you want to block them, and let the search/user bots through:

    User-agent: GPTBot
    User-agent: ClaudeBot
    User-agent: Google-Extended
    User-agent: CCBot
    Disallow: /
    
    User-agent: OAI-SearchBot
    User-agent: ChatGPT-User
    User-agent: Claude-SearchBot
    User-agent: Claude-User
    User-agent: PerplexityBot
    Allow: /
    

    robots.txt is virtual unless a file exists at the web root. Edit it under your SEO plugin (Yoast: Tools › File editor; Rank Math: General Settings › Edit robots.txt) or upload a physical file. Some security plugins (Wordfence, All In One Security) and hosts (SiteGround, WP Engine) inject AI-bot blocks; check their settings too.

    Shopify generates robots.txt from templates/robots.txt.liquid. Edit it in the theme code editor; remove the group that disallows the search bots. Note that Shopify's own default blocks some bots on cart/checkout paths only, which is fine.

    Edit the static file or the route that serves it; remember that the first matching group wins per user agent, so a specific User-agent: OAI-SearchBot group overrides *. Deploy, then verify with curl https://example.com/robots.txt.

  3. Publish and clear caches. Save the file, then purge /robots.txt (and your sitemap URLs) at the CDN: Cloudflare and most CDNs cache them for up to a day. Confirm with curl -s https://your-domain/robots.txt that the live file shows the change.

  4. Verify in Glimana. Click Re-test on the AI access page, or wait for the weekly test (Sunday 03:00). The task closes when no search or user bot is disallowed.

How the fix is verified

Glimana re-tests on the next weekly run or when you click Re-test on the AI access page; the task closes when no search or user bot is disallowed.

FAQ

I blocked GPTBot. Does that stop ChatGPT from citing me?

No. GPTBot is OpenAI's training crawler; ChatGPT search uses OAI-SearchBot and ChatGPT-User. Blocking GPTBot alone keeps you out of training data while leaving you citable.

Why does this rule mention user bots too?

User bots (ChatGPT-User, Claude-User, Perplexity-User) fetch a page when a person asks about it in a chat. The former separate rule for them was merged into this one, because the fix and the effect are the same.

Sources