glimanaDocs Open Glimana →

Issue guide · Indexability

Crawl blocked

SeverityCritical
CategoryIndexability
ScopeWhole site
EffortSmall
Verified byRe-crawl after you mark the task done
Rule keycrawl_blocked

Short answer

This is not a rule from the catalogue but a state raised by the audit engine itself when the crawl cannot produce a usable result. Health and structure scores are not calculated while it is active, because there is nothing valid to measure. The title and fix you see in the panel depend on the cause: the home page returned an error to our crawler (most often HTTP 503 or 403 from bot protection), the site is marked as in development, every page is noindex, or no crawlable page was found. Clear the cause and the next crawl resumes measurement on its own.

Why it matters

Everything else in Glimana (issues, scores, tasks, impact reports) is built on the crawl. When the crawl cannot start, the panel would otherwise show a misleading "100% healthy" or a wall of false errors. The crawl-blocked state makes the gap explicit and tells you what to do.

How Glimana detects it

At the start of each audit Glimana fetches the home page and checks the result. If the home page returns an error status, the site is flagged as not live, no indexable page exists, or no crawlable page is found, the audit stops at that gate and records this issue with the matching variant. It is re-evaluated on every crawl.

The four causes and their fixes

Home page returns an error (HTTP 503 / 403)

Your server's bot or WAF protection (Cloudflare, LiteSpeed, Imunify, Wordfence and similar) answered the crawler's request with an error or a challenge page. No measurement is possible while the home page returns an error to this identity. Allow the Glimana crawler:

  • User-Agent: GlimanaBot/1.0 (+https://glimana.com/bot)
  • IP address: shown on glimana.com/bot and in the issue's evidence

Wordfence › Firewall › Blocking (remove the match) and Rate Limiting (allowlist). Hosts with Imunify360: whitelist the IP in the host panel or ask support. Cloudflare: Security › WAF › Custom rules › Skip for the user agent or IP; turn off Bot Fight Mode or add the exception in Super Bot Fight Mode.

Stores served directly by Shopify do not 503 crawlers; check a Cloudflare zone or firewall app in front of the store, or whether the store is password-protected (which returns the password page).

Nginx/Apache deny rules, limit_req, fail2ban and CDN bot management: add an exception for the user agent and IP. A 503 from the application itself usually means maintenance mode; take the site out of maintenance.

Site is in development mode

You marked this site as "not live yet". Because its pages can't be indexed, health and structure scores aren't calculated; this is expected. When the site goes live, remove this flag in the site settings and measurement starts by itself.

The whole site is noindex

If the site is live, remove the noindex now. When it comes from the X-Robots-Tag HTTP header, it is usually a line in .htaccess or the server configuration left over from staging; when it comes from the robots meta tag, check the CMS setting that discourages search engines (in WordPress: Settings → Reading). If the site is intentionally not launched yet, ignore this issue until launch day. The full guide is The whole site is noindex.

No crawlable page

The home page answered, but the crawler could not find any page it is allowed to fetch: robots.txt disallows everything for the crawler, every link is blocked or points off-site, or the home page is an empty shell rendered entirely by JavaScript that failed. Check robots.txt for a Disallow: / that applies to GlimanaBot or *, make sure the home page contains real links in the HTML, and that the main content renders without user interaction.

How the fix is verified

The state clears automatically when the next crawl gets a usable home page and at least one indexable, crawlable page. Use Crawl now on the Crawls page after fixing the cause instead of waiting for the daily run.

Sources