Server returns 403/429 to the crawler
crawler_blocked_4xxShort answer
Some pages returned 403 (forbidden) or 429 (too many requests) to Glimana's crawler. The page exists but our crawler can't read it, so nothing can be measured on these pages. If the same block applies to Googlebot, the page may drop out of the index. Your bot protection (Cloudflare, LiteSpeed, Imunify, Wordfence, etc.) is blocking our crawler: a page that returns 403 can't be read, and 429 means "slow down". Allow our crawler's identity, and check whether the same block applies to Googlebot with the live test in Search Console's URL Inspection; bot protection often lets Googlebot through separately.
Why it matters
A 403 to a crawler means the page is invisible to that crawler; a 429 means it is being rate-limited and will be partially crawled. For Glimana it means no data on those pages; for search engines, if the same rule hits them, it means deindexing. WAFs usually treat verified Googlebot specially, so the two cases must be checked separately.
How Glimana detects it
The rule lists up to 500 crawled URLs with status 403 or 429, with the status for each. When the home page itself is blocked, the audit cannot start at all and the crawl blocked state is reported instead.
How to fix it
- Find the affected pages in Glimana. Open Issues › Server returns 403/429 to the crawler. This is a site-level check, so the single row carries the full evidence: the URLs, values or days involved. The task under Tasks shows the same list and the expected gain.
-
Apply the fix on your platform.
Glimana's crawler identifies itself as
GlimanaBot/1.0 (+https://glimana.com/bot)and crawls from a fixed IP address shown on the bot page. Allow that identity:Then open Search Console › URL Inspection on one of the affected pages and run Test live URL to confirm Googlebot gets a 200.
Wordfence › Firewall › Blocking: remove rules that match the user agent; Wordfence › Rate Limiting: add the crawler to "Allowlisted" or raise the limit for verified crawlers. Hosts with Imunify360/LiteSpeed: add the user agent or IP to the whitelist in the host panel, or ask support.
Shopify may throttle aggressive crawling with 429; Glimana respects crawl delay and retries. Persistent 403s come from a Cloudflare zone or firewall app in front of the store.
Cloudflare: Security › WAF › Custom rules › Skip for
http.user_agent contains "GlimanaBot"(or the IP); Nginxlimit_reqzones: add an exception or raise the burst; fail2ban: whitelist the IP. -
Publish and clear caches. Apply the server or CDN change (reload Nginx/Apache, save the Cloudflare rule), purge the cache, then check the live response with
curl -sI https://your-domain/page— the status line and headers you see there are exactly what Glimana and Googlebot get. - Verify in Glimana. Open the task under Tasks and click Mark as done. The affected pages are queued for a verification crawl within a few hours, and the task closes when the listed pages return 200 to the crawler. If the check still fails, the task returns to New with a note; when it passes, the task is listed under Resolved technical issues on Impact reports. To check sooner, use Re-crawl affected pages on the issue.
How the fix is verified
Re-crawl; the task closes when the listed pages return 200 to the crawler. You can also lower the crawl speed in Site settings › Crawl if 429 was the only problem.