glimanaDocs Open Glimana →

Data

Crawls

Everything Glimana knows about your pages comes from its crawler. This page is the log: what ran, what it found, what it cost, and what the search engines' own crawlers saw on your server.

Crawl history with type, health, open issues, pages visited, duration and speed.
Crawl history with type, health, open issues, pages visited, duration and speed.

Crawl types

Type When What it fetches
Full Weekly, when the sitemap changes, or on demand Every URL reachable from the home page and the sitemap, up to the page limit.
Incremental Nightly Pages changed since the last crawl (by sitemap lastmod, ETag and content hash), new URLs, and the pages behind open tasks.
Single page Recrawl on a page, or after a task is marked done Only the listed URLs. Does not update the health score; it updates the affected issues.

A crawl that was stopped by the budget or by a server problem is listed too, with the reason, so the history never looks like a crawl silently vanished.

The re-analysed note under a crawl means the rules were run again on the same crawl data without fetching — for example after a sitemap change or a rule update. The crawl time stays the original one.

Starting a crawl

From the All sites row menu: Incremental crawl or Full crawl. From the Issues page: Re-crawl affected pages on an issue. From a page: Recrawl. On-demand crawls queue behind the nightly job if one is running.

Compare

Compare picks two crawls and lists every rule whose affected count changed, with the pages that became affected or clean. Use it after a release to see exactly what the deploy changed — it is the same data the Change column on the Issues page uses, for any two points in time.

Crawled pages

Crawled pages lists the URLs visited in a specific crawl with status, response time, size, depth, render mode and whether they came from the sitemap or from links. Filter by status code to find the 4xx pages of one crawl, or by source to see pages that exist only in the sitemap.

Search engine crawler · Bing

If Bing Webmaster is connected, the bottom table shows, per day, how many pages Bing crawled, how many it holds in its index, and the status codes it received: 2xx, 301, 4xx, 5xx, robots-blocked, errors. This is a third party hitting your server every day, so it sees intermittent 5xx errors that Glimana's once-a-night crawl can miss. A run of 5xx days raises the search engine server errors rule. Bing's statistics can stop a few days short; the table says so when they do.

Budget

Every completed crawl deducts its page count from the monthly budget shown under Plan & billing. Incremental crawls are small; full crawls on large sites are what spends it. If the budget is close to its end, a crawl is not cancelled — it finishes with the pages that remain and the history shows how many it fetched. When the budget is gone, new crawls wait for the 1st of the month; nothing already collected is deleted.