Issue guides
Every rule Glimana checks, explained: what it means, why it matters, how we detect it and how to fix it. The same text opens from “Why it matters” next to each issue in the panel.
Indexability 36
Can Google find, crawl and index the page?
- Crawl blockedThe audit could not produce a usable result — the home page returned an error to the crawler, the site is in development mode, every page is noindex, or no crawlable page was found. What each cause means and how to clear it.
- Page that gets clicks is blocked by robots.txtA page with search clicks is disallowed in robots.txt. Google cannot read it, cannot build a snippet and cannot see updates.
- Page that gets clicks is noindexA page that received clicks from Google in the last 28 days carries noindex. Google will drop it and the clicks disappear.
- Server error (5xx)The page answers with a server error. If it persists, Google drops the URL from its index and reduces the crawl rate for the host.
- Server returns 403/429 to the crawlerYour bot protection blocks or throttles Glimana's crawler on some pages, so nothing can be measured there. Check whether Googlebot is affected too.
- The whole site is noindexEvery crawled page carries noindex. Google drops the whole site from search results within days to weeks.
- Canonical is outside <head> (Google ignores it)The rel=canonical link is in the body, often because an element inside head closed it early. Google only reads canonicals from head.
- Canonical points to a noindex pageThe page says its main version is a page that asks not to be indexed. Point the canonical to an indexable page or remove the target's noindex.
- Canonical points to another domainThe page declares a canonical URL on a different domain, which tells Google the real version lives elsewhere. Unless this is deliberate syndication, point it back to this site.
- Canonical target returns an error or redirectsThe page's canonical points at a URL that is not a working 200 page, so Google cannot use it to choose the indexed version.
- Conflicting canonicals (different URLs)The page declares more than one canonical with different URLs, so Google ignores them all. Keep a single canonical.
- Google treats the page as a soft 404The page returns 200 but Google reads its content as "not found" and does not index it. Fix the content or return a real 404/410.
- JavaScript changes the canonical to a different URLThe canonical in the raw HTML and the one after JavaScript runs point to different URLs. Google sees two signals and the outcome is unpredictable.
- Page did not respond (timeout or connection error)The crawler could not get a response at all. Google treats this like a server error; if it persists, the crawl rate is lowered.
- Page listed in the sitemap is noindexThe sitemap asks Google to index a page that its own robots meta or X-Robots-Tag says not to index. Decide which you mean and remove the other signal.
- Redirect loopThe URL redirects back into its own chain and never reaches a page. Fix the conflicting redirect rules.
- Sitemap URL returns an error (4xx/5xx)A URL in your sitemap answers with an error. Remove deleted pages from the sitemap or redirect them, and fix pages that should exist.
- URL returning 404 still shown in GoogleGoogle still shows these URLs in search results, but everyone who clicks sees an error page. Redirect them or let them drop.
- noindex in the raw HTML is removed by JavaScriptThe server sends a noindex tag and JavaScript removes it. Google may skip rendering when it sees noindex, so the page can stay out of the index.
- Crawled by Google but not indexedSearch Console reports that Google fetched these pages and chose not to index them. Google does not say why; look at whether the page offers an original answer.
- Google chose a different canonicalGoogle indexed a different URL than the canonical the page declares. Signals are split between two addresses until the two agree.
- Links use URL fragments (#/ orViews reached only through hash links are not separate URLs to Google. Give each view a real URL with the History API.
- No XML sitemap foundGlimana could not find a working sitemap for the site. Create one, declare it in robots.txt and submit it in Search Console.
- Possible soft 404 (empty or error page returning 200)The page returns 200 but looks like a "not found" page. Return a real 404/410, or add the missing content.
- Same page served at URLs that differ only in letter case/Page and /page both return 200 with identical content. Google treats them as different URLs and splits the signals.
- Sitemap URL redirectsA URL in your sitemap redirects somewhere else. List the final destination instead, so Google does not have to follow a hop for every entry.
- Sitemap file returns an errorA declared sitemap URL does not return 200 or is not valid XML. Search engines cannot read it.
- Canonical uses a relative URLA relative canonical works but is fragile — if a staging copy is crawled, it points at the wrong host. Use an absolute URL.
- Meta refresh redirectThe page redirects with a meta refresh tag instead of a server-side 301. Replace it.
- Non-canonical URL in the sitemapThe sitemap lists URLs whose canonical points elsewhere. Google treats the sitemap as a list of canonicals; conflicting entries weaken both signals.
- Page is noindex and also canonicalised to an indexable URLnoindex and rel=canonical contradict each other. Keep the one you mean.
- Redirect chain (2+ hops)A URL passes through two or more redirects before reaching a page. Point every redirect straight to the final destination.
- Canonical declared in both HTML and the HTTP headerThe page sends the same canonical twice, in the head and in a Link header. Harmless while they agree; consider keeping one.
- Internal links contain non-ASCII characters that aren't percent-encodedLinks with raw accented or non-Latin characters usually work, but Google recommends percent-encoding them. A tidy-up, not an emergency.
- No canonical tagThe page has no rel=canonical. That is allowed — Google picks the canonical itself — but a self-referencing canonical protects you against parameter and host duplicates.
- URLs use underscores instead of hyphensGoogle recommends hyphens to separate words in URLs. A readability convention with no meaningful ranking impact; change it only for new URLs or with redirects.
Content 19
Titles, descriptions, headings, text and duplicates.
- Missing title tagThe page has no title, so Google writes one from other text on the page. Give every page a unique, descriptive title.
- Content depends on JavaScript renderingHalf or more of the page's text exists only after JavaScript runs. Google renders, but crawlers that don't see an empty page. Render on the server.
- Near-duplicate pagesPages that are almost identical and get no impressions. No penalty, but signals split between copies. Merge, canonicalise or make them distinct.
- Same title on more than one pageSeveral pages share one title, so searchers cannot tell them apart. Give each page its own title, or consolidate the pages.
- Title contains a past yearA title such as "Best X 2024" on a page published in 2024 looks stale once the year changes. Update or remove the year.
- Title is only the site nameThe title says the brand and nothing about the page. Google rewrites such titles; write one that describes the page.
- Article page has no authorAn article without a named author lacks an E-E-A-T signal. Add the author's name, an author page and author.name in the schema.
- Article page has no dateAn article without a visible or structured date loses freshness signals and Discover eligibility. Add datePublished and dateModified and show the date.
- Missing H1The page has no H1 heading. Add one that states the page's main topic.
- Missing meta descriptionWithout a description, Google writes the snippet itself. Not a ranking factor, but the snippet you write is often used and affects clicks.
- Same meta description on more than one pageOne description on several pages means the snippet does not distinguish them. Write a description specific to each page.
- Meta description is too longGoogle truncates descriptions that do not fit; keep the key sentence under about 920 px (~155 characters).
- Missing HTML lang attributeGoogle works out the language from content and does not use lang; screen readers and browsers do. Add it.
- More than one H1Several H1s on a page. Google's John Mueller says you can rank fine with multiple H1s; this is informational.
- Site name appears more than once in the titleThe brand is repeated in the title, wasting the space searchers see. Keep it once.
- Thin contentFewer than 300 words and no Google impressions in the last 28 days. Word count is not a ranking factor; the question is whether the page has a purpose.
- Title is long enough to be truncated in search resultsGoogle has no title length limit; it only cuts what does not fit. Keep the important words first and the title under about 580 px.
- Title is very shortA title under 20 characters is not a problem in itself, but it rarely describes the page or matches a query.
- Very low text-to-HTML ratioGoogle does not measure text-to-HTML ratio. If the page is slow, the cause is its size, not the ratio.
Links 11
Internal and external links, anchors, depth and orphans.
- Broken internal link (4xx/5xx target)Your own pages link to addresses that return an error. Fix the link or redirect the target.
- Broken images, stylesheets or scriptsA file the page loads returns 404, fails DNS or has answered 5xx repeatedly. Visitors see broken images or an unstyled page, and Googlebot renders what it can load.
- Orphan page (no internal links)No page on the site links to this one. Google struggles to find it and reads the absence of links as a lack of importance. Link to it or remove it.
- Broken external linkA link on your page points to an outside address that no longer works. Update the link, point to an archived copy, or remove it.
- Internal link to a redirecting pageYour pages link to URLs that redirect. No value is lost, but every click and crawl pays an extra hop. Link to the final address.
- Internal links with no anchor textLinks with no text, alt or title give Google nothing to understand the target. Add descriptive anchor text, or alt text on image links.
- Nofollow on internal linksInternal links marked nofollow stop link value flowing to your own pages. Use robots rules for pages you do not want crawled.
- Page has no internal links (dead end)The page links to nothing else on the site, so visitors and crawlers stop there. Add links to related content, the category and the home page.
- Page is 4+ clicks from the homepageDeep pages are crawled less often and receive less internal link equity. Give important pages a shorter path through category, menu or related-content links.
- Internal links with generic anchor text"Click here", "read more" and the like say nothing about the target page. Replace them with words that describe where the link goes.
- Too many links on the page (300+)More than 300 links on one page dilute each link's value and usually mean oversized navigation blocks. Trim and simplify.
Performance 10
Speed, Core Web Vitals, compression and caching.
- HTML exceeds the 2 MB limitGooglebot processes only the first 2 MB of an HTML file. Text and links after the limit are never seen.
- Poor LCP in field data (p75 > 4 s)Real users wait more than 4 seconds for the largest element on the page. LCP is a Core Web Vital and part of Google's page-experience signal.
- Search engine crawler got 5xx from your serverBing's crawler reported server errors on recent days. Glimana's daily crawl may have missed the outage, but the search engine saw it.
- Poor CLS in field data (p75 > 0.25)The page visibly shifts while loading for real users. CLS is a Core Web Vital; the usual causes are images without dimensions, late-loading ads or embeds and web-font swaps.
- Poor INP in field data (p75 > 500 ms)Real users wait more than half a second for the page to respond to taps and clicks. INP replaced FID as a Core Web Vital in 2024.
- Response is not compressed (no gzip/br)The server sends HTML uncompressed. Every visitor and every crawler downloads several times more data than necessary.
- Slow server response (TTFB > 800 ms)The server takes more than 800 ms to start sending the page. TTFB is not a Core Web Vital, but it is the first clue when a page feels slow.
- Heavy inline script (> 200 KB)More than 200 KB of JavaScript is written directly into the HTML, so it cannot be cached and must be downloaded on every visit.
- No Cache-Control headerThe HTML response has no Cache-Control header, so browsers and CDNs have to guess how long they may keep it.
- Too many external scripts (25+)The page loads more than 25 external JavaScript files. Each one costs a request and often blocks rendering.
AI access 8
Whether AI search and training bots can read the site.
- robots.txt blocks AI search botsYour robots.txt disallows the bots that fetch pages to cite them in ChatGPT, Claude, Perplexity or Copilot answers. Blocked pages can never appear as a source.
- The server blocks or challenges automated visitorsA generic bot-protection rule answers non-browser requests with a block or verification page. AI crawlers cannot solve challenges, so they probably cannot read your pages.
- The server or WAF blocks AI bots (403 or a challenge page)Requests with an AI bot's user agent get a 403 or a verification page from your firewall. Search bots that cannot fetch a page cannot cite it.
- AI bots acting on user requests are blockedThis rule has been merged into "robots.txt blocks AI search bots"; user-triggered fetchers are now covered there.
- AI bots are served different contentA request with an AI bot's user agent gets a different title, much less text or no main content than a normal visitor. Usually a cache, security or A/B layer; it can also look like cloaking.
- AI training bots are blocked (may be intentional)robots.txt blocks bots that collect data for model training. That is a valid choice and does not affect whether you appear in AI answers.
- Gemini training and grounding are turned off (Google-Extended)robots.txt blocks Google-Extended. That affects Gemini training and grounding only, not Google Search, AI Overviews or AI Mode.
- These AI bots can still access the siteYour preference is to block all AI bots, but some can still fetch your pages. robots.txt handles most; user-triggered fetchers need a firewall rule.
Structured data 6
JSON-LD, schema types and rich results.
- JSON-LD parse errorA JSON-LD block on the page is not valid JSON, so search engines ignore all of it and the page loses its rich results.
- Article lacks max-image-preview:largeThe article does not allow large image previews, so Google Discover shows it with a small thumbnail or not at all.
- Structured data is missing properties Google requiresThe page has schema markup, but an item lacks properties Google lists as required, so it is not eligible for a rich result.
- Article-like page without Article schemaThe page looks like an article (publish date, blog path, article og:type) but has no Article/NewsArticle markup.
- No BreadcrumbList schemaA page at depth 2 or deeper has no breadcrumb markup. Breadcrumbs help Google understand the site structure and can replace the URL in the result.
- Structured data for a rich result Google no longer showsThe page has valid markup for a rich result type that Google has retired (FAQ, HowTo). Harmless, but it brings nothing anymore.
Security 5
HTTPS, HSTS and mixed content.
- Page is not on HTTPSThe page is served over plain HTTP. Browsers mark it "Not secure", and Google has preferred HTTPS URLs since 2014.
- Both the www and non-www versions return 200Your site answers on two hostnames without redirecting one to the other. Pick one canonical host and 301 the other.
- Mixed content (HTTP resource on an HTTPS page)An HTTPS page loads scripts, styles, images or iframes over HTTP. Browsers block or upgrade them, and the page can look broken.
- No HTTP → HTTPS redirectThe HTTP version of your site returns 200 instead of redirecting. Two copies of the site exist, and visitors on old links stay on the insecure one.
- No HSTS headerThe site does not send Strict-Transport-Security. It is a security best practice; it has no effect on search rankings.
Images 3
Alt text, dimensions and formats.
- Images missing alt textImages without alt text do not exist for screen-reader users and are harder for Google to understand. Describe meaningful images; mark decorative ones with an empty alt.
- Images without dimensions (CLS risk)Images without width and height make the page jump while loading. Add width/height attributes or a CSS aspect-ratio.
- Image URLs use older formats (jpg/png)WebP and AVIF are typically 25–35% smaller than JPEG/PNG at the same quality. Serve modern formats, unless your server already converts them under the same URL.
International 3
hreflang and language declarations.
- Invalid hreflang language/region codeAn hreflang value is not a valid ISO 639-1 language code with an optional ISO 3166-1 region. Google ignores invalid entries.
- hreflang is not reciprocalPage A lists B as a language version, but B does not list A back. Google only trusts hreflang pairs that confirm each other.
- hreflang present but no x-defaultThe page declares language versions but no x-default, so Google has no fallback for users who match none of them.
Mobile 1
Viewport and mobile rendering.
Accessibility 1
Lighthouse accessibility findings.