<- Blog.SEO Basics

How to Read the Google Search Console Page Indexing Report

The Page Indexing report shows which URLs Google has indexed and which are excluded. Learn how to diagnose status reasons, soft 404s, and validate fixes.

Sep 23, 2026.8 min read
Updated on: Sep 23, 2026
How to Read the Google Search Console Page Indexing Report

Google cannot rank a web page that it has not stored in its search database. The Page Indexing report in Google Search Console is the definitive dashboard showing exactly which pages on your website Google has successfully indexed, which pages have been excluded, and the specific technical reasons behind those decisions.

For many site owners, opening this report creates immediate anxiety. Seeing thousands of URLs categorized under "Not indexed" often leads beginners to believe that their website is penalized or broken.

In reality, having unindexed URLs is standard behavior for any healthy website. Search engines are designed to exclude administrative paths, faceted navigation, filtered parameters, and obsolete pages. Understanding how to read this report allows you to separate normal crawler maintenance from critical technical errors that hurt your organic traffic.

What is the Page Indexing report?

The Page Indexing report[1]Source 1Google Search Central. Page Indexing Report Documentation.View source ↗, located under the Indexing section in the Search Console navigation sidebar, tracks the crawl and storage lifecycle of every URL Google has discovered on your domain.

The report serves three main functions:

  1. Index Health Tracking: It charts the total number of valid indexed URLs against unindexed URLs over a rolling 16-month timeline.
  2. Issue Categorization: It groups unindexed URLs into distinct technical reasons, distinguishing between intentional webmaster directives and unexpected crawler obstacles.
  3. Verification Management: It provides a direct interface to trigger the Validate Fix cycle after you repair server errors or update technical tags.

Indexed vs Not Indexed: Understanding the top-level chart

At the top of the report, a summary timeline presents two primary metrics: Indexed (displayed in green) and Not indexed (displayed in gray or amber).

Many webmasters mistakenly assume that a successful website should have zero unindexed pages. This assumption is incorrect. Large websites with ecommerce filtering, user profile pages, tag archives, or historical URL migrations will frequently have more unindexed URLs than indexed ones.

Your primary objective is not to force every discovered URL into Google's index. Your objective is to ensure that every canonical, high-value page you intended for searchers is stably counted under Indexed, while low-value, duplicate, or private pages remain properly excluded.

The status triage matrix: What needs fixing vs what to ignore

To avoid wasting time on harmless crawler notices, use this operational matrix to sort Search Console statuses into actionable categories.

Status Reason Typical Severity What It Means Recommended Action
Page with redirect Harmless The URL correctly forwards visitors and crawlers to another page via an HTTP 301 or 302 redirect. None required if the destination URL is indexed and the redirect is intentional.
Excluded by 'noindex' tag Harmless / Expected The page contains an explicit noindex directive instructing Google not to index it. None, unless you accidentally placed a noindex tag on an important article.
Not found (404) Harmless / Expected The URL returns an HTTP 404 status because the page was deleted or never existed. None required for dead content. Add a 301 redirect only if an equivalent page exists.
Soft 404 High Priority The page appears blank, thin, or displays a "Not found" message while returning an HTTP 200 success code. Ensure missing pages return a true 404 or 410 header, or add substantial helpful content.
Server error (5xx) High Priority Googlebot encountered a server timeout, gateway crash, or database error while requesting the page. Inspect web server logs, check hosting resource limits, and resolve backend script crashes.
Crawled - currently not indexed Medium Priority Google fetched and evaluated the content, but decided not to index it based on quality, intent, or duplication. Improve content quality, eliminate internal duplicate text, and build contextual internal links.
Discovered - currently not indexed Medium Priority Google added the URL to its crawl queue, but has not yet fetched it due to crawl demand or server load limits. Check internal linking structure, audit sitemap quality, and verify server responsiveness.
Duplicate without user-selected canonical High Priority The URL has duplicate versions, but lacks an explicit rel="canonical" tag. Google chose an arbitrary canonical. Add an explicit self-referential or master rel="canonical" tag to the document head.

Deep-dive into common Not Indexed statuses

Understanding the technical mechanics behind the most frequent status codes prevents panic and directs your attention to genuine problems.

1. Crawled - currently not indexed

This status indicates that Googlebot successfully made an HTTP request to the URL, downloaded the HTML, and rendered the document. However, Google's indexing systems chose not to add the page to the search database.

This status is rarely caused by server configurations. Instead, it reflects an evaluation of content utility, intent overlap, or topical depth. If multiple pages on your site answer the same search question with slight variations in wording, Google often indexes only one version and holds the others under this status.

To resolve this issue for valuable pages, review the search intent. Ensure the article provides original value, first-hand data, or clear practical guidance that competing search results do not offer.

2. Discovered - currently not indexed

Unlike the crawled status, this label means Google knows the URL exists, but Googlebot has not yet fetched the HTML. The URL is sitting in Google's crawl queue.

This status commonly occurs on:

  • Brand new websites where Google has not yet developed high crawl demand.
  • Websites that recently published large batches of new pages simultaneously.
  • URLs located deep in the site architecture with few internal links pointing to them.

If a critical page remains discovered but uncrawled for weeks, inspect your internal link structure. Adding links from established, high-traffic pages helps Googlebot discover and prioritize the URL.

3. Page with redirect and Excluded by 'noindex' tag

These two statuses represent Google respecting your technical configurations.

When you migrate a page and set up a 301 redirect, Googlebot follows the redirect to the new URL. The old address is properly marked as "Page with redirect." Attempting to "fix" this row is unnecessary.

Similarly, if you intentionally place a noindex tag on thank-you pages, checkout steps, or author archives, Search Console will record them under "Excluded by 'noindex' tag." This confirms that your robots directives are functioning correctly.

4. Not found (404) vs Soft 404

An HTTP 404 status code tells search engines that a resource has been permanently removed. If you deleted an old, obsolete blog post with no modern replacement, returning a 404 is the correct protocol. Google will gradually drop the URL from its reporting.

A Soft 404, by contrast, is a configuration flaw. It occurs when a web server serves an empty page, an "Item not found" notice, or an error template, but returns an HTTP 200 (Success) header instead of an HTTP 404. Googlebot recognizes that the page contains no meaningful content and flags it as a soft error. Ensure your server returns an authentic 404 status header for non-existent resources.

A 4-step diagnostic workflow for troubled URLs

When an important page appears in the Not Indexed table, follow this structured diagnostic routine to identify the root cause.

  1. Step 1: Click the status row: Select the specific reason code in Search Console to view the list of impacted URLs and review the trend graph.
  2. Step 2: Sample representative URLs: Click an individual URL in the examples table to open the slide-out inspector tray, then click Inspect URL[2]Source 2Google Search Central. URL Inspection Tool Guide.View source ↗.
  3. Step 3: Test the Live URL: On the inspection summary screen, click Test Live URL. This executes an instant real-time fetch using Googlebot, verifying whether the page is currently accessible, can be rendered, and returns valid canonical tags.
  4. Step 4: Review crawler data: Examine the Coverage details. Check the User-declared canonical against the Google-selected canonical, inspect the HTTP response code, and confirm that robots.txt is not blocking critical CSS or JavaScript assets.

How Validate Fix actually works

When you fix a systematic error, such as removing an accidental noindex tag or correcting server timeout issues, Search Console allows you to initiate validation by clicking Validate Fix.

Starting validation does not instantly re-index your pages. Instead, it triggers a multi-stage verification cycle:

  • Immediate check: Google immediately checks a small sample of the affected URLs. If the sample still exhibits the error, validation fails immediately.
  • Queueing verification: If the sample passes, Google marks the status as Pending and schedules the remaining URLs for re-crawling as crawl capacity allows.
  • Continuous re-crawling: Depending on the size of your website and Googlebot's crawl rate, validation can take anywhere from several days to several weeks.
  • Resolution: If all URLs pass inspection, Search Console changes the status to Passed and moves the valid pages into the Indexed bucket.

Do not click Validate Fix before verifying that your technical repairs are live and returning correct server headers. Clicking validation prematurely simply delays subsequent re-crawling cycles.

Expert recommendations for healthy indexation

Maintaining a healthy indexation profile[3]Source 3Google Search Central. Introduction to Crawling and Indexing.View source ↗ requires disciplined architectural management rather than reactive troubleshooting.

  1. Align XML sitemaps with indexing intent: Only include canonical, indexable URLs in your XML sitemaps[4]Source 4Google Search Central. Sitemaps Report Documentation.View source ↗. Never submit redirected, noindexed, or broken 404 URLs in a sitemap.
  2. Prune thin content proactively: If your site generates hundreds of thin parameter pages or auto-generated archives, apply noindex directives or canonicalize them to primary hub pages.
  3. Monitor crawl spikes in Crawl Stats: Review the Crawl Stats report (located under Settings) to identify whether Googlebot is spending bandwidth on unnecessary system directories or query strings.
  4. Audit internal links regularly: Ensure that important articles are never more than three clicks away from your homepage. Strong internal linking is the most effective way to prevent pages from stalling under discovered or crawled exclusions.

Sources

  1. Google Search Central. Page Indexing Report Documentation.

  2. Google Search Central. URL Inspection Tool Guide.

  3. Google Search Central. Introduction to Crawling and Indexing.

  4. Google Search Central. Sitemaps Report Documentation.

Share

Frequently Asked Questions (FAQs)

What is the Google Search Console Page Indexing report?+

The Page Indexing report is a diagnostic tool in Google Search Console that details which pages on your website are stored in Google's search index and provides specific technical reasons why other pages are excluded.

Why are my pages marked 'Not indexed'?+

Pages are classified as Not Indexed when Google cannot or should not include them in search results. Reasons range from standard configurations (such as 301 redirects and noindex tags) to technical errors (such as 404s, 5xx server issues, or content quality constraints).

What is the difference between 'Crawled - currently not indexed' and 'Discovered - currently not indexed'?+

Discovered means Google added the URL to its crawl queue but has not yet fetched the page. Crawled means Google successfully downloaded and evaluated the HTML, but decided not to add the page to its search index at this time.

How long does Google take to index a page?+

Indexation timing varies from a few hours to several weeks. It depends on your website's domain authority, crawl budget, publishing frequency, internal link structure, and content quality.

Should I try to fix every 'Not found (404)' error?+

No. If a deleted page has no relevant replacement, returning an HTTP 404 status code is completely correct. You only need to fix 404s that occur due to broken internal links or deleted pages that have a direct, equivalent replacement.

What does 'Validate Fix' do in Search Console?+

Validate Fix initiates an automated verification cycle where Google re-crawls a sample of affected URLs to confirm that a technical issue has been resolved. If the sample passes, Google gradually re-evaluates the remaining URLs over several days or weeks.

Can unindexed pages hurt my site's rankings?+

Harmless exclusions (like redirects and noindex tags) do not hurt your rankings. However, widespread crawl errors, server crashes, or large volumes of low-quality thin pages can drain crawl capacity and delay the discovery of valuable content.

Why does Search Console say 'Duplicate, Google chose different canonical than user'?+

This happens when you declare a canonical URL using a rel="canonical" tag, but Google's automated systems conclude that a different URL on your site is a better, more authoritative representative for that content.

What is a Soft 404 error?+

A Soft 404 occurs when a web page displays a "Page Not Found" message or has virtually no content, but returns an HTTP 200 (Success) response code instead of an HTTP 404 or 410 error code.

How do I check if a specific URL is indexed right now?+

Paste the full URL into the search bar at the very top of Google Search Console and press Enter. The URL Inspection tool will immediately display whether the URL is on Google and when it was last crawled.