Submitting an XML sitemap to Google Search Console is supposed to be straightforward: enter your file path, click submit, and receive a green "Success" confirmation. Instead, many site owners are greeted by a red warning message: Sitemap could not be read, frequently paired with a status of Couldn't fetch.
The situation becomes especially confusing when you paste the exact sitemap URL into your browser address bar and the file opens instantly, displaying cleanly formatted XML lines without any visible errors.
This disconnect occurs because a standard web browser and Googlebot process web resources through completely different network paths, user agents, and parsing engines. When Google Search Console displays Sitemap could not be read, it means Googlebot encountered an obstacle while attempting to fetch, download, or parse the raw XML file.
This guide provides an exhaustive technical walkthrough of why this error occurs, how to diagnose the underlying failure point, and the exact steps required to resolve it.
Understanding the "Sitemap Could Not Be Read" Error
In Google Search Console, the Sitemaps[1]Source 1Google Search Central. Sitemaps Overview and Guidelines.View source ↗ report tracks whether Googlebot can retrieve and process your submitted XML sitemaps. When Googlebot attempts to process your sitemap, it completes three separate stages:
- Network Retrieval: Googlebot establishes a TCP and TLS connection to your domain, requests the sitemap URL using its dedicated user agent, and expects a direct HTTP 200 response.
- Header and Content Evaluation: Googlebot verifies that the HTTP response headers conform to web standards (specifically confirming a valid XML Content-Type header).
- XML Schema Parsing: Googlebot parses the XML structure, verifying that tags (
<urlset>,<url>,<loc>,<lastmod>) comply with the official Sitemaps.org schema specifications.
If a failure occurs during any of these stages, Google Search Console flags the file as Sitemap could not be read or Couldn't fetch.
Why the Sitemap Works in Your Browser but Fails in Search Console
When you open a sitemap in Google Chrome, Edge, or Safari, several conditions mask underlying issues that block automated crawlers:
- Browser Header Negotiation: Browsers send standard consumer request headers and accept diverse content types, automatically executing visual stylesheets (XSLT) attached to the sitemap.
- Automated Redirect Following: Browsers automatically resolve multi-hop redirects, trailing slash corrections, and HTTP-to-HTTPS upgrades seamlessly.
- Firewall Whitelisting: Cloudflare, AWS WAF, or hosting security suites recognize your browser session as human, while Googlebot's automated crawler IP addresses may trigger automated challenge pages or rate limits.
The 6 Common Causes Behind Sitemap Read Errors

Diagnosing this issue requires checking both server-side delivery and document-level formatting. Below are the six primary reasons Google Search Console fails to read sitemaps.
1. Web Application Firewall (WAF) and Bot Fight Mode Blocks
Security layers such as Cloudflare, Sucuri, Wordfence, or AWS WAF are designed to intercept automated scrapers and distributed denial of service attacks. However, overly aggressive security rules frequently block legitimate search engine crawlers.
Features such as Cloudflare Bot Fight Mode or Super Bot Fight Mode issue JavaScript verification challenges or HTTP 403 Forbidden responses to requests matching automated patterns. Because Googlebot cannot solve interactive verification challenges when requesting static XML files, the request fails, resulting in a Sitemap could not be read error.
2. Redirect Chains and Canonical URL Mismatches
Google requires sitemaps to be submitted using their final, direct, canonical URL. If your submitted sitemap triggers any of the following redirect scenarios, Googlebot may abort the fetch:
- Redirecting from HTTP to HTTPS (e.g.,
http://example.com/sitemap.xml->https://example.com/sitemap.xml). - Redirecting between www and non-www versions.
- Trailing slash discrepancies (e.g., submitting
example.com/sitemap.xml/when the real file isexample.com/sitemap.xml, or vice versa). - Soft 404 redirects that route failed URLs back to the homepage.
Your sitemap must return an immediate HTTP 200 OK status without a single intermediate redirect hop.
3. Incorrect HTTP Content-Type Headers
When your web server responds to Googlebot's request, it includes a Content-Type header declaring the MIME type of the resource.
For an XML sitemap, the server must return:
Content-Type: application/xml; charset=UTF-8
or:
Content-Type: text/xml; charset=UTF-8
If your server misclassifies the file and returns Content-Type: text/html or Content-Type: text/plain, Googlebot's XML parser will encounter a MIME-type mismatch and fail to process the document.
4. XML Formatting, Syntax, and Character Encoding Errors
XML is an unforgiving data format. While web browsers often render broken HTML gracefully, an XML parser halts execution at the very first syntax error.
Common XML syntax mistakes include:
- Unescaped Special Characters: The ampersand symbol (
&) in URLs must be escaped as&. Writinghttps://example.com/page?cat=1&sort=2instead ofhttps://example.com/page?cat=1&sort=2invalidates the entire document. - Unclosed XML Tags: A missing
</url>,</loc>, or</urlset>closing tag. - BOM (Byte Order Mark) Artifacts: Invisible Unicode characters inserted at the beginning of the file by text editors.
- Whitespace Preceding XML Declaration: Any space or blank line before the opening
<?xml version="1.0" encoding="UTF-8"?>declaration causes parser rejection.
5. Sitemap File Size and URL Limit Violations
The official Sitemaps.org protocol[2]Source 2Sitemaps.org. XML Sitemap Protocol Specification.View source ↗ establishes strict capacity limits for single sitemap files:
- Maximum uncompressed file size: 50 MB.
- Maximum URL count per sitemap: 50,000 URLs.
If your sitemap exceeds either threshold, Googlebot will reject the file. Websites with large URL inventories must utilize a Sitemap Index file that links to multiple smaller sub-sitemaps.
6. Temporary Google Search Console Processing Delays
In many instances, the "Couldn't fetch" or "Sitemap could not be read" warning is simply a temporary reporting state. When you submit a new sitemap, Search Console queues the request. During the brief window between submission and physical crawling, the dashboard frequently displays "Couldn't fetch" as a placeholder status before updating to "Success".
If your sitemap is technically flawless and accessible, this status often resolves automatically within 24 to 72 hours.
Step-by-Step Troubleshooting and Resolution Guide

Follow this sequential diagnostic workflow to isolate and fix the exact breakdown point.
Step 1: Inspect the Sitemap in Google Search Console's Live Test
Do not rely on the Sitemaps tab alone. Use the URL Inspection[4]Source 4Google Search Central. URL Inspection Tool Documentation.View source ↗ tool to evaluate how Googlebot interacts with your sitemap in real time:
- Paste your full sitemap URL (e.g.,
https://example.com/sitemap.xml) into the top search bar in Google Search Console and press Enter. - Click Test Live URL in the top-right corner.
- Once the test finishes, check the status:
- URL is available to Google (Green checkmark): Your server and firewall are accessible to Googlebot. The issue is likely a temporary console delay or an internal XML syntax error.
- URL is not available to Google (Red error): Click View Tested Page and inspect the HTTP response code and Page fetch field. If it reports "Failed: Blocked by robots.txt" or "Failed: Crawl error (403 / 500)", your server or security layer is actively rejecting Googlebot.
Step 2: Test Server Headers with a Terminal curl Command
Open your terminal or command prompt and execute a curl request impersonating Googlebot. This reveals the exact status code and headers your server delivers to search crawlers:
curl -I -A "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)" https://example.com/sitemap.xml
Analyze the terminal output:
- Look at the first line: It must state
HTTP/1.1 200 OKorHTTP/2 200. If it states301 Moved Permanently, you submitted the wrong URL variation. - Check the Content-Type line: Verify that it states
Content-Type: application/xmlortext/xml. - Check for 403 Forbidden: If you receive an HTTP 403 error, your firewall is blocking requests that carry the Googlebot user agent header.
Step 3: Configure Cloudflare or Firewall Rules to Whitelist Googlebot
If your site is routed through Cloudflare and your live tests return 403 or challenge screens:
- Log into your Cloudflare Dashboard and select your domain.
- Navigate to Security -> WAF -> Custom Rules[3]Source 3Cloudflare. Cloudflare WAF Custom Rules Documentation.View source ↗.
- Create a rule named "Allow Verified Googlebot":
- Field:
Verified Bot - Operator:
equals - Value:
true - Action:
Skip(Select all security features including WAF and Managed Challenge).
- Field:
- If you have "Bot Fight Mode" enabled under Security -> Bots, verify whether disabling it resolves the Search Console fetch error.
Google provides cryptographically verifiable IP ranges and reverse DNS lookups, allowing Cloudflare to identify genuine Googlebot crawlers while blocking malicious spoofed bots.
Step 4: Validate XML Syntax and Escape Special Characters
Verify your XML code using an official schema validator:
- Copy your raw sitemap URL and run it through the W3C XML Validation Service or an XML sitemap validator.
- Ensure all special characters in query strings or post titles are properly converted into XML entity references:
| Character | XML Entity Name | Correct XML Usage |
|---|---|---|
& |
& |
https://example.com/item?a=1&b=2 |
< |
< |
<title>Safe Title</title> |
> |
> |
> Description text |
" |
" |
title="Keywords" |
' |
' |
author='Search Engine Basics' |
Step 5: Force a Fresh Fetch Using Cache-Busting Submission Techniques
Google Search Console caches failed sitemap fetch attempts. If you fixed an error on your server, Search Console may continue showing "Sitemap could not be read" for days because it has not yet re-queried your URL.
Use these proven techniques to force an immediate re-fetch:
- Submit with a Trailing Parameter: In Search Console, submit your sitemap with a harmless cache-busting query parameter (e.g.,
sitemap.xml?v=2). This treats the URL as a new submission, forcing Googlebot to fetch the live document immediately. - Rename the Sitemap Index File: If your CMS permits, rename
sitemap.xmltositemap_index.xmlorsitemap-index.xmland submit the new filename. - Reference the Sitemap in Robots.txt: Ensure your robots.txt file contains a dedicated directive pointing to your sitemap:
Googlebot reads robots.txt regularly and discovers updated sitemaps organically.User-agent: * Allow: / Sitemap: https://example.com/sitemap.xml
Platform-Specific Fixes

Depending on your content management system or web stack, apply these platform-specific adjustments:
WordPress (Yoast SEO / Rank Math)
- 404 or Blank White Page: In WordPress Admin, navigate to Settings -> Permalinks and simply click Save Changes without altering any settings. This flushes rewrite rules and regenerates dynamic XML sitemaps.
- Nginx Rewrite Directives: If using Nginx, ensure your configuration includes the required fastcgi rewrite directives for WordPress XML sitemaps rather than attempting to serve them as static files.
Next.js (App Router)
- Static Generation vs Dynamic Route Handlers: Ensure your
sitemap.tsorsitemap.xml/route.tsexports the proper HTTP response headers:return new Response(xmlString, { headers: { 'Content-Type': 'application/xml; charset=utf-8', }, }); - Ensure the route is configured with
dynamic = 'force-static'or has revalidation caching enabled to prevent rendering timeouts when Googlebot requests the endpoint.
Shopify
- Shopify automatically generates and manages your sitemap at
example.com/sitemap.xml. You cannot edit this file manually. - If Shopify displays "Sitemap could not be read", the cause is almost always domain verification or primary domain redirection. Verify that your custom domain is designated as the primary domain and that all alternative domains redirect properly to the primary HTTPS address.
What to Do If the Status Remains Stuck

If you validated your sitemap with curl, confirmed HTTP 200 status, verified clean XML schema compliance, and passed the URL Inspection Live Test, yet the Sitemaps tab still shows "Sitemap could not be read":
- Delete and Re-add the Sitemap: Click on the failed sitemap in Search Console, click the three-dot menu in the upper right corner, select Remove sitemap, and submit the URL again.
- Wait 72 Hours: Google Search Console's dashboard reporting engine runs asynchronously from Google's actual crawling infrastructure. Googlebot frequently crawls and indexes URLs listed in a sitemap days before the reporting interface updates to "Success". Check your Page Indexing report to verify whether your pages are being actively crawled regardless of the sitemap status message.
Sources
Google Search Central. Sitemaps Overview and Guidelines.
Sitemaps.org. XML Sitemap Protocol Specification.
Cloudflare. Cloudflare WAF Custom Rules Documentation.
Google Search Central. URL Inspection Tool Documentation.
