Back to Blog

Which URLs Should Not Be in an XML Sitemap

Written by SeLinkPro
•
October 03, 2026
Sitemap URLs That Should Not Be Submitted

An XML sitemap serves as an explicit declaration of a website’s preferred indexing intent. Submitting a URL in this file signals to search engines that the page represents primary content meant to be crawled and indexed. However, sitemaps frequently suffer from a lack of curation, inadvertently pulling in URLs that actively contradict this request. Knowing which URLs should not be in an XML sitemap is a baseline requirement for maintaining a clear and consistent line of communication with search engine crawlers.

A mechanical conflict arises when a sitemap submits URLs that carry restrictive directives or invalid status codes. If a sitemap requests indexation for a page that simultaneously returns a 4xx or 5xx error, redirects to another location, contains a meta noindex tag, points to an alternate canonical URL, or is blocked by robots.txt, it forces search engines to resolve competing instructions. The sitemap advocates for indexation, while the page or server level mandates the exact opposite.

Broadcasting these mixed signals can degrade the perceived reliability of the sitemap. When crawlers consistently encounter a high volume of disallowed, broken, or duplicate URLs, they may begin to treat the XML file as an untrustworthy discovery source, which can reduce indexing efficiency at scale. Purging these contradictory URLs ensures the sitemap functions strictly as a curated list of valid, accessible, and indexable pages.

The impact of contradictory sitemap signals

When an XML sitemap submits a URL that carries a conflicting directive, it initiates a sequence where the search engine must evaluate competing inputs. The sitemap submission serves as a non-binding hint requesting indexation, while page-level, header-level, and server-level directives act as binding rules. Because binding rules override hints, crawlers like Googlebot must process the URL to discover the final instruction. This forces the crawler to request a page that the server or HTML ultimately prevents from being indexed.

The resource requirement to resolve this discrepancy depends on where the contrary directive is located. A conflicting HTTP status code or an x-robots-tag in the HTTP header can be processed early in the network request. However, if the restrictive directive is a meta tag located in the HTML document, the crawler must download and parse the page before discovering that the sitemap's indexation request is invalid. In cases where directives depend on client-side execution, the search engine may even expend rendering resources before identifying the conflict.

Search engines evaluate the overall quality of a sitemap by observing the ratio of valid, indexable pages to those returning errors or restrictive directives. While transient discrepancies are expected due to the delay between content updates and sitemap generation, a persistently high rate of contradictory signals alters how the file is processed. If a sitemap frequently presents disallowed, redirected, or non-indexable URLs, crawlers begin to view the file as an unreliable data source rather than a high-confidence discovery queue.

This loss of perceived reliability directly impacts indexing efficiency. On sites with extensive URL inventories, crawlers must prioritize which URLs to request. If the sitemap is deemed untrustworthy, the search engine may reduce the frequency with which it polls the file. As a result, when legitimately new or updated content is added, discovery and indexation may be delayed because the crawler no longer relies on the sitemap as an authoritative indicator of priority.

Technical SEO site audit tool

Run a deep technical crawl to identify 4xx errors, missing meta tags, and indexation blockers.

Excluding noindex directives and disallowed URLs

An XML sitemap must omit any URL configured with a noindex directive. This applies equally to restrictions declared within the HTML document via a meta robots tag and those delivered in the HTTP response via an x-robots-tag header. Including these URLs creates a direct contradiction, as the sitemap requests indexation while the page or header explicitly forbids it. Because sitemaps act as a queue for canonical, indexable content, any URL deliberately removed from the search index must also be removed from the sitemap.

URLs blocked by a robots.txt file present a distinct structural conflict. While a noindex directive prevents indexation, a robots.txt disallow rule prevents crawling. Submitting a disallowed URL in a sitemap instructs the search engine to discover and index the page, while the robots.txt file simultaneously forbids the crawler from accessing the server to retrieve the document. Consequently, the crawler cannot evaluate the page content, process canonical tags, or read any page-level noindex directives that might be present.

This situation forces search engines to register the URL's existence without being able to verify its contents. When a search engine processes a sitemap containing a disallowed URL, it frequently flags the anomaly with a 'Submitted URL blocked by robots.txt' warning. The sitemap submission indicates the URL is important, but the crawl restriction prevents validation.

If the search engine discovers the blocked URL through other internal or external links, it may decide to index the URL based solely on anchor text and external signals, despite being unable to crawl the page text. When a sitemap actively submits a URL that enters this state, it results in an 'Indexed, though blocked by robots.txt' condition. A URL indexed in this manner can appear in search results, but it typically lacks a standard description snippet because the crawler was never permitted to extract text from the HTML.

To avoid generating these warnings and degrading the sitemap's utility, dynamic sitemap generation logic must evaluate both page-level indexation states and server-level crawl directives. If a page requires a noindex tag to stay out of search results, or falls under a robots.txt disallow path to preserve server resources, the URL must be systematically excluded from the XML sitemap.

Removing Non-Canonical and duplicate content

An XML sitemap serves as a definitive list of a site's preferred URLs. To maintain this utility, the sitemap must be strictly limited to canonical URLs. Submitting alternate versions of a page introduces ambiguity and contradicts the fundamental purpose of the file.

Duplicate content frequently emerges through URL parameters and dynamic routing. Common non-canonical variations include URLs appending tracking parameters, session IDs, or display settings. Similarly, faceted navigation and sorting filters generate unique URLs that display the same core content in a different order. Structural duplication also occurs when a content management system creates alternate directory paths for the same asset, such as a single product accessible through multiple distinct category hierarchies.

None of these variant URLs belong in a sitemap. When a sitemap submits non-canonical duplicates alongside or instead of the primary URL, it requires search engines to expend resources evaluating redundant content. The crawler must fetch the duplicate document, parse the HTML, process the canonical tag, and execute deduplication algorithms to consolidate the variants. This process consumes parsing capacity that should be directed toward discovering unique content.

Furthermore, submitting multiple variations of a single page dilutes indexation signals. Search engines treat sitemap inclusion as a strong indicator of canonical preference. When the sitemap actively submits a parameterized or filtered URL while the page-level HTML canonical tag points to a clean URL, the crawler receives conflicting directives. Persistent canonicalization conflicts can reduce the search engine's reliance on the sitemap as an authoritative signal, increasing the likelihood that the algorithm will independently select an unintended URL as the canonical version for search results.

Bulk Google and Yandex index checker

Verify agency reports and track live SERP status in Google and Yandex to protect your SEO ROI.

Purging redirects and non-200 status codes

An XML sitemap functions as an inventory of active, accessible content. To fulfill this purpose, every URL included in the file must return a direct HTTP 200 OK status code. Search engine crawlers expect the sitemap to provide the immediate location of a page. Submitting URLs that return other status codes contradicts this expectation, creating unnecessary network requests and reducing the efficiency of the crawling process.

Including URLs that return 301 or 302 redirects transforms a direct discovery tool into a map of historical routing. When a crawler processes a redirected URL from a sitemap, it must execute an initial network request, process the 3xx status code, extract the location header, and then queue the new destination URL for a future crawl. This sequence delays the discovery of the actual content. Crawlers rely on sitemaps to bypass legacy routing chains; the file must contain only the final destination URL.

Client errors present a different operational failure. Submitting URLs that return a 404 Not Found or a 410 Gone status code forces the search engine to request pages that no longer exist. When crawlers persistently encounter 4xx status codes within a sitemap, the perceived reliability of the file drops. A high volume of dead links can degrade the trust a search engine places in the sitemap, which may result in a lower crawl priority for the file and slower discovery of valid URLs.

URLs returning 5xx server errors, such as a 500 Internal Server Error or a 503 Service Unavailable, also disrupt the discovery pipeline. While 5xx codes can sometimes indicate temporary infrastructure strain, their persistent presence in a sitemap suggests a disconnect between the database generating the XML file and the actual availability of the frontend pages. The sitemap should reflect the functional state of the website, containing only URLs capable of rendering accessible content.

Maintaining a sitemap free of non-200 status codes requires synchronous updates between a site's routing logic and its XML generation rules. When a page is permanently moved, the old URL must be purged from the sitemap and the new HTTP 200 destination URL evaluated for inclusion. When a page is deleted, its corresponding URL must be removed from the sitemap entirely, rather than leaving a legacy entry that returns a 404 or 410 status code.

Evaluating utility, archive, and Low-Quality pages

A URL returning a 200 OK status code and lacking a noindex directive is technically eligible for indexing. However, technical validity does not automatically qualify a page for sitemap inclusion. An XML sitemap serves as a curated list of a website's most valuable organic entry points. Submitting utility pages, thin archives, or temporary campaign URLs dilutes this focus, instructing search engines to evaluate low-value URLs alongside primary content.

Internal search results and utility pages

Pages designed solely for operational site functions rarely satisfy general organic search intent. Account login screens, shopping cart views, password reset forms, and user profiles offer no standalone informational or commercial value to search engine users. While these pages are necessary for user experience, they do not belong in a sitemap.

Similarly, internal search result pages generate an essentially infinite number of dynamic URLs based on user queries. Submitting internal search pages to a sitemap signals that these dynamic pathways hold primary value. This can inflate the site's perceived size with duplicate or thin content and prompt search engines to process endless query variations rather than crawling stable category and product pages. These dynamic URL patterns should be explicitly excluded from the sitemap generation logic.

Thin tag and archive pages

Many content management systems automatically generate archive pages based on tags, authors, or publication dates. While these pages can aid internal navigation, they frequently present duplicated excerpts of primary articles without adding unique context. If a tag page lists the exact same articles as a main category page, submitting both to the sitemap forces search engines to process redundant content.

Unless an archive page is deliberately optimized with unique, authoritative content targeting a specific taxonomy search intent, it provides little organic value. Automated taxonomy pages that merely reorder existing content should be kept out of the XML sitemap to prioritize the discovery of the individual articles and primary categories.

Temporary landing pages

Landing pages built for short-term paid advertising campaigns, email promotions, or limited-time events usually focus on immediate conversions rather than long-term organic discovery. These pages often duplicate core service or product pages but use modified copy tailored to a specific audience segment.

Submitting temporary campaign URLs to the sitemap risks indexing redundant content and complicates canonicalization signals. Furthermore, because these pages are often removed or redirected once a campaign ends, including them introduces temporary volatility into the sitemap. When a page is designed for an isolated, non-organic traffic source, its URL should be omitted.

Establishing exclusion criteria

To maintain a focused sitemap, organizations can implement programmatic filters based on page templates, CMS categories, or URL patterns. Decision criteria for excluding a technically valid URL should include:

  • Does the page lack unique, primary content, such as thin auto-generated archives or paginated lists with no distinct text?
  • Is the page intended exclusively for users already navigating the site, such as internal search results or checkout flows?
  • Is the content's lifespan shorter than a standard organic indexing cycle, such as a flash sale landing page?
  • Is the page an isolated variant built specifically for a paid traffic channel?

Applying these criteria ensures the sitemap strictly represents the high-quality, canonical pages the business actually wants surfaced in organic search results, keeping crawler attention focused on the content that drives organic acquisition.

SEO structure and reciprocal link analyzer

Detect orphan pages, deep click depths, and toxic reciprocal links built by careless agencies.

Diagnosing sitemap errors in Google search console

Google Search Console provides direct visibility into sitemap conflicts through the Page Indexing report. When search engines process an XML file, they cross-reference the submitted URLs against their own crawling and indexing findings. Discrepancies between the explicit indexation request and the actual page state surface as precise error flags.

To isolate these conflicts, navigate to the Page Indexing report and use the filter dropdown positioned above the main chart. Selecting "All submitted pages" or filtering by a specific sitemap URL restricts the data exclusively to URLs provided in the sitemap, filtering out organic discovery issues. Evaluating the "Why pages aren't indexed" table under this filtered view reveals the exact mechanical conflicts requiring attention.

Submitted URL marked noindex

This status appears when the sitemap submits a URL for indexation, but upon crawling, Googlebot encounters a meta robots noindex tag or an X-Robots-Tag HTTP header. The sitemap and the page-level directive provide opposing instructions. Search engines honor the restrictive directive and drop the page. Diagnosing this requires determining the true intent for the URL: if the page should be indexed, the noindex directive must be removed from the page template; if the page is intentionally excluded, the URL must be stripped from the sitemap.

Submitted URL blocked by robots.txt

This flag triggers when a sitemap includes a URL that matches a disallow rule in the site's robots.txt file. Search engines register the URL from the sitemap but are forbidden from crawling it to verify its contents, status code, or canonical tags. Resolution involves auditing the robots.txt paths against the sitemap output to ensure disallowed directories, parameters, or internal search paths are completely excluded from the sitemap generation logic.

Submitted URL seems to be a soft 404

A Soft 404 occurs when a URL returns a successful HTTP 200 OK status code, but Googlebot evaluates the content and interprets it as an error page, an empty template, or a missing item. When this flag appears for submitted URLs, it frequently points to dynamically generated paths for out-of-stock products, expired job listings, or empty taxonomy categories that fail to return a proper 404 Not Found or 410 Gone HTTP status. Correcting this requires updating the server or CMS to return the appropriate 4xx client error code for missing content, which subsequently disqualifies the URL from sitemap inclusion.

Submitted URL has crawl issue

This is a broader diagnostic flag indicating that Googlebot could not fully fetch the submitted URL due to an error not categorized under standard HTTP status codes or directives. Common triggers include server timeouts, database query overloads, network routing blocks, or unhandled exceptions during page rendering. Diagnosing a crawl issue on a submitted URL usually requires running the specific URL through the Search Console URL Inspection tool to view the live fetch status. If the live test fails, cross-referencing the URL and timestamp in server access logs is necessary to identify the exact transaction failure preventing the crawl.

Establishing dynamic sitemap generation rules

Manual curation of XML sitemaps is impractical for websites with constantly changing inventory, user-generated content, or daily editorial publishing. Preventing conflicting submissions at scale requires a sitemap generation process governed by strict, automated inclusion logic. Whether using a native CMS feature, a headless architecture middleware, or a custom script, the system must evaluate a precise set of criteria before appending any absolute URL to the sitemap file.

The sitemap controller or generation script should function as a rigid filter. Instead of simply pulling all permalinks from a database table, the script must query the exact publication state, canonical configuration, routing rules, and indexability status of each URL.

Sequential validation logic

A robust dynamic sitemap generation routine executes a series of validation checks. A URL should only be written to the XML output if it passes every condition in the sequence:

  • Active publication status: The database query must verify that the content entity is marked as published and accessible to the public. Drafts, archived posts, scheduled content, and disabled products must be excluded to prevent 404 Not Found or 410 Gone responses from entering the sitemap.
  • Absence of redirect flags: The system must check the routing table or redirection manager to ensure the URL is not the source of a 301 or 302 redirect. Only the final destination URL of an active redirect rule should be eligible for inclusion.
  • Self-referencing canonical tag matching: The generation script must compare the requested URL path against the defined canonical URL for that page. If the system generates URLs with alternate taxonomy paths, session IDs, or tracking parameters, those variants must fail the check. The URL appended to the sitemap must exactly match the canonical tag output on the page.
  • Noindex exclusion: The script must query page-level SEO metadata and global directory rules. If a URL carries a database flag triggering a meta noindex tag, or if the path matches a rule applying an x-robots-tag noindex HTTP header, the generation logic must bypass the URL.

Implementation and performance considerations

Executing sequential validation checks for tens of thousands of URLs requires careful resource management. Querying multiple database tables for publication status, canonical definitions, and SEO directives on the fly can cause severe server load if the sitemap is dynamically generated every time a crawler requests the file.

To maintain performance, sitemap generation should be decoupled from live HTTP requests. A common and reliable approach is to use a background process or cron job that compiles the XML sitemap during low-traffic windows. The script processes the validation logic, generates a static XML file, and saves it to the server. Search engine crawlers then access this cached static file, ensuring fast response times without triggering repeated database queries.

For systems that update frequently, event-driven generation can supplement scheduled tasks. When a database transaction updates a page from published to draft, adds a redirect, or applies a noindex directive, the CMS can trigger a localized update to the cached sitemap. This removes the conflicting URL immediately and keeps the sitemap accurate between full regeneration cycles.

Keep Reading

Explore more insights and technical guides from our blog.

Sitemap and Search Console Validation

Sitemap and Search Console Validation

Show how to reconcile submitted sitemap data, discovered URLs, errors, and the live sitemap contents.

XML Sitemap Index Files

XML Sitemap Index Files

Cover sitemap indexes, child sitemaps, URL limits, response status, and consistency across large sites.

Incorrect lastmod Values in XML Sitemaps

Incorrect lastmod Values in XML Sitemaps

Explain when lastmod is useful, what it should represent, and why automated fake dates reduce its usefulness.

Audit technical issues, analyze backlinks and donors, and monitor the signals that matter to your SEO work

Create Account