LW IT Solutions
« Blog Overview /Web Development / Sitemap Structures Measured at 24 Sites: Indexes,...
This post in other languages:

Sitemap Structures Measured at 24 Sites: Indexes, Splitting, Extensions and Violations

Sitemap Structures Measured at 24 Sites: Indexes, Splitting, Extensions and Violations
Contents
  1. How the measurement was done
  2. Six domains without a readable sitemap
  3. robots.txt as a table of contents
  4. Five ways to build a sitemap
  5. The limit reached first
  6. Where the violations sit
  7. lastmod: one date for everything
  8. The location rule and the XSD
  9. Checking a sitemap before it is published
  10. Context and limits
  11. Questions and answers
  12. Sources

On a small site a sitemap is a list. On a large one it is a building made of files: an index that points to child files, a split by content type, language or running number, and extensions for images, videos, news and language versions. Each of these decisions has a limit in the protocol, and each can go wrong without any page visibly breaking.

For this article the sitemaps of 24 domains were read on 1 October 2026 and checked with the core of the new Sitemap Checker & Validator: search, infrastructure, news, public administration, retail and travel, in three countries, plus this site.

Five ways sitemaps are built, with examples from the measurement, next to six violations reported by the validator, below a bar with an Ikea file at 95.3 percent of the 50 MB limit and three figures on lastmod, priority and changefreq
Structures, violations and limits in the sitemaps of 18 readable domains, measured on 1 October 2026.

How the measurement was done

The sitemap was located the way a crawler finds it: first the first Sitemap: line in robots.txt, otherwise /sitemap.xml and /sitemap_index.xml. Where the file was an index, three child files were fetched: the first, one from the middle and the last. Compressed files (.gz) were decompressed. The check was run by the same code that runs in the tool in the browser, here under Node on this site’s own server.

Whether the addresses in the sitemaps are reachable was not checked. That is a different question with its own traps, and it is what the XML Sitemap Auditor is for, which fetches a sample of the addresses. This article is about the files themselves: structure, format and the rules of sitemaps.org and Google.

Six domains without a readable sitemap

Of 24 domains, 18 delivered a readable sitemap. Four name no sitemap in robots.txt and answer 404 at /sitemap.xml, one of them only after a redirect. Two turned the fetching client away with 403, one of them on the very file its own robots.txt names.

One case lies in between. heise.de names seven sitemaps in robots.txt, and the first of them, /sitemapindex.xml, answers 404. The standard address /sitemap.xml, on the other hand, delivers a file with 10,000 entries. A crawler that reads every line loses nothing; a tool that takes only the first reports a missing sitemap.

robots.txt as a table of contents

The protocol allows any number of Sitemap: lines, and that is exactly how several large sites work: instead of an index, the list of files sits directly in robots.txt.

Domain Sitemap: lines
booking.com 434
nytimes.com 25
microsoft.com 22
otto.de 19
bbc.co.uk 13
wordpress.org 9
heise.de 7
seven more 1 to 5

This has an advantage that no specification spells out: every line is also proof that the operator of the host wants this file. Under sitemaps.org, a sitemap may contain addresses of another host only if that host’s robots.txt points to it.

Five ways to build a sitemap

The 18 sitemaps fall into five structures, and some sites mix two of them.

Structure Examples from the measurement
one file heise.de (10,000 entries), lukaswojcik.com (1,006), Cloudflare (933); the news sitemaps of the New York Times, Spiegel and the Guardian, each named first in its robots.txt
index by content type wordpress.org (pages, images, videos), Otto (49 categories), Airbnb (homes, experiences, destinations, help pages)
index by language or market Ikea with 2,166 files following the pattern prod-et-EE_1.xml, MDN with ten languages, Booking with 45 languages per content type, Shopify with 568 files
index by running number gov.uk with 35 files of up to 25,000 addresses each, Airbnb with 140 files for homes alone
index of indexes Apple (47 country indexes, one of them on apple.com.cn), Microsoft (five sub-indexes), one child file at Google

The last structure deserves a second look. sitemaps.org describes an index as a list of sitemaps, and Google’s guide to large sitemaps says nothing about nesting. An index that points to further indexes therefore relies on behaviour nobody promises. Apple’s root index also names files on a second host, which under the protocol is covered only by that host’s robots.txt.

The limit reached first

A sitemap may contain at most 50,000 addresses and may be at most 50 MB uncompressed. gov.uk stays at half with 25,000 addresses per file; one Airbnb file sits exactly on the limit with 50,000. At Ikea the other limit applies.

The file prod-et-EE_1.xml contains only 4,617 addresses but weighs 47.6 MB, 95.3 percent of what is allowed. Each address carries on average 69 xhtml:link entries for its language versions, 319,635 in total, plus 22,797 images and 1,585 videos. Around 10.8 KB per entry is no longer the address but its description.

This is the arithmetic hreflang in a sitemap brings with it: every version of a page lists all other versions and itself. Across all files the number of links therefore grows with the square of the language versions, not with the number of pages. For sitemaps with hreflang the byte limit is the better yardstick for splitting than the number of addresses.

Where the violations sit

In the core format of loc, lastmod, changefreq and priority the validator found little: a few duplicate entries, three addresses with a fragment (#) and two on another host at heise.de, and empty child files at Shopify and Ikea. Almost all violations of weight sit in the namespaces and the extensions.

  • Google root index: it uses the old Google namespace http://www.google.com/schemas/sitemap/0.84 instead of http://www.sitemaps.org/schemas/sitemap/0.9. Namespaces are compared character by character; to a strict reader this is not a sitemap index.
  • Gmail sitemap: 166 addresses with 9,075 hreflang links, 165 of them with the code es-419 for Latin American Spanish. Google’s help page on localized versions names exactly this code as an example of one that is not supported, because only regions per ISO 3166-1 count. The help page itself lists a version with hreflang="es-419" in its HTML head.
  • Cloudflare: the namespaces for the sitemap and for xhtml are written with https://. A namespace is an identifier, not an address; with https it is a different identifier, and the 10,263 hreflang links sit in a namespace no reader expects.
  • wordpress.org: none of the 62 videos in the video sitemap has a video:description, which Google lists as a required field. The image sitemap of the same site lists pages several times, once per image, the home page alone 91 times; 521 of the 693 entries repeat an address already listed, although one url entry may carry up to 1,000 images.
  • New York Times: 13 of the 745 entries of the news sitemap carry news:language as en-US. An ISO 639 code of two or three letters is required, with exceptions only for zh-cn and zh-tw.
  • BBC: the index names three sitemaps on the 2014 elections, one of them with 204 addresses, all still on http://.

Whether Google discards these cases one by one or reads them generously cannot be measured from outside. What can be shown is that they contradict the formats’ own rules, and that a strict reader fails on them.

lastmod: one date for everything

Of the optional fields Google uses only lastmod, and only if it is verifiably accurate. Google ignores priority and changefreq; even so, they appear at 9 of the 18 sites.

At 7 sites at least one file carries the same date on every entry, among them Cloudflare (933 entries, one date), an Ikea file (4,617 entries, one date) and Booking (3,032 entries per language file, one date). Such a value describes the generator run, not the pages. At 6 sites no measured file contains a lastmod at all. The values look credible at heise.de (8,913 distinct among 10,000), the New York Times (736 among 745) and wordpress.org (51 among 52).

A case no validator sees within one file turned up on this site. On the day of measurement the Yoast index named 28 September as the last change of post-sitemap.xml, while the newest entry in that file carried 1 October. The lastmod in the index is meant to describe the change of the child file; checking it requires both files side by side.

The location rule and the XSD

Large sites break two rules of the protocol almost across the board. The first is the location rule: under sitemaps.org, a sitemap at /sitemaps/sitemap.xml may contain only addresses below /sitemaps/. At ten of the 18 sites at least one measured file violates this rule, mostly because it sits in a subfolder such as /sitemaps/ and lists addresses above it. At gov.uk, Ikea, Otto, MDN and the three news sites this affects every address of the measured files. Google still uses such files if they are submitted in Search Console. The validator therefore reports the location rule as a note, and only when the address of the file is given.

The second is the schema. The sitemaps.org XSD admits elements from other namespaces only with processContents="strict". A correct image sitemap with a single image:image therefore fails a plain XSD check; libxml2 2.9.14 reported “No matching global element declaration available, but demanded by the strict wildcard”. Checking against the XSD means loading the schemas of the extensions as well. The XSD is just as strict about time: 2026-10-01T14:05 without seconds is invalid there, with seconds but without a time zone it is valid, but not in the W3C format that sitemaps.org refers to.

Checking a sitemap before it is published

No single command covers the chain of well-formedness, protocol and extensions. The Sitemap Checker & Validator in the toolbox takes pasted XML or a file, compressed too, and checks in the browser:

  • well-formedness with line and column, such as an unescaped & or a truncated end of file,
  • namespace, encoding and the limits of 50,000 entries and 50 MB,
  • loc, lastmod, changefreq, priority and their order as in the XSD,
  • hreflang with valid codes, self-reference and return links within the file,
  • the required fields and limits of the image, video and news extensions, including the deprecated tags,
  • with the address of the file given, the location rule.

The file does not leave the browser. What the addresses answer live, whether they redirect, carry a noindex or a canonical pointing elsewhere, is what the XML Sitemap Auditor checks on the published version.

Context and limits

24 domains were measured on one day, three child files per index. Ikea alone has 2,166 files; the measurement says nothing about the rest. Six domains remained unreadable, two of them because of a refusal that a crawler with a known name may not meet. And the measurement describes violations of the formats, not their consequences in search: how lenient Google is with a wrong namespace or an en-US is not documented.

Questions and answers

Why does a correct image sitemap fail a check against the sitemaps.org XSD?

The schema admits elements from other namespaces only with processContents="strict". A strictly checking parser therefore has to find a declaration for image:image, and that declaration is not in the sitemap schema but in a separate schema for the extension. Loading only sitemap.xsd produces an error for every extension, even when the file is perfectly fine. The remedy is to load the schemas of the extensions in use as well, or to check the extensions against their documented rules instead of the XSD.

Does hreflang belong in the sitemap or in the head of the page?

Google accepts language versions in three ways: as a link element in the HTML head, as an HTTP header and in the sitemap. For the evaluation the three are equivalent; what differs is the cost.

In the sitemap the number of entries grows with the square of the versions, because every version lists all the others and itself. That is how the measured Ikea file reached 95.3 percent of the 50 MB limit with 4,617 addresses. In return the pages themselves stay lean, and the annotations sit in one place where a validator can check self-references and return links in a single pass.

In the HTML head each page carries only its own group, and the annotations change together with the page. With a few languages that is usually the simpler route; with dozens of markets the sitemap has more in its favour, but then split by bytes rather than by addresses. Maintaining both at once doubles the sources of error, because two lists of the same group can drift apart.

How is a sitemap with more than 50,000 addresses split cleanly?

With an index that points to several files. Four rules help:

  • The child files sit in the index’s folder or below, as Google requires for indexes.
  • The split follows a property that does not change constantly, such as content type or language; a running number moves addresses into other files with every rebuild.
  • The lastmod in the index describes the last change of the respective child file and is rewritten together with it.
  • An index points to sitemaps, not to further indexes, because neither specification promises nesting.
Lukas Wojcik

Lukas Wojcik

Systems architect and technology enthusiast specializing in scalable tracking solutions, GMP Stack (GA4 & GTM), and robust backend architectures. Advocate for clean code and privacy-first design.

Get in Touch

Briefly describe your project or inquiry for a tailored response. This site is protected by reCAPTCHA.

Write a comment

Edge cases the article misses and questions about the code are welcome here.

The email address is not published. Required fields are marked with an asterisk.

ALL ARTICLES & CATEGORIES

CCTV

Follow this category by RSS

Cloud & AI

Follow this category by RSS

Data Privacy

All 17 articles in this category Follow this category by RSS

Digital Analytics

All 54 articles in this category Follow this category by RSS

Digital Marketing

All 36 articles in this category Follow this category by RSS

IT & Networks

All 17 articles in this category Follow this category by RSS

Music Production

All 14 articles in this category Follow this category by RSS

Raspberry PI

Follow this category by RSS

Smart Home

All 18 articles in this category Follow this category by RSS

Web Development

All 11 articles in this category Follow this category by RSS

WordPress Plugins & Tricks

Follow this category by RSS