{"id":15567,"date":"2026-09-05T14:38:55","date_gmt":"2026-09-05T12:38:55","guid":{"rendered":"https:\/\/www.lukaswojcik.com\/blog\/toolbox\/xml-sitemap-auditor-lastmod-index-sample-check\/"},"modified":"2026-09-05T14:38:55","modified_gmt":"2026-09-05T12:38:55","slug":"xml-sitemap-auditor-lastmod-index-sample-check","status":"publish","type":"page","link":"https:\/\/www.lukaswojcik.com\/blog\/en\/toolbox\/xml-sitemap-auditor-lastmod-index-sample-check\/","title":{"rendered":"XML Sitemap Auditor: Index, lastmod and a Sample Check"},"content":{"rendered":"<div class=\"gtm-analyser-container\" style=\"background: var(--bg-panel, #1e1e24); padding: 25px; border-radius: 8px; border: 1px solid var(--border, #2a2a35);\">\n<p style=\"color: var(--text-secondary, #a0a0b0); margin-bottom: 20px;\">A sitemap is a promise to crawlers: these addresses exist, they are meant to be indexed, and this is when they last changed. Broken promises cost crawl budget and trust: listed pages that redirect or answer 404, noindex pages in the list, a lastmod that is the same for every entry, or a sitemap that no robots.txt announces. This check finds the sitemap the way a crawler does, reads an index and up to ten child sitemaps, judges the entries, and then fetches twelve of the listed pages spread across the list to compare promise and reality.<\/p>\n<div style=\"margin-bottom: 14px;\">\n        <label for=\"sm-url\" style=\"color: var(--text-secondary, #a0a0b0); display: block; font-size: 0.85rem; margin-bottom: 5px;\">Domain or sitemap address<\/label><br \/>\n        <input id=\"sm-url\" type=\"text\" value=\"www.lukaswojcik.com\" class=\"form-control\" style=\"width: 100%; padding: 10px; background: var(--bg-body, #14141a); border: 1px solid var(--border, #2a2a35); color: var(--text-primary, #e8e8ee); border-radius: 6px; box-sizing: border-box; font-family: monospace;\">\n    <\/div>\n<p style=\"color: var(--text-secondary, #a0a0b0); font-size: 0.82rem; margin-bottom: 16px;\">From this server: robots.txt, the sitemap and up to ten child sitemaps (5 MB each at most), then twelve listed pages of at most 64 KB each. Hosts in private, loopback and link-local networks are refused. The query is protected by reCAPTCHA v3; the calling IP address and the target are stored for one hour to limit the rate.<\/p>\n<p>    <button id=\"sm-btn\" class=\"button\" style=\"background: var(--accent, #7ee787); color: var(--on-accent, #0b1114); border: none; padding: 12px 24px; border-radius: 6px; font-weight: 700; cursor: pointer;\">Audit the sitemap<\/button><\/p>\n<div id=\"sm-ausgabe\" style=\"margin-top: 22px;\"><\/div>\n<p style=\"color: var(--text-secondary, #a0a0b0); font-size: 0.8rem; margin: 24px 0 0;\">Limits worth knowing: only the first ten child sitemaps of an index are read, and only the first 20,000 entries are kept for the statistics; the sample is twelve pages spread evenly through the list, which finds systematic problems and misses single broken entries; noindex is read from the HTML only, not from an X-Robots-Tag header; and image, video and news extensions are counted as ordinary entries.<\/p>\n<\/div>\n<p><script>\n(function () {\n    'use strict';<\/p>\n<p>    const T = {\"einleitung\":\"A sitemap is a promise to crawlers: these addresses exist, they are meant to be indexed, and this is when they last changed. Broken promises cost crawl budget and trust: listed pages that redirect or answer 404, noindex pages in the list, a lastmod that is the same for every entry, or a sitemap that no robots.txt announces. This check finds the sitemap the way a crawler does, reads an index and up to ten child sitemaps, judges the entries, and then fetches twelve of the listed pages spread across the list to compare promise and reality.\",\"label_url\":\"Domain or sitemap address\",\"datenschutz\":\"From this server: robots.txt, the sitemap and up to ten child sitemaps (5 MB each at most), then twelve listed pages of at most 64 KB each. Hosts in private, loopback and link-local networks are refused. The query is protected by reCAPTCHA v3; the calling IP address and the target are stored for one hour to limit the rate.\",\"knopf\":\"Audit the sitemap\",\"laeuft\":\"Finding the sitemap, reading the index, sampling pages...\",\"grenze\":\"Limits worth knowing: only the first ten child sitemaps of an index are read, and only the first 20,000 entries are kept for the statistics; the sample is twelve pages spread evenly through the list, which finds systematic problems and misses single broken entries; noindex is read from the HTML only, not from an X-Robots-Tag header; and image, video and news extensions are counted as ordinary entries.\",\"h_uebersicht\":\"Overview\",\"zeile_start\":\"Sitemap\",\"zeile_quelle\":\"Found via\",\"zeile_robots\":\"Announced in robots.txt\",\"zeile_gesamt\":\"Listed addresses\",\"zeile_dauer\":\"Took\",\"zeile_limit\":\"Checks left this hour\",\"quelle_robots\":\"Sitemap line in robots.txt\",\"quelle_genannt\":\"address entered\",\"quelle_standard\":\"default path \\\/sitemap.xml\",\"quelle_alternative\":\"an alternative default path\",\"h_sitemaps\":\"Sitemap files\",\"spalte_sitemap\":\"File\",\"spalte_art\":\"Type\",\"spalte_n\":\"Entries\",\"spalte_groesse\":\"Size\",\"spalte_status\":\"Status\",\"art_index\":\"index\",\"art_urlset\":\"URL list\",\"art_unbekannt\":\"not a sitemap\",\"fehlgeschlagen\":\"failed\",\"h_proben\":\"Sample of {n} listed pages\",\"spalte_url\":\"Address\",\"spalte_befund\":\"Result\",\"spalte_hinweis\":\"Detail\",\"probe_ok\":\"answers 200\",\"probe_fehler\":\"not reachable\",\"probe_status\":\"wrong status\",\"probe_weitergeleitet\":\"redirects\",\"probe_noindex\":\"noindex\",\"probe_canonical\":\"canonical elsewhere\",\"h_befunde\":\"Findings\",\"stufe_hoch\":\"High\",\"stufe_mittel\":\"Medium\",\"stufe_niedrig\":\"Low\",\"stufe_info\":\"Information\",\"stufe_gut\":\"In order\",\"keine_befunde\":\"Nothing to report.\",\"urteil_gut\":\"The sitemap is findable, well-formed and the sample keeps its promise.\",\"urteil_warnung\":\"The sitemap works, but part of what it promises is not true or not usable.\",\"urteil_kritisch\":\"No usable sitemap was found.\",\"fehler_kein_token\":\"The check could not be started because reCAPTCHA did not load.\",\"fehler_captcha\":\"reCAPTCHA classified the request as automated. Reloading the page normally helps.\",\"fehler_zu_viele\":\"The limit of {limit} checks per hour for this address has been reached.\",\"fehler_adresse\":\"That address cannot be used.\",\"fehler_zaehler\":\"The rate counter is unavailable, so nothing was checked.\",\"fehler_pruefdienst\":\"The reCAPTCHA service could not be reached.\",\"fehler_aufbau\":\"The service is not configured correctly. The fault is on this side.\",\"fehler_eingabe\":\"The request could not be read.\",\"fehler_methode\":\"Wrong request method.\",\"fehler_netz\":\"The service could not be reached.\",\"fehler_antwort\":\"The answer could not be read.\",\"fehler_unbekannt\":\"Something went wrong that has no message of its own.\",\"grund_leer\":\"Nothing was entered.\",\"grund_zu_lang\":\"The input is too long.\",\"grund_unlesbar\":\"It does not parse as an address.\",\"grund_schema\":\"Only http and https addresses are checked.\",\"grund_benutzerinfo\":\"Inputs with credentials in them are refused.\",\"grund_kein_punkt\":\"The host name needs at least one dot.\",\"grund_hostform\":\"Enter a host name, not an IP address.\",\"grund_port\":\"Only the ports 80, 443, 8080 and 8443 are checked.\",\"grund_gesperrter_bereich\":\"That host resolves into a private or reserved network. Those are never contacted.\",\"grund_kein_dns\":\"The host name does not resolve.\",\"befund_robots_sitemap\":\"robots.txt announces {n} sitemap(s); checked: {url} || Crawlers read the Sitemap line without any other hint. This is how a sitemap should be found.\",\"befund_genannt\":\"Checked the address entered: {url} || The sitemap was taken from the input, not discovered.\",\"befund_robots_ohne_sitemap\":\"robots.txt does not announce a sitemap || Google finds \\\/sitemap.xml by convention and via Search Console; Bing and others rely on the Sitemap line in robots.txt. One line closes the gap.\",\"befund_robots_andere_sitemap\":\"robots.txt announces a different sitemap: {liste} || The file checked is not the one robots.txt points crawlers to.\",\"befund_alternative\":\"Found at an alternative path: {url} || \\\/sitemap.xml did not answer; a WordPress or plugin default path did. Announcing it in robots.txt removes the guessing.\",\"befund_sitemap_fehlt\":\"No sitemap found ({url}: {status} {fehler}) || Neither robots.txt nor the default paths lead to a sitemap. Crawlers then rely on links alone; new and deep pages are found later or not at all.\",\"befund_start_weitergeleitet\":\"The sitemap address redirects to {ziel} || Crawlers follow it, but the announced address should be the final one. A redirect chain in front of every sitemap fetch is wasted work.\",\"befund_kein_sitemap_xml\":\"The address answers, but not with a sitemap ({typ}) || No urlset or sitemapindex element was found. Often an HTML error page under the sitemap path.\",\"befund_nicht_wohlgeformt\":\"Not well-formed XML: {url} || The entries could still be counted, but a strict parser stops at the first error, and Google reports the file as unreadable. Unescaped ampersands in URLs are the usual cause.\",\"befund_namespace_fehlt\":\"Sitemap namespace missing || The root element should declare xmlns=\\u0022http:\\\/\\\/www.sitemaps.org\\\/schemas\\\/sitemap\\\/0.9\\u0022. Most crawlers cope, validators do not.\",\"befund_content_type\":\"Served as {typ} || A sitemap should be served as application\\\/xml or text\\\/xml. Some crawlers refuse other types.\",\"befund_gz\":\"Compressed with gzip || Allowed and useful for large files; the 50 MB limit applies to the uncompressed size.\",\"befund_index\":\"Sitemap index with {n} child sitemap(s), {gelesen} read || An index bundles several sitemaps. Each child is fetched and judged on its own.\",\"befund_index_gekuerzt\":\"Only the first {max} of {n} child sitemaps were read || The statistics below cover the sitemaps read, not the whole set.\",\"befund_kind_fehler\":\"{n} child sitemap(s) not readable: {liste} || An index that points to missing or failing files makes crawlers skip those parts entirely.\",\"befund_kind_leer\":\"{n} child sitemap(s) contain no entries || Empty child sitemaps are harmless but pointless; removing them keeps the index clean.\",\"befund_index_verschachtelt\":\"{n} child sitemap(s) are themselves indexes || A sitemap index must not point to another index. Google ignores nested indexes.\",\"befund_zu_viele_urls\":\"{n} entries in {url} || A single sitemap may list at most 50,000 URLs. Anything beyond is ignored; a sitemap index with several files is the fix.\",\"befund_zu_gross\":\"{mb} MB uncompressed: {url} || A single sitemap may be at most 50 MB uncompressed. Larger files are refused; split them via an index.\",\"befund_leer\":\"The sitemap lists no addresses || A sitemap without entries tells crawlers nothing. Either the generator produced nothing, or the wrong file is announced.\",\"befund_urls_gesamt\":\"{n} addresses in {sitemaps} sitemap file(s) || The set of addresses the site asks crawlers to index.\",\"befund_lastmod_fehlt\":\"No entry carries lastmod || Google uses lastmod to decide what to recrawl and has said it only trusts it when it is consistently accurate. Without it, every entry looks equally old.\",\"befund_lastmod_teilweise\":\"lastmod on {mit} of {n} entries || Mixed use makes the value hard to trust. Either every entry gets a real modification date, or none does.\",\"befund_lastmod_ungueltig\":\"{n} lastmod value(s) are not W3C dates || The format is YYYY-MM-DD or a full date-time with time zone. Anything else is ignored.\",\"befund_lastmod_zukunft\":\"{n} lastmod value(s) lie in the future || A modification date after now is impossible and marks the field as unreliable for the whole file.\",\"befund_lastmod_alle_gleich\":\"All {n} entries carry the same lastmod ({datum}) || The generator writes the build time into every entry. That is not a modification date; Google has stated it ignores lastmod that is set on every URL to the same value.\",\"befund_lastmod_ok\":\"lastmod on {mit} of {n} entries with {verschieden} distinct dates || The dates vary, which is what real modification dates do.\",\"befund_host_fremd\":\"{n} address(es) on a host other than {host} || A sitemap may only list addresses on its own host (protocol and www variant included), unless robots.txt of the other host announces it. Foreign entries are ignored.\",\"befund_http_urls\":\"{n} http:\\\/\\\/ address(es) in an https sitemap || The listed addresses should be the canonical https ones; http entries send crawlers through a redirect for every page.\",\"befund_relativ\":\"{n} relative address(es) || Sitemap entries must be absolute URLs. Relative ones are ignored.\",\"befund_fragment\":\"{n} address(es) with a #fragment || Fragments are never sent to the server; such entries describe the same document twice.\",\"befund_query\":\"{n} address(es) with query parameters || Fine when those are canonical pages; a sign of parameter noise when they are filters or tracking parameters.\",\"befund_duplikate\":\"{n} duplicate address(es) || The same URL appears more than once. Harmless, but it inflates the file and hints at a generator issue.\",\"befund_priority_changefreq\":\"priority on {priority}, changefreq on {changefreq} entries || Google ignores both fields; Bing largely too. They do no harm, but they do no work either.\",\"befund_stichprobe\":\"Sample: {ok} of {n} answer 200 without issues, {weiter} redirect, {fehler} fail, {noindex} noindex, {canonical} canonical elsewhere || Twelve pages spread through the list, fetched like a crawler would.\",\"befund_probe_fehler\":\"{n} of {von} sampled pages do not answer 200 || Listed pages that answer 404, 410 or 5xx waste crawl budget and, at scale, make Google distrust the file.\",\"befund_probe_weitergeleitet\":\"{n} of {von} sampled pages redirect || The sitemap should list the final address, not one that redirects to it.\",\"befund_probe_noindex\":\"{n} of {von} sampled pages carry noindex || A page that asks not to be indexed does not belong in the sitemap. The two signals contradict each other.\",\"befund_probe_canonical\":\"{n} of {von} sampled pages point their canonical elsewhere || The sitemap lists an address that the page itself says is not the preferred one. The canonical wins, and the entry is wasted.\",\"befund_probe_sauber\":\"All {n} sampled pages answer 200, without redirect, noindex or a foreign canonical || Promise and reality match in the sample.\"};<\/p>\n<p>    var ENDPUNKT = '\/lw-sitemap.php';\n    var SITEKEY = '6LcrPkEtAAAAAPo1QCOf-IIM2fCL0UfdJz4y2iSY';\n    var STUFEN = ['hoch', 'mittel', 'niedrig', 'info', 'gut'];\n    var QUELLEN = ['robots', 'genannt', 'standard', 'alternative'];\n    var PROBEN = ['ok', 'fehler', 'status', 'weitergeleitet', 'noindex', 'canonical'];<\/p>\n<p>    function liste(w) { return Array.isArray(w) ? w : []; }\n    function zahl(w) { return (typeof w === 'number' && isFinite(w)) ? w : null; }\n    function text(schluessel, daten) {\n        var t = T[schluessel];\n        if (typeof t !== 'string') { return null; }\n        return t.replace(\/\\{([a-z_0-9]+)\\}\/g, function (m, k) {\n            return (daten && daten[k] !== undefined && daten[k] !== null) ? String(daten[k]) : m;\n        });\n    }<\/p>\n<p>    function bewerten(antwort) {\n        var a = antwort || {};\n        if (a.fehler) { return { fehler: String(a.fehler), grund: a.grund || null, limit: zahl(a.limit) }; }\n        var befunde = liste(a.befunde).map(function (f) {\n            return { key: String(f.key || ''), stufe: STUFEN.indexOf(f.stufe) === -1 ? 'info' : f.stufe, daten: (f.daten && typeof f.daten === 'object') ? f.daten : {} };\n        });\n        var gruppen = {};\n        STUFEN.forEach(function (s) { gruppen[s] = befunde.filter(function (f) { return f.stufe === s; }); });\n        return {\n            fehler: null, basis: a.basis ? String(a.basis) : '', start: a.start ? String(a.start) : '', quelle: QUELLEN.indexOf(a.quelle) === -1 ? 'standard' : a.quelle, robots: liste(a.robots).map(String),\n            urteil: (a.urteil === 'gut' || a.urteil === 'warnung' || a.urteil === 'kritisch') ? a.urteil : 'kritisch',\n            gesamt: zahl(a.gesamt) || 0,\n            sitemaps: liste(a.sitemaps).map(function (s) { s = s || {}; return { url: String(s.url || ''), art: s.art ? String(s.art) : null, n: zahl(s.n) || 0, kb: zahl(s.kb), gz: !!s.gz, status: zahl(s.status), fehler: s.fehler ? String(s.fehler) : null, kind: !!s.kind, lastmod: s.lastmod ? String(s.lastmod) : null }; }),\n            proben: liste(a.proben).map(function (p) { p = p || {}; return { url: String(p.url || ''), status: zahl(p.status), spruenge: zahl(p.spruenge) || 0, endurl: p.endurl ? String(p.endurl) : null, noindex: !!p.noindex, canonical: p.canonical ? String(p.canonical) : null, fehler: p.fehler ? String(p.fehler) : null, befund: PROBEN.indexOf(p.befund) === -1 ? 'fehler' : p.befund }; }),\n            befunde: befunde, gruppen: gruppen, dauer: zahl(a.dauer_ms), limitRest: zahl(a.limit_rest)\n        };\n    }<\/p>\n<p>    function skriptLaden() {\n        return new Promise(function (auf) {\n            if (window.grecaptcha) { auf(true); return; }\n            var s = document.createElement('script');\n            s.src = 'https:\/\/www.google.com\/recaptcha\/api.js?render=' + SITEKEY;\n            s.onload = function () { auf(true); }; s.onerror = function () { auf(false); };\n            document.head.appendChild(s);\n            setTimeout(function () { auf(!!window.grecaptcha); }, 8000);\n        });\n    }\n    function tokenHolen() {\n        return skriptLaden().then(function (da) {\n            if (!da || !window.grecaptcha || !window.grecaptcha.ready) { return null; }\n            return new Promise(function (auf) {\n                var fertig = false;\n                function einmal(w) { if (!fertig) { fertig = true; auf(w); } }\n                setTimeout(function () { einmal(null); }, 12000);\n                try {\n                    window.grecaptcha.ready(function () {\n                        if (fertig) { return; }\n                        if (!window.grecaptcha.execute) { einmal(null); return; }\n                        window.grecaptcha.execute(SITEKEY, { action: 'sitemap' }).then(function (t) { einmal(t || null); }, function () { einmal(null); });\n                    });\n                } catch (e) { einmal(null); }\n            });\n        }, function () { return null; });\n    }\n    function abfragen(url) {\n        return tokenHolen().then(function (token) {\n            return fetch(ENDPUNKT, { method: 'POST', headers: { 'Content-Type': 'application\/json' }, body: JSON.stringify({ url: url, token: token || '' }) })\n                .then(function (r) { return r.json().then(function (j) { return j; }, function () { return { fehler: 'antwort' }; }); }, function () { return { fehler: 'netz' }; });\n        });\n    }<\/p>\n<p>    window.LW_TEST = window.LW_TEST || {};\n    window.LW_TEST.SM = { bewerten: bewerten, darstellen: null, text: text, ENDPUNKT: ENDPUNKT, STUFEN: STUFEN, QUELLEN: QUELLEN, PROBEN: PROBEN, T: T };<\/p>\n<p>    var btn = document.getElementById('sm-btn');\n    var out = document.getElementById('sm-ausgabe');\n    if (!btn || !out) { return; }<\/p>\n<p>    function el(tag, stil, txt) { var e = document.createElement(tag); if (stil) { e.setAttribute('style', stil); } if (txt !== undefined && txt !== null) { e.textContent = txt; } return e; }\n    function wert(id) { var e = document.getElementById(id); return e ? e.value : ''; }\n    var UEBERSCHRIFT = 'font-family: \"Nunito Sans\", sans-serif; font-weight: 700; color: var(--text-primary, #e8e8ee); font-size: 0.95rem; margin: 22px 0 8px;';\n    var ZELLE = 'padding: 5px 12px 5px 0; color: var(--text-secondary, #a0a0b0); font-size: 0.87rem;';\n    var ZELLE_WERT = 'padding: 5px 12px 5px 0; color: var(--text-primary, #e8e8ee); font-size: 0.87rem; font-family: monospace; word-break: break-all;';\n    var KASTEN = 'background: var(--bg-body, #14141a); border: 1px solid var(--border, #2a2a35); border-radius: 8px; padding: 14px 16px; margin: 0 0 10px;';\n    var GUT = ' color: #7ee787;', WARN = ' color: #ffa94d;', ROT = ' color: #ff7b72;', GRAU = ' color: #8a8a99;';\n    var FARBEN = { hoch: '#ff7b72', mittel: '#ffa94d', niedrig: '#e3b341', info: '#8a8a99', gut: '#7ee787' };\n    var URTEIL = { gut: GUT, warnung: WARN, kritisch: ROT };\n    var PROBE_FARBE = { ok: GUT, fehler: ROT, status: ROT, weitergeleitet: WARN, noindex: WARN, canonical: WARN };\n    function gitter(sp) { return el('div', 'display: grid; grid-template-columns: ' + sp + '; gap: 0 18px; align-items: baseline;'); }\n    function paar(g, n, w, stil) { g.appendChild(el('div', ZELLE, n)); g.appendChild(el('div', ZELLE_WERT + (stil || ''), w)); }<\/p>\n<p>    function darstellen(b, ziel) {\n        ziel.innerHTML = '';\n        if (b.fehler) {\n            var txt = T['fehler_' + b.fehler] || T.fehler_unbekannt;\n            if (b.grund && T['grund_' + b.grund]) { txt = txt + ' ' + T['grund_' + b.grund]; }\n            if (b.limit !== null) { txt = txt.replace('{limit}', String(b.limit)); }\n            ziel.appendChild(el('p', ZELLE + WARN + ' margin: 0;', txt));\n            return;\n        }\n        ziel.appendChild(el('div', 'font-size: 1.05rem; font-weight: 700; margin: 0 0 12px;' + URTEIL[b.urteil], T['urteil_' + b.urteil]));<\/p>\n<p>        ziel.appendChild(el('div', UEBERSCHRIFT, T.h_uebersicht));\n        var g = gitter('auto auto');\n        paar(g, T.zeile_start, b.start || '-');\n        paar(g, T.zeile_quelle, T['quelle_' + b.quelle], GRAU);\n        if (b.robots.length) { paar(g, T.zeile_robots, b.robots.join(', '), GRAU); }\n        paar(g, T.zeile_gesamt, String(b.gesamt), b.gesamt ? GUT : ROT);\n        if (b.dauer !== null) { paar(g, T.zeile_dauer, b.dauer + ' ms', GRAU); }\n        if (b.limitRest !== null) { paar(g, T.zeile_limit, String(b.limitRest), GRAU); }\n        ziel.appendChild(g);<\/p>\n<p>        if (b.sitemaps.length) {\n            ziel.appendChild(el('div', UEBERSCHRIFT, T.h_sitemaps));\n            var huelle = el('div', 'overflow-x: auto;');\n            var gp = gitter('auto auto auto auto auto');\n            [T.spalte_sitemap, T.spalte_art, T.spalte_n, T.spalte_groesse, T.spalte_status].forEach(function (s) { gp.appendChild(el('div', ZELLE + ' font-weight: 700; color: var(--text-primary, #e8e8ee); white-space: nowrap;', s)); });\n            b.sitemaps.forEach(function (s) {\n                gp.appendChild(el('div', ZELLE_WERT + (s.kind ? ' padding-left: 16px;' : ''), s.url));\n                gp.appendChild(el('div', ZELLE_WERT + GRAU, s.art ? T['art_' + s.art] || s.art : '-'));\n                gp.appendChild(el('div', ZELLE_WERT, s.fehler ? '-' : String(s.n)));\n                gp.appendChild(el('div', ZELLE_WERT + GRAU, s.kb === null ? '-' : s.kb + ' KB' + (s.gz ? ' (gz)' : '')));\n                gp.appendChild(el('div', ZELLE_WERT + (s.fehler ? ROT : GUT), s.fehler ? (s.status !== null ? String(s.status) : T.fehlgeschlagen) : '200'));\n            });\n            huelle.appendChild(gp);\n            ziel.appendChild(huelle);\n        }<\/p>\n<p>        if (b.proben.length) {\n            ziel.appendChild(el('div', UEBERSCHRIFT, text('h_proben', { n: b.proben.length })));\n            var hp = el('div', 'overflow-x: auto;');\n            var gq = gitter('auto auto auto auto');\n            [T.spalte_url, T.spalte_status, T.spalte_befund, T.spalte_hinweis].forEach(function (s) { gq.appendChild(el('div', ZELLE + ' font-weight: 700; color: var(--text-primary, #e8e8ee); white-space: nowrap;', s)); });\n            b.proben.forEach(function (p) {\n                gq.appendChild(el('div', ZELLE_WERT, p.url));\n                gq.appendChild(el('div', ZELLE_WERT + GRAU, p.status === null ? (p.fehler || '-') : String(p.status)));\n                gq.appendChild(el('div', ZELLE_WERT + PROBE_FARBE[p.befund] + ' white-space: nowrap;', T['probe_' + p.befund]));\n                var hinweis = p.befund === 'weitergeleitet' && p.endurl ? '\u2192 ' + p.endurl : (p.befund === 'canonical' && p.canonical ? 'canonical: ' + p.canonical : '');\n                gq.appendChild(el('div', ZELLE_WERT + GRAU, hinweis));\n            });\n            hp.appendChild(gq);\n            ziel.appendChild(hp);\n        }<\/p>\n<p>        ziel.appendChild(el('div', UEBERSCHRIFT, T.h_befunde));\n        var irgendwas = false;\n        STUFEN.forEach(function (s) {\n            var gr = b.gruppen[s];\n            if (!gr.length) { return; }\n            irgendwas = true;\n            ziel.appendChild(el('div', 'font-weight: 700; font-size: 0.82rem; text-transform: uppercase; letter-spacing: 0.5px; margin: 14px 0 6px; color: ' + FARBEN[s] + ';', T['stufe_' + s] + ' (' + gr.length + ')'));\n            gr.forEach(function (f) {\n                var k = el('div', KASTEN + ' border-left: 3px solid ' + FARBEN[s] + ';');\n                var t = text('befund_' + f.key, f.daten) || f.key;\n                var teile = t.split(' || ');\n                k.appendChild(el('div', 'color: var(--text-primary, #e8e8ee); font-size: 0.9rem; font-weight: 700;', teile[0]));\n                if (teile[1]) { k.appendChild(el('div', ZELLE + ' padding: 6px 0 0; line-height: 1.55;', teile[1])); }\n                ziel.appendChild(k);\n            });\n        });\n        if (!irgendwas) { ziel.appendChild(el('p', ZELLE + ' margin: 0;', T.keine_befunde)); }\n    }\n    window.LW_TEST.SM.darstellen = darstellen;<\/p>\n<p>    var laeuft = false;\n    btn.addEventListener('click', function () {\n        if (laeuft) { return; }\n        laeuft = true; btn.disabled = true;\n        out.innerHTML = '';\n        out.appendChild(el('p', ZELLE + GRAU + ' margin: 0;', T.laeuft));\n        abfragen(wert('sm-url')).then(function (a) { darstellen(bewerten(a), out); laeuft = false; btn.disabled = false; },\n            function () { darstellen(bewerten({ fehler: 'netz' }), out); laeuft = false; btn.disabled = false; });\n    });\n})();\n<\/script><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Finds a site&#8217;s XML sitemap via robots.txt or the default paths, follows a sitemap index to its children, checks the entries for lastmod quality, foreign hosts, duplicates and size limits, and fetches a sample of twelve listed pages to see whether they answer 200, redirect, carry noindex or point their canonical elsewhere.<\/p>\n","protected":false},"author":1,"featured_media":0,"parent":38,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"template-tool-base.php","meta":{"footnotes":""},"tags":[91144,91094,91096],"class_list":["post-15567","page","type-page","status-publish","hentry","tag-devops","tag-seo","tag-wordpress"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v28.1 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>XML Sitemap Auditor: Index, lastmod and a Sample Check - Lukas Wojcik - Blog<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/www.lukaswojcik.com\/blog\/en\/toolbox\/xml-sitemap-auditor-lastmod-index-sample-check\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"XML Sitemap Auditor: Index, lastmod and a Sample Check - Lukas Wojcik - Blog\" \/>\n<meta property=\"og:description\" content=\"Finds a site&#039;s XML sitemap via robots.txt or the default paths, follows a sitemap index to its children, checks the entries for lastmod quality, foreign hosts, duplicates and size limits, and fetches a sample of twelve listed pages to see whether they answer 200, redirect, carry noindex or point their canonical elsewhere.\" \/>\n<meta property=\"og:url\" content=\"https:\/\/www.lukaswojcik.com\/blog\/en\/toolbox\/xml-sitemap-auditor-lastmod-index-sample-check\/\" \/>\n<meta property=\"og:site_name\" content=\"Lukas Wojcik - Blog\" \/>\n<meta property=\"og:image\" content=\"https:\/\/www.lukaswojcik.com\/blog\/wp-content\/uploads\/2026\/08\/og-default.jpg\" \/>\n\t<meta property=\"og:image:width\" content=\"1200\" \/>\n\t<meta property=\"og:image:height\" content=\"630\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/jpeg\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data1\" content=\"1 minute\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/en\\\/toolbox\\\/xml-sitemap-auditor-lastmod-index-sample-check\\\/\",\"url\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/en\\\/toolbox\\\/xml-sitemap-auditor-lastmod-index-sample-check\\\/\",\"name\":\"XML Sitemap Auditor: Index, lastmod and a Sample Check - Lukas Wojcik - Blog\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/#website\"},\"datePublished\":\"2026-09-05T12:38:55+00:00\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/en\\\/toolbox\\\/xml-sitemap-auditor-lastmod-index-sample-check\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/en\\\/toolbox\\\/xml-sitemap-auditor-lastmod-index-sample-check\\\/\"]}]},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/en\\\/toolbox\\\/xml-sitemap-auditor-lastmod-index-sample-check\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Toolbox\",\"item\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/en\\\/toolbox\\\/\"},{\"@type\":\"ListItem\",\"position\":3,\"name\":\"XML Sitemap Auditor: Index, lastmod and a Sample Check\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/#website\",\"url\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/\",\"name\":\"Lukas Wojcik - Blog\",\"description\":\"\",\"publisher\":{\"@id\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/#\\\/schema\\\/person\\\/895f7604f9b6b71aad9bba33af28d0f9\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":[\"Person\",\"Organization\"],\"@id\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/#\\\/schema\\\/person\\\/895f7604f9b6b71aad9bba33af28d0f9\",\"name\":\"luky\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/07\\\/lw-x2.jpg\",\"url\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/07\\\/lw-x2.jpg\",\"contentUrl\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/07\\\/lw-x2.jpg\",\"width\":424,\"height\":636,\"caption\":\"luky\"},\"logo\":{\"@id\":\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/07\\\/lw-x2.jpg\"},\"sameAs\":[\"https:\\\/\\\/www.lukaswojcik.com\\\/blog\"]}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"XML Sitemap Auditor: Index, lastmod and a Sample Check - Lukas Wojcik - Blog","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/www.lukaswojcik.com\/blog\/en\/toolbox\/xml-sitemap-auditor-lastmod-index-sample-check\/","og_locale":"en_US","og_type":"article","og_title":"XML Sitemap Auditor: Index, lastmod and a Sample Check - Lukas Wojcik - Blog","og_description":"Finds a site's XML sitemap via robots.txt or the default paths, follows a sitemap index to its children, checks the entries for lastmod quality, foreign hosts, duplicates and size limits, and fetches a sample of twelve listed pages to see whether they answer 200, redirect, carry noindex or point their canonical elsewhere.","og_url":"https:\/\/www.lukaswojcik.com\/blog\/en\/toolbox\/xml-sitemap-auditor-lastmod-index-sample-check\/","og_site_name":"Lukas Wojcik - Blog","og_image":[{"width":1200,"height":630,"url":"https:\/\/www.lukaswojcik.com\/blog\/wp-content\/uploads\/2026\/08\/og-default.jpg","type":"image\/jpeg"}],"twitter_card":"summary_large_image","twitter_misc":{"Est. reading time":"1 minute"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"WebPage","@id":"https:\/\/www.lukaswojcik.com\/blog\/en\/toolbox\/xml-sitemap-auditor-lastmod-index-sample-check\/","url":"https:\/\/www.lukaswojcik.com\/blog\/en\/toolbox\/xml-sitemap-auditor-lastmod-index-sample-check\/","name":"XML Sitemap Auditor: Index, lastmod and a Sample Check - Lukas Wojcik - Blog","isPartOf":{"@id":"https:\/\/www.lukaswojcik.com\/blog\/#website"},"datePublished":"2026-09-05T12:38:55+00:00","breadcrumb":{"@id":"https:\/\/www.lukaswojcik.com\/blog\/en\/toolbox\/xml-sitemap-auditor-lastmod-index-sample-check\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/www.lukaswojcik.com\/blog\/en\/toolbox\/xml-sitemap-auditor-lastmod-index-sample-check\/"]}]},{"@type":"BreadcrumbList","@id":"https:\/\/www.lukaswojcik.com\/blog\/en\/toolbox\/xml-sitemap-auditor-lastmod-index-sample-check\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/www.lukaswojcik.com\/blog\/"},{"@type":"ListItem","position":2,"name":"Toolbox","item":"https:\/\/www.lukaswojcik.com\/blog\/en\/toolbox\/"},{"@type":"ListItem","position":3,"name":"XML Sitemap Auditor: Index, lastmod and a Sample Check"}]},{"@type":"WebSite","@id":"https:\/\/www.lukaswojcik.com\/blog\/#website","url":"https:\/\/www.lukaswojcik.com\/blog\/","name":"Lukas Wojcik - Blog","description":"","publisher":{"@id":"https:\/\/www.lukaswojcik.com\/blog\/#\/schema\/person\/895f7604f9b6b71aad9bba33af28d0f9"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/www.lukaswojcik.com\/blog\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":["Person","Organization"],"@id":"https:\/\/www.lukaswojcik.com\/blog\/#\/schema\/person\/895f7604f9b6b71aad9bba33af28d0f9","name":"luky","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.lukaswojcik.com\/blog\/wp-content\/uploads\/2026\/07\/lw-x2.jpg","url":"https:\/\/www.lukaswojcik.com\/blog\/wp-content\/uploads\/2026\/07\/lw-x2.jpg","contentUrl":"https:\/\/www.lukaswojcik.com\/blog\/wp-content\/uploads\/2026\/07\/lw-x2.jpg","width":424,"height":636,"caption":"luky"},"logo":{"@id":"https:\/\/www.lukaswojcik.com\/blog\/wp-content\/uploads\/2026\/07\/lw-x2.jpg"},"sameAs":["https:\/\/www.lukaswojcik.com\/blog"]}]}},"_links":{"self":[{"href":"https:\/\/www.lukaswojcik.com\/blog\/wp-json\/wp\/v2\/pages\/15567","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.lukaswojcik.com\/blog\/wp-json\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/www.lukaswojcik.com\/blog\/wp-json\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/www.lukaswojcik.com\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.lukaswojcik.com\/blog\/wp-json\/wp\/v2\/comments?post=15567"}],"version-history":[{"count":0,"href":"https:\/\/www.lukaswojcik.com\/blog\/wp-json\/wp\/v2\/pages\/15567\/revisions"}],"up":[{"embeddable":true,"href":"https:\/\/www.lukaswojcik.com\/blog\/wp-json\/wp\/v2\/pages\/38"}],"wp:attachment":[{"href":"https:\/\/www.lukaswojcik.com\/blog\/wp-json\/wp\/v2\/media?parent=15567"}],"wp:term":[{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.lukaswojcik.com\/blog\/wp-json\/wp\/v2\/tags?post=15567"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}