Chrome extension

Sitemap ExplorerFind every URL on a site

Sitemap Explorer lists every page a website publishes — up to 50,000 — as a folder tree you can search and tick. Send the ones you want straight into a scrape.

Works in Chrome, Edge & Brave · No account needed

100,000+ users

4.6 on the Chrome Web Store

A whole domain as a tree you can search — tick the branches you want and they go straight into a scrape.
How it works

For sites with nothing to click

Some sites have no index page worth scraping, or bury their catalogue behind a search box. Their sitemap has everything, and this is how you read it.

  1. Open it on the site

    It finds the site’s sitemaps by itself, checking the usual places sites publish them. You can also paste one in if you already know where it lives.

  2. Browse the tree

    Addresses are deduplicated, tidied and grouped by their path, so all the product pages sit in one branch and all the blog posts in another. Search across the whole tree, with matches expanded for you.

  3. Send them somewhere

    Tick whole branches, then carve out the bits you don’t want. Continue hands the selection straight to the Page Extractor; Export CSV downloads it as a one-column file for anything else.

How the scan works

It runs in your browser, not on our servers

That is not a small detail. The scan uses your own browser session, which is why it reaches sites that turn away anonymous requests — and why nothing about the site you are looking at leaves your machine.

  • Your session, your machine

    The whole scan happens inside your active tab. That helps it past bot protection that would block a server, and means nothing is uploaded anywhere.

  • It follows indexes down

    Big sites split their sitemap into a sitemap of sitemaps. Those are walked automatically, up to five levels deep, so a large multi-file site is covered in one pass.

  • Compressed files handled

    Sitemaps published as compressed .gz files are unpacked as it goes — you don’t download anything or unzip anything.

  • Grouped into a tree

    Addresses are grouped by their path, so /collections/shoes becomes a branch you can take in one click rather than 400 links you have to read.

  • Searchable, with partial ticks

    Search across the tree with several words at once. Tick a whole group, then untick the branches you don’t want — the parent shows a partial state so you can see what you’ve done.

  • Stop early, keep results

    Stop a scan at any point and work with what has been found so far. Switching to a different site rescans automatically.

What you get

What you can do with the result

The point of a list of 12,000 addresses is what happens next, and there are two things that happen next.

  • Straight into a scrape

    Continue hands your selection to the Page Extractor, which visits every one and pulls the same details off each — with no address list to build by hand.

  • Out as a spreadsheet

    Export CSV downloads the selected addresses as a one-column file, for a workflow that lives somewhere else entirely.

  • Up to 50,000 per scan

    A single scan collects up to fifty thousand addresses; anything beyond that is marked as capped rather than silently dropped.

  • Free for everyone

    Discovery, selection and the spreadsheet export are all free, with no account needed.

Use cases

What people use it for

It is the step before a scrape, whenever the hard part is knowing which pages exist.

  • Scraping a catalogue with no category page

    A shop whose products are only reachable through search still lists them all in its sitemap. Tick the product branch and hand it to the Page Extractor.

  • SEO and content audits

    Every published page as a tree, grouped by section, is the fastest inventory of a site you will get — and it feeds straight into a metadata scrape.

  • Finding a section that isn’t linked

    Old landing pages, regional variants and archived sections are usually still in the sitemap long after they leave the navigation.

  • Building an AI text set

    Take the article branch, hand it to the Page Text Extractor, and you have every post on a site as clean text.

Get started

Two minutes to your first spreadsheet.

Nothing to configure, and no account needed to start.

Works in Chrome, Edge & Brave · No account needed

Not in Chrome right now? The Sitemap URL Extractor reads a sitemap in the browser, with nothing to install.

Questions about the Sitemap Explorer

Up to 50,000 in a single scan. Sitemap indexes are followed several levels deep, so a large site that splits its sitemap across many files is covered in one pass. Anything beyond the limit is marked as capped rather than silently dropped.