Page ExtractorScrape product pages in bulk
The Page Extractor visits a list of web addresses and pulls the same details off every one — one row per page. On a shop it needs nothing picked at all.
Works in Chrome, Edge & Brave · No account needed
100,000+ users
4.6 on the Chrome Web Store
For when you have the pages and need the details
This is the tool for “I have a list of pages and I want the same three things off each of them”. The addresses can come from a file, from a scrape you already ran, or from the site’s own index.
Give it the addresses
Upload a spreadsheet and pick the column, reuse a column of links from a list you already scraped, or let Sitemap Explorer find every page on the site. The first address opens in your tab so you have a real page to point at.
Say what to take
Stack up the steps you want: click elements on the loaded page to define fields, or add a ready-made step that reads product details, page metadata, phone numbers or Google Maps listings with nothing to pick.
Start it and walk away
It works through the list in the background with a live count. A page that fails is reported and skipped — it never stops the batch — and there is no cap on how many addresses you give it.
Five ways to say what you want off a page
Each step you add writes its own columns, and they combine — run Automatic Extract for what the page is about and Page Metadata for its title and description, and you get both, side by side.
Pick it by clicking
Click elements on the loaded page to define your own fields. Each one is remembered several different ways, so it still resolves when the next page is laid out slightly differently.
Product and article details
Automatic Extract reads what the page already publishes for search engines and turns it into columns. No picking, no setup.
Page metadata
Title, description, share image, author, publish and update dates, language, canonical address — the same columns for every page, whatever each one is about. Free.
Phone numbers
Finds numbers written in the page text as well as the ones behind a call link. Free.
Google Maps listings
Place name, rating, address, phone, website and opening hours off Maps listings. Offered automatically when Maps addresses are spotted in your list, and it needs nothing picked.
Failures don’t sink the run
Each address gets two attempts, with a pause between. Timeouts and network blips retry; a 404 doesn’t. Whatever fails is reported per address and the batch carries on.
On a shop, it already knows what a product is
Most shops and news sites publish a machine-readable description of the page for Google — the name, the price, whether it’s in stock, the brand, the images, the rating. Automatic Extract reads that and lays it out as columns, so on a catalogue you often pick nothing at all. If the first page you load has it, the step is added for you.
- Name, price, availability, currency, brand, colour, size, material, rating and review count, category, SKU, barcode and images
- Works on any shop that publishes it — WooCommerce, Magento, custom builds — not just one platform
- Articles, jobs, events, recipes and local businesses are recognised too, each with its own columns
- A product with twelve variants stays one row, with the variants joined into a single cell
It reads what the page publishes and invents nothing: a detail the page never declared comes back empty rather than guessed. On a Shopify store, the Shopify Extractor goes further still — one row per variant, and an import-ready file.
What you can tune
The defaults are polite and work on most sites. Everything below is adjustable when a site needs more patience — or when you need the run finished sooner.
- Pages at once
One
Turn on Faster Extraction to run several tabs in parallel for a large batch
- Pause between pages
1 second
Anywhere from none to five seconds
- Give-up time per page
30 seconds
Adjustable from 5 to 120 seconds
- Randomised timing
Off
Varies the pause and adds occasional extra ones, for sites that notice a steady rhythm
- Retries
2 per address
Transient failures back off and retry; permanent ones don’t
- Address limit
None
Process as many pages as you like
Everything here travels with the setup, so a saved recipe and a scheduled cloud run behave exactly like the run you tested.
What people use it for
Anywhere you already know which pages matter and just need what’s on them.
Price and stock tracking
Hand it your competitors’ product pages and get a dated price list. Save it as a recipe and it’s a one-click job every Monday.
Turning a list of links into details
Scrape a directory with the List Extractor, then feed its link column in here to visit every entry and collect what the listing page didn’t show.
Content and SEO audits
Every title, description, canonical address and publish date across a site, in one table — feed the addresses straight in from Sitemap Explorer.
Building a lead list
Company pages in, contact details and business facts out, one row each, ready to hand to a CRM.
Two minutes to your
first spreadsheet.
Nothing to configure, and no account needed to start.
Works in Chrome, Edge & Brave · No account needed
Not in Chrome right now? The Meta Tag Extractor reads titles, descriptions and share tags from a list of pages in the browser.
Questions about the Page Extractor
The List Extractor works on one page that contains many items — a search results page, a product grid. The Page Extractor works on many pages that each contain one thing — a product page, a profile, an article. They chain: extract a list of links with the first, then hand that column to the second to visit each link.
Usually not. Most shops publish a machine-readable description of each product for search engines, and Automatic Extract reads it — name, price, availability, brand, images, rating — and lays it out as columns. If the first page you load has it, the step is added for you automatically.
There is no cap. You can process as many pages as you like when the run happens on your own computer, and there is no monthly allowance to use up.
Five places: a spreadsheet you upload, a column of links from a scrape you already ran, the site's own index via Sitemap Explorer, a cloud table if you're signed in, or a saved recipe. There's no free-text paste box in this tool — the addresses always come from a real list.
It's retried twice if the failure looks temporary, like a timeout, and skipped if it doesn't, like a 404. Either way it's reported against that address and the rest of the batch carries on. A run never aborts because one page went wrong.
Yes. A setup you build here runs unchanged on our Cloud Platform, on whatever schedule you choose, from whichever country you pick — and it can drop the fresh rows into a Google Sheet while your computer is off.