· 13 min read
Common Crawl Scraper: Up to 1,000 Free Results a Month (2026)
Every record carries 20 fields for URL captures, including the direct WARC location and byte offset required to fetch raw historical page content. This scraper queries the Common Crawl URL Index to return archived URLs, capture dates, and MIME types without requiring proxies or authentication. The free-plan price of $5.00 per 1,000 results applies to items written to your dataset, which can be tried using the $5.00 monthly usage included in Apify's free plan. It is designed for SEO auditors and data engineers who need to map a domain's historical footprint rather than users seeking real-time contact details or live page metadata, which the index does not provide.
Try it: open Common Crawl Scraper on Apify, sign in on the free plan and run the prefilled example.
Can you try Common Crawl Scraper before paying?
Yes. Apify's free plan includes $5.00 of prepaid usage every month and asks for no credit card. At $0.005 per result, that covers up to 1,000 results of Common Crawl Scraper a month, before run-start charges and platform usage.
The example request further down caps maxItems at 15, so a first run returns at most 15 results and costs at most $0.075 in result charges. That is enough to see the real shape of the data before deciding anything.
Common Crawl Scraper was last updated on 2026-07-02. It is one of 1,724 Actors CrawlerBros publishes on Apify, which together have 716,703 lifetime public runs and an average rating of 4.63 out of 5 across 416 reviews.
What does it cost to run Common Crawl Scraper?
Each result costs $0.005 on Apify's free plan, which is $5.00 per 1,000 results. Starting a run is charged separately at $0.005 per GB of Actor memory. Apify also bills the platform usage each run consumes, at the rates of your Apify plan, on top of these charges.
| Apify plan | Per result | Per 1,000 results |
|---|---|---|
| FREE | $0.005 | $5.00 |
| BRONZE | $0.00433 | $4.33 |
| SILVER | $0.00367 | $3.67 |
| GOLD | $0.003 | $3.00 |
| PLATINUM | $0.003 | $3.00 |
| DIAMOND | $0.003 | $3.00 |
The maxItems control has the most direct impact on your bill because charges are calculated per dataset item written. To verify your URL pattern without a large spend, run the example input which caps results at 15 items to keep result charges at $0.075 or less.
How do you run Common Crawl Scraper from the API?
The schema marks 1 of its 9 controls as required: mode. The payload below uses the schema's own prefilled values, so it runs as written once you substitute your API token.
Call the synchronous endpoint to start a run and receive dataset items in one request:
curl -X POST "https://api.apify.com/v2/acts/crawlerbros~common-crawl-scraper/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"mode":"urlCaptures","urlPattern":"en.wikipedia.org/wiki/*","crawl":"latest","matchType":"auto","maxItems":15}'
The same run from Python, using the official client:
from apify_client import ApifyClient
client = ApifyClient("<YOUR_APIFY_TOKEN>")
run_input = {
"mode": "urlCaptures",
"urlPattern": "en.wikipedia.org/wiki/*",
"crawl": "latest",
"matchType": "auto",
"maxItems": 15
}
run = client.actor("crawlerbros~common-crawl-scraper").call(run_input=run_input)
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
print(item)
And from Node.js:
import { ApifyClient } from 'apify-client'
const client = new ApifyClient({ token: '<YOUR_APIFY_TOKEN>' })
const input = {
"mode": "urlCaptures",
"urlPattern": "en.wikipedia.org/wiki/*",
"crawl": "latest",
"matchType": "auto",
"maxItems": 15
}
const run = await client.actor('crawlerbros~common-crawl-scraper').call(input)
const { items } = await client.dataset(run.defaultDatasetId).listItems()
console.log(items)
The synchronous endpoint holds the connection open until the run finishes, which is convenient for small batches and wrong for large ones. For anything long running, start the run asynchronously and poll, or attach a webhook, so a dropped connection does not cost you the results.
Which Common Crawl Scraper inputs matter, and which can you skip?
The urlPattern is the primary control for targeting data; use wildcards like example.com/* to capture a site's entire directory. Most practitioners should leave statusFilter and mimeFilter empty on the first run to see the full variety of archived data before narrowing the scope.
mode(string): What to fetch. Default:"urlCaptures".urlPattern(string): Domain or URL to look up. Use*wildcards, e.g.example.com,example.com/*,en.wikipedia.org/wiki/*, or*.example.com. Default:"en.wikipedia.org/wiki/*".crawl(string): Which monthly crawl to query. Uselatestfor the newest, or a crawl id likeCC-MAIN-2024-10. Run mode=listCrawls to discover valid ids. Default:"latest".matchType(string): How to interpret the pattern.autolets the*wildcards decide. Default:"auto".statusFilter(integer): Only keep captures with this HTTP status (e.g. 200). Leave empty for all.mimeFilter(string): Only keep captures whose MIME type contains this text (e.g.text/html,pdf,json).fromDate(string): Only include captures on or after this date, e.g.20240101.toDate(string): Only include captures on or before this date, e.g.20241231.maxItems(integer): Hard cap on emitted records. Default:100.
Fixed-choice controls: mode accepts urlCaptures (lookup a domain / URL pattern), listCrawls (List available crawls); matchType accepts auto (use wildcards in the pattern), exact (Exact URL), prefix (path under a URL), host (one hostname), domain (host + all subdomains).
What does Common Crawl Scraper return?
The returned records are excellent for identifying deleted pages and mapping site structure across historical crawls. They do not contain the actual HTML body of the page, but they provide the warcUrl and offset needed to download the raw response from the archive.
Output - urlCaptures (one row per archived capture)
url- the archived URLurlKey- Common Crawl's canonical (SURT) keytimestamp- capture time,YYYYMMDDHHMMSScaptureDate- the same time as ISO 8601status- HTTP status at capture timemime- declared MIME typemimeDetected- MIME type detected from contentdigest- content digest (dedupe identical pages)length- record byte lengthoffset- byte offset within the WARC filefilename- WARC file path in the archivelanguages- detected language codesencoding- character encodingredirectUrl- redirect target (for 3xx captures)truncated- truncation reason (when present)crawlId- which crawl this came fromwarcUrl- direct link to the WARC file ondata.commoncrawl.orgrecordType: "capture",sourceUrl,scrapedAt
Output - listCrawls (one row per crawl)
crawlId- e.g.CC-MAIN-2024-10name- human-readable name (e.g.February/March 2024 Index)fromDate,toDate- crawl time windowcdxApiUrl- the crawl's index API endpointtimegateUrl- the crawl's timegaterecordType: "crawl",sourceUrl,scrapedAt
These are the documented fields. Optional ones can be empty on a given record, so measure how often each field your deliverable depends on is populated across a real sample before automating the handoff.
How do you build the workflow end to end?
Open Common Crawl Scraper and work through these in order. Each step ends with something to check, so a bad configuration surfaces on a small run rather than a scheduled one.
- Set the mode control to listCrawls to see all available monthly index identifiers.
- Identify a crawlId like CC-MAIN-2024-10 from the results to target a specific historical period.
- Switch the mode to urlCaptures and input your target domain into urlPattern using a wildcard like example.com/*.
- Choose the matchType that fits your scope, such as domain to include every subdomain found in the crawl.
- Use statusFilter set to 200 to exclude broken links and redirects from your final dataset.
- Perform a trial run with maxItems set to 15 to verify the url and captureDate fields align with your expectations.
- Check the warcUrl and offset fields to ensure you have the coordinates needed for raw content retrieval.
- Increase the maxItems value to your desired volume once the record structure is confirmed.
How do you apply it? Three worked playbooks
These are Common Crawl Scraper's own documented use cases, each worked through as an operating pattern rather than a description.
Use case 1: SEO & site audits
Outcome: Discover every URL a domain has ever exposed to crawlers
Configure: mode="urlCaptures", urlPattern="example.com/*", matchType="domain", statusFilter=200
Working method: Run a broad crawl across the latest index and compare the returned url list against your current XML sitemap to find orphaned pages.
Deliverable: A spreadsheet of historical URLs including the first seen captureDate and the content digest for each record.
Stop condition: The run completes with a result count matching the maxItems cap without reaching the end of the index.
Use case 2: Domain intelligence
Outcome: Profile a competitor's URL structure and content types
Configure: mode="urlCaptures", urlPattern="competitor.com/*", mimeFilter="text/html"
Working method: Query the URL index for a competitor's domain and group the output by urlKey to map their internal directory structure.
Deliverable: A dataset mapping the competitor's URL hierarchy and the distribution of their content types across monthly crawls.
Stop condition: The urlPattern returns zero results for three consecutive monthly crawl IDs.
Use case 3: Historical URL discovery
Outcome: Recover old / removed pages for migration or research
Configure: mode="urlCaptures", urlPattern="*.legacy-site.org", toDate="20211231"
Working method: Search for all captures preceding a specific date and use the record offset to prepare for bulk retrieval of deleted page HTML.
Deliverable: A CSV file containing the original URLs, their historical status codes, and exact WARC file locations.
Stop condition: The returned captureDate values exceed the date specified in the toDate filter.
What breaks, and how do you design around it?
When a search returns zero results, use the listCrawls mode to ensure the crawl ID you are querying is currently active in the index. For extremely large domains where you hit the 5,000 item limit, refine your search by applying more specific urlPattern wildcards or specific MIME types.
When should you not use Common Crawl Scraper?
Do not use this Actor if you need the most recent version of a page that was updated within the last few days, as Common Crawl operates on a monthly refresh cycle. If you require snapshots from a specific day or need to see how a page looked to a human user in a browser, Wayback Machine Search is the better choice. For researchers who only need to find all current URLs without historical context, Sitemap Sniffer will be faster and more cost-effective. This tool provides coordinates to content rather than the text itself; if your goal is audience research and reviews, consider Common Sense Media Scraper for structured entertainment data.
What should you check before trusting the output?
- Verify the urlKey follows the SURT format for reliable sorting and comparison across crawls.
- Confirm that the timestamp contains exactly 14 digits in the YYYYMMDDHHMMSS format.
- Check that the warcUrl points to data.commoncrawl.org for all records intended for raw data fetching.
- Ensure the digest field is populated for content deduplication between different monthly snapshots.
- Inspect the mimeDetected field to resolve ambiguity when the primary mime type is application/octet-stream.
None of this proves a record is correct. It gives a scheduled Common Crawl Scraper run defined points where it should stop instead of quietly passing bad data downstream.
Frequently asked questions
What is the free-plan price for 1,000 results?
The free-plan price is $5.00 per 1,000 results. Each individual result written to the dataset costs $0.005. While Apify's free plan provides $5.00 of monthly usage, you should account for the platform usage and the run-start fee that apply to every run. These additional costs mean that the $5.00 monthly credit will cover slightly fewer than 1,000 results in practice depending on the platform resources consumed during the execution.
How can I find out which monthly crawl IDs to use?
You should set the mode control to listCrawls. This run returns 9 fields per record, including the crawlId and a human-readable name. This allows you to see the exact date range covered by each monthly snapshot before you start querying for specific domain patterns using the urlCaptures mode.
Why am I getting zero results for a specific domain?
Common Crawl respects robots.txt directives, so some websites may not be included in the archive. If you get no results for a major site, try changing the crawl parameter to an older ID or check if the site blocks crawlers. You can also try matchType set to domain to include all subdomains that might have different crawl histories.
Does this Actor return the actual text of the archived pages?
No, it returns the metadata and location coordinates. Each result includes a warcUrl, filename, offset, and length. These four fields are the industry-standard way to locate and download the raw archived response directly from Common Crawl's public store at data.commoncrawl.org using a separate HTTP range request.
Can I filter results by content type in this Actor?
Yes, you can use the mimeFilter to only keep captures whose MIME type contains specific text like pdf or html. While this filter is applied to the data before it is written to your dataset, it helps ensure your final output is limited to the items you need for your specific research or audit task.
Where to go next
When you are ready to run it, open Common Crawl Scraper on Apify; the free plan covers up to 1,000 results a month.
Start with the Common Crawl Scraper Actor page for the current input schema, pricing tier, and run history.
Other Actors we maintain for related data:
- Wayback Machine Search: Query Internet Archive's Wayback Machine for historical snapshots of any URL or domain.
- Sitemap Sniffer: Discover every sitemap file for a website.
- Common Sense Media Scraper: Scrape CommonSenseMedia.org with search reviews, browse by content type (movies/games/books/apps/TV shows/websites), or fetch specific review pages.
- Contact Info Scraper Pro: Crawl any website and extract emails, phones, and social media profiles.
Related guides:
- Historical Web Data Extraction: Wayback Machine Search Playbooks
- Darkweb Scraper: Guide and 3 Core Use Cases
- Website Image Scraper: 13 Data Fields, Up to 2,500 Free Results/Month
- GitHub Repository Intelligence: Up to 2,500 Free Results a Month
- LinkedIn Events Scraper: Operating Playbooks and Workflows
Resources
Actor documentation, input schema, and pricing: verified against the published Actor on 2026-10-01.
Actor last updated by its maintainers on 2026-07-02.
Run outcome figures cover the 30 day public window ending 2026-10-01.
Featured actors
Common Crawl Scraper
Query the Common Crawl URL Index for any domain or URL pattern. Discover a site's archived pages, historical URLs, capture dates, HTTP statuses and MIME types for SEO, domain intelligence and research. Also lists the available monthly crawls.
Run on Apify ↗