Skip to content
    ↑↓ to choose · Enter to open

    · 12 min read

    Substack Scraper: 16 Data Fields, Up to 1,000 Free Results/Month

    By CrawlerBros Engineering Team

    Each post record carries 16 fields, including post title, full body HTML, author, publication date, categories, and cover image URLs extracted directly from public RSS feeds. A thousand results costs $5.00 on the free plan, which includes $5.00 of monthly prepaid usage so you can process up to 1,000 posts each month without entering a credit card. The Actor resolves full URLs, custom domains, and bare Substack subdomains automatically. This tool is built for content analysts, market researchers, and media monitoring teams tracking public newsletter posts; it is not for anyone needing full paywalled post text, as Substack RSS feeds only include free posts and public previews.

    Try it: open Substack Scraper on Apify, sign in on the free plan and run the prefilled example.

    Can you try Substack Scraper before paying?

    Yes. Apify's free plan includes $5.00 of prepaid usage every month and asks for no credit card. At $0.005 per result, that covers up to 1,000 results of Substack Scraper a month, before run-start charges and platform usage.

    The example request further down caps maxItems at 50, so a first run returns at most 50 results and costs at most $0.25 in result charges. That is enough to see the real shape of the data before deciding anything.

    Substack Scraper was last updated on 2026-05-06. It is one of 1,725 Actors CrawlerBros publishes on Apify, which together have 686,270 lifetime public runs and an average rating of 4.63 out of 5 across 416 reviews.

    What does it cost to run Substack Scraper?

    Each result costs $0.005 on Apify's free plan, which is $5.00 per 1,000 results. Starting a run is charged separately at $0.005 per GB of Actor memory. Apify also bills the platform usage each run consumes, at the rates of your Apify plan, on top of these charges.

    Apify plan Per result Per 1,000 results
    FREE $0.005 $5.00
    BRONZE $0.00433 $4.33
    SILVER $0.00367 $3.67
    GOLD $0.003 $3.00
    PLATINUM $0.003 $3.00
    DIAMOND $0.003 $3.00

    Result charges are driven directly by the total number of dataset items written, which you control using the maxItems setting and the number of publications in your input. The publications array and filters like containsKeyword dictate how many total post records are matched and written. To verify your setup cheaply before scaling, run a test with maxItems capped at 50, which limits result charges to $0.25 on the free tier.

    How do you run Substack Scraper from the API?

    The schema marks 1 of its 6 controls as required: publications. Nothing in the payload below is illustrative. Those are the schema's prefilled defaults for Substack Scraper, so the request works once your token is in place.

    Call the synchronous endpoint to start a run and receive dataset items in one request:

    curl -X POST "https://api.apify.com/v2/acts/crawlerbros~substack-scraper/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
      -H "Content-Type: application/json" \
      -d '{"publications":["platformer.news"],"categoryAnyOf":[],"includeBody":true,"maxItems":50}'
    

    The same run from Python, using the official client:

    from apify_client import ApifyClient
    
    client = ApifyClient("<YOUR_APIFY_TOKEN>")
    
    run_input = {
      "publications": [
        "platformer.news"
      ],
      "categoryAnyOf": [],
      "includeBody": True,
      "maxItems": 50
    }
    
    run = client.actor("crawlerbros~substack-scraper").call(run_input=run_input)
    
    for item in client.dataset(run["defaultDatasetId"]).iterate_items():
        print(item)
    

    And from Node.js:

    import { ApifyClient } from 'apify-client'
    
    const client = new ApifyClient({ token: '<YOUR_APIFY_TOKEN>' })
    
    const input = {
      "publications": [
        "platformer.news"
      ],
      "categoryAnyOf": [],
      "includeBody": true,
      "maxItems": 50
    }
    
    const run = await client.actor('crawlerbros~substack-scraper').call(input)
    const { items } = await client.dataset(run.defaultDatasetId).listItems()
    console.log(items)
    

    Because the call is synchronous, your client waits for the whole run. Keep it for exploration. For scheduled work, start the run without waiting and collect the dataset afterwards, so network trouble costs you a retry rather than the results.

    Which Substack Scraper inputs matter, and which can you skip?

    The publications field is the only required input and accepts full URLs, custom domains like platformer.news, or bare slugs like noahpinion. Optional controls like publishedAfter and containsKeyword let you narrow the returned post set by date and content strings before writing to the dataset. For your initial run, keep default settings for includeBody and set maxItems to a small number to review sample output.

    • publications (array): Substack publication URLs or subdomains (e.g. https://platformer.news, noahpinion.substack.com, just noahpinion). Default: ["platformer.news","noahpinion.substack.com"].
    • categoryAnyOf (array): Only emit posts tagged with at least one of these categories (matches RSS <category> tags case-insensitively). Default: [].
    • publishedAfter (string): Drop posts published before this date.
    • containsKeyword (string): Only emit posts whose title or description contains this substring (case-insensitive).
    • includeBody (boolean): Include the full post body HTML (content:encoded from RSS). Default: true.
    • maxItems (integer): Hard cap on emitted records. Default: 50.

    What does Substack Scraper return?

    Returned records provide structured post data including body HTML, plain-text summaries, reading time estimates, categories, and original publication metadata. They are ideal for topic modeling, media monitoring, and market research across independent publications. Note that for paid-only Substack posts, the RSS output contains only public preview snippets rather than the complete subscriber text.

    • title, url, guid
    • author - from <dc:creator>
    • publishedAt - ISO 8601 UTC (parsed from RFC 822 pubDate)
    • publishedAtRaw - original RFC 822 string
    • summary - plain-text version of <description> (capped at 500 chars)
    • bodyHtml - full HTML body from <content:encoded> (when includeBody=true)
    • wordCount, readingTimeMinutes
    • categories[]
    • coverImage - from <enclosure> URL
    • publication, publicationUrl
    • recordType: "post", scrapedAt

    These are the documented fields. Optional ones can be empty on a given record, so measure how often each field your deliverable depends on is populated across a real sample before automating the handoff.

    How do you build the workflow end to end?

    Open Substack Scraper and work through these in order. Each step ends with something to check, so a bad configuration surfaces on a small run rather than a scheduled one.

    1. Populate the publications array with target domains, subdomains, or bare slugs such as platformer.news or noahpinion.
    2. Set the publishedAfter filter using the YYYY-MM-DD format to skip archived posts older than your scope.
    3. Specify target keywords in containsKeyword if you only want posts matching particular terms in the title or summary.
    4. Keep includeBody set to true if you need full HTML content:encoded output, or set it to false to return lighter metadata records.
    5. Set maxItems to cap the total emitted records across all publications, starting with 50 for initial tests.
    6. Run the Actor and inspect the dataset to confirm that output fields like title, bodyHtml, and categories contain expected values.
    7. Verify that empty fields are omitted properly and check that publicationUrl accurately reflects the source feed.

    How do you apply it? Three worked playbooks

    These are Substack Scraper's own documented use cases, each worked through as an operating pattern rather than a description.

    Use case 1: Newsletter intel

    Outcome: Track competitor publications, harvest content

    Configure: Set publications to ["platformer.news", "noahpinion.substack.com"], set includeBody to true, set publishedAfter to "2024-01-01", and set maxItems to 100.

    Working method: Start by targeting two major competitor publication domains in the publications array. Execute the run and review the output dataset to ensure fields like title, author, and bodyHtml populate cleanly across both feeds. Use the publishedAfter filter to maintain a rolling window of recent market coverage.

    Deliverable: A structured JSON dataset containing post records with full body HTML, author metadata, and publication timestamps from competitor newsletters.

    Stop condition: Stop execution if a targeted publication feed consistently returns 403 or 429 status codes after exponential backoff retries.

    Use case 2: Market research

    Outcome: Newsletters in your domain (analyst notes, sector reports)

    Configure: Set publications to ["platformer.news"], set containsKeyword to "antitrust", set includeBody to true, and set maxItems to 50.

    Working method: Define a set of industry newsletters and apply the containsKeyword control to isolate target terms. Verify that emitted items match the keyword in either the title or summary string. Filter down the results to relevant sector reports and analysis notes without harvesting unrelated posts.

    Deliverable: A dataset of targeted newsletter posts filtered by domain keywords, containing plain-text summaries, categories, and post URLs.

    Stop condition: Halt the run if containsKeyword filters out all results across every listed publication, indicating overly restrictive search terms.

    Use case 3: RSS aggregation

    Outcome: Consolidate multiple Substacks into a single feed

    Configure: Set publications to ["noahpinion", "thedailyupside"], set categoryAnyOf to [], set includeBody to false, and set maxItems to 500.

    Working method: Provide bare slugs in the publications array to test auto-resolution to standard Substack subdomains. Collect records across multiple feeds in a single batch while leaving includeBody set to false for fast metadata extraction. Ensure post URLs and GUIDs are unique across the consolidated output dataset.

    Deliverable: A unified feed of multi-publication post metadata including titles, URLs, published dates, and word counts without full HTML payloads.

    Stop condition: Stop if the output dataset contains duplicate post URLs across different publication entries.

    What breaks, and how do you design around it?

    Substack limits public RSS feeds to recent posts, so older archive items may not appear in the feed. When harvesting historical content, set publishedAfter to focus strictly on available dates. If a target feed returns 403, 429, or 5xx HTTP status codes, the Actor retries up to 3 times with exponential backoff before skipping that publication to ensure the rest of your run completes.

    When should you not use Substack Scraper?

    Do not use this Actor if you need full text from paid subscriber-only newsletters, because Substack RSS feeds only expose public post previews. If you need to monitor general tech blogging or developer articles beyond Substack, Dev.to Scraper provides native API access to developer posts, tags, and comments. If you are scraping platform-agnostic news websites using HTML discovery and sitemaps rather than RSS feeds, News Source Crawler is the better choice. Avoid this tool if you require social interaction metrics like comments or subscriber counts, as public RSS XML does not track engagement numbers.

    What should you check before trusting the output?

    • Confirm that publishedAt is formatted as a valid ISO 8601 UTC string rather than falling back to publishedAtRaw.
    • Check that summary text contains plain text capped at 500 characters without residual unparsed RSS tags.
    • Verify that coverImage contains a valid URL string when enclosure tags exist in the target feed.
    • Set an automated alert to halt execution if bodyHtml returns null while includeBody is explicitly set to true.
    • Monitor maxItems execution output to verify that total returned items do not exceed the set integer limit.

    None of this proves a record is correct. It gives a scheduled Substack Scraper run defined points where it should stop instead of quietly passing bad data downstream.

    Frequently asked questions

    How much does it cost to run Substack Scraper?

    Substack Scraper costs $0.005 per result written to your dataset, which equals $5.00 per 1,000 results on Apify's free plan. Paid Apify tiers offer lower per-result charges, down to $3.00 per 1,000 results on Gold, Platinum, and Diamond plans. Platform usage is billed separately by Apify for each run.

    Can I extract paid Substack newsletter posts?

    No. The scraper reads public RSS feeds exposed by Substack publications. These feeds contain the complete text for free public posts, but only provide short public previews for subscriber-only paid posts. Accessing full paid content requires account authentication, which this scraper does not support.

    Does Substack Scraper work with custom domains?

    Yes. You can pass custom domains such as platformer.news directly into the publications array. The scraper automatically appends /feed to the root domain to locate and parse the underlying RSS feed regardless of whether it uses a standard substack.com address.

    How does the scraper handle Substack rate limits?

    The scraper uses HTTP requests with curl_cffi Chrome TLS impersonation to bypass basic edge blocks. If a publication feed returns 403, 429, or 5xx HTTP status codes, the Actor retries up to 3 times with exponential backoff. If it still fails, it logs a warning and proceeds to the next publication.

    Can I filter posts by keyword or publish date?

    Yes. You can pass a date string in YYYY-MM-DD format to publishedAfter to ignore older posts. You can also supply a substring in containsKeyword to restrict the output to posts where the title or plain-text summary matches your target term case-insensitively.

    Where to go next

    When you are ready to run it, open Substack Scraper on Apify; the free plan covers up to 1,000 results a month.

    Start with the Substack Scraper Actor page for the current input schema, pricing tier, and run history.

    Other Actors we maintain for related data:

    • Medium Article Scraper: Scrape Medium articles by tag/topic, user, publication, or search query.
    • Yahoo News Scraper: Scrape Yahoo News articles across categories, trending stories, and keyword search.
    • News Source Crawler: Given a news website URL, discover and extract articles with full metadata with title, authors, publish date, body text, top image, keywords, and summary.
    • Beehiiv Newsletter Discovery Scraper: Discover and scrape newsletters from Beehiiv's public directory.
    • Dev.to Scraper: Scrape Dev.to, the popular blogging platform for developers (forem.com).
    • TikTok Explore/Trending Scraper: Scrape TikTok's Explore/Trending feed across categories.
    • LinkedIn Post Scraper: Scrape posts from any LinkedIn personal profile activity feed.
    • Ghost Blog Scraper: Scrape posts, pages, tags, authors, and site info from any Ghost CMS-powered blog via Ghost's public Content API - including ghost.org's own /resources, /changelog, and /help sections.

    Related guides:

    Resources

    • Actor documentation, input schema, and pricing: verified against the published Actor on 2026-09-28.

    • Actor last updated by its maintainers on 2026-05-06.

    • Run outcome figures cover the 30 day public window ending 2026-09-28.

    • Substack Scraper on Apify

    Featured actors

    Substack Scraper

    Scrape Substack publications via the public RSS feed of any newsletter. Extract post title, URL, author, publication date, body HTML, categories, and enclosures. HTTP-only with TLS impersonation (no auth, no proxy).

    Run on Apify ↗