Skip to content
    ↑↓ to choose · Enter to open

    · 12 min read

    Hacker News Scraper: Up to 2,500 Free Results a Month (2026)

    By CrawlerBros Engineering Team

    Each record carries 19 fields for stories, 14 fields for comments, and 10 fields for user profiles, pulled directly from official public APIs without requiring logins or proxies. A thousand results costs $2.00 on the free plan, and you can test it using Apify's free plan which includes $5.00 of monthly usage with no credit card. The scraper supports ten distinct modes ranging from top feeds to full-text search and direct ID lookups. It suits developers and analysts building digests, trend monitors, or datasets, but is not for anyone who needs buyer contact details, which the records do not include.

    Try it: open Hacker News Scraper on Apify, sign in on the free plan and run the prefilled example.

    Can you try Hacker News Scraper before paying?

    Yes. Apify's free plan includes $5.00 of prepaid usage every month and asks for no credit card. At $0.002 per result, that covers up to 2,500 results of Hacker News Scraper a month, before run-start charges and platform usage.

    Hacker News Scraper was last updated on 2026-05-02. It is one of 1,724 Actors CrawlerBros publishes on Apify, which together have 710,197 lifetime public runs and an average rating of 4.63 out of 5 across 416 reviews.

    What does it cost to run Hacker News Scraper?

    Each result costs $0.002 on Apify's free plan, which is $2.00 per 1,000 results. Starting a run is charged separately at $0.005 per GB of Actor memory. Apify also bills the platform usage each run consumes, at the rates of your Apify plan, on top of these charges.

    Apify plan Per result Per 1,000 results
    FREE $0.002 $2.00
    BRONZE $0.00167 $1.67
    SILVER $0.00133 $1.33
    GOLD $0.001 $1.00
    PLATINUM $0.001 $1.00
    DIAMOND $0.001 $1.00

    The main cost driver is maxItems, which places a hard cap on emitted records across stories, comments, and users combined. Increasing maxComments or maxDepth will fetch deeper reply trees per story, which increases total item count and raises the final result charge. The cheapest way to test this Actor is to run a small batch with maxItems set to a low value like ten before scaling up your collection.

    How do you run Hacker News Scraper from the API?

    The schema marks 1 of its 18 controls as required: mode. Nothing in the payload below is illustrative. Those are the schema's prefilled defaults for Hacker News Scraper, so the request works once your token is in place.

    Call the synchronous endpoint to start a run and receive dataset items in one request:

    curl -X POST "https://api.apify.com/v2/acts/crawlerbros~hackernews-scraper/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
      -H "Content-Type: application/json" \
      -d '{"mode":"topStories"}'
    

    The same run from Python, using the official client:

    from apify_client import ApifyClient
    
    client = ApifyClient("<YOUR_APIFY_TOKEN>")
    
    run_input = {
      "mode": "topStories"
    }
    
    run = client.actor("crawlerbros~hackernews-scraper").call(run_input=run_input)
    
    for item in client.dataset(run["defaultDatasetId"]).iterate_items():
        print(item)
    

    And from Node.js:

    import { ApifyClient } from 'apify-client'
    
    const client = new ApifyClient({ token: '<YOUR_APIFY_TOKEN>' })
    
    const input = {
      "mode": "topStories"
    }
    
    const run = await client.actor('crawlerbros~hackernews-scraper').call(input)
    const { items } = await client.dataset(run.defaultDatasetId).listItems()
    console.log(items)
    

    Because the call is synchronous, your client waits for the whole run. Keep it for exploration. For scheduled work, start the run without waiting and collect the dataset afterwards, so network trouble costs you a retry rather than the results.

    Which Hacker News Scraper inputs matter, and which can you skip?

    The mode control determines what data the Actor fetches, offering options like topStories, search, item, and user. Most users should start with topStories or search and leave advanced filters like commentAuthorFilter and domainBlocklist empty on a first run.

    • mode (string): What to scrape. topStories / newStories / bestStories / askStories / showStories / jobStories walk the public feeds. item fetches specific item IDs (itemIds). user fetches profiles (usernames). search uses the Algolia HN API (searchQuery). Default: "topStories".
    • itemIds (array): Direct Hacker News item IDs (numeric strings or integers). Used when mode=item. Default: [].
    • usernames (array): Hacker News usernames. Used when mode=user. Default: [].
    • searchQuery (string): Free-text query passed to the Algolia HN search API.
    • startUrls (array): Hacker News URLs to scrape directly. Supports news.ycombinator.com/item?id=N and news.ycombinator.com/user?id=USER. Each URL is converted to an item or user fetch. Default: [].
    • enableCommentHierarchy (boolean): When true, attach a nested comments[] array to each story (recursive replies). When false, comments are emitted as flat sibling records (one per row, with parentId and depth). Default: false.
    • maxItems (integer): Hard cap on emitted records (stories + comments + users combined). Default: 100.
    • maxComments (integer): Cap on comments fetched per story. 0 = no comments. Default 0 (skip comments). Default: 0.
    • maxDepth (integer): Maximum reply depth when fetching comments (only applies when maxComments > 0). Root comments are depth 0. Default: 5.
    • minScore (integer): Drop stories with fewer points than this. Comments are unaffected.
    • domainAllowlist (array): Only emit stories whose URL host contains one of these substrings (case-insensitive). e.g. ['github.com','arxiv.org']. Default: [].
    • domainBlocklist (array): Drop stories whose URL host contains one of these substrings. e.g. ['twitter.com','x.com']. Default: [].

    The other 6 controls, with their defaults, are listed in the input schema on Hacker News Scraper on Apify.

    Fixed-choice controls: mode accepts 9 values (default topStories), including topStories (Top stories), newStories (New stories), bestStories (Best stories), askStories (Ask HN).

    What does Hacker News Scraper return?

    The returned records are ideal for building newsletters, content archives, or sentiment analysis datasets from real-time Hacker News activity. They conspicuously do not contain private user email addresses, external enrichment metadata, or purchase contact details.

    Output per story

    • id, hnUrl, type, title
    • url, domain (parsed host)
    • score, numComments, author
    • createdAt (ISO-8601 UTC), createdAtEpoch, ageHours
    • text (Ask/Show/Job posts only - HN-flavored markdown converted to plain text)
    • kids (list of top-level comment IDs)
    • comments[] (nested replies - only when enableCommentHierarchy=true)
    • dead / deleted (boolean - only when true)
    • recordType: "story", scrapedAt

    Output per comment

    • id, hnUrl, storyId, parentId, depth
    • author, text (HTML → plain text)
    • createdAt, createdAtEpoch, ageHours
    • kids (list of reply IDs), replies[] (nested - only when enableCommentHierarchy=true)
    • recordType: "comment", scrapedAt

    Output per user

    • id, profileUrl
    • karma, createdAt, createdAtEpoch, about
    • submittedCount, submitted (capped at 50 most-recent)
    • recordType: "user", scrapedAt

    These are the documented fields. Optional ones can be empty on a given record, so measure how often each field your deliverable depends on is populated across a real sample before automating the handoff.

    How do you build the workflow end to end?

    Open Hacker News Scraper and work through these in order. Each step ends with something to check, so a bad configuration surfaces on a small run rather than a scheduled one.

    1. Set the mode control to topStories, newStories, bestStories, askStories, showStories, jobStories, item, user, or search.
    2. Adjust maxItems to bound the total emitted records across stories, comments, and users.
    3. Configure maxComments and maxDepth if you want to pull comment threads alongside stories.
    4. Apply minScore, domainAllowlist, or domainBlocklist to filter out unwanted story submissions.
    5. Set dateRangeFrom and dateRangeTo using ISO YYYY-MM-DD format to isolate specific date windows.
    6. Add commentAuthorFilter or commentMinScore if you need to refine comment output before executing the run.

    How do you apply it? Three worked playbooks

    These are Hacker News Scraper's own documented use cases, each worked through as an operating pattern rather than a description.

    Use case 1: Trend monitoring

    Outcome: Track which domains hit the front page each week

    Configure: Set mode to topStories, set maxItems to 100, and supply domainAllowlist or domainBlocklist as needed.

    Working method: Run once with a broad topStories feed, then check that the returned records contain populated domain and score fields before applying blocklists.

    Deliverable: A dataset of top story records containing id, title, domain, score, and createdAt fields.

    Stop condition: Zero records return a recognized domain attribute or the dataset comes back empty.

    Use case 2: Comment intelligence

    Outcome: Pull every comment for an Ask HN thread to study reactions

    Configure: Set mode to item, supply target numeric identifiers in itemIds, set maxComments to 500, and enableCommentHierarchy to true.

    Working method: Execute a test fetch on a single item ID first, inspect the nested comments array for depth integrity, and then scale up maxComments.

    Deliverable: A structured JSON dataset containing story records with fully nested comment threads and author details.

    Stop condition: Returned comment counts drop to zero despite active discussion on the source thread.

    Use case 3: YC job-ads digest

    Outcome: Weekly extract of jobStories for the careers newsletter

    Configure: Set mode to jobStories, set maxItems to 50, and specify dateRangeFrom to filter recent posts.

    Working method: Run the jobStories mode, verify that the text field contains the full listing description, and review the createdAt timestamps.

    Deliverable: A clean dataset of current startup job postings with titles, links, and posting dates.

    Stop condition: Job records return missing text bodies or duplicate stale identifiers from prior weeks.

    What breaks, and how do you design around it?

    Hacker News rarely exposes per-comment scores, so commentMinScore will skip items without scores rather than filtering them reliably. When fetching deep comment hierarchies, keep maxDepth reasonable to avoid large payloads and memory constraints.

    When should you not use Hacker News Scraper?

    Do not use this Actor if you require official historical database dumps with sub-second analytical querying across millions of archived items, as live API walking has natural throughput bounds. If you only need startup job board postings and nothing else, use Hacker News Jobs Scraper to fetch direct listings without processing general community feeds or comment trees. Similarly, if your primary goal is broad multi-platform community mining that includes comment threads and keyword filters across different sites, consider Reddit MCP Scraper Pro (Multi-Mode) instead.

    What should you check before trusting the output?

    • Verify that the recordType field correctly matches story, comment, or user for every emitted dataset item.
    • Check that null or empty fields are omitted from the output rather than written as empty strings.
    • Confirm that url and text fields align correctly with story types, noting that Ask HN and Show HN posts carry text instead of a url.
    • Stop scheduled runs if commentMinScore returns unexpected results due to Hacker News sporadically exposing comment scores.
    • Inspect nested comments[] or replies[] structures when enableCommentHierarchy is set to true to ensure depth boundaries hold.

    None of this proves a record is correct. It gives a scheduled Hacker News Scraper run defined points where it should stop instead of quietly passing bad data downstream.

    Frequently asked questions

    Does Hacker News Scraper require a login or cookies?

    No. Both Firebase and Algolia HN APIs are fully public, meaning you can execute runs without supplying login credentials, session cookies, or proxy configurations.

    Why are some stories missing a url field in the dataset?

    Ask HN and Show HN posts are self-text submissions. They have a text field instead of an external url, and the Actor's omit-empty contract drops the url field entirely on these records.

    What is the difference between maxItems and maxComments?

    maxItems places a hard cap on total emitted records combining stories, comments, and users. maxComments caps how many comments the Actor fetches into each story's comment tree per story.

    How fresh is the data returned by the Actor?

    The data is retrieved in real-time. Both underlying APIs serve the live Hacker News database directly without caching delays.

    How much does a run cost on the free tier?

    On the free tier, each result costs $0.002, which equals $2.00 per 1,000 results. Apify's free plan includes $5.00 of monthly usage with no credit card required, covering up to 2,500 results before run-start charges.

    Where to go next

    When you are ready to run it, open Hacker News Scraper on Apify; the free plan covers up to 2,500 results a month.

    Start with the Hacker News Scraper Actor page for the current input schema, pricing tier, and run history.

    If you are comparing approaches rather than committing to one Actor, these category pages list every option we publish:

    Other Actors we maintain for related data:

    Related guides:

    Resources

    • Actor documentation, input schema, and pricing: verified against the published Actor on 2026-09-30.

    • Actor last updated by its maintainers on 2026-05-02.

    • Run outcome figures cover the 30 day public window ending 2026-09-30.

    • Hacker News Scraper on Apify

    Featured actors

    Hacker News Scraper

    Scrape Hacker News stories, comments, jobs, and user profiles. Modes: top/new/best/ask/show/jobs/past/item/user/search. Filters: minScore, domainFilter, dateRange, commentMinScore. No proxy, no auth.

    Run on Apify ↗