Skip to content
    ↑↓ to choose · Enter to open

    · 15 min read

    Hacker News Stories, Comments & Users Scraper: $5.00 per 1,000 Results

    By CrawlerBros Engineering Team

    Hacker News Stories, Comments & Users Scraper provides 11 fields for story records and 13 fields for comment records, enabling detailed analysis of content and user activity. This Actor works by utilizing the official Algolia HN Search API and Hacker News Firebase API, requiring no authentication. It's suitable for practitioners who need to gather specific discussions, track trends, or identify influential users from Hacker News. This Actor is not for those who need private user data or discussions from non-public channels, as it only accesses publicly available information.

    Try it before you read further. Apify's free plan includes $5.00 of usage every month with no credit card, enough for up to 1,000 results at $0.005 each before platform usage. Open Hacker News Stories, Comments & Users Scraper on Apify and run the prefilled example.

    How reliable is Hacker News Stories, Comments & Users Scraper in production?

    Across the last 30 days of public runs on the Apify platform, Hacker News Stories, Comments & Users Scraper recorded 57 runs with the following outcomes.

    Outcome Runs Share
    Succeeded 57 100.0%
    Failed 0 0.0%
    Aborted by the user 0 0.0%
    Timed out 0 0.0%
    Total 57 100.0%

    No run failed or timed out in the last 30 days. Keep a retry and an alert on scheduled runs all the same: a clean month is a record, not a guarantee.

    What does it cost to run Hacker News Stories, Comments & Users Scraper?

    Each result costs $0.005 on Apify's free plan, which is $5.00 per 1,000 results. Starting a run is charged separately at $0.005 per GB of Actor memory. Apify also bills the platform usage each run consumes, at the rates of your Apify plan, on top of these charges.

    Apify plan Per result Per 1,000 results
    FREE $0.005 $5.00
    BRONZE $0.00433 $4.33
    SILVER $0.00367 $3.67
    GOLD $0.003 $3.00
    PLATINUM $0.003 $3.00
    DIAMOND $0.003 $3.00

    Worked example: collecting 10,000 results costs $50.00 in result charges before run-start fees and platform usage. No run failed or timed out in the last 30 days, so the list price is a fair budget; keep a retry in place all the same.

    The primary factor affecting your bill is the "maxItems" input, which caps the total number of records returned. Each result written to your dataset incurs a charge. To estimate costs and determine if this Actor meets your needs before committing significant spend, start with a low "maxItems" value (such as the example input's default of 50). This will return a small sample of data for a minimal cost.

    How do you run Hacker News Stories, Comments & Users Scraper from the API?

    The schema marks 1 of its 10 controls as required: mode. Nothing in the payload below is illustrative. Those are the schema's prefilled defaults for Hacker News Stories, Comments & Users Scraper, so the request works once your token is in place.

    Call the synchronous endpoint to start a run and receive dataset items in one request:

    curl -X POST "https://api.apify.com/v2/acts/crawlerbros~hacker-news-scraper/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
      -H "Content-Type: application/json" \
      -d '{"mode":"searchStories","searchQuery":"AI startup","startUrls":[],"storyType":"all","sortBy":"relevance","maxItems":50}'
    

    The same run from Python, using the official client:

    from apify_client import ApifyClient
    
    client = ApifyClient("<YOUR_APIFY_TOKEN>")
    
    run_input = {
      "mode": "searchStories",
      "searchQuery": "AI startup",
      "startUrls": [],
      "storyType": "all",
      "sortBy": "relevance",
      "maxItems": 50
    }
    
    run = client.actor("crawlerbros~hacker-news-scraper").call(run_input=run_input)
    
    for item in client.dataset(run["defaultDatasetId"]).iterate_items():
        print(item)
    

    And from Node.js:

    import { ApifyClient } from 'apify-client'
    
    const client = new ApifyClient({ token: '<YOUR_APIFY_TOKEN>' })
    
    const input = {
      "mode": "searchStories",
      "searchQuery": "AI startup",
      "startUrls": [],
      "storyType": "all",
      "sortBy": "relevance",
      "maxItems": 50
    }
    
    const run = await client.actor('crawlerbros~hacker-news-scraper').call(input)
    const { items } = await client.dataset(run.defaultDatasetId).listItems()
    console.log(items)
    

    Because the call is synchronous, your client waits for the whole run. Keep it for exploration. For scheduled work, start the run without waiting and collect the dataset afterwards, so network trouble costs you a retry rather than the results.

    Which Hacker News Stories, Comments & Users Scraper inputs matter, and which can you skip?

    The mode control is required and dictates the type of data to scrape, such as searching stories or comments, or fetching user submissions. For most users, begin by setting mode and then providing a searchQuery or username relevant to your goal. Many users can leave controls like sortBy and storyType at their defaults for an initial run.

    • mode (string): What to scrape. Default: "searchStories".
    • searchQuery (string): Keyword or phrase to search for (modes: searchStories, searchComments).
    • username (string): Hacker News username to fetch submissions for (mode: byUser).
    • startUrls (array): Hacker News item IDs (numeric) or URLs like https://news.ycombinator.com/item?id=12345 (mode: getItem). Default: [].
    • storyType (string): Filter by story type (applies to searchStories). Default: "all".
    • sortBy (string): Sort order for search results. Default: "relevance".
    • dateFrom (string): Only return items created on or after this date (ISO format, e.g. 2024-01-01). Applies to search modes.
    • dateTo (string): Only return items created on or before this date (ISO format, e.g. 2024-12-31). Applies to search modes.
    • minPoints (integer): Only emit stories/comments with at least this many points.
    • maxItems (integer): Maximum number of records to emit. Default: 50.

    Fixed-choice controls: mode accepts searchStories (Search stories by keyword), searchComments (Search comments by keyword), topStories (front page), newStories (latest), bestStories (all-time top), byUser (All submissions by a user), getItem (Get specific item(s) by ID or URL); storyType accepts all (All types), story (Regular story), ask_hn (Ask HN), show_hn (Show HN), job (Job posting); sortBy accepts relevance, date (newest first).

    What does Hacker News Stories, Comments & Users Scraper return?

    The output records are structured to provide comprehensive data for either stories or comments. Story records include fields like title, url, author, points, and commentsCount, making them suitable for trend analysis or content research. Comment records include text, author, points, and links back to their storyId and storyTitle, which is useful for sentiment analysis or identifying engaged users. These records do not contain direct contact information for authors.

    Story record

    • storyId (e.g. 39876543)
    • type (e.g. story)
    • title (e.g. Show HN: I built a faster way to index large ...)
    • url
    • hnUrl (e.g. https://news.ycombinator.com/item?id=39876543)
    • author (e.g. johndoe)
    • points (e.g. 312)
    • commentsCount (e.g. 87)
    • createdAt (e.g. 2024-03-15T14:22:00+00:00)
    • recordType (e.g. story)
    • scrapedAt (e.g. 2024-03-15T15:00:00+00:00)

    Comment record

    • commentId (e.g. 39876999)
    • type (e.g. comment)
    • text
    • hnUrl (e.g. https://news.ycombinator.com/item?id=39876999)
    • author (e.g. janedoe)
    • points (e.g. 42)
    • storyId (e.g. 39876543)
    • storyTitle (e.g. Show HN: I built a faster way to index large ...)
    • storyUrl
    • storyHnUrl (e.g. https://news.ycombinator.com/item?id=39876543)
    • createdAt (e.g. 2024-03-15T14:45:00+00:00)
    • recordType (e.g. comment)
    • scrapedAt (e.g. 2024-03-15T15:00:00+00:00)

    These are the documented fields. Optional ones can be empty on a given record, so measure how often each field your deliverable depends on is populated across a real sample before automating the handoff.

    How do you build the workflow end to end?

    Open Hacker News Stories, Comments & Users Scraper and work through these in order. Each step ends with something to check, so a bad configuration surfaces on a small run rather than a scheduled one.

    1. Set the "mode" to "searchStories" or "searchComments" for keyword-based extraction, or to "byUser" for user-specific data, considering the type of information you need.
    2. Populate the "searchQuery" field with your keywords if using search modes, or the "username" field if fetching user submissions.
    3. Specify a "dateFrom" and "dateTo" in YYYY-MM-DD format to narrow down results for search modes, which can help manage the volume of data.
    4. For story searches, consider applying a "storyType" filter (e.g., "ask_hn", "show_hn", "job") to focus on specific discussion formats.
    5. Set "minPoints" to filter for more impactful or popular stories and comments, starting with a value like 10 or 20 to see relevant items.
    6. Limit the total number of records by setting "maxItems" to a reasonable value like 100 or 200 for your initial runs.
    7. After the run, check the dataset to ensure the type field correctly identifies whether records are story or comment and that relevant fields like title and text are populated.
    8. Verify that scrapedAt timestamps are recent and that author fields are present, indicating complete record extraction.

    How do you apply it? Three worked playbooks

    These are Hacker News Stories, Comments & Users Scraper's own documented use cases, each worked through as an operating pattern rather than a description.

    Use case 1: Market research

    Outcome: Track mentions of products, technologies, or companies

    Configure: Set the "mode" to "searchStories" or "searchComments". Set "searchQuery" to a product or company name, e.g., "AI chip startups". Set "dateFrom" and "dateTo" to define a relevant period.

    Working method: Begin by searching for a single, specific product or company name. Review the initial results for relevance and completeness. Then, broaden your search by adding related keywords or expanding the date range.

    Deliverable: A dataset containing stories or comments mentioning the specified product, technology, or company, with associated metadata like title, url, author, and createdAt dates.

    Stop condition: Search results contain a high percentage of irrelevant mentions or the title or text fields are consistently empty for relevant items.

    Use case 2: Competitor intelligence

    Outcome: Search for mentions of competitors or alternatives

    Configure: Set "mode" to "searchStories" and "searchQuery" to a competitor's name, e.g., "AcmeCorp AI". Adjust "sortBy" to "date" for recent mentions.

    Working method: Start with a precise search for one competitor. Examine the hnUrl and title fields to see how and where they are being discussed. Expand to alternative names or product lines, or compare mentions across a few key competitors.

    Deliverable: A structured list of Hacker News stories where competitors or their products are mentioned, including author, points, and direct links to the discussions.

    Stop condition: The commentsCount or points for competitor mentions are consistently low, indicating limited discussion volume or impact, or the search returns too many false positives.

    Use case 3: Lead generation

    Outcome: Identify users active in your target domain

    Configure: Set "mode" to "searchStories" or "searchComments". Use a "searchQuery" related to your target domain, e.g., "SaaS founder" or "backend developer". Set "minPoints" to a value like 50 to find active contributors.

    Working method: First, run a broad search for relevant keywords in stories or comments. Review the author fields in the results to identify potential users of interest. Then, for specific promising authors, switch to byUser mode with their username to get their full submission history.

    Deliverable: A list of Hacker News usernames who have actively engaged in discussions or submitted stories within your target domain, along with their related story or comment data.

    Stop condition: Identified users consistently have very few storyId or commentId entries in their byUser history, suggesting low activity or an incorrect target domain definition.

    What breaks, and how do you design around it?

    • Over the last 30 days, 0.0% of public runs failed and 0.0% timed out. Build retries and alerting around those rates rather than assuming every run completes.

    When dealing with the 5,000-item maximum, consider splitting your task into multiple runs by adjusting the dateFrom and dateTo parameters to cover different time periods. If you need to search a very broad range of dates, it's best to iterate through smaller date windows rather than attempting one massive query. For very long lists of startUrls or usernames, process them in batches to prevent individual run timeouts.

    When should you not use Hacker News Stories, Comments & Users Scraper?

    This Actor is not suitable if your primary need is real-time, low-latency data for individual posts or comments, as its design is geared towards batch data extraction. For single-item lookups that require immediate responses, directly querying the Algolia HN Search API or Hacker News Firebase API might be a more efficient approach. If you are trying to acquire data from very obscure or niche discussions that are unlikely to garner significant points or comments, the minPoints filter might exclude all relevant results, in which case a more general search Actor or manual browsing might be necessary. Also, if your goal is to monitor job postings exclusively, the Hacker News Jobs Scraper offers a more specialized solution focused solely on the "Who is Hiring" board.

    What should you check before trusting the output?

    • Check for storyId or commentId values that are null or malformed, as this indicates a failed item extraction.
    • Inspect title fields in story records and text fields in comment records for unexpected characters or truncated content.
    • Confirm that points and commentsCount (for stories) or points (for comments) are numeric and within expected ranges; unusually low or zero values might indicate incomplete data.
    • Look for createdAt timestamps that are outside your specified dateFrom and dateTo ranges, or are missing entirely.
    • Validate url fields in story records for accessibility and hnUrl fields for both story and comment records for correctness.
    • If using minPoints, ensure that all returned records indeed meet or exceed the specified minimum.

    None of this proves a record is correct. It gives a scheduled Hacker News Stories, Comments & Users Scraper run defined points where it should stop instead of quietly passing bad data downstream.

    Frequently asked questions

    What is the cost for 1,000 results?

    On Apify's free plan, this Actor costs $5.00 for every 1,000 results. This allows you to collect a substantial amount of data before needing to consider a paid plan. The platform also includes $5.00 of monthly usage without requiring a credit card, which covers up to 1,000 results. Be aware that run-start fees and platform usage are billed separately.

    How reliable is this Actor for unattended runs?

    The Actor has a 100.0% success rate across 57 runs in the last 30 days, indicating high reliability. This means you can schedule it for unattended execution with confidence that it will complete successfully. You should not expect any runs to fail or time out based on recent telemetry, making it suitable for automated workflows without frequent intervention.

    Can I filter search results by a specific date range?

    Yes, you can filter search results by date. Use the dateFrom and dateTo input fields to specify your desired range in YYYY-MM-DD format. This applies to searchStories and searchComments modes, allowing you to focus on content published within a particular timeframe.

    Is it possible to scrape user profiles and their submission history?

    Yes, you can scrape all submissions by a specific Hacker News user. Set the mode input to byUser and provide the target username in the corresponding field. This will return a dataset containing all stories and comments submitted by that user.

    How far back does the available search history go?

    The Hacker News Search API, which this Actor utilizes, indexes all content from Hacker News since its inception in 2006. This means you can retrieve historical data spanning many years, making it suitable for extensive trend analysis and historical research, provided you adjust your dateFrom parameter accordingly.

    Where to go next

    When you are ready to run it, open Hacker News Stories, Comments & Users Scraper on Apify; the free plan covers up to 1,000 results a month.

    Start with the Hacker News Stories, Comments & Users Scraper Actor page for the current input schema, pricing tier, and run history.

    If you are comparing approaches rather than committing to one Actor, these category pages list every option we publish:

    Other Actors we maintain for related data:

    Related guides:

    Resources

    • Actor documentation, input schema, and pricing: verified against the published Actor on 2026-10-05.

    • Actor last updated by its maintainers on 2026-06-11.

    • Run outcome figures cover the 30 day public window ending 2026-10-05.

    • Hacker News Stories, Comments & Users Scraper on Apify

    Featured actors

    Hacker News Stories, Comments & Users Scraper

    Scrape Hacker News - search stories and comments, fetch top/new/best stories, get user profiles and submission history. Uses the official Algolia HN Search API and Hacker News Firebase API.

    Run on Apify ↗