· 15 min read
Hacker News Stories, Comments & Users Scraper: $5.00 per 1,000 Results
Hacker News Stories, Comments & Users Scraper provides 11 fields for story records and 13 fields for comment records, enabling detailed analysis of content and user activity. This Actor works by utilizing the official Algolia HN Search API and Hacker News Firebase API, requiring no authentication. It's suitable for practitioners who need to gather specific discussions, track trends, or identify influential users from Hacker News. This Actor is not for those who need private user data or discussions from non-public channels, as it only accesses publicly available information.
Try it before you read further. Apify's free plan includes $5.00 of usage every month with no credit card, enough for up to 1,000 results at $0.005 each before platform usage. Open Hacker News Stories, Comments & Users Scraper on Apify and run the prefilled example.
How reliable is Hacker News Stories, Comments & Users Scraper in production?
Across the last 30 days of public runs on the Apify platform, Hacker News Stories, Comments & Users Scraper recorded 57 runs with the following outcomes.
| Outcome | Runs | Share |
|---|---|---|
| Succeeded | 57 | 100.0% |
| Failed | 0 | 0.0% |
| Aborted by the user | 0 | 0.0% |
| Timed out | 0 | 0.0% |
| Total | 57 | 100.0% |
No run failed or timed out in the last 30 days. Keep a retry and an alert on scheduled runs all the same: a clean month is a record, not a guarantee.
What does it cost to run Hacker News Stories, Comments & Users Scraper?
Each result costs $0.005 on Apify's free plan, which is $5.00 per 1,000 results. Starting a run is charged separately at $0.005 per GB of Actor memory. Apify also bills the platform usage each run consumes, at the rates of your Apify plan, on top of these charges.
| Apify plan | Per result | Per 1,000 results |
|---|---|---|
| FREE | $0.005 | $5.00 |
| BRONZE | $0.00433 | $4.33 |
| SILVER | $0.00367 | $3.67 |
| GOLD | $0.003 | $3.00 |
| PLATINUM | $0.003 | $3.00 |
| DIAMOND | $0.003 | $3.00 |
Worked example: collecting 10,000 results costs $50.00 in result charges before run-start fees and platform usage. No run failed or timed out in the last 30 days, so the list price is a fair budget; keep a retry in place all the same.
The primary factor affecting your bill is the "maxItems" input, which caps the total number of records returned. Each result written to your dataset incurs a charge. To estimate costs and determine if this Actor meets your needs before committing significant spend, start with a low "maxItems" value (such as the example input's default of 50). This will return a small sample of data for a minimal cost.
How do you run Hacker News Stories, Comments & Users Scraper from the API?
The schema marks 1 of its 10 controls as required: mode. Nothing in the payload below is illustrative. Those are the schema's prefilled defaults for Hacker News Stories, Comments & Users Scraper, so the request works once your token is in place.
Call the synchronous endpoint to start a run and receive dataset items in one request:
curl -X POST "https://api.apify.com/v2/acts/crawlerbros~hacker-news-scraper/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"mode":"searchStories","searchQuery":"AI startup","startUrls":[],"storyType":"all","sortBy":"relevance","maxItems":50}'
The same run from Python, using the official client:
from apify_client import ApifyClient
client = ApifyClient("<YOUR_APIFY_TOKEN>")
run_input = {
"mode": "searchStories",
"searchQuery": "AI startup",
"startUrls": [],
"storyType": "all",
"sortBy": "relevance",
"maxItems": 50
}
run = client.actor("crawlerbros~hacker-news-scraper").call(run_input=run_input)
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
print(item)
And from Node.js:
import { ApifyClient } from 'apify-client'
const client = new ApifyClient({ token: '<YOUR_APIFY_TOKEN>' })
const input = {
"mode": "searchStories",
"searchQuery": "AI startup",
"startUrls": [],
"storyType": "all",
"sortBy": "relevance",
"maxItems": 50
}
const run = await client.actor('crawlerbros~hacker-news-scraper').call(input)
const { items } = await client.dataset(run.defaultDatasetId).listItems()
console.log(items)
Because the call is synchronous, your client waits for the whole run. Keep it for exploration. For scheduled work, start the run without waiting and collect the dataset afterwards, so network trouble costs you a retry rather than the results.
Which Hacker News Stories, Comments & Users Scraper inputs matter, and which can you skip?
The mode control is required and dictates the type of data to scrape, such as searching stories or comments, or fetching user submissions. For most users, begin by setting mode and then providing a searchQuery or username relevant to your goal. Many users can leave controls like sortBy and storyType at their defaults for an initial run.
mode(string): What to scrape. Default:"searchStories".searchQuery(string): Keyword or phrase to search for (modes: searchStories, searchComments).username(string): Hacker News username to fetch submissions for (mode: byUser).startUrls(array): Hacker News item IDs (numeric) or URLs like https://news.ycombinator.com/item?id=12345 (mode: getItem). Default:[].storyType(string): Filter by story type (applies to searchStories). Default:"all".sortBy(string): Sort order for search results. Default:"relevance".dateFrom(string): Only return items created on or after this date (ISO format, e.g. 2024-01-01). Applies to search modes.dateTo(string): Only return items created on or before this date (ISO format, e.g. 2024-12-31). Applies to search modes.minPoints(integer): Only emit stories/comments with at least this many points.maxItems(integer): Maximum number of records to emit. Default:50.
Fixed-choice controls: mode accepts searchStories (Search stories by keyword), searchComments (Search comments by keyword), topStories (front page), newStories (latest), bestStories (all-time top), byUser (All submissions by a user), getItem (Get specific item(s) by ID or URL); storyType accepts all (All types), story (Regular story), ask_hn (Ask HN), show_hn (Show HN), job (Job posting); sortBy accepts relevance, date (newest first).
What does Hacker News Stories, Comments & Users Scraper return?
The output records are structured to provide comprehensive data for either stories or comments. Story records include fields like title, url, author, points, and commentsCount, making them suitable for trend analysis or content research. Comment records include text, author, points, and links back to their storyId and storyTitle, which is useful for sentiment analysis or identifying engaged users. These records do not contain direct contact information for authors.
Story record
storyId(e.g.39876543)type(e.g.story)title(e.g.Show HN: I built a faster way to index large ...)urlhnUrl(e.g.https://news.ycombinator.com/item?id=39876543)author(e.g.johndoe)points(e.g.312)commentsCount(e.g.87)createdAt(e.g.2024-03-15T14:22:00+00:00)recordType(e.g.story)scrapedAt(e.g.2024-03-15T15:00:00+00:00)
Comment record
commentId(e.g.39876999)type(e.g.comment)texthnUrl(e.g.https://news.ycombinator.com/item?id=39876999)author(e.g.janedoe)points(e.g.42)storyId(e.g.39876543)storyTitle(e.g.Show HN: I built a faster way to index large ...)storyUrlstoryHnUrl(e.g.https://news.ycombinator.com/item?id=39876543)createdAt(e.g.2024-03-15T14:45:00+00:00)recordType(e.g.comment)scrapedAt(e.g.2024-03-15T15:00:00+00:00)
These are the documented fields. Optional ones can be empty on a given record, so measure how often each field your deliverable depends on is populated across a real sample before automating the handoff.
How do you build the workflow end to end?
Open Hacker News Stories, Comments & Users Scraper and work through these in order. Each step ends with something to check, so a bad configuration surfaces on a small run rather than a scheduled one.
- Set the "mode" to "searchStories" or "searchComments" for keyword-based extraction, or to "byUser" for user-specific data, considering the type of information you need.
- Populate the "searchQuery" field with your keywords if using search modes, or the "username" field if fetching user submissions.
- Specify a "dateFrom" and "dateTo" in YYYY-MM-DD format to narrow down results for search modes, which can help manage the volume of data.
- For story searches, consider applying a "storyType" filter (e.g., "ask_hn", "show_hn", "job") to focus on specific discussion formats.
- Set "minPoints" to filter for more impactful or popular stories and comments, starting with a value like 10 or 20 to see relevant items.
- Limit the total number of records by setting "maxItems" to a reasonable value like 100 or 200 for your initial runs.
- After the run, check the dataset to ensure the
typefield correctly identifies whether records arestoryorcommentand that relevant fields liketitleandtextare populated. - Verify that
scrapedAttimestamps are recent and thatauthorfields are present, indicating complete record extraction.
How do you apply it? Three worked playbooks
These are Hacker News Stories, Comments & Users Scraper's own documented use cases, each worked through as an operating pattern rather than a description.
Use case 1: Market research
Outcome: Track mentions of products, technologies, or companies
Configure: Set the "mode" to "searchStories" or "searchComments". Set "searchQuery" to a product or company name, e.g., "AI chip startups". Set "dateFrom" and "dateTo" to define a relevant period.
Working method: Begin by searching for a single, specific product or company name. Review the initial results for relevance and completeness. Then, broaden your search by adding related keywords or expanding the date range.
Deliverable: A dataset containing stories or comments mentioning the specified product, technology, or company, with associated metadata like title, url, author, and createdAt dates.
Stop condition: Search results contain a high percentage of irrelevant mentions or the title or text fields are consistently empty for relevant items.
Use case 2: Competitor intelligence
Outcome: Search for mentions of competitors or alternatives
Configure: Set "mode" to "searchStories" and "searchQuery" to a competitor's name, e.g., "AcmeCorp AI". Adjust "sortBy" to "date" for recent mentions.
Working method: Start with a precise search for one competitor. Examine the hnUrl and title fields to see how and where they are being discussed. Expand to alternative names or product lines, or compare mentions across a few key competitors.
Deliverable: A structured list of Hacker News stories where competitors or their products are mentioned, including author, points, and direct links to the discussions.
Stop condition: The commentsCount or points for competitor mentions are consistently low, indicating limited discussion volume or impact, or the search returns too many false positives.
Use case 3: Lead generation
Outcome: Identify users active in your target domain
Configure: Set "mode" to "searchStories" or "searchComments". Use a "searchQuery" related to your target domain, e.g., "SaaS founder" or "backend developer". Set "minPoints" to a value like 50 to find active contributors.
Working method: First, run a broad search for relevant keywords in stories or comments. Review the author fields in the results to identify potential users of interest. Then, for specific promising authors, switch to byUser mode with their username to get their full submission history.
Deliverable: A list of Hacker News usernames who have actively engaged in discussions or submitted stories within your target domain, along with their related story or comment data.
Stop condition: Identified users consistently have very few storyId or commentId entries in their byUser history, suggesting low activity or an incorrect target domain definition.
What breaks, and how do you design around it?
- Over the last 30 days, 0.0% of public runs failed and 0.0% timed out. Build retries and alerting around those rates rather than assuming every run completes.
When dealing with the 5,000-item maximum, consider splitting your task into multiple runs by adjusting the dateFrom and dateTo parameters to cover different time periods. If you need to search a very broad range of dates, it's best to iterate through smaller date windows rather than attempting one massive query. For very long lists of startUrls or usernames, process them in batches to prevent individual run timeouts.
When should you not use Hacker News Stories, Comments & Users Scraper?
This Actor is not suitable if your primary need is real-time, low-latency data for individual posts or comments, as its design is geared towards batch data extraction. For single-item lookups that require immediate responses, directly querying the Algolia HN Search API or Hacker News Firebase API might be a more efficient approach. If you are trying to acquire data from very obscure or niche discussions that are unlikely to garner significant points or comments, the minPoints filter might exclude all relevant results, in which case a more general search Actor or manual browsing might be necessary. Also, if your goal is to monitor job postings exclusively, the Hacker News Jobs Scraper offers a more specialized solution focused solely on the "Who is Hiring" board.
What should you check before trusting the output?
- Check for
storyIdorcommentIdvalues that are null or malformed, as this indicates a failed item extraction. - Inspect
titlefields in story records andtextfields in comment records for unexpected characters or truncated content. - Confirm that
pointsandcommentsCount(for stories) orpoints(for comments) are numeric and within expected ranges; unusually low or zero values might indicate incomplete data. - Look for
createdAttimestamps that are outside your specifieddateFromanddateToranges, or are missing entirely. - Validate
urlfields in story records for accessibility andhnUrlfields for both story and comment records for correctness. - If using
minPoints, ensure that all returned records indeed meet or exceed the specified minimum.
None of this proves a record is correct. It gives a scheduled Hacker News Stories, Comments & Users Scraper run defined points where it should stop instead of quietly passing bad data downstream.
Frequently asked questions
What is the cost for 1,000 results?
On Apify's free plan, this Actor costs $5.00 for every 1,000 results. This allows you to collect a substantial amount of data before needing to consider a paid plan. The platform also includes $5.00 of monthly usage without requiring a credit card, which covers up to 1,000 results. Be aware that run-start fees and platform usage are billed separately.
How reliable is this Actor for unattended runs?
The Actor has a 100.0% success rate across 57 runs in the last 30 days, indicating high reliability. This means you can schedule it for unattended execution with confidence that it will complete successfully. You should not expect any runs to fail or time out based on recent telemetry, making it suitable for automated workflows without frequent intervention.
Can I filter search results by a specific date range?
Yes, you can filter search results by date. Use the dateFrom and dateTo input fields to specify your desired range in YYYY-MM-DD format. This applies to searchStories and searchComments modes, allowing you to focus on content published within a particular timeframe.
Is it possible to scrape user profiles and their submission history?
Yes, you can scrape all submissions by a specific Hacker News user. Set the mode input to byUser and provide the target username in the corresponding field. This will return a dataset containing all stories and comments submitted by that user.
How far back does the available search history go?
The Hacker News Search API, which this Actor utilizes, indexes all content from Hacker News since its inception in 2006. This means you can retrieve historical data spanning many years, making it suitable for extensive trend analysis and historical research, provided you adjust your dateFrom parameter accordingly.
Where to go next
When you are ready to run it, open Hacker News Stories, Comments & Users Scraper on Apify; the free plan covers up to 1,000 results a month.
Start with the Hacker News Stories, Comments & Users Scraper Actor page for the current input schema, pricing tier, and run history.
If you are comparing approaches rather than committing to one Actor, these category pages list every option we publish:
- Comment scrapers covers 37 Actors in this family.
Other Actors we maintain for related data:
- Hacker News Scraper: Scrape Hacker News stories, comments, jobs, and user profiles via the official Firebase API.
- Hacker News Scraper: Scrape Hacker News stories, comments, jobs, and user profiles.
- Hacker News Jobs Scraper: Scrape the Hacker News job board - the famous 'Who is Hiring' startup job posts on news.ycombinator.com.
- Snapchat User Stories Scraper: Scrape curated highlight stories from public Snapchat profiles.
Related guides:
- Hacker News Scraper: Up to 2,500 Free Results a Month (2026)
- Yandex News Scraper: Operational Guide, Configurations, and Playbooks
- Facebook Comments Scraper: 24 Data Fields per Record (2026)
- Indie Hackers Scraper: 3 Practical Use Cases
- Nexus Mods Scraper: Up to 1,000 Free Results a Month (2026)
- Crunchbase News Scraper: Data Extraction and Integration Guide
- TikTok Comments Scraper: 3 Practical Use Cases
Resources
Actor documentation, input schema, and pricing: verified against the published Actor on 2026-10-05.
Actor last updated by its maintainers on 2026-06-11.
Run outcome figures cover the 30 day public window ending 2026-10-05.
Featured actors
Hacker News Stories, Comments & Users Scraper
Scrape Hacker News - search stories and comments, fetch top/new/best stories, get user profiles and submission history. Uses the official Algolia HN Search API and Hacker News Firebase API.
Run on Apify ↗