Skip to content
    ↑↓ to choose · Enter to open

    · 11 min read

    StartEngine Scraper: 10 Data Fields, Up to 1,000 Free Results/Month

    By CrawlerBros Engineering Team

    Each record carries 10 output fields, returning active Reg CF, Reg A+, and Reg D campaign metadata at $5.00 per 1,000 results on the free tier. The scraper pulls company slugs, direct offering URLs, regulatory types, investment minimums, and company descriptions straight from the public StartEngine sitemap without requiring an account. It is built for angel syndicates, venture analysts, and market intelligence teams who need clean deal flow feeds. It is not for anyone who needs founder contact emails or deep cap table financials, which the public records do not include.

    Try it before you read further. Apify's free plan includes $5.00 of usage every month with no credit card, enough for up to 1,000 results at $0.005 each before platform usage. Open StartEngine Scraper on Apify and run the prefilled example.

    How reliable is StartEngine Scraper in production?

    Across the last 30 days of public runs on the Apify platform, StartEngine Scraper recorded 103 runs with the following outcomes.

    Outcome Runs Share
    Succeeded 103 100.0%
    Failed 0 0.0%
    Aborted by the user 0 0.0%
    Timed out 0 0.0%
    Total 103 100.0%

    No run failed or timed out in the last 30 days. Keep a retry and an alert on scheduled runs all the same: a clean month is a record, not a guarantee.

    What does it cost to run StartEngine Scraper?

    Each result costs $0.005 on Apify's free plan, which is $5.00 per 1,000 results. Starting a run is charged separately at $0.005 per GB of Actor memory. Apify also bills the platform usage each run consumes, at the rates of your Apify plan, on top of these charges.

    Apify plan Per result Per 1,000 results
    FREE $0.005 $5.00
    BRONZE $0.00433 $4.33
    SILVER $0.00367 $3.67
    GOLD $0.003 $3.00
    PLATINUM $0.003 $3.00
    DIAMOND $0.003 $3.00

    Worked example: collecting 10,000 results costs $50.00 in result charges before run-start fees and platform usage. No run failed or timed out in the last 30 days, so the list price is a fair budget; keep a retry in place all the same.

    The primary cost driver is maxItems, which directly controls the count of dataset rows generated and charged. To evaluate output quality cheaply before scaling, leave maxItems at a low number like 5 to verify the fields against your pipeline.

    How do you run StartEngine Scraper from the API?

    The schema marks 1 of its 3 controls as required: mode. Nothing in the payload below is illustrative. Those are the schema's prefilled defaults for StartEngine Scraper, so the request works once your token is in place.

    Call the synchronous endpoint to start a run and receive dataset items in one request:

    curl -X POST "https://api.apify.com/v2/acts/crawlerbros~startengine-scraper/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
      -H "Content-Type: application/json" \
      -d '{"mode":"browseOfferings","searchQuery":"ai","maxItems":49}'
    

    The same run from Python, using the official client:

    from apify_client import ApifyClient
    
    client = ApifyClient("<YOUR_APIFY_TOKEN>")
    
    run_input = {
      "mode": "browseOfferings",
      "searchQuery": "ai",
      "maxItems": 49
    }
    
    run = client.actor("crawlerbros~startengine-scraper").call(run_input=run_input)
    
    for item in client.dataset(run["defaultDatasetId"]).iterate_items():
        print(item)
    

    And from Node.js:

    import { ApifyClient } from 'apify-client'
    
    const client = new ApifyClient({ token: '<YOUR_APIFY_TOKEN>' })
    
    const input = {
      "mode": "browseOfferings",
      "searchQuery": "ai",
      "maxItems": 49
    }
    
    const run = await client.actor('crawlerbros~startengine-scraper').call(input)
    const { items } = await client.dataset(run.defaultDatasetId).listItems()
    console.log(items)
    

    Because the call is synchronous, your client waits for the whole run. Keep it for exploration. For scheduled work, start the run without waiting and collect the dataset afterwards, so network trouble costs you a retry rather than the results.

    Which StartEngine Scraper inputs matter, and which can you skip?

    The input schema exposes 3 controls, with mode being the single required setting. Most runs rely on browseOfferings to fetch all active opportunities; use searchCompanies alongside searchQuery only when targeting specific industry sectors or tags.

    • mode (string): What to fetch. Default: "browseOfferings".
    • searchQuery (string): Keyword to filter company names (e.g. AI, fintech, health).
    • maxItems (integer): Maximum number of offerings to return. Default: 49.

    Fixed-choice controls: mode accepts browseOfferings (Browse all active private offerings), searchCompanies (Search companies by keyword).

    What does StartEngine Scraper return?

    Returned records provide structured campaign discovery fields including companyName, offeringUrl, minInvestment, and offeringType. They do not contain founder personal contact details, cap table history, or real-time investor ledgers.

    • companyName: String - Company/offering name
    • companySlug: String - URL slug identifier
    • offeringUrl: String - Direct URL to the StartEngine offering page
    • offeringType: String - Regulation type: Reg CF, Reg A+, Reg D, etc.
    • minInvestment: String - Minimum investment amount (e.g. $250)
    • description: String - Company/offering description
    • imageUrl: String - Company logo or offering cover image
    • platform: String - Always StartEngine
    • source: String - Always startengine.com
    • scrapedAt: String - ISO 8601 timestamp when scraped

    These are the documented fields. Optional ones can be empty on a given record, so measure how often each field your deliverable depends on is populated across a real sample before automating the handoff.

    How do you build the workflow end to end?

    Open StartEngine Scraper and work through these in order. Each step ends with something to check, so a bad configuration surfaces on a small run rather than a scheduled one.

    1. Run an initial probe with mode set to browseOfferings and maxItems capped at 5.
    2. Inspect the resulting dataset to verify that offeringUrl, companySlug, and offeringType resolve to non-empty strings.
    3. Decide whether your pipeline needs catalog-wide monitoring or keyword tracking; if tracking specific sectors, switch mode to searchCompanies and pass terms like AI or biotech into searchQuery.
    4. Increase maxItems up to the platform sitemap ceiling (typically 40 to 60 items) when ready for full batch ingestion.
    5. Validate that scrapedAt returns valid ISO 8601 timestamps and minInvestment contains properly formatted values such as $250.
    6. Deduplicate records in your target store against companySlug across consecutive executions to avoid duplicate entries.
    7. Direct raw records into your downstream pipeline while monitoring dataset count against the run-start and per-result charges.

    How do you apply it? Three worked playbooks

    These are StartEngine Scraper's own documented use cases, each worked through as an operating pattern rather than a description.

    Use case 1: Private equity opportunity tracking

    Outcome: Track active private equity investment opportunities on StartEngine

    Configure: Set mode to browseOfferings and maxItems to 49.

    Working method: Execute an initial run to capture current active listings, store companySlug as your primary key, and schedule recurrent weekly executions to identify new listings by comparing slug sets.

    Deliverable: A structured database table of active equity campaigns with company names, offering types, and direct URLs.

    Stop condition: Zero records returned when mode is browseOfferings or the sitemap responds with empty batches.

    Use case 2: Reg CF and Reg A+ research

    Outcome: Research emerging startups raising capital via Reg CF or Reg A+

    Configure: Set mode to searchCompanies, searchQuery to Reg CF, and maxItems to 49.

    Working method: Execute runs across specific offering types or sector terms, comparing minimum investments and descriptions across the returned cohort to assess capital requirements.

    Deliverable: A clean CSV export summarizing funding types, company summaries, and entry minimums across active raises.

    Stop condition: Returned items show missing offeringType values across more than 5% of dataset records.

    Use case 3: Deal flow aggregator feed

    Outcome: Build an investment opportunity aggregator or deal flow pipeline

    Configure: Set mode to browseOfferings and maxItems to 200.

    Working method: Ingest the full sitemap batch, filter against existing deals in your aggregator pipeline by companySlug, and push newly discovered opportunities into your staging queue.

    Deliverable: An automated JSON feed updating a deal sourcing CRM with live offering links and target investment floors.

    Stop condition: Output records missing offeringUrl or companyName completely.

    What breaks, and how do you design around it?

    • Over the last 30 days, 0.0% of public runs failed and 0.0% timed out. Build retries and alerting around those rates rather than assuming every run completes.

    StartEngine's private offerings sitemap typically lists 40 to 60 active offerings at any given time, so asking for large result sets will still yield bounded batches. If you need historical campaign metrics or closed raises, you must store snapshots over time yourself rather than expecting archived records.

    When should you not use StartEngine Scraper?

    Do not use this scraper if your mandate requires historical public market equity metrics, balance sheets, or SEC filing text. For public market assets, use Morningstar Scraper to pull exchange-listed equity quotes, market caps, and trading volume instead. Similarly, if your focus is product-based reward crowdfunding rather than equity, run Kickstarter Project Scraper to retrieve backer tallies, funding goals, and campaign deadlines. Skip this scraper entirely if you need real-time cap tables, investor contact databases, or closed deals, as the StartEngine sitemap only exposes active equity offerings.

    What should you check before trusting the output?

    • Verify that companySlug is populated and unique within the current batch; halt the ingest pipeline if duplicates emerge in a single run.
    • Check that offeringUrl begins with the expected startengine.com URL structure; alert immediately if missing or malformed.
    • Ensure minInvestment contains currency symbols and numeric values (e.g. $250); flag records returning blank strings.
    • Confirm offeringType matches expected regulatory exemptions like Reg CF, Reg A+, or Reg D; trigger an alert if unexpected categorizations appear.
    • Check that scrapedAt parses cleanly as an ISO 8601 date string and matches the current UTC day.

    None of this proves a record is correct. It gives a scheduled StartEngine Scraper run defined points where it should stop instead of quietly passing bad data downstream.

    Frequently asked questions

    What does scraping StartEngine cost on the free tier?

    Results are billed at $0.005 per result, which is $5.00 per 1,000 results on the free tier. Apify provides $5.00 of monthly prepaid platform usage on the free plan, which covers up to 1,000 results before run-start and platform consumption charges.

    Do I need a StartEngine login or API token?

    No. The scraper reads public sitemap endpoints and public offering pages directly, meaning no StartEngine user account, credentials, or session cookies are required to extract active offerings.

    How many active offerings can I expect per run?

    The StartEngine private offerings sitemap typically lists 40-60 active offerings at a given moment. Setting maxItems higher than 60 will simply collect all currently live listings present on the sitemap.

    Can I filter listings by specific categories or keywords?

    Yes. Switch the mode setting to searchCompanies and provide a keyword in searchQuery, such as AI or fintech. The scraper will then filter listings matching the target terms.

    Are closed or expired offerings included in the output?

    No. The scraper gathers active offerings surfaced through the live sitemap. To maintain visibility into expired campaigns or historic raises, persist each run's output to an external database.

    Where to go next

    When you are ready to run it, open StartEngine Scraper on Apify; the free plan covers up to 1,000 results a month.

    Start with the StartEngine Scraper Actor page for the current input schema, pricing tier, and run history.

    Other Actors we maintain for related data:

    • Kickstarter Project Scraper: Extract crowdfunding project data from Kickstarter like name, blurb, goal, pledged amount, backers, deadline, creator, category, country, and more.
    • Morningstar Scraper: Scrape Morningstar - stock and ETF quotes with intraday and yearly price data, market cap, volume, pre/post-market prices, plus top gainers/losers/actives.
    • Private Property Scraper: Scrape PrivateProperty.co.za - a major South African property portal.

    Related guides:

    Resources

    • Actor documentation, input schema, and pricing: verified against the published Actor on 2026-09-30.

    • Actor last updated by its maintainers on 2026-06-06.

    • Run outcome figures cover the 30 day public window ending 2026-09-30.

    • StartEngine Scraper on Apify

    Featured actors

    StartEngine Scraper

    Scrape StartEngine (startengine.com) - an equity crowdfunding platform. Browse active private investment offerings. Extracts company name, offering slug, type, valuation cap, minimum investment, deadline, and offering URL.

    Run on Apify ↗