Skip to content
    ↑↓ to choose · Enter to open

    · 15 min read

    Legacy.com Obituary Scraper: 23 Data Fields per Record (2026)

    By CrawlerBros Engineering Team

    Each record carries 23 output fields, including the full obituary text, birth and death dates, age, and the guest book visit counter. This scraper operates without a login or paid proxy, capturing obituary details from the largest online archive in the United States. You can try it using Apify's free plan, which provides $5.00 of monthly usage to cover up to 1,000 results. This tool is built specifically for researchers, family historians, and legal professionals who need structured proof-of-death data; it is not for those seeking overseas listings, as it is limited to US obituaries.

    Try it before you read further. Apify's free plan includes $5.00 of usage every month with no credit card, enough for up to 1,000 results at $0.005 each before platform usage. Open Legacy.com Obituary Scraper on Apify and run the prefilled example.

    How reliable is Legacy.com Obituary Scraper in production?

    Across the last 30 days of public runs on the Apify platform, Legacy.com Obituary Scraper recorded 174 runs with the following outcomes.

    Outcome Runs Share
    Succeeded 168 96.6%
    Failed 0 0.0%
    Aborted by the user 6 3.4%
    Timed out 0 0.0%
    Total 174 100.0%

    No run failed or timed out in the last 30 days; the 6 that did not finish were stopped by the people who started them. Keep a retry and an alert on scheduled runs all the same: a clean month is a record, not a guarantee.

    What does it cost to run Legacy.com Obituary Scraper?

    Each result costs $0.005 on Apify's free plan, which is $5.00 per 1,000 results. Starting a run is charged separately at $0.005 per GB of Actor memory. Apify also bills the platform usage each run consumes, at the rates of your Apify plan, on top of these charges.

    Apify plan Per result Per 1,000 results
    FREE $0.005 $5.00
    BRONZE $0.00433 $4.33
    SILVER $0.00367 $3.67
    GOLD $0.003 $3.00
    PLATINUM $0.003 $3.00
    DIAMOND $0.003 $3.00

    Worked example: collecting 10,000 results costs $50.00 in result charges before run-start fees and platform usage. No run failed or timed out in the last 30 days, so the list price is a fair budget; keep a retry in place all the same.

    The price of your run is heavily influenced by the maxPages and maxItems controls, as they dictate the total number of records written to your dataset. If you set maxPages to its ceiling of 2000, you will scan up to 20,000 items and scale your cost accordingly. To test your queries without spending platform resources unnecessarily, configure a low maxItems limit of 10 to check the data structure first.

    How do you run Legacy.com Obituary Scraper from the API?

    The schema marks 1 of its 10 controls as required: mode. The payload below uses the schema's own prefilled values, so it runs as written once you substitute your API token.

    Call the synchronous endpoint to start a run and receive dataset items in one request:

    curl -X POST "https://api.apify.com/v2/acts/crawlerbros~legacy-com-scraper/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
      -H "Content-Type: application/json" \
      -d '{"mode":"search"}'
    

    The same run from Python, using the official client:

    from apify_client import ApifyClient
    
    client = ApifyClient("<YOUR_APIFY_TOKEN>")
    
    run_input = {
      "mode": "search"
    }
    
    run = client.actor("crawlerbros~legacy-com-scraper").call(run_input=run_input)
    
    for item in client.dataset(run["defaultDatasetId"]).iterate_items():
        print(item)
    

    And from Node.js:

    import { ApifyClient } from 'apify-client'
    
    const client = new ApifyClient({ token: '<YOUR_APIFY_TOKEN>' })
    
    const input = {
      "mode": "search"
    }
    
    const run = await client.actor('crawlerbros~legacy-com-scraper').call(input)
    const { items } = await client.dataset(run.defaultDatasetId).listItems()
    console.log(items)
    

    The synchronous endpoint holds the connection open until the run finishes, which is convenient for small batches and wrong for large ones. For anything long running, start the run asynchronously and poll, or attach a webhook, so a dropped connection does not cost you the results.

    Which Legacy.com Obituary Scraper inputs matter, and which can you skip?

    The mode control defines your retrieval strategy, allowing you to choose between "search", "browseByState", or "byUrl". If you are starting out, leave proxyConfiguration at its default setting. The actor first tries a direct connection and only engages the proxy if the source blocks the request.

    • mode (string): What to fetch. Default: "search".
    • searchQuery (string): Free-text query matching the deceased name or obituary text (mode=search). Default: "smith".
    • state (string): US state filter. Required for mode=browseByState (scans recent obituaries and keeps only those whose address is in this state). In mode=search, also narrows results to this state when set - leave unset for an unfiltered, nationwide search.
    • obituaryUrls (array): Full obituary URLs, e.g. https://www.legacy.com/legacy/osea-ocasio (mode=byUrl). Default: [].
    • dateRangeFrom (string): Only keep obituaries whose death date is on or after this date (ISO YYYY-MM-DD). Default: "".
    • dateRangeTo (string): Only keep obituaries whose death date is on or before this date (ISO YYYY-MM-DD). Default: "".
    • containsKeyword (string): Only keep obituaries whose name, headline, or obituary text contains this substring (case-insensitive). Default: "".
    • maxItems (integer): Hard cap on emitted records. Default: 50.
    • maxPages (integer): Maximum number of result/feed pages to scan (10 results per page). Higher values find more results but take longer. The recent-obituaries feed used by browseByState/search holds roughly 20,000 obituaries (~7-8 months of US obituaries across all states) - set near 2000 to reach the full feed. Default: 10.
    • proxyConfiguration (object): Optional Apify proxy. The actor first tries a direct connection and only engages the proxy if the source blocks the request. Default: {"useApifyProxy":false}.

    Fixed-choice controls: mode accepts search (Search obituaries (name/keyword)), browseByState (Browse recent obituaries by state), byUrl (Fetch obituaries by URL); state accepts 51 values, including AL (Alabama), AK (Alaska), AZ (Arizona), AR (Arkansas).

    What does Legacy.com Obituary Scraper return?

    The extracted dataset is ideal for building genealogy records, legal verification pipelines, and historical databases, as it outputs clean ISO-formatted dates and unformatted text blocks. However, it does not contain the individual guest book messages or the underlying image files, providing only the raw photo URLs and visit counters instead.

    • obituaryId - Legacy.com's internal obituary UUID
    • name, givenName, familyName - deceased's full name and parts
    • birthDate, deathDate - ISO dates (YYYY-MM-DD)
    • age - age at death, derived from birth/death dates when both are present
    • city, state - location from the obituary's address (2-letter US state code)
    • headline - e.g. Jean Smith Obituary 2026
    • description - short obituary description
    • genre, articleSection - content category labels from the embedded article metadata
    • obituaryText - the full obituary text (paragraphs separated by blank lines)
    • datePublished, dateCreated, dateModified - publication timestamps (date part)
    • photoUrl - the obituary's main photo
    • galleryPhotoUrls[] - all photos from the obituary's photo gallery
    • guestBookVisits - the guest book visit counter shown on the page
    • sourceUrl - canonical Legacy.com URL of the obituary
    • recordType: "obituary", scrapedAt

    These are the documented fields. Optional ones can be empty on a given record, so measure how often each field your deliverable depends on is populated across a real sample before automating the handoff.

    How do you build the workflow end to end?

    Open Legacy.com Obituary Scraper and work through these in order. Each step ends with something to check, so a bad configuration surfaces on a small run rather than a scheduled one.

    1. Select your target retrieval method by configuring the mode control, choosing "search" to query names, "browseByState" to scan the recent feed, or "byUrl" to pass specific obituary links.
    2. Define your target subject by entering a name or keyword into the searchQuery field, or leave it set to the default surname if you are running a broad verification pass.
    3. If you selected "browseByState" mode, you must set the state control to a two-letter US state code such as TX or NY to isolate that region's recent listings.
    4. Establish chronological boundaries for your search using the dateRangeFrom and dateRangeTo fields, ensuring you format both as ISO YYYY-MM-DD strings.
    5. Apply string-matching filters to the obituary text and headlines by inserting your target term into the containsKeyword control.
    6. Prevent unexpected platform usage on large runs by setting maxItems to a conservative figure, remembering that the actor enforces a hard cap between 1 and 1000 records.
    7. Adjust the depth of your crawl by changing maxPages, noting that each page contains 10 listings and you will need a value near 2000 to scan the entire recent-obituaries feed.
    8. Click start to run the scraper, then open the dataset tab to inspect the returned records and verify that the obituaryText and deathDate fields contain valid data.

    How do you apply it? Three worked playbooks

    These are Legacy.com Obituary Scraper's own documented use cases, each worked through as an operating pattern rather than a description.

    Use case 1: Genealogy research

    Outcome: Bulk-export obituary records with dates, locations, and full text for family-history projects

    Configure: mode is set to "search", searchQuery is set to "Patricia", dateRangeFrom is set to "2026-01-01", dateRangeTo is set to "2026-12-31", and maxItems is set to 30.

    Working method: Begin by executing a targeted search using a common family name and historical date constraints. Inspect the returned dataset to ensure that the birthDate, deathDate, and age fields are populated. If they are missing, look for these values inside the raw text of the obituaryText field.

    Deliverable: A bulk-exported JSON dataset containing full names, birth and death dates, family relationship details, and complete biographical text.

    Stop condition: The returned records contain empty obituaryText fields, indicating the scraper has fallen back to listing-only snippets.

    Use case 2: Death-notice monitoring

    Outcome: Watch for obituaries of specific people or in specific states

    Configure: mode is set to "browseByState", state is set to "TX", maxItems is set to 10, and maxPages is set to 20.

    Working method: Set up a daily schedule to poll the recent obituary feed for your chosen state. Use a downstream notification script to parse the output for specific names or high-priority keywords, comparing the scrapedAt timestamp to your previous run.

    Deliverable: A daily status report or automated alert system highlighting new death notices matched against a monitored watchlist.

    Stop condition: The browse feed returns zero new records across multiple state-specific runs, indicating a regional feed block or layout change.

    Use case 3: Legal & probate

    Outcome: Collect proof-of-death notices published by families and funeral homes

    Configure: mode is set to "byUrl", and obituaryUrls is set to ["https://www.legacy.com/legacy/osea-ocasio", "https://www.legacy.com/legacy/karen-todd"].

    Working method: Compile your target list of official Legacy.com URLs from your probate case files. Input the absolute URLs into the array control to bypass the search index and directly request the official published notice from the platform.

    Deliverable: An audit-ready collection of structured death records, including the published date, hosting newspaper details, and canonical sourceUrl.

    Stop condition: An input URL returns a blank record or a connection error, indicating that the target obituary has been taken down or moved.

    What breaks, and how do you design around it?

    • US obituaries only - the actor targets Legacy.com's standard US obituary pages. Legacy.com's newspaper-branded and international properties use different page structures and aren't currently supported.
    • No direct state search - the obituary search index has no state filter, so browseByState scans the most recent obituaries (up to maxPages × 10) and keeps those in the chosen state. Older obituaries may require more pages to find.
    • Detail-page fallback - if an obituary detail page cannot be loaded (rare - page not yet on the obituary platform), the record is emitted with the listing fields only (name, deathDate, photoUrl, snippet, sourceUrl).
    • Per-record completeness varies - not every obituary publishes birth dates, photos, or full text; missing values are omitted, never filled with placeholders.
    • Guest book visits - only the visit counter shown on the page is collected; guest book entries themselves are not scraped.

    When you hit Legacy.com's limit of 200 results per search query, you should segment your target names by state or combine keyword searches to slice the index. If you need historical records beyond the 20,000-item limit of the recent feed, switch the mode to direct URL input and feed the scraper list files from external indexing services.

    When should you not use Legacy.com Obituary Scraper?

    Do not use this Actor if you need to gather historical death notices from outside the United States or from localized, newspaper-branded portals, as it only targets Legacy.com's standard US layout. If your research is focused on global historical figures or ancestry verification through wider open-linked databases, you should instead query WikiData using the WikiData Notable Persons Scraper. Additionally, if your target dataset requires scraping commercial classifieds or pet listings rather than human obituaries, specialized tools like the Lancaster Puppies Scraper or NextDayPets Scraper (Pawrade) are much better choices.

    What should you check before trusting the output?

    • Verify that the deathDate field conforms to the ISO YYYY-MM-DD format and is not empty on records critical for compliance.
    • Check if the state field returns the expected two-letter US postal code, particularly when filtering national search results.
    • Monitor the presence of the obituaryText field, checking if a record has reverted to listing-only fields due to a detail page failing to load.
    • Ensure that guestBookVisits values are numeric when you intend to use them as an engagement metric in your downstream analytics pipeline.

    None of this proves a record is correct. It gives a scheduled Legacy.com Obituary Scraper run defined points where it should stop instead of quietly passing bad data downstream.

    Frequently asked questions

    How much does it cost to scrape Legacy.com obituaries?

    Pricing is set at $0.005 per result, which is $5.00 per 1,000 results on the free plan. Paid Apify plans reduce this per-result cost further. You also need to account for Apify's standard platform usage fees which are billed on top of the result costs.

    Does this scraper require a proxy configuration?

    No. The public pages are served without cookies or authentication. The actor first tries a direct connection and only engages the proxy if the source blocks the request, meaning you do not need to buy expensive proxies upfront.

    Why does the browseByState mode require a high maxPages value?

    Legacy.com does not have a direct search filter for states. The Actor scans the shared nationwide feed of recent obituaries and filters out matching states. You must set maxPages near 2000 to search the full 20,000-item feed.

    Are the actual photo files downloaded to my dataset?

    No. The scraper collects the public photo URLs and photo gallery URLs directly from the obituary page. Downloading, archiving, and hosting the actual image files is your responsibility.

    Can I retrieve historical obituaries from several years ago?

    The search index is limited, and the browse feed only holds the last 7 to 8 months of data. To scrape older archives, you must collect the specific obituary URLs first and run the Actor in byUrl mode.

    Where to go next

    When you are ready to run it, open Legacy.com Obituary Scraper on Apify; the free plan covers up to 1,000 results a month.

    Start with the Legacy.com Obituary Scraper Actor page for the current input schema, pricing tier, and run history.

    Other Actors we maintain for related data:

    Related guides:

    Resources

    • Actor documentation, input schema, and pricing: verified against the published Actor on 2026-10-05.

    • Actor last updated by its maintainers on 2026-10-01.

    • Run outcome figures cover the 30 day public window ending 2026-10-05.

    • Legacy.com Obituary Scraper on Apify

    Featured actors

    Legacy.com Obituary Scraper

    Scrape obituaries from Legacy.com - the largest online obituary archive in the US. Search obituaries by name or keyword, browse recent obituaries by state, fetch full obituary pages by URL, and get deceased name, birth/death dates, age, city/state, full obituary text, photos, and guest book visits.

    Run on Apify ↗