Skip to content
    ↑↓ to choose · Enter to open

    · 14 min read

    H-1B LCA Disclosure Data Scraper: 91 Data Fields per Record (2026)

    By CrawlerBros Engineering Team

    Each record carries 91 output fields detailing US Department of Labor H-1B, H-1B1, and E-3 Labor Condition Applications directly from official quarterly disclosures. The free-plan price is $5.00 per 1,000 results, and the tool captures employer details, offered wages, prevailing wage determinations, and secondary worksites without web scraping fragility. It serves immigration attorneys, talent recruiters, and corporate compliance officers benchmarking wage compliance. It is not for teams seeking real-time USCIS petition approval statuses or candidate identity records, which DOL disclosure files never include.

    Try it: open H-1B LCA Disclosure Data Scraper on Apify, sign in on the free plan and run the prefilled example.

    Can you try H-1B LCA Disclosure Data Scraper before paying?

    Yes. Apify's free plan includes $5.00 of prepaid usage every month and asks for no credit card. At $0.005 per result, that covers up to 1,000 results of H-1B LCA Disclosure Data Scraper a month, before run-start charges and platform usage.

    H-1B LCA Disclosure Data Scraper was last updated on 2026-08-29. It is one of 1,725 Actors CrawlerBros publishes on Apify, which together have 692,561 lifetime public runs and an average rating of 4.63 out of 5 across 416 reviews.

    What does it cost to run H-1B LCA Disclosure Data Scraper?

    Each result costs $0.005 on Apify's free plan, which is $5.00 per 1,000 results. Starting a run is charged separately at $0.005 per GB of Actor memory. Apify also bills the platform usage each run consumes, at the rates of your Apify plan, on top of these charges.

    Apify plan Per result Per 1,000 results
    FREE $0.005 $5.00
    BRONZE $0.00433 $4.33
    SILVER $0.00367 $3.67
    GOLD $0.003 $3.00
    PLATINUM $0.003 $3.00
    DIAMOND $0.003 $3.00

    The primary cost driver is maxItems, which directly controls how many dataset records are written from the streamed disclosure file. Setting tight filters like employerFein or socCode while leaving maxItems capped low keeps result fees minimal. To validate your schema mapping before incurring larger result fees, run with maxItems set to a handful of items.

    How do you run H-1B LCA Disclosure Data Scraper from the API?

    The schema marks 2 of its 37 controls as required: mode, proxyConfiguration. The payload below uses the schema's own prefilled values, so it runs as written once you substitute your API token.

    Call the synchronous endpoint to start a run and receive dataset items in one request:

    curl -X POST "https://api.apify.com/v2/acts/crawlerbros~h1b-lca-disclosure-scraper/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
      -H "Content-Type: application/json" \
      -d '{"mode":"search","proxyConfiguration":{"useApifyProxy":true,"apifyProxyGroups":["RESIDENTIAL"]}}'
    

    The same run from Python, using the official client:

    from apify_client import ApifyClient
    
    client = ApifyClient("<YOUR_APIFY_TOKEN>")
    
    run_input = {
      "mode": "search",
      "proxyConfiguration": {
        "useApifyProxy": True,
        "apifyProxyGroups": [
          "RESIDENTIAL"
        ]
      }
    }
    
    run = client.actor("crawlerbros~h1b-lca-disclosure-scraper").call(run_input=run_input)
    
    for item in client.dataset(run["defaultDatasetId"]).iterate_items():
        print(item)
    

    And from Node.js:

    import { ApifyClient } from 'apify-client'
    
    const client = new ApifyClient({ token: '<YOUR_APIFY_TOKEN>' })
    
    const input = {
      "mode": "search",
      "proxyConfiguration": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
          "RESIDENTIAL"
        ]
      }
    }
    
    const run = await client.actor('crawlerbros~h1b-lca-disclosure-scraper').call(input)
    const { items } = await client.dataset(run.defaultDatasetId).listItems()
    console.log(items)
    

    The synchronous endpoint holds the connection open until the run finishes, which is convenient for small batches and wrong for large ones. For anything long running, start the run asynchronously and poll, or attach a webhook, so a dropped connection does not cost you the results.

    Which H-1B LCA Disclosure Data Scraper inputs matter, and which can you skip?

    The two required controls are mode and proxyConfiguration. The mode selector toggles between querying the full dataset and looking up specific records by case numbers, while fiscalYear controls whether you scan current or historical files. Leave secondary filters like totalWorkerPositionsMin and receivedDateFrom blank until your initial search scope proves too broad.

    • mode (string): What to fetch. Default: "search".
    • proxyConfiguration (object): Required. dol.gov's Akamai edge blocks Apify's own egress IPs and the default Apify datacenter proxy pool with a 403 on the large XLSX disclosure files. The actor always tries a direct connection first (fast to fail) and falls back to this proxy - it must be a Residential group for the download to succeed reliably from Apify's cloud infrastructure. Default: {"useApifyProxy":true,"apifyProxyGroups":["RESIDENTIAL"]}.
    • fiscalYear (string): Which DOL disclosure file to scan. latest auto-discovers the newest quarterly file for the current fiscal year (currently FY2026 Q2). Older fiscal years are closed/final (through Q4). Default: "latest".
    • caseNumbers (array): Exact LCA case numbers to look up, e.g. I-200-25181-143022. Default: [].
    • employerName (string): Case-insensitive substring match against employer legal name OR trade name / DBA.
    • jobTitle (string): Case-insensitive substring match against the job title.
    • socCode (string): Exact match or prefix match against the Standard Occupational Classification code, e.g. 19-2031 or 19-2031.00.
    • caseStatus (string): Filter to a specific case status. Default: "".
    • visaClass (string): Filter to a specific visa classification. Default: "".
    • worksiteState (string): Filter to a specific US worksite state/territory. Default: "".
    • worksiteCity (string): Case-insensitive substring match against the worksite city.
    • naicsCode (string): Exact match or prefix match against the employer's NAICS industry code, e.g. 611310 or 6113.

    The other 25 controls, with their defaults, are listed in the input schema on H-1B LCA Disclosure Data Scraper on Apify.

    Fixed-choice controls: mode accepts search (Search disclosures), byCaseNumbers (Lookup by case number(s)); fiscalYear accepts latest (current fiscal year, auto-discovered), FY2025 (full year, Oct 2024-Sep 2025), FY2024 (full year, Oct 2023-Sep 2024), FY2023 (full year, Oct 2022-Sep 2023), FY2022 (full year, Oct 2021-Sep 2022), FY2021 (full year, Oct 2020-Sep 2021), FY2020 (full year, Oct 2019-Sep 2020); caseStatus accepts "" (Any status), Certified, Certified - Withdrawn, Denied, Withdrawn; visaClass accepts "" (Any visa class), H-1B, H-1B1 Chile, H-1B1 Singapore, E-3 Australian; worksiteState accepts 56 values (default ""), including "" (Any state), AK (Alaska), AL (Alabama), AR (Arkansas).

    What does H-1B LCA Disclosure Data Scraper return?

    Returned items provide comprehensive docket intelligence, including wage rates, prevailing wage levels, worksite addresses, and legal counsel names. The records do not contain employee identities, passport numbers, or USCIS adjudication outcomes. They reflect purely what employers filed with the Department of Labor.

    • caseNumber, caseStatus, visaClass
    • receivedDate, decisionDate, originalCertDate, beginDate, endDate
    • jobTitle, socCode, socTitle, fullTimePosition
    • totalWorkerPositions, newEmploymentCount, continuedEmploymentCount, changePreviousEmploymentCount, newConcurrentEmploymentCount, changeEmployerCount, amendedPetitionCount
    • employerName, tradeNameDba, employerAddress, employerCity, employerState, employerPostalCode, employerCountry, employerProvince, employerPhone, employerFein, naicsCode
    • employerPocName, employerPocJobTitle, employerPocAddress, employerPocCity, employerPocState, employerPocPostalCode, employerPocCountry, employerPocProvince, employerPocPhone, employerPocEmail
    • agentRepresentingEmployer, agentAttorneyName, agentAttorneyAddress, agentAttorneyCity, agentAttorneyState, agentAttorneyPostalCode, agentAttorneyCountry, agentAttorneyProvince, agentAttorneyPhone, agentAttorneyEmail, lawfirmName, lawfirmFein, attorneyBarState, attorneyBarCourt
    • worksiteWorkers, secondaryEntity, secondaryEntityBusinessName, worksiteAddress, worksiteCity, worksiteCounty, worksiteState, worksitePostalCode
    • wageRateFrom, wageRateTo, wageUnitOfPay, prevailingWage, pwUnitOfPay, pwTrackingNumber, pwWageLevel, pwOesYear, pwOtherSource, pwOtherYear, pwSurveyPublisher, pwSurveyName
    • totalWorksiteLocations, agreeToLcStatement, h1bDependent, willfulViolator, supportH1b, statutoryBasis, appendixAAttached, publicDisclosure
    • masterExemptionWorkerCount, masterExemptionDegrees (array of {institutionName, fieldOfStudy, dateOfDegree}) - advanced-degree exemption detail from DOL's Appendix A file, present only for the subset of H-1B-dependent-employer cases that claim the masters/advanced-degree exemption
    • additionalWorksiteLocations (array of {worksiteWorkers, secondaryEntity, secondaryEntityBusinessName, worksiteAddress, worksiteCity, worksiteCounty, worksiteState, worksitePostalCode, wageRateFrom, wageRateTo, wageUnitOfPay, prevailingWage, pwUnitOfPay, pwWageLevel}) - every worksite location DOL's separate LCA Worksites file discloses for the case, joined in by case number automatically; present only for the minority of cases that report more than one worksite (a filing can disclose up to 10)
    • preparerName, preparerBusinessName, preparerEmail
    • sourceUrl - the DOL disclosure file this record came from
    • recordType: "lcaDisclosure", scrapedAt

    These are the documented fields. Optional ones can be empty on a given record, so measure how often each field your deliverable depends on is populated across a real sample before automating the handoff.

    How do you build the workflow end to end?

    Open H-1B LCA Disclosure Data Scraper and work through these in order. Each step ends with something to check, so a bad configuration surfaces on a small run rather than a scheduled one.

    1. Select mode as search and specify fiscalYear as latest, or set it to an archival year like FY2024 to target historical baselines.
    2. Target specific entities using employerName, or supply an exact employerFein to avoid entity name ambiguity across subsidiaries.
    3. Refine the occupation scope by entering socCode or jobTitle, and restrict visaClass to H-1B if non-immigrant specialty workers are your sole interest.
    4. Ensure proxyConfiguration retains a residential proxy group so large quarterly file downloads from dol.gov do not return an edge 403 status.
    5. Set maxItems to a conservative test number like 10, execute the run, and verify that the run starts and outputs valid JSON records.
    6. Inspect the resulting dataset to verify core fields such as caseNumber, caseStatus, prevailingWage, and wageRateFrom are populated.
    7. Widen maxItems toward your desired sample size, or swap mode to byCaseNumbers with a caseNumbers array to pull specific docket items.

    How do you apply it? Three worked playbooks

    These are H-1B LCA Disclosure Data Scraper's own documented use cases, each worked through as an operating pattern rather than a description.

    Use case 1: Immigration law firms

    Outcome: Research an employer's H-1B sponsorship history and prevailing wage levels before filing

    Configure: Set mode to "search", fiscalYear to "latest", employerName to "Google", visaClass to "H-1B", and maxItems to 250.

    Working method: Execute an initial run on the employer name, inspect prevailingWage against wageRateFrom on certified cases, and widen across prior years using fiscalYear set to FY2024 for multi-year trend analysis.

    Deliverable: A historical dataset of certified wage distributions, job titles, and attorney representation details for the target sponsor.

    Stop condition: Stop if zero records return across multiple fiscal years or if caseStatus returns Denied for the entire sample.

    Use case 2: Recruiting & talent sourcing

    Outcome: Identify which employers actively sponsor visas for a given job title or occupation

    Configure: Set mode to "search", socCode to "15-1252", fullTimePositionOnly to true, wageMin to 120000, and maxItems to 500.

    Working method: Run the occupation query across the latest file, group emitted records by employerName, and calculate which companies maintain the highest volume of approved filings.

    Deliverable: A ranked list of sponsoring companies hiring for the specific occupational code alongside offered compensation ranges.

    Stop condition: Stop if wageUnitOfPay does not match Year, distorting minimum salary filtering thresholds.

    Use case 3: Compliance monitoring

    Outcome: Track your own or a competitor's LCA filings, case statuses, and worksite locations

    Configure: Set mode to "search", employerFein to "12-3456789", caseStatus to "Certified", and maxItems to 100.

    Working method: Schedule periodic queries against the latest quarterly disclosure, evaluate additionalWorksiteLocations arrays for unlisted branches, and map changes in caseStatus.

    Deliverable: A structured audit report detailing newly certified dockets, secondary worksite addresses, and total worker positions.

    Stop condition: Stop if the employer FEIN returns mismatching legal entity names inside employerName.

    What breaks, and how do you design around it?

    Large quarterly DOL disclosure files must be streamed in their entirety, so runs can take several minutes when filtering against rare search terms. To prevent timeouts, constrain your queries using broad initial parameters before layering multiple date filters. When scanning older periods, specify closed fiscal years individually rather than relying on automatic quarterly detection.

    When should you not use H-1B LCA Disclosure Data Scraper?

    Do not use this Actor if you require general occupational labor market trends rather than employer-specific visa dockets; use O*NET Occupation Data Scraper to evaluate standard SOC task requirements and baseline industry wage bands instead. If your objective is monitoring macro-level state or industry wage metrics without employer identity, BLS Labor Statistics Scraper is better suited. Avoid this tool entirely if you need live job vacancy postings, where an active job board scraper provides actual open requisitions rather than regulatory visa pre-filings.

    What should you check before trusting the output?

    • Confirm that caseNumber matches standard DOL filing formats, such as starting with an I-200 prefix.
    • Verify prevailingWage contains a positive numeric value and wageUnitOfPay matches an expected duration like Year or Hour.
    • Check that worksiteState contains a valid two-letter US state code when regional targeting is specified.
    • Ensure employerName is present; halt downstream ingestion if raw null values appear across required entity fields.
    • Inspect additionalWorksiteLocations to confirm secondary location structures parse cleanly when multi-worksite filings are processed.

    None of this proves a record is correct. It gives a scheduled H-1B LCA Disclosure Data Scraper run defined points where it should stop instead of quietly passing bad data downstream.

    Frequently asked questions

    What is the pricing model for running this Actor?

    Billing includes a run-start fee for every run initiated, platform usage billed separately by Apify, and dataset results priced at $0.005 per result. That equals $5.00 per 1,000 results on the free tier, with lower per-result costs available on paid plans.

    Why is a residential proxy required for execution?

    The Department of Labor blocks standard cloud datacenter egress IP addresses with 403 Forbidden errors when attempting to download large quarterly XLSX files. Residential proxies route through legitimate residential IP addresses, ensuring reliable downloads from government file servers.

    Does this tool provide individual worker names?

    No. The Department of Labor public LCA disclosure files omit personal beneficiary identities to protect privacy. The scraper returns filing metadata, job titles, wage figures, employer identifiers, and worksite locations, but no employee names or visa holder biographical details.

    How do I fetch historical filings prior to the current fiscal year?

    Set the fiscalYear parameter to an earlier fiscal year such as FY2025, FY2024, or back to FY2020. Closed fiscal years scan the full cumulative Q4 file containing all annual submissions in a single dataset stream.

    Can I query filings by an employer's tax ID?

    Yes. Set employerFein to the exact Federal Employer Identification Number. The Actor strips hyphens and spaces automatically, allowing precise entity matching without substring false positives from shared corporate names.

    Where to go next

    When you are ready to run it, open H-1B LCA Disclosure Data Scraper on Apify; the free plan covers up to 1,000 results a month.

    Start with the H-1B LCA Disclosure Data Scraper Actor page for the current input schema, pricing tier, and run history.

    Other Actors we maintain for related data:

    Related guides:

    Resources

    • Actor documentation, input schema, and pricing: verified against the published Actor on 2026-09-28.

    • Actor last updated by its maintainers on 2026-08-29.

    • Run outcome figures cover the 30 day public window ending 2026-09-28.

    • H-1B LCA Disclosure Data Scraper on Apify

    Featured actors

    H-1B LCA Disclosure Data Scraper

    Scrape US Department of Labor (DOL) H-1B / H-1B1 / E-3 Labor Condition Application (LCA) disclosure data. Search by employer, job title, SOC code, wage, worksite state, visa class, and more, or look up exact case numbers.

    Run on Apify ↗