Skip to content
    ↑↓ to choose · Enter to open

    · 13 min read

    GitHub Repository Intelligence: Up to 2,500 Free Results a Month

    By CrawlerBros Engineering Team

    Each record carries 9 output fields, including stars, forks, language byte counts, topic tags, and decoded README text up to 500 KB. Results cost $2.00 per 1,000 results on the free plan, with up to 1,000 repositories returned per search query in search mode or unlimited target lists fetched in parallel via direct mode. This tool is for software engineers, product managers, and data analysts building technology maps or tracking project activity. It is not for anyone who needs complete commit histories, issue comments, or contributor pull requests, which the output does not include.

    Try it before you read further. Apify's free plan includes $5.00 of usage every month with no credit card, enough for up to 2,500 results at $0.002 each before platform usage. Open GitHub Repository Intelligence on Apify and run the prefilled example.

    How reliable is GitHub Repository Intelligence in production?

    Across the last 30 days of public runs on the Apify platform, GitHub Repository Intelligence recorded 56 runs with the following outcomes.

    Outcome Runs Share
    Succeeded 56 100.0%
    Failed 0 0.0%
    Aborted by the user 0 0.0%
    Timed out 0 0.0%
    Total 56 100.0%

    No run failed or timed out in the last 30 days. Keep a retry and an alert on scheduled runs all the same: a clean month is a record, not a guarantee.

    What does it cost to run GitHub Repository Intelligence?

    Each result costs $0.002 on Apify's free plan, which is $2.00 per 1,000 results. Starting a run is charged separately at $0.005 per GB of Actor memory. Apify also bills the platform usage each run consumes, at the rates of your Apify plan, on top of these charges.

    Apify plan Per result Per 1,000 results
    FREE $0.002 $2.00
    BRONZE $0.00167 $1.67
    SILVER $0.00133 $1.33
    GOLD $0.001 $1.00
    PLATINUM $0.001 $1.00
    DIAMOND $0.001 $1.00

    Worked example: collecting 10,000 results costs $20.00 in result charges before run-start fees and platform usage. No run failed or timed out in the last 30 days, so the list price is a fair budget; keep a retry in place all the same.

    Your output dataset size is determined by maxResults in search mode or the number of items passed in repositoryUrls in direct mode. Enabling or disabling includeReadme, includeLanguages, or includeTopics adjusts internal API requests per repository but does not alter the dataset result charge. To test your parameters cheaply, set maxResults to 30 on a first run to cap result charges at $0.06.

    How do you run GitHub Repository Intelligence from the API?

    None of its 9 controls is strictly required, so the defaults below produce a valid run on their own. Nothing in the payload below is illustrative. Those are the schema's prefilled defaults for GitHub Repository Intelligence, so the request works once your token is in place.

    Call the synchronous endpoint to start a run and receive dataset items in one request:

    curl -X POST "https://api.apify.com/v2/acts/crawlerbros~github-repo-intelligence/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
      -H "Content-Type: application/json" \
      -d '{"mode":"search","searchQuery":"stars:>10000 language:python","repositoryUrls":["https://github.com/apify/crawlee"],"sortBy":"stars","maxResults":30,"includeReadme":true,"includeTopics":true,"includeLanguages":true}'
    

    The same run from Python, using the official client:

    from apify_client import ApifyClient
    
    client = ApifyClient("<YOUR_APIFY_TOKEN>")
    
    run_input = {
      "mode": "search",
      "searchQuery": "stars:>10000 language:python",
      "repositoryUrls": [
        "https://github.com/apify/crawlee"
      ],
      "sortBy": "stars",
      "maxResults": 30,
      "includeReadme": True,
      "includeTopics": True,
      "includeLanguages": True
    }
    
    run = client.actor("crawlerbros~github-repo-intelligence").call(run_input=run_input)
    
    for item in client.dataset(run["defaultDatasetId"]).iterate_items():
        print(item)
    

    And from Node.js:

    import { ApifyClient } from 'apify-client'
    
    const client = new ApifyClient({ token: '<YOUR_APIFY_TOKEN>' })
    
    const input = {
      "mode": "search",
      "searchQuery": "stars:>10000 language:python",
      "repositoryUrls": [
        "https://github.com/apify/crawlee"
      ],
      "sortBy": "stars",
      "maxResults": 30,
      "includeReadme": true,
      "includeTopics": true,
      "includeLanguages": true
    }
    
    const run = await client.actor('crawlerbros~github-repo-intelligence').call(input)
    const { items } = await client.dataset(run.defaultDatasetId).listItems()
    console.log(items)
    

    Because the call is synchronous, your client waits for the whole run. Keep it for exploration. For scheduled work, start the run without waiting and collect the dataset afterwards, so network trouble costs you a retry rather than the results.

    Which GitHub Repository Intelligence inputs matter, and which can you skip?

    The main setting is mode, which toggles between evaluating a searchQuery string or retrieving a list of repositoryUrls. Most runs should leave includeReadme, includeTopics, and includeLanguages set to true unless you are running a large scan and want to minimise internal GitHub API calls.

    • mode (string): Choose 'search' to query GitHub's search API with a keyword / qualifier expression, or 'direct' to fetch a known list of repository URLs. Default: "search".
    • searchQuery (string): GitHub search query (used when mode = search). Supports qualifiers like 'stars:>1000', 'language:python', 'topic:llm', 'user:apify'. See GitHub's search syntax docs for the full grammar. Default: "stars:>10000 language:python".
    • repositoryUrls (array): List of full GitHub repository URLs to scrape (used when mode = direct). Accepts https://github.com/owner/repo, ending with .git, trailing slash, or the git@github.com:owner/repo.git SSH form.
    • sortBy (string): Sort order for search results. Ignored in direct mode. Default: "stars".
    • maxResults (integer): Hard cap on the number of repositories to return. GitHub's search API itself is capped at 1000 results; direct mode is only limited by this value. Default: 30.
    • includeReadme (boolean): Fetch and include the decoded README text for each repository (capped at 500 KB). Default: true.
    • includeTopics (boolean): Include the repository's topic tags (e.g. 'machine-learning', 'web-scraping'). Default: true.
    • includeLanguages (boolean): Include the per-language byte counts (e.g. {"Python": 123456, "JavaScript": 2345}). Default: true.
    • githubToken (string): Optional GitHub personal access token. Mark this field as Secret when pasting. Increases the API rate limit from 60 requests/hour (unauthenticated) to 5000 requests/hour. A classic token with 'public_repo' scope is enough for public repositories.

    Fixed-choice controls: mode accepts search (Search GitHub), direct (Direct repository URLs); sortBy accepts stars, forks, updated (Recently updated), help-wanted-issues (Help-wanted issues).

    What does GitHub Repository Intelligence return?

    Output records supply repository identity fields, star and fork counts, license SPDX keys, and byte counts per programming language. They omit commit histories, issue discussions, and contributor email addresses.

    • Identity - id, name, fullName, htmlUrl, owner (login / id / avatar / URL / type)
    • Description - description, homepage, primaryLanguage
    • Engagement - stars, forks, watchers, openIssues, size (in KB)
    • Tags and license - topics array, license block with SPDX identifier
    • Timestamps - createdAt, updatedAt, pushedAt
    • Status flags - isFork, isArchived, isDisabled, isTemplate
    • Branch - defaultBranch
    • Content - readme (decoded, truncated to 500 KB), languages (bytes per language)
    • scrapedAt - ISO-8601 timestamp of this run

    These are the documented fields. Optional ones can be empty on a given record, so measure how often each field your deliverable depends on is populated across a real sample before automating the handoff.

    How do you build the workflow end to end?

    Open GitHub Repository Intelligence and work through these in order. Each step ends with something to check, so a bad configuration surfaces on a small run rather than a scheduled one.

    1. Set mode to search with a searchQuery like stars:>1000 language:python and maxResults set to 5 to verify repository fields on a tiny run.
    2. Inspect the dataset to verify that returned records include primary keys like fullName, stars, and languages, rather than error objects.
    3. Paste a classic personal access token into githubToken if your batch exceeds roughly 20 repositories to increase the GitHub API limit to 5000 requests per hour.
    4. Disable includeReadme, includeTopics, or includeLanguages if you only need core metrics, dropping internal API calls per repository down to one.
    5. Switch mode to direct and add repository URLs to repositoryUrls if you already have a predefined list of target repositories.
    6. Increase maxResults up to 1000 or partition your searchQuery into distinct star or date ranges if your query hits the 1000-result search limit.

    How do you apply it? Three worked playbooks

    These are GitHub Repository Intelligence's own documented use cases, each worked through as an operating pattern rather than a description.

    Use case 1: Ecosystem mapping

    Outcome: Enumerate every repo tagged with topic:llm or topic:web3 to build a competitive map

    Configure: Set mode to search, searchQuery to topic:llm stars:>100, sortBy to stars, maxResults to 250, includeTopics to true, and provide a valid githubToken.

    Working method: Run an initial search on primary ecosystem topic tags, inspect the returned topics list on high-star repositories to extract secondary tags, and execute follow-up queries on those discovered terms.

    Deliverable: A deduplicated catalog of GitHub repositories containing repository names, star counts, primary languages, and associated topic tags.

    Stop condition: The search query returns 1000 repositories, indicating the query must be split into star-range brackets.

    Use case 2: Leaderboards and dashboards

    Outcome: Rank a set of repos by stars, forks, or recent activity for internal dashboards

    Configure: Set mode to direct, add repository target URLs to repositoryUrls, set includeReadme to false, and populate githubToken.

    Working method: Supply a curated list of target project URLs to direct mode on a recurring run schedule and sort output items by engagement parameters.

    Deliverable: A structured tabular feed of engagement metrics including stars, forks, openIssues, watchers, and updatedAt timestamps.

    Stop condition: Output records show type equal to github_repo_intelligence_error with reason set to not_found or rate_limit.

    Use case 3: Market research

    Outcome: Study language mix, topic distribution, and growth across a population of related projects

    Configure: Set mode to search, searchQuery to topic:web3, includeLanguages to true, includeTopics to true, and maxResults to 500.

    Working method: Query category qualifiers, aggregate the byte counts returned in the languages key, and calculate relative adoption rates across programming stacks.

    Deliverable: A dataset showing programming language byte breakdowns, project sizes, and topic distributions across a technology sector.

    Stop condition: Zero records are returned or searchQuery fails to match repositories due to invalid qualifier syntax.

    What breaks, and how do you design around it?

    • 1000-result search cap. GitHub itself caps any search query at 1000 results. For larger spaces, slice the query into ranges.
    • Anonymous rate limit is tight. Without a githubToken you get about 60 API calls per hour. Each enriched repo is up to 3 calls, so runs over ~20 repositories need a token.
    • README truncation at 500 KB. Very long READMEs (rare) are cut at 500 KB with a marker.
    • No commit, issue, or PR data. This actor focuses on repository-level metadata; commit history, issues, and pull requests are out of scope.
    • Private repos need explicit access. The githubToken must have repo scope for private repositories; public-only tokens return a not_found error record for private URLs.

    GitHub enforces a 1000-result limit on search queries; split your searchQuery using explicit star ranges like stars:100..500 to extract larger populations. Standard unauthenticated runs hit GitHub API caps quickly, so pass a classic personal access token with public_repo scope in githubToken to raise the limit from 60 to 5000 requests per hour. README content is capped at 500 KB per repository, so save original file links separately if unclipped text is necessary.

    When should you not use GitHub Repository Intelligence?

    Do not use this Actor if you require commit history logs, pull request threads, or user issue comments. This Actor extracts high-level repository metadata and omits detailed event histories. If you want historical metrics, contributor velocity, and curated project rankings across millions of repositories without writing query strings, use OSS Insight Scraper instead. Furthermore, if you are auditing local repositories or internal projects where you already have local git clones, running git commands locally or making direct requests to the official GitHub GraphQL API is faster and avoids platform usage charges.

    What should you check before trusting the output?

    • Check dataset records for type matching github_repo_intelligence_error and inspect the reason field for rate_limit or not_found values.
    • Verify that the stars field contains an integer value of zero or greater across all repository records.
    • Check that the languages record contains expected language byte counts when includeLanguages is set to true.
    • Confirm that license contains a non-null spdxId string when verifying open-source compliance terms.

    None of this proves a record is correct. It gives a scheduled GitHub Repository Intelligence run defined points where it should stop instead of quietly passing bad data downstream.

    Frequently asked questions

    What does it cost to scrape 1,000 GitHub repositories?

    Results cost $2.00 per 1,000 results on the free plan, which equals $0.002 per result written to the dataset. Apify's free plan includes $5.00 of monthly usage with no credit card, covering up to 2,500 results of this Actor before run-start and platform resource charges. Paid Apify plans pay lower rates per result.

    Why should I supply a GitHub token?

    Without a token, requests run under GitHub's unauthenticated rate limit of 60 requests per hour. Fetching detailed metadata, languages, and README files takes up to three API calls per repo, restricting unauthenticated runs to roughly 20 repos per hour. Supplying a personal access token in githubToken lifts the API rate limit to 5000 requests per hour.

    Can I return more than 1,000 results from a single search query?

    No. GitHub's search API imposes a hard limit of 1000 results per search query. To capture more repositories, split your searchQuery into multiple ranges using numerical qualifiers such as stars:100..500 and stars:501..1000.

    What happens when a repository URL is deleted or private?

    The Actor continues running if a single repository fails. It writes an error record to the dataset containing the repository path and a reason code like not_found or invalid_url, then proceeds with the remaining URLs in the batch.

    Are private GitHub repositories supported?

    Yes. Private repositories can be scraped in direct mode if the token provided in githubToken has repo scope permissions granting access. If the token lacks permission, the Actor writes an error record with a not_found reason code.

    Where to go next

    When you are ready to run it, open GitHub Repository Intelligence on Apify; the free plan covers up to 2,500 results a month.

    Other Actors we maintain for related data:

    • OSS Insight Scraper: Scrape OSS Insight - open-source intelligence on 5M+ GitHub repos.
    • PubMed Search Scraper: Search PubMed (NCBI E-utilities) for biomedical articles by keyword, date range, and article type.
    • ImportYeti Trade Intelligence Scraper: Scrape US import/export trade data from ImportYeti: companies, suppliers, shipments, top trading partners, trademarks, countries, and shipment recency.
    • B2B Sales Trigger Intelligence: Detect high-intent B2B sales triggers hiring surges, funding rounds, executive changes, and news momentum from a list of company names.

    Related guides:

    Resources

    • Actor documentation, input schema, and pricing: verified against the published Actor on 2026-09-29.

    • Actor last updated by its maintainers on 2026-08-18.

    • Run outcome figures cover the 30 day public window ending 2026-09-29.

    • GitHub Repository Intelligence on Apify

    Featured actors

    GitHub Repository Intelligence

    Fetch rich metadata (stars, forks, README, languages, topics, license) from GitHub repositories. Search by query or provide direct URLs. Optional GitHub token for 80x higher rate limit.

    Run on Apify ↗