Cross-Platform AI Trending Scraper avatar

Cross-Platform AI Trending Scraper

Pricing

Pay per usage

Go to Apify Store
Cross-Platform AI Trending Scraper

Cross-Platform AI Trending Scraper

Scrapes trending AI repositories, tweets, and papers across GitHub, Twitter, and ArXiv.

Pricing

Pay per usage

Rating

0.0

(0)

Developer

CQ

CQ

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

10 days ago

Last modified

Share

Scrapes recent trending AI content from two public sources and filters it by your topics:

  • GitHub – most-starred repositories created in the last 30 days that match your topics, via the public GitHub Search API.
  • ArXiv – most recent papers in cs.AI / cs.LG / cs.CL matching your topics, via the public ArXiv Atom API.

Both sources are queried with no API key required. A GitHub token can be supplied optionally to raise the GitHub rate limit, but the actor runs fine without one.

Twitter/X is not included. There is no free, key-less Twitter trending API, so this actor does not claim to scrape Twitter. (An earlier version listed Twitter and returned fabricated data; that has been removed.)

Input

FieldTypeDefaultNotes
topicsstring[]["AI","LLM","DeepSeek","Agent"]Terms to match in repo/paper name & text.
maxItemsPerSourceinteger101–50 items per platform.
githubTokenstring (secret)emptyOptional. Raises GitHub rate limit.

Output

Each dataset item:

{
"source": "GitHub",
"title": "owner/repo",
"description": "...",
"url": "https://github.com/owner/repo",
"stars": 1234,
"language": "Python",
"published": "2026-06-01",
"scrapedAt": "2026-06-20T12:00:00.000Z"
}

ArXiv items use authors instead of stars/language. A run summary is also written to the key-value store under SUMMARY.

Reliability

  • Each platform runs independently (Promise.allSettled). One source failing (network error, rate limit, downtime) does not abort the run; remaining sources still produce output and the failure is logged.
  • Each source request is attempted up to 4 times (1 initial try + 3 retries) with a growing backoff (3s / 6s / 9s). Public APIs like GitHub and ArXiv intermittently rate-limit or block shared datacenter IPs, so a single attempt is flaky from the cloud even though it works from a home IP.
  • Every request has a 15s timeout (AbortController) so a hung endpoint fails fast instead of stalling the run.
  • A 4-minute wall-clock budget caps all retries so the run always finishes inside Apify's 5-minute quality-test window.
  • GitHub is queried with exactly one request per run (plus retries on failure) to stay well under the unauthenticated rate limit (10 requests/minute).
  • If every source is temporarily unavailable, the run still finishes successfully with a single marker record (carrying a note explaining the temporary outage) rather than an empty dataset or a failed run.

Usage example

Run with custom topics:

{
"topics": ["Rust", "WebAssembly"],
"maxItemsPerSource": 5
}

This returns up to 5 GitHub repos and 5 ArXiv papers matching Rust or WebAssembly. Leave topics empty (or omit it) to use the defaults.

Limitations

  • Two sources only. GitHub and ArXiv. No Twitter/X, Reddit, Hacker News, Product Hunt, blogs, or other platforms. There is no free, key-less Twitter API, so Twitter is deliberately excluded.
  • No AI / semantic analysis. Filtering is plain keyword matching performed by GitHub's and ArXiv's own search APIs — your topic terms are OR'd into a single query per source. The actor does not rank, summarize, cluster, embed, or use any ML/LLM. "AI" in the name refers to the subject matter (AI repos and papers), not the method.
  • GitHub "trending" is an approximation. Results are repos created in the last 30 days (fixed window, by creation date) whose name or description matches your topic terms, sorted by star count. This is not GitHub's official trending feed and it excludes older repos that are gaining stars now. Terms are matched literally, not semantically.
  • ArXiv "trending" = newest, not most-cited. ArXiv results are the most recently submitted papers in cs.AI / cs.LG / cs.CL matching your terms, sorted by submission date. There is no citation or popularity signal — recency only. Other ArXiv categories are not searched.
  • Volume caps. At most maxItemsPerSource items per source (hard maximum 50). One page of results per source — no pagination. GitHub is queried once per run to stay under the unauthenticated rate limit (10 requests/minute; 1000 results/query).
  • Shared-IP rate limits. From Apify's shared datacenter IPs, GitHub and ArXiv intermittently return 403/429 or time out. The actor retries (up to 4 attempts each, 3s/6s/9s backoff, 15s per-request timeout, 4-minute total budget) and never fails the run, but a given run can still return fewer items than requested. If every source is blocked — or no items match your topics — the run finishes cleanly with a single marker record (source: "none", null title/url/published, plus a note) instead of an empty dataset. Re-run shortly, or supply a GitHub token / enable a proxy for higher reliability.
  • Field shape varies by source. GitHub items include stars and language; ArXiv items include authors instead. Not every field is present on every record.
  • Minimal ArXiv parsing. ArXiv Atom XML is parsed with lightweight, dependency-free regex targeting ArXiv's standard response format — not a general XML parser.
  • No history or dedup. Each run is independent. There is no cross-run deduplication or persistence beyond the dataset and the per-run SUMMARY key-value record.
  • Blank-topic fallback. If you supply topics but every entry is blank, the actor falls back to the default topics rather than failing the run.