Github AI Coding Task Dataset Builder
Pricing
from $3.00 / 1,000 results
Github AI Coding Task Dataset Builder
Turn GitHub issues into structured coding-agent tasks with repository context, relevant files, build commands, and ready-to-use prompts.
Github AI Coding Task Dataset Builder
Pricing
from $3.00 / 1,000 results
Turn GitHub issues into structured coding-agent tasks with repository context, relevant files, build commands, and ready-to-use prompts.
Use the safe built-in demo or scan live GitHub repositories and organizations.
Repositories as owner/repo or github.com URLs. You can combine these with organizations and a source dataset.
[]Organization logins or github.com organization URLs. Repositories are expanded automatically.
[]Optional Apify dataset containing GitHub repository strings.
Field containing owner/repo or a GitHub repository URL in each source dataset item.
Optional for small public scans, strongly recommended for large batches, and required for private repositories. Read-only repository access is sufficient.
Open issues are best for current agent work. Closed/all issues are useful for benchmark and historical datasets.
Only scan issues updated within this many days. Set 0 to disable the time filter.
Optional label allow-list. If set, an issue must have at least one listed label.
[]Issues with any of these labels are skipped before scoring.
[ "duplicate", "invalid", "wontfix", "won't fix", "question", "support", "discussion", "not planned"]Keep all issues, only assigned issues, or only unassigned issues.
0-100 heuristic confidence that an issue is an actionable coding task. 55 is a balanced default.
Maximum unique repositories expanded and processed in one run.
Maximum non-pull-request issues scanned per repository, ordered by most recently updated.
Stop emitting new task rows after this many accepted coding tasks.
Include archived repositories.
Maximum high-value manifests/config/instruction files fetched per repository for stack and command detection.
Selected context files larger than this are skipped. Secret-like files are never fetched.
Maximum likely-relevant repository paths attached to each coding task.
Maximum issue-body characters stored in each task and generated prompt.
Generate a ready-to-use task prompt for Codex, Cursor, Claude Code, and other coding agents.
Save compact per-repository context JSON files in the run key-value store.
Repositories processed in parallel. Keep conservative to reduce GitHub secondary-rate-limit risk.
Selected repository blobs fetched in parallel inside each repository.