Open Canada Dataset Catalogue Scraper avatar

Open Canada Dataset Catalogue Scraper

Pricing

$1.00 / 1,000 results

Go to Apify Store
Open Canada Dataset Catalogue Scraper

Open Canada Dataset Catalogue Scraper

Open Canada scraper and API: export Government of Canada open-data dataset records (title, publisher, licence, file formats, keywords, last update) from open.canada.ca to JSON, CSV or Excel. Official public CKAN API, no login, keyword and organisation search.

Pricing

$1.00 / 1,000 results

Rating

0.0

(0)

Developer

COMPASSLAB

COMPASSLAB

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

4 hours ago

Last modified

Categories

Share

Get Open Canada dataset records as clean JSON, CSV or Excel, from the Apify API or on a schedule. No login, about $1.00 for 1,000 datasets.

What does Open Canada Dataset Catalogue Scraper do?

Open Canada Dataset Catalogue Scraper extracts structured data from open.canada.ca. Collect dataset metadata from the Government of Canada Open Data portal (open.canada.ca) via its public CKAN API, with filters and automatic pagination, for research, data discovery and monitoring of open government data. It works as an API for open.canada.ca data: run it from Apify Console, on a schedule, or from your own code, and get clean, typed JSON with numbers as numbers and dates in ISO 8601.

What you get

Data12 fields per item: url, id, name, title, organization, datasetUrl, ...
FormatsJSON, CSV, Excel, HTML, or the Apify API
Price$1.00 per 1,000 datasets, pay per result
AccessPublic open.canada.ca data only: no login, no cookies, robots.txt respected
LicenceOpen Government Licence – Canada

Why use Open Canada Dataset Catalogue Scraper?

  • Data journalists and researchers: find every dataset on a topic, with its licence, publisher and files.
  • Govtech and civic tech: watch a portal for new or updated datasets and their download links.
  • Data catalogue aggregation: merge several open-data portals into one searchable index.
  • AI agents and RAG: give an LLM the portal's metadata and file links as clean JSON.

Main features:

  • Follows pagination up to maxPages pages per start URL and stops at maxItems results.
  • Filters: query (search query), organization (organization), pageSize (page size), so you only get (and pay for) the results you need.
  • Polite by default: respects robots.txt, at most maxConcurrency parallel requests and a delay between requests.
  • Checks every result against field validators, so layout changes show up as clear data-quality warnings.
  • Runs on the Apify platform: scheduling, API access, integrations, monitoring and datasets you can export.

What data can Open Canada Dataset Catalogue Scraper extract?

FieldTypeDescription
urlstringAPI page URL the dataset was scraped from
idstringUnique dataset identifier (UUID)
namestringDataset URL slug
titlestringDataset title (English, falls back to French if missing)
organizationstringPublishing organization name
datasetUrlstringPublic dataset page on open.canada.ca
notesstringDataset description (may be truncated to 900 characters)
licensestringLicense title
resourceCountintegerNumber of resources (files/links) in the dataset
formatsarrayDistinct resource formats (e.g. CSV, JSON, XLSX)
keywordsarrayDataset tags/keywords
metadataModifiedISO 8601 dateLast modified date of the metadata (ISO 8601 date)

How to scrape open.canada.ca

  1. Open Open Canada Dataset Catalogue Scraper in Apify Console and go to the Input tab.
  2. Enter what to scrape (see the Input section below), for example the start URLs.
  3. Set Max items to the number of results you need.
  4. Click Start and wait for the run to finish.
  5. Download the results from the Output tab, or fetch them with the API.

How much will it cost to scrape open.canada.ca?

This Actor is priced per result: $1.00 per 1,000 results, with no extra charge for platform usage. That is about $1.00 for 1,000 datasets: 100 results cost $0.10 and 10,000 results cost $10.00. Set a maximum cost per run and the Actor stops when it is reached.

Input

See the Input tab for full configuration options.

FieldTypeRequiredDescription
querystringnoOptional free-text query (CKAN 'q' parameter) added to each start URL, e.g. 'climate'.
organizationstringnoOptional organization name (slug) to filter datasets, e.g. 'statcan'. Applied as an fq filter.
maxItemsintegernoMaximum number of items to return (0 = unlimited).
startUrlsarraynopackage_search API URLs to scrape, e.g. https://open.canada.ca/data/api/action/package_search?rows=50. Query parameters such as q, fq and sort are preserved; start is advanced automatically for pagination.
pageSizeintegernoNumber of datasets requested per API call (rows parameter, 1-1000).
maxPagesintegernoMaximum listing pages to follow per start URL (pagination).
maxConcurrencyintegernoMaximum parallel requests (politeness; 1-10).
requestDelayMsintegernoMinimum delay between requests, in milliseconds (at least 250).
proxyTypestringnonone (direct connection), datacenter (Apify Proxy, cheapest) or residential (opt-in, billed per GB, fewer blocks). The actor never switches by itself.
proxyCountrystringnoTwo-letter country code for the proxy IP (optional).

Example input:

{
"startUrls": [
{
"url": "https://open.canada.ca/data/api/action/package_search?rows=50"
}
],
"maxItems": 75,
"maxPages": 3,
"maxConcurrency": 2,
"requestDelayMs": 1000,
"proxyType": "none",
"pageSize": 50
}

Output

You can download the dataset in various formats such as JSON, HTML, CSV, or Excel. Example results from a real run:

[
{
"url": "https://open.canada.ca/data/api/action/package_search?rows=50&start=0",
"id": "7c03f039-3753-4093-af60-74b0f7b2385d",
"name": "7c03f039-3753-4093-af60-74b0f7b2385d",
"title": "Government of Canada - Consultations",
"organization": "Privy Council Office | Bureau du Conseil privé",
"datasetUrl": "https://open.canada.ca/data/en/dataset/7c03f039-3753-4093-af60-74b0f7b2385d",
"notes": "**This dataset consolidates all the consultations submitted by departments and agencies in the Government of Canada.** Please note that consultations that were entered before January 2018 may be missing information such as description and subject(s). In spring of 2018, consultation information from…",
"license": "Open Government Licence - Canada",
"resourceCount": 4,
"formats": [
"CSV",
"JSON",
"XLSX"
],
"keywords": [
"Public engagement",
"Consultations",
"Open dialogue"
],
"metadataModified": "2026-10-01"
},
{
"url": "https://open.canada.ca/data/api/action/package_search?rows=50&start=0",
"id": "b27ec3ef-7338-4e76-a6fd-128339a92df5",
"name": "b27ec3ef-7338-4e76-a6fd-128339a92df5",
"title": "Who we regulate",
"organization": "Office of the Superintendent of Financial Institutions Canada | Bureau du surintendant des institutions financières Canada",
"datasetUrl": "https://open.canada.ca/data/en/dataset/b27ec3ef-7338-4e76-a6fd-128339a92df5",
"notes": "We regulate and supervise approximately 350 financial institutions. In this dataset you’ll find a list of these institutions. We regulate and supervise approximately 1200 private pension plans for employees working in federally regulated industries. Here you’ll find a list of these plans, or see our…",
"license": "Open Government Licence - Canada",
"resourceCount": 10,
"formats": [
"CSV",
"DOCX"
],
"keywords": [
"FRFI",
"Federally Regulated Financial Institutions",
"Domestic Banks",
"Foreign Banks"
],
"metadataModified": "2026-10-01"
},
{
"url": "https://open.canada.ca/data/api/action/package_search?rows=50&start=0",
"id": "ededff77-a021-48d6-89a5-cdbcd75fb4ff",
"name": "ededff77-a021-48d6-89a5-cdbcd75fb4ff",
"title": "Pest Management Regulatory Agency (PMRA) List of Formulants",
"organization": "Health Canada | Santé Canada",
"datasetUrl": "https://open.canada.ca/data/en/dataset/ededff77-a021-48d6-89a5-cdbcd75fb4ff",
"notes": "This dataset contains a list of formulants that are found in pest control products currently registered in Canada under the Pest Control Products Act and Regulations. This list is updated twice a year to reflect the addition of new formulants and the deletion of formulants no longer found in registe…",
"license": "Open Government Licence - Canada",
"resourceCount": 5,
"formats": [
"HTML",
"XLSX",
"CSV"
],
"keywords": [
"Regulation of Formulants; formulant",
"pest control product"
],
"metadataModified": "2026-10-01"
}
]

Integrations and API

  • Apify API: start a run and get the results in one HTTP request:
curl -X POST "https://api.apify.com/v2/acts/compass_lab~canada-general-scraper/run-sync-get-dataset-items?token=<YOUR_APIFY_TOKEN>" \
-H "Content-Type: application/json" -d '{"startUrls": [{"url": "https://open.canada.ca/data/api/action/package_search?rows=50"}], "maxItems": 75, "pageSize": 50}'
  • Python (pip install apify-client):
from apify_client import ApifyClient
client = ApifyClient("<YOUR_APIFY_TOKEN>")
run = client.actor("compass_lab/canada-general-scraper").call(run_input={"startUrls": [{"url": "https://open.canada.ca/data/api/action/package_search?rows=50"}], "maxItems": 75, "pageSize": 50})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
print(item)
  • JavaScript (npm install apify-client):
import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: '<YOUR_APIFY_TOKEN>' });
const run = await client.actor('compass_lab/canada-general-scraper').call({"startUrls": [{"url": "https://open.canada.ca/data/api/action/package_search?rows=50"}], "maxItems": 75, "pageSize": 50});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
  • Make, Zapier, n8n, Google Sheets, webhooks: use the Apify integrations (Integrations tab) to send each run's results where you need them, or to start a run from your workflow.
  • Schedules: run it hourly, daily or weekly from Apify Console (Schedules) and always have fresh datasets.

Tips and advanced options

  • Keep Max items and Max pages as low as you need: fewer pages means a faster, cheaper run.
  • Raise Delay between requests if the site responds slowly; keep Max concurrency low to stay polite.
  • Missing values are null. Fields that often come back empty are listed in the run log as data-quality warnings.

FAQ, disclaimers and support

Vetted by the autonomy policy (green tier) on 2026-10-01: Green tier (autonomy policy, 2026-10-01): robots.txt exists and permits /data/api/action/package_search (no matching Disallow); public endpoint; platform family 'open-data-ckan' confirmed by Mouad on 2026-10-01: Open-source portal software, government open-data licences

The data is published under the Open Government Licence – Canada: you may reuse it, including commercially, provided you credit the source as the licence requires.

Our Actors are ethical and do not extract any private user data, such as email addresses, gender, or location. They only extract what the user has chosen to share publicly. We therefore believe that our Actors, when used for ethical purposes by Apify users, are safe. However, you should be aware that your results could contain personal data. Personal data is protected by the GDPR in the European Union and by other regulations around the world. You should not scrape personal data unless you have a legitimate reason to do so. If you're unsure whether your reason is legitimate, consult your lawyers.

How many results can I get?

Up to maxItems per run (0 means no limit), as many as the source lists. Each result is one dataset item, and you are only charged for items that are saved.

Can I run it on a schedule or from my own code?

Yes. Schedule it in Apify Console (Schedules), or call it with the run-sync-get-dataset-items endpoint or the Python/JavaScript clients shown in Integrations and API above.

What are the limitations?

  • Only dataset metadata is returned, not the contents of the data files themselves
  • The portal is a government service, so very large runs should be kept moderate to avoid rate limiting or temporary errors
  • Many datasets are bilingual; the Actor returns English text first and falls back to French when English is missing
  • Descriptions are truncated to 900 characters to keep results compact
  • The CKAN API may cap deep pagination, so very broad searches should be split using the query or organization filters
  • Some datasets have no formats, tags or license, so those fields can be empty

Where can I get help?

Report problems or ideas on the Issues tab. To call this Actor from your own code, see the API tab.

Open Data Suite: the same clean, typed output across sources, so you can combine them in one dataset.

ActorWhat it scrapesPrice
data.gouv.fr Dataset Catalogue ScraperDataset records from data.gouv.fr$1.00 / 1,000