# Creative Commons Media Search - Openverse Scraper (`jungle_synthesizer/openverse-media-scraper`) Actor

Search and export Creative Commons and public domain images and audio from Openverse's 800M+ item media index (Flickr, Wikimedia Commons, NASA, Smithsonian, Europeana, Jamendo and 50+ more). Filter by license, category and provider; each record includes the file URL, creator and attribution text.

- **URL**: https://apify.com/jungle\_synthesizer/openverse-media-scraper.md
- **Developed by:** [BowTiedRaccoon](https://apify.com/jungle_synthesizer) (community)
- **Categories:** AI, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.80 / 1,000 record scrapeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Creative Commons Media Search — Openverse Scraper

Search and export Creative Commons and public domain images and audio from [Openverse](https://openverse.org/), the open media index spanning 800M+ works. Returns the direct file URL, creator, license code, ready-to-use attribution text, tags, and format details for every record across 50+ sources including Flickr, Wikimedia Commons, NASA, and the Smithsonian.

***

### Openverse Scraper Features

- Searches images and audio in one run, or narrows to a single media type
- Filters by license — CC BY, CC0, public domain mark, and 7 more — plus category, provider, and (for images) aspect ratio and size
- Returns full license metadata and pre-formatted attribution text, not just a license code
- Covers 50+ providers spanning Flickr, Wikimedia Commons, NASA, the Smithsonian, Europeana, and Jamendo
- Includes direct file URLs, dimensions or duration, tags, and file type — everything needed to actually use the media, not just find it

***

### Who Uses Openverse Media Data?

- **Content marketers** — source properly licensed imagery for blog posts and social without a stock-photo subscription
- **Educators and course creators** — pull public domain photographs and CC-licensed audio for slide decks and course materials, with the attribution text already written
- **App and product teams** — populate a media library with openly licensed assets and keep the license trail for legal review
- **Researchers** — build datasets of openly licensed images or audio filtered by exact license, so provenance is never in question
- **Archivists and librarians** — track what a specific institution — Smithsonian, NYPL, Europeana — has released under open licenses

***

### How Openverse Scraper Works

1. Pick a media type — images, audio, or both — and an optional search phrase.
2. Narrow the results with license, category, provider, or (for images) size and aspect ratio filters.
3. The scraper pages through Openverse's index until it hits your `maxItems` cap, capturing every field the source returns.
4. Each record lands in the dataset with a direct file URL, creator, and pre-formatted attribution text — ready to drop into a CMS or dataset.

***

### Input

```json
{
  "mediaTypes": ["images"],
  "searchQuery": "mountain landscape",
  "licenses": ["cc0", "by"],
  "imageCategories": ["photograph"],
  "maxItems": 200
}
```

| Field             | Type    | Default      | Description |
|-------------------|---------|--------------|-------------|
| `mediaTypes`      | Array   | `["images"]` | Which Openverse media types to search — images, audio, or both. Empty selects both. |
| `searchQuery`     | String  | `""`         | Full-text search phrase, matched against title, description and tags. Leave blank to browse by filters only. |
| `licenses`        | Array   | `[]`         | Restrict to one or more specific licenses (CC BY, CC0, public domain mark, and 7 more). Empty includes every license. |
| `imageCategories` | Array   | `[]`         | Restrict image results to one or more categories (photograph, illustration, digitized artwork). Images only. |
| `audioCategories` | Array   | `[]`         | Restrict audio results to one or more categories (music, podcast, sound effect, and 3 more). Audio only. |
| `sources`         | Array   | `[]`         | Restrict to one or more of Openverse's 50+ provider sources (Flickr, NASA, Jamendo, and more). A provider only applies to the media type it actually serves. |
| `aspectRatio`     | Array   | `[]`         | Restrict image results to square, tall, or wide. Images only. |
| `imageSize`       | Array   | `[]`         | Restrict image results to large, medium, or small. Images only. |
| `includeMature`   | Boolean | `false`      | Include results Openverse flags as sensitive/mature content. |
| `maxItems`        | Integer | `10`         | Maximum number of media records to return. |

Searching only audio, filtered to a couple of providers:

```json
{
  "mediaTypes": ["audio"],
  "searchQuery": "jazz",
  "sources": ["jamendo", "freesound"],
  "maxItems": 100
}
```

***

### Openverse Scraper Output Fields

Every record carries the same 24 fields. `width`/`height` populate on images; `durationMs` and `genres` populate on audio — the fields that don't apply to a record's media type come back `null`.

#### Example — Image Record

```json
{
  "id": "1c5442f6-6bb6-4ab7-b603-f598e7579dd2",
  "mediaType": "image",
  "title": "Cat Fish 2",
  "creator": "admiller",
  "creatorUrl": "https://www.flickr.com/photos/32426194@N00",
  "source": "flickr",
  "license": "by",
  "licenseVersion": "2.0",
  "licenseUrl": "https://creativecommons.org/licenses/by/2.0/",
  "attribution": "\"Cat Fish 2\" by admiller is licensed under CC BY 2.0. To view a copy of this license, visit https://creativecommons.org/licenses/by/2.0/.",
  "foreignLandingUrl": "https://www.flickr.com/photos/32426194@N00/3481540500",
  "fileUrl": "https://live.staticflickr.com/3313/3481540500_c846c62863_b.jpg",
  "thumbnailUrl": "https://api.openverse.org/v1/images/1c5442f6-6bb6-4ab7-b603-f598e7579dd2/thumb/",
  "category": null,
  "tags": ["cat", "glass", "one"],
  "fileType": null,
  "fileSize": null,
  "mature": false,
  "width": 716,
  "height": 1024,
  "durationMs": null,
  "genres": [],
  "detailUrl": "https://api.openverse.org/v1/images/1c5442f6-6bb6-4ab7-b603-f598e7579dd2/",
  "scraped_at": "2026-09-30T02:15:00.000Z"
}
```

#### Example — Audio Record

```json
{
  "id": "8457fac8-84be-48b9-9b57-143ec5e5fbd9",
  "mediaType": "audio",
  "title": "latino-jazz-cash",
  "creator": "Les Oreilles en Ballades",
  "creatorUrl": "https://www.jamendo.com/artist/2210/les.oreilles.en.ballades",
  "source": "jamendo",
  "license": "by-nc-nd",
  "licenseVersion": "2.5",
  "licenseUrl": "https://creativecommons.org/licenses/by-nc-nd/2.5/",
  "attribution": "\"latino-jazz-cash\" by Les Oreilles en Ballades is licensed under CC BY-NC-ND 2.5.",
  "foreignLandingUrl": "https://www.jamendo.com/track/15767",
  "fileUrl": "https://prod-1.storage.jamendo.com/?trackid=15767&format=mp32",
  "thumbnailUrl": null,
  "category": "music",
  "tags": ["energetic", "instrumental"],
  "fileType": "mp32",
  "fileSize": null,
  "mature": false,
  "width": null,
  "height": null,
  "durationMs": 202000,
  "genres": ["funk", "pop", "rnb"],
  "detailUrl": "https://api.openverse.org/v1/audio/8457fac8-84be-48b9-9b57-143ec5e5fbd9/",
  "scraped_at": "2026-09-30T02:15:00.000Z"
}
```

| Field               | Type            | Description                                                                 |
|---------------------|-----------------|-----------------------------------------------------------------------------|
| `id`                | String          | Openverse UUID for this media item.                                         |
| `mediaType`         | String          | `image` or `audio`.                                                         |
| `title`             | String          | Title of the work.                                                          |
| `creator`           | String          | Name of the creator/artist, where credited.                                 |
| `creatorUrl`        | String          | Link to the creator's profile on the source platform.                       |
| `source`            | String          | Machine-readable provider code (e.g. `flickr`, `wikimedia`, `jamendo`).     |
| `license`           | String          | License code (e.g. `by`, `cc0`, `pdm`).                                     |
| `licenseVersion`    | String          | License version (e.g. `4.0`, `2.0`).                                        |
| `licenseUrl`        | String          | Link to the full legal license text.                                        |
| `attribution`       | String          | Ready-to-use attribution text for this work.                                |
| `foreignLandingUrl` | String          | Link to the work's original page on the source platform.                    |
| `fileUrl`           | String          | Direct URL to the media file.                                               |
| `thumbnailUrl`      | String          | Direct URL to a thumbnail/preview of the media file.                        |
| `category`          | String          | Openverse category (e.g. `photograph`, `illustration`, `music`, `podcast`). |
| `tags`              | Array           | Tags/keywords associated with the work.                                     |
| `fileType`          | String          | File extension/format (e.g. `jpg`, `mp3`), where known.                     |
| `fileSize`          | Integer | null | File size in bytes, where known.                                            |
| `mature`            | Boolean         | Whether Openverse flags this item as sensitive/mature content.              |
| `width`             | Integer | null | Image width in pixels. Images only.                                         |
| `height`            | Integer | null | Image height in pixels. Images only.                                        |
| `durationMs`        | Integer | null | Audio duration in milliseconds. Audio only.                                 |
| `genres`            | Array           | Music genres, where tagged. Audio only.                                     |
| `detailUrl`         | String          | Openverse API detail endpoint for this item.                                |
| `scraped_at`        | String          | ISO 8601 timestamp of when the record was emitted.                          |

***

### FAQ

#### How do I search Openverse for Creative Commons images?

Set `mediaTypes` to `["images"]`, add a `searchQuery`, and optionally narrow with `licenses` or `imageCategories`. Leave `searchQuery` blank to browse a filtered slice of the index instead of running a text search.

#### Do I need an account or API key for Openverse?

No. Openverse Scraper needs no account, no API key, and no login on your end — point it at a query and filters, and it returns records.

#### Can I filter results to a specific license?

Yes. The `licenses` field accepts any combination of the ten licenses Openverse tracks — CC BY, CC BY-SA, CC0, the Public Domain Mark, and the rest. Leave it empty to pull every license.

#### What's the difference between images and audio results?

Both media types share the same 24-field output shape. Image records populate `width` and `height`; audio records populate `durationMs` and `genres` instead — the fields that don't apply come back `null` rather than being omitted.

#### How many records can I pull in one run?

As many as your `maxItems` allows. Openverse Scraper paginates automatically until it reaches that cap or the source runs out of matching records, whichever comes first.

***

### Need More Features?

Need custom fields, filters, or a different target site? [File an issue](https://console.apify.com/actors/issues) or get in touch.

### Why Use Openverse Scraper?

- **Accurate licensing, every time** — full license code, version, URL, and pre-formatted attribution text on every record, not a best-guess string.
- **One actor, two media types** — search images and audio together or separately, with filters tuned to what each media type actually supports (aspect ratio and size for images, genres and category for audio).
- **Broad provider coverage** — 50+ sources in one query, from Flickr and NASA to the Smithsonian and Europeana, instead of hitting each source separately.

# Actor input Schema

## `sp_intended_usage` (type: `string`):

What will this data feed? E.g. lead lists, KYB checks, price tracking.

## `sp_improvement_suggestions` (type: `string`):

Provide any feedback or suggestions for improvements.

## `sp_contact` (type: `string`):

We'll personally help with your use case. No spam.

## `mediaTypes` (type: `array`):

Which Openverse media types to search. Select both to search images and audio in the same run. Leave empty to search both.

## `searchQuery` (type: `string`):

Full-text search phrase, matched against title, description and tags. Leave blank to browse by the filters below only.

## `licenses` (type: `array`):

Restrict results to one or more specific licenses. Leave empty to include every license Openverse indexes.

## `imageCategories` (type: `array`):

Restrict image results to one or more categories. Only applies when Media Types includes Images. Leave empty to include all image categories.

## `audioCategories` (type: `array`):

Restrict audio results to one or more categories. Only applies when Media Types includes Audio. Leave empty to include all audio categories.

## `sources` (type: `array`):

Restrict results to one or more specific source providers. A provider only applies to the media type(s) it serves. Leave empty to include every provider.

## `aspectRatio` (type: `array`):

Restrict image results to one or more aspect ratios. Only applies when Media Types includes Images.

## `imageSize` (type: `array`):

Restrict image results to one or more file size classes. Only applies when Media Types includes Images.

## `includeMature` (type: `boolean`):

Include results Openverse flags as sensitive/mature. Off by default, matching Openverse's own default.

## `maxItems` (type: `integer`):

Maximum number of media records to return.

## Actor input object example

```json
{
  "sp_intended_usage": "Describe your intended use...",
  "sp_improvement_suggestions": "Share your suggestions here...",
  "sp_contact": "Share your email here...",
  "mediaTypes": [
    "images"
  ],
  "licenses": [],
  "imageCategories": [],
  "audioCategories": [],
  "sources": [],
  "aspectRatio": [],
  "imageSize": [],
  "includeMature": false,
  "maxItems": 10
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "sp_intended_usage": "Describe your intended use...",
    "sp_improvement_suggestions": "Share your suggestions here...",
    "sp_contact": "Share your email here...",
    "mediaTypes": [
        "images"
    ],
    "searchQuery": "",
    "licenses": [],
    "imageCategories": [],
    "audioCategories": [],
    "sources": [],
    "aspectRatio": [],
    "imageSize": [],
    "maxItems": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("jungle_synthesizer/openverse-media-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "sp_intended_usage": "Describe your intended use...",
    "sp_improvement_suggestions": "Share your suggestions here...",
    "sp_contact": "Share your email here...",
    "mediaTypes": ["images"],
    "searchQuery": "",
    "licenses": [],
    "imageCategories": [],
    "audioCategories": [],
    "sources": [],
    "aspectRatio": [],
    "imageSize": [],
    "maxItems": 10,
}

# Run the Actor and wait for it to finish
run = client.actor("jungle_synthesizer/openverse-media-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "sp_intended_usage": "Describe your intended use...",
  "sp_improvement_suggestions": "Share your suggestions here...",
  "sp_contact": "Share your email here...",
  "mediaTypes": [
    "images"
  ],
  "searchQuery": "",
  "licenses": [],
  "imageCategories": [],
  "audioCategories": [],
  "sources": [],
  "aspectRatio": [],
  "imageSize": [],
  "maxItems": 10
}' |
apify call jungle_synthesizer/openverse-media-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,jungle_synthesizer/openverse-media-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/Abt2owobQB4dU8WC5/builds/0CBH1k3xpYa4arpPc/openapi.json
