# Coles Scraper (`mlg14/coles-scraper`) Actor

Scrape public Coles Australian grocery products, prices, specials, and categories. Export structured results, source URLs, IDs, and available details to CSV, JSON, or Excel for Australian grocery assortment and price research.

- **URL**: https://apify.com/mlg14/coles-scraper.md
- **Developed by:** [MLG Data](https://apify.com/mlg14) (community)
- **Categories:** E-commerce, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$1.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Coles Scraper

Scrape Coles Australian grocery products, prices, specials, and categories from public pages and export Coles data to CSV, JSON, or Excel. This Coles API alternative turns repeatable inputs into structured results for Australian grocery assortment and price research.

Each dataset row represents a public result or a documented alternate record type. The output includes source identifiers and links where available, so you can verify a row, compare later runs, and distinguish missing source data from an omitted column.

### What data can you extract from Coles?

The dataset schema defines every field below. Examples come from one successful published run; `null` means that field was not exposed for that particular record. Availability may differ by input mode, page, and record type.

| Field | Description | Example |
| --- | --- | --- |
| `recordType` | Product or category row. | `product` |
| `operation` | Operation that produced the row. | `category` |
| `query` | Search keyword. | `null` |
| `page` | Native Coles page number. | `1` |
| `productId` | Coles product identifier. | `8150288` |
| `productUrl` | Canonical product page URL. | `https://www.coles.com.au/product/coles-full-cream-milk-3l-8150288` |
| `name` | Product name. | `Full Cream Milk` |
| `brand` | Brand. | `Coles` |
| `description` | Short product description. | `COLES FULL CREAM MILK 3L` |
| `longDescription` | Detailed product description. | `null` |
| `size` | Pack size. | `3L` |
| `imageUrl` | Product image URL. | `https://cdn.productimages.coles.com.au/productimages/8/8150288.jpg` |
| `price` | Regular price in Australian dollars; current price if no reduction. | `4.95` |
| `currentPrice` | Current price in Australian dollars. | `4.95` |
| `discountPrice` | Current reduced price, if a was price is shown. | `null` |
| `wasPrice` | Previous price when reduced. | `null` |
| `currency` | ISO currency code. | `AUD` |
| `unitPrice` | Comparable unit price. | `1.65` |
| `unitMeasure` | Unit for comparable pricing. | `l` |
| `unitPriceText` | Displayed comparable price label. | `$1.65/ 1L` |
| `promotionType` | Promotion classification. | `null` |
| `specialType` | Special classification. | `null` |
| `offerDescription` | Displayed offer text. | `null` |
| `savePercent` | Discount percentage when supplied. | `null` |
| `isOnlineSpecial` | Whether the offer is online special. | `false` |
| `available` | Availability in the anonymous shopping context. | `true` |
| `availabilityType` | Online or store availability type. | `InStoreAndOnline` |
| `availableQuantity` | Reported available quantity. | `2440` |
| `retailLimit` | Maximum retail quantity. | `20` |
| `category` | Product category. | `Milk` |
| `subCategory` | Top level department. | `Dairy, Eggs & Fridge` |
| `aisle` | Product aisle. | `Full Cream Milk` |
| `categoryId` | Coles category identifier. | `8881800` |
| `gtin` | Global trade item number from detail page. | `null` |
| `ingredients` | Ingredient text from detail page. | `null` |
| `allergens` | Allergen text from detail page. | `null` |
| `servingSize` | Nutrition serving size. | `null` |
| `energyKjPer100` | Energy per 100 g or ml. | `null` |
| `proteinPer100` | Protein per 100 g or ml. | `null` |
| `fatPer100` | Total fat per 100 g or ml. | `null` |
| `sugarsPer100` | Total sugar per 100 g or ml. | `null` |
| `sodiumPer100` | Sodium per 100 g or ml. | `null` |
| `countryOfOrigin` | Product origin statement. | `null` |
| `lastUpdated` | Source product update timestamp. | `null` |
| `categoryUrl` | Category browse URL. | `null` |
| `categorySlug` | Category path slug. | `null` |
| `categoryName` | Category name derived from path. | `null` |

Use the identifier and URL fields as your join keys before comparing snapshots. Fields that describe a page, search, category, author, or tournament establish where the record came from; keep them when you export a subset. Numeric values and boolean flags reflect the public page at collection time, not a permanent claim about the underlying item.

### How to scrape Coles

1. Open the actor input form and choose a narrow public source: a search phrase, page URL, item URL, or ID supported by the input fields below.
2. Set the relevant per-source page or result limit and the total item limit. Start with a small sample to check which optional fields the public source exposes.
3. Run the actor. Its dataset contains one structured row per saved result; inspect the first rows and their source links for the chosen input.
4. Download the dataset as JSON, CSV, or Excel. Preserve identifiers when combining runs so repeated results can be deduplicated.

For this actor, choose search, category, specials, product, or category-discovery operation according to the input schema. Product records and category records have different populated fields; inspect recordType before comparing prices.

### Input

Use only the parameters relevant to your collection mode. An omitted optional filter uses the schema default shown here; an empty array, zero, and an omitted value can have different meanings, so keep intentional settings in your saved input.

| Parameter | Type | Default | Description |
| --- | --- | --- |
| `operation` | `string` | `search` | Choose product search, specials, a category shelf, individual items, or category discovery. |
| `query` | `string` | No default; example `milk` | Product keyword for search. |
| `special` | `string` | `all` | All specials or half-price offers when operation is specials. |
| `categoryUrl` | `string` | No default; example `https://www.coles.com.au/browse/dairy-eggs-fridge/milk` | Full Coles browse URL for category operation. |
| `productUrls` | `array` | No default; example `["https://www.coles.com.au/product/coles-full-cream-milk-3l-8150288"]` | Full Coles product URLs for item operation. |
| `startPage` | `integer` | `1` | First native Coles page for search, specials, or category. |
| `maxPages` | `integer` | `1` | Maximum native pages to fetch from the start page. |
| `maxItems` | `integer` | `100` | Stop after this many dataset rows; 0 means no cap. |
| `details` | `boolean` | `false` | Fetch each listing product page for barcode, ingredients, allergens, and nutrition. |
| `proxyConfiguration` | `object` | `{"useApifyProxy":true}` | Proxy settings for remote access; the actor escalates to an Australian residential proxy when blocked. |

Example input based on the published golden run (long URL lists are shortened):

```json
{
  "operation": "search",
  "query": "milk",
  "startPage": 1,
  "maxPages": 1,
  "maxItems": 50,
  "details": false
}
```

The example is a starting shape, not a guarantee of a particular result count. Source inventory and page accessibility change. When you need repeatable comparisons, save the exact input JSON with the run date and inspect the returned source or record-type field.

### Output example

The following is one real item from a successful published dataset. Long text and media arrays are shortened for readability; the actual dataset keeps the original values and all schema fields.

```json
{
  "recordType": "product",
  "operation": "category",
  "page": 1,
  "productId": "8150288",
  "productUrl": "https://www.coles.com.au/product/coles-full-cream-milk-3l-8150288",
  "name": "Full Cream Milk",
  "brand": "Coles",
  "description": "COLES FULL CREAM MILK 3L",
  "size": "3L",
  "imageUrl": "https://cdn.productimages.coles.com.au/productimages/8/8150288.jpg",
  "price": 4.95,
  "currentPrice": 4.95,
  "currency": "AUD",
  "unitPrice": 1.65,
  "unitMeasure": "l",
  "unitPriceText": "$1.65/ 1L",
  "isOnlineSpecial": false,
  "available": true,
  "availabilityType": "InStoreAndOnline",
  "availableQuantity": 2440,
  "retailLimit": 20,
  "category": "Milk",
  "subCategory": "Dairy, Eggs & Fridge",
  "aisle": "Full Cream Milk",
  "categoryId": "8881800"
}
```

This row illustrates the observed output structure, including its identifiers and public links. Empty values elsewhere in the dataset should be interpreted field by field; a field shown in this example is not promised for every result.

### Use cases

- Grocery analysts can compare product and unit prices across recurring category snapshots.
- Brands can monitor public assortment, promotion labels, and shelf visibility.
- Nutrition researchers can compile available ingredients, allergens, and per-100 values for selected products.
- Retail operations can reconcile product IDs, barcodes, and availability signals.

The strongest analyses keep source context. A field such as price, rating, engagement, or rank has meaning only with its associated item, query, date, and public URL. Keep the raw export and create a separate cleaned view for charts or alerts.

### How much does it cost to scrape Coles?

The price is **$1.00 per 1,000 saved results**. Platform usage is included. The charge scales with output rows, so a restrictive filter or inaccessible page can produce fewer billable results than the requested maximum.

- 100 results: **$0.10**. This is useful for checking a small cohort and confirming which optional fields are present.
- 1,000 results: **$1.00**. This is the reference price for a larger export.
- 5,000 results: **$5.00**. Reaching this size may require multiple focused sources or scheduled runs, depending on public inventory and source caps.

Compute any other estimate as saved result count × $1.00 / 1,000. A maximum input is a ceiling, not a purchase of that many rows. For planning, use the actual saved-item count from an initial representative run.

### Tips for best results

For comparable prices, keep a consistent operation and query across runs. Use productUrls for known items, and details when ingredients or nutrition matter. Compare unitPrice with unitMeasure rather than raw package price alone.

Collect a small baseline first, record the exact input and date, and inspect both a typical row and a sparse row. Expand by adding focused sources instead of assuming one broad input can reveal the full public inventory. When comparing two runs, match stable IDs or canonical URLs and use the same filters so changes reflect the source rather than a changed query.

Export the full JSON when nested arrays or objects matter. CSV and Excel are convenient for sorting and joins, but nested structures may need flattening before spreadsheet analysis. Keep numeric fields numeric and preserve source URLs as text; do not infer a zero from a null.

### Limits

Retail prices and availability can depend on storefront context and change after collection. Nutrition, allergens, offers, and origin are only populated where the public product data exposes them.

Public pages can change, disappear, or expose different fields for different records. The schema is a list of possible output columns, not a promise that each column is filled in each row. The `maxItems`-style input limits cap saved results; they do not bypass source pagination, public visibility, or a site-specific result ceiling.

Treat a saved result as a snapshot. If a later run returns fewer rows, first compare the input, source access, and public inventory before concluding that the underlying market changed. If you need an audit trail, retain the source URL, stable identifier, and run timestamp with the export.

### Use with AI agents (MCP)

An agent can supply the documented input JSON, run this actor, and work from its dataset. Ask it to keep source links and identify null values explicitly when summarizing results. Example prompts:

> Run Coles Scraper for the sample input above. Return the first 20 results with their source URLs and the fields needed for Australian grocery assortment and price research. Mark unavailable fields as null.

> Schedule a repeat Coles collection with the same filters. Compare records by stable ID or canonical URL and report only new, removed, or changed public values.

### FAQ

#### Is it legal to scrape Coles?

This actor collects public data. Check the site’s terms, applicable law, and the rights of people whose data appears in your export. Follow GDPR and other privacy rules where relevant; do not use personal data for misuse, intrusive profiling, or unauthorized contact.

#### Do I need to configure proxies?

The input includes proxyConfiguration and defaults to Apify Proxy. Usually the default is enough to start. Availability can vary by page and region; changing network settings cannot make private or login-only information public.

#### How fast will a run finish?

Duration depends on the number of inputs, pages, detail requests, and source responses. A small sample is the best way to measure your workload. Increase the limits gradually, then use the observed run duration for scheduling and monitoring.

#### Can I schedule and monitor recurring runs?

Yes. Save the input and schedule recurring runs in Apify. Monitor run status, item count, and any missing-field changes; store the run date alongside exports when comparing snapshots.

#### Can I export to Google Sheets or Excel?

Yes. Download CSV or Excel from the dataset, or pass JSON to a spreadsheet integration. For nested fields, flatten the specific child values you need rather than losing the original JSON.

#### What if a field is empty?

An empty or null field means it was not available from that public source record in that run. Check the source URL, input mode, and record type. Do not replace missing prices, counts, dates, or flags with zero unless your own analysis has a documented rule for doing so.

#### Are search and category rows identical?

No. recordType and operation identify the source. Category discovery rows describe taxonomy, while product rows carry prices and product attributes.

### Integrations

Use the Apify API to start runs and retrieve the dataset, webhooks to react when a run finishes, or Zapier, Make, and n8n to route records into other systems. Google Sheets supports lightweight review, while scheduled runs provide repeat snapshots. Keep the raw JSON when downstream workflows need nested fields or exact null values.

### Support

Open an issue on the Issues tab; we reply within 24h and add fields on request.

# Actor input Schema

## `operation` (type: `string`):

Choose product search, specials, a category shelf, individual items, or category discovery.

## `query` (type: `string`):

Product keyword for search.

## `special` (type: `string`):

All specials or half-price offers when operation is specials.

## `categoryUrl` (type: `string`):

Full Coles browse URL for category operation.

## `productUrls` (type: `array`):

Full Coles product URLs for item operation.

## `startPage` (type: `integer`):

First native Coles page for search, specials, or category.

## `maxPages` (type: `integer`):

Maximum native pages to fetch from the start page.

## `maxItems` (type: `integer`):

Stop after this many dataset rows; 0 means no cap.

## `details` (type: `boolean`):

Fetch each listing product page for barcode, ingredients, allergens, and nutrition.

## `proxyConfiguration` (type: `object`):

Proxy settings for remote access; the actor escalates to an Australian residential proxy when blocked.

## Actor input object example

```json
{
  "operation": "search",
  "query": "milk",
  "special": "all",
  "categoryUrl": "https://www.coles.com.au/browse/dairy-eggs-fridge/milk",
  "productUrls": [
    "https://www.coles.com.au/product/coles-full-cream-milk-3l-8150288"
  ],
  "startPage": 1,
  "maxPages": 1,
  "maxItems": 100,
  "details": false,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `results` (type: `string`):

Product and category rows in the default dataset.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "query": "milk",
    "categoryUrl": "https://www.coles.com.au/browse/dairy-eggs-fridge/milk",
    "productUrls": [
        "https://www.coles.com.au/product/coles-full-cream-milk-3l-8150288"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("mlg14/coles-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "query": "milk",
    "categoryUrl": "https://www.coles.com.au/browse/dairy-eggs-fridge/milk",
    "productUrls": ["https://www.coles.com.au/product/coles-full-cream-milk-3l-8150288"],
}

# Run the Actor and wait for it to finish
run = client.actor("mlg14/coles-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "query": "milk",
  "categoryUrl": "https://www.coles.com.au/browse/dairy-eggs-fridge/milk",
  "productUrls": [
    "https://www.coles.com.au/product/coles-full-cream-milk-3l-8150288"
  ]
}' |
apify call mlg14/coles-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,mlg14/coles-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/vUXTzZhut7zrcEmpP/builds/KrifGTwa4JEzA7NWp/openapi.json
