# Globo Esporte Scraper - Sports News, Videos & Matches (`abotapi/globo-ge`) Actor

Scrape public sports content from ge.globo.com, including news, videos, matches and feed records. Extract clean, structured data for sports coverage, content monitoring, match tracking and analysis.

- **URL**: https://apify.com/abotapi/globo-ge.md
- **Developed by:** [Abot API](https://apify.com/abotapi) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 1 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.50 / 1,000 ge sports records

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Globo GE Sports Scraper

Globo GE Sports Scraper collects public sports content from GE (ge.globo.com), including news, videos, matches, and standings. Search by keyword, or paste content links -- full URLs and bare section paths both work. Results share one dataset with a record type so you can separate news from match and table records.

### Why This Scraper?

- Search public GE content by one or more keywords.
- Collect the homepage or named sections.
- Read direct article, match, video, and public search URLs.
- Preserve source IDs, titles, summaries, sections, dates, images, and match metadata.
- Keep article cards and event records in one dataset with a stable recordType.
- Track recurring runs as NEW, UPDATED, UNCHANGED, REAPPEARED, and EXPIRED.

### Data You Get

| Field | Description |
| --- | --- |
| `id` | Stable GE or search-result identifier |
| `recordType` | article, video, match, or standing |
| `title` | Headline or content title |
| `summary` | Public description, summary, or video body |
| `section` | GE section or publisher |
| `url` | Canonical public content URL when supplied; some video search cards have no link |
| `imageUrl` | Source image when available |
| `thumbnailUrl` | Search thumbnail when available |
| `publishedAt` | Publication timestamp when available |
| `modifiedAt` | Source modification timestamp when available |
| `source` | Public source label |
| `fullText` | Public article text for direct article URLs or enriched article cards |
| `author` | Public article byline when available |
| `homeTeam` | Home team for match records |
| `awayTeam` | Away team for match records |
| `homeScore` | Home score when available |
| `awayScore` | Away score when available |
| `team` | Team name for standings records |
| `position` | Table position when available |
| `points` | Table points when available |
| `changeType` | Incremental classification |
| `changedFields` | Fields changed since the prior incremental run |
| `firstSeenAt` | First observation in the incremental state |
| `lastSeenAt` | Latest observation in the incremental state |

### How to Use

#### Search mode

```json
{
  "mode": "search",
  "queries": ["flamengo", "Libertadores"],
  "maxItems": 20,
  "maxPages": 2
}
```

#### Link mode (URLs or section paths)

```json
{
  "mode": "urls",
  "urls": ["/", "/futebol/", "https://ge.globo.com/futebol/libertadores/"],
  "fetchDetails": true,
  "maxItems": 10
}
```

A bare section path and a full section URL are the same page -- use whichever
is handier. (The old `mode: "sections"` and the `sections` field keep working
for saved tasks; both fold into this mode.)

#### Recurring monitoring

```json
{
  "mode": "search",
  "queries": ["Libertadores"],
  "incrementalMode": true,
  "stateKey": "libertadores-news",
  "emitUnchanged": false,
  "emitExpired": false,
  "maxItems": 50
}
```

### Input Parameters

| Parameter | Type | Default | Description |
| --- | --- | --- | --- |
| `mode` | string | search | Select search or urls. |
| `queries` | string\[] | \["flamengo"] | Keywords used in search mode. |
| `urls` | string\[] | \[] | Public GE links used in URL mode: full URLs or bare section paths. The legacy `sections` field still works and folds in. |
| `fetchDetails` | boolean | false | Add author and full text to article cards. Successfully enriched saved cards add a surcharge. Direct article URLs include available details without an extra charge. |
| `maxItems` | integer | 20 | Maximum rows to save. 0 means no item cap. |
| `maxPages` | integer | 0 | Maximum pages per source. 0 walks until the source ends or the item cap is reached. |
| `resumeFromRunId` | string | empty | Previous run or dataset ID whose saved IDs should be skipped. |
| `incrementalMode` | boolean | false | Enable cross-run change classification. |
| `stateKey` | string | empty | Optional name for the incremental campaign. |
| `emitUnchanged` | boolean | false | Return and bill extra unchanged rows in incremental mode. |
| `emitExpired` | boolean | false | Return and bill extra missing-record rows after a complete nonempty scan. |
| `mcpConnectors` | string\[] | \[] | Optional connector IDs for exporting the first results. |
| `notionParentPageUrl` | string | empty | Optional parent page for a connector export. |
| `maxNotifyListings` | integer | 50 | Maximum records offered to connector exports. |
| `proxyConfiguration` | object | Apify Proxy | Optional Apify Proxy configuration. |

#### Resume and recurring updates

`resumeFromRunId` skips records already saved in one previous run or dataset. `incrementalMode` stores a baseline for the same search and labels later observations. Use one of these modes at a time. Unchanged records are suppressed unless enabled. Missing records are marked EXPIRED only after a complete, nonempty scan with no item or page truncation. A smaller output cap saves only that many new records; later runs can collect the remaining unseen records.

### Send results into your apps (MCP connectors)

Connector export is optional. Add connector IDs when an approved connector is configured for the run. Use `notionParentPageUrl` when the selected connector needs a parent page, and `maxNotifyListings` to limit exported records. Connector failures leave the dataset available.

### Output Example

> Sample shape: values are illustrative placeholders, not from a live match.

```json
{
  "id": "event-0001",
  "recordType": "match",
  "title": "Sample Home Team x Sample Away Team",
  "summary": "Match card from the public GE sports feed.",
  "section": "sample-championship",
  "url": "https://example.com/match/0001",
  "publishedAt": "2026-01-01T00:00:00Z",
  "homeTeam": "Sample Home Team",
  "awayTeam": "Sample Away Team",
  "homeScore": 2,
  "awayScore": 0,
  "source": "ge.globo.com",
  "changeType": "NEW",
  "changedFields": [],
  "firstSeenAt": "2026-09-18T08:00:00+00:00",
  "lastSeenAt": "2026-09-18T08:00:00+00:00"
}
```

### Plan Requirement

An Apify account is required to run this actor. Available charges are shown on the actor's pricing page.

### 🔗 Want more sports data?

Pair this actor with these related scrapers from the same team:

<table>
<tr><td>⚽ <a href="https://apify.com/abotapi/onefootball-com-scraper"><b>OneFootball Scraper</b></a><br>Scrape public OneFootball data including football news, upcoming fixtures, match results...</td><td>⚽ <a href="https://apify.com/abotapi/sofascore-scraper"><b>SofaScore Scraper ⚽ Live Scores, Stats, Players &amp; Odds</b></a><br>From $1/1K. Pull structured sports data from SofaScore across football, basketball...</td></tr>
<tr><td>⚽ <a href="https://apify.com/abotapi/hotstar-com-scraper"><b>JioHotstar Scraper</b></a><br>Scrape the JioHotstar catalog by content type or URL. Extract shows, movies, episodes...</td><td>⚽ <a href="https://apify.com/abotapi/sportsbook-odds-scraper"><b>Sportsbook Odds Scraper (1xBet, Melbet, Linebet, Paripulse)</b></a><br>Collect live and prematch betting odds from 1xBet, Melbet, Linebet and Paripulse. Give it...</td></tr>
<tr><td>⚽ <a href="https://apify.com/abotapi/netshoes-scraper"><b>Netshoes Brazil</b></a><br>Scrape sportswear and footwear from netshoes.com.br. Search by keyword with the store's...</td><td>⚽ <a href="https://apify.com/abotapi/bookmyshow-scraper"><b>BookMyShow Scraper</b></a><br>Scrape BookMyShow movies and live events by city, category or URL. Extract genres...</td></tr>
</table>

👉 [Browse all abotapi scrapers](https://apify.com/abotapi)

### 💬 Support & custom scrapers

- 🐞 **Found a bug or a missing field?** Open a ticket on the [Issues tab](https://apify.com/abotapi/globo-ge/issues/open). We usually reply within hours.
- 🛠️ **Need another site, extra fields or a private build?** Email <abotapi@proton.me> or message [Telegram @abotapi](https://t.me/abotapi).
- ⭐ **Enjoying it?** A quick review on the actor page helps other users find it.

# Actor input Schema

## `mode` (type: `string`):

Search by keywords, or paste public GE links (full URLs or bare section paths).

## `queries` (type: `array`):

One or more public GE search terms.

## `urls` (type: `array`):

Public ge.globo.com links -- full URLs (homepage, section, search, article, match, or video) and bare section paths such as / or /futebol/brasileirao-serie-a/ both work. Used only in URL mode.

## `fetchDetails` (type: `boolean`):

Add author and full text to article cards when available. Successfully enriched saved cards add a surcharge. Direct article URLs include available details without an extra charge.

## `maxItems` (type: `integer`):

Maximum rows to save. Set 0 for unlimited within the source page budget.

## `maxPages` (type: `integer`):

Optional page bound per source. Leave at 0 to walk until the public source ends or Maximum records is reached.

## `mcpConnectors` (type: `array`):

Optional connector IDs for a best-effort side-channel export. The full dataset remains unchanged.

## `notionParentPageUrl` (type: `string`):

Optional parent page ID or URL for Notion export.

## `maxNotifyListings` (type: `integer`):

Maximum records sent to each connector. Does not affect the dataset.

## `resumeFromRunId` (type: `string`):

Paste a previous run or dataset ID. Already-saved record IDs are skipped.

## `incrementalMode` (type: `boolean`):

Classify later observations as NEW, UPDATED, UNCHANGED, REAPPEARED, or EXPIRED.

## `stateKey` (type: `string`):

Optional name for the recurring monitoring campaign. Leave empty for an automatic scope hash.

## `emitUnchanged` (type: `boolean`):

Incremental mode only. Return and bill extra rows for records that did not change.

## `emitExpired` (type: `boolean`):

Incremental mode only. Return and bill extra rows for missing records after a complete uncapped nonempty scan.

## `proxyConfiguration` (type: `object`):

Direct public requests are normally sufficient. Apify Proxy can be enabled when needed.

## Actor input object example

```json
{
  "mode": "search",
  "queries": [
    "flamengo"
  ],
  "urls": [
    "https://ge.globo.com/"
  ],
  "fetchDetails": false,
  "maxItems": 20,
  "maxPages": 0,
  "mcpConnectors": [],
  "maxNotifyListings": 50,
  "incrementalMode": false,
  "emitUnchanged": false,
  "emitExpired": false,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `overview` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "mode": "search",
    "queries": [
        "flamengo"
    ],
    "urls": [
        "https://ge.globo.com/"
    ],
    "fetchDetails": false,
    "maxItems": 20,
    "maxPages": 0,
    "incrementalMode": false,
    "emitUnchanged": false,
    "emitExpired": false,
    "proxyConfiguration": {
        "useApifyProxy": true
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("abotapi/globo-ge").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "mode": "search",
    "queries": ["flamengo"],
    "urls": ["https://ge.globo.com/"],
    "fetchDetails": False,
    "maxItems": 20,
    "maxPages": 0,
    "incrementalMode": False,
    "emitUnchanged": False,
    "emitExpired": False,
    "proxyConfiguration": { "useApifyProxy": True },
}

# Run the Actor and wait for it to finish
run = client.actor("abotapi/globo-ge").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "mode": "search",
  "queries": [
    "flamengo"
  ],
  "urls": [
    "https://ge.globo.com/"
  ],
  "fetchDetails": false,
  "maxItems": 20,
  "maxPages": 0,
  "incrementalMode": false,
  "emitUnchanged": false,
  "emitExpired": false,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}' |
apify call abotapi/globo-ge --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,abotapi/globo-ge"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/ZhCsfWs7zOXM5qJlu/builds/lB75Kl3d88OxPIhZH/openapi.json
