Indonesian News Scraper
Pricing
Pay per event
Indonesian News Scraper
Search Kompas and Detik by topic and export normalized public Indonesian news for recurring media monitoring.
Pricing
Pay per event
Rating
0.0
(0)
Developer
Stas Persiianenko
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
3 days ago
Last modified
Categories
Share
Collect public Indonesian news from Kompas and Detik in one normalized feed. Search several topics, select either publisher or both, and export headlines, summaries, article links, publication times, sections, images, and collection provenance for media monitoring.
The Actor uses the publishers' public topic-search pages. It does not require a publisher account, browser automation, or a residential proxy.
What Indonesian News Scraper does
Indonesian News Scraper turns two different publisher result formats into one stable dataset.
For every topic, it:
- searches the selected Kompas and Detik public surfaces;
- paginates until the requested limit is reached or results end;
- alternates records across selected publishers for balanced coverage;
- removes duplicate canonical article URLs;
- normalizes the remaining records into one output schema;
- stores the results in the run's default Apify dataset.
Results are shared fairly across multiple topics, so a large first topic does not consume the entire result limit before later topics run.
Who is it for
- PR and communications teams tracking coverage of a company, executive, campaign, or issue.
- Media-monitoring analysts comparing Kompas and Detik reporting.
- Market researchers collecting Indonesian economic, technology, policy, or consumer-news signals.
- Newsrooms and academics building a repeatable article-discovery dataset.
- Data teams sending current public news metadata to a spreadsheet, warehouse, or alerting pipeline.
This Actor is for article discovery and metadata monitoring. It does not download full article bodies or generate sentiment, emotion, or factuality scores.
Why use this Indonesian news feed
- One schema across Kompas and Detik.
- Topic and publisher controls instead of an unfocused site crawl.
- Balanced multi-source and multi-topic collection.
- Stable canonical URLs for deduplication between scheduled runs.
- Public metadata only; no login or account handling.
- Lightweight direct HTTP extraction with 256 MB of memory.
- Per-result charging: rejected, duplicate, and empty records are not item-charged.
Data you can extract
| Field | Meaning |
|---|---|
title | Article headline shown by the publisher |
publisher | Normalized Kompas or Detik name |
section | Publisher section or channel, when available |
articleUrl | Canonical public article URL |
publishedAt | Best-effort ISO 8601 publication date/time |
publishedText | Date/time text exactly as exposed by the source |
summary | Public search-result excerpt, when available |
imageUrl | Search-result thumbnail URL, when available |
query | Input topic that produced the result |
rank | One-based accepted-result rank for that topic |
page | One-based publisher result page |
sourceUrl | Search page used to collect the record |
fetchedAt | ISO 8601 collection timestamp |
Nullable fields remain null when a publisher does not expose them. The Actor does not invent missing summaries or exact publication times.
Get started
- Open the Actor in Apify Console.
- Add one or more Indonesian topics under News topics.
- Keep both publishers selected, or choose only Kompas or Detik.
- Set Maximum articles for the whole run.
- Click Start.
- Open the News feed dataset view to inspect or export the records.
A useful first run is:
{"queries": ["kecerdasan buatan"],"publishers": ["kompas", "detik"],"maxItems": 20}
Input parameters
queries
Required array of 1–20 non-empty topic strings. Each topic can contain up to 150 characters.
Use concrete Indonesian terms such as:
kecerdasan buatanekonomi Indonesiaenergi terbarukankebijakan pemerintah
Multiple topics receive a fair share of maxItems. If an earlier topic has too few results, later topics can use the remaining capacity.
publishers
Required source selection with one or both values:
kompasdetik
The default includes both. Selecting one publisher is useful for source-specific editorial monitoring.
maxItems
Maximum number of unique records saved across all topics and publishers.
- minimum:
1 - default:
100 - maximum:
1000
A lower value makes smoke tests faster and cheaper. A higher value may require more publisher result pages.
Output example
This record shape comes from the current public Kompas search behavior; live headlines change over time:
{"title": "Mengurai Kemacetan Rantai Pasok Sumatera lewat Kecerdasan Buatan","publisher": "Kompas","section": "Properti","articleUrl": "https://www.kompas.com/properti/read/2026/07/31/082820121/mengurai-kemacetan-rantai-pasok-sumatera-lewat-kecerdasan-buatan","publishedAt": "2026-07-31T00:00:00.000Z","publishedText": "31 Juli 2026","summary": "KIM mengelola dua kawasan industri utama yang menampung ratusan aktivitas manufaktur dan logistik skala regional maupun internasional.","imageUrl": "https://asset.kompas.com/example-image.jpg","query": "kecerdasan buatan","rank": 1,"page": 1,"sourceUrl": "https://search.kompas.com/search/?q=kecerdasan+buatan&page=1","fetchedAt": "2026-08-14T14:50:00.000Z"}
The example image URL is shortened for documentation. Dataset records retain the source image URL.
How much does it cost to collect Indonesian news?
Pricing combines one small Actor-start charge with one item event for each accepted dataset record. Duplicate URLs, malformed cards, and valid no-result searches do not produce item charges.
The one-time start price is $0.0005. Item prices decrease across the six Apify tiers:
| Tier | Price per accepted article |
|---|---|
| FREE | $0.0007636 |
| BRONZE | $0.000664 |
| SILVER | $0.00051792 |
| GOLD | $0.0003984 |
| PLATINUM | $0.0002656 |
| DIAMOND | $0.00018592 |
For planning, multiply your tier's item price by the requested record ceiling, then add the single start charge. At the BRONZE price:
| Accepted records | Example total |
|---|---|
| 10 | $0.00714 |
| 100 | $0.06690 |
| 500 | $0.33250 |
A run may cost less than that ceiling if the selected topics have fewer available results. These examples do not promise that every query returns the requested maximum.
Media-monitoring workflows
Monitor a brand or organization
Schedule a daily run with the organization name and both publishers. Use articleUrl as the stable key when comparing today's dataset with yesterday's.
Compare publisher coverage
Search one policy or market topic across both publishers. Group the output by publisher, then compare headline language, section placement, and publication timing.
Build a multi-topic policy feed
Submit several related policy terms in one run. The Actor allocates result capacity across topics rather than allowing the first topic to consume all output.
Send results to a spreadsheet or warehouse
Connect the default dataset to Make, Zapier, Google Sheets, or an Apify webhook. Store query, publisher, publishedAt, and fetchedAt alongside each canonical URL for auditable monitoring.
Scheduling and change detection
Apify schedules can run this Actor hourly, daily, or weekly. For change detection:
- schedule the same input;
- export each run's default dataset;
- compare records by
articleUrl; - treat unseen URLs as newly discovered coverage;
- use
fetchedAtas collection provenance, not as publication time.
The Actor itself does not persist a cross-run alert state or send notifications. Use an integration, webhook, or downstream dataset comparison for alerts.
API usage with cURL
Start a run and wait for its dataset:
curl -X POST \"https://api.apify.com/v2/acts/automation-lab~indonesian-news-aggregation-feed/run-sync-get-dataset-items?token=$APIFY_TOKEN" \-H "Content-Type: application/json" \-d '{"queries": ["ekonomi Indonesia"],"publishers": ["kompas", "detik"],"maxItems": 20}'
Keep your Apify token in an environment variable. Do not commit it to source control.
API usage with JavaScript
import { ApifyClient } from 'apify-client';const client = new ApifyClient({ token: process.env.APIFY_TOKEN });const run = await client.actor('automation-lab/indonesian-news-aggregation-feed').call({queries: ['energi terbarukan'],publishers: ['kompas', 'detik'],maxItems: 50,});const { items } = await client.dataset(run.defaultDatasetId).listItems();console.log(items);
API usage with Python
import osfrom apify_client import ApifyClientclient = ApifyClient(os.environ['APIFY_TOKEN'])run = client.actor('automation-lab/indonesian-news-aggregation-feed').call(run_input={'queries': ['kebijakan pemerintah'],'publishers': ['kompas', 'detik'],'maxItems': 50,})items = client.dataset(run['defaultDatasetId']).list_items().itemsprint(items)
Use with Apify MCP
Add this Actor to Claude Code:
claude mcp add --transport http apify \"https://mcp.apify.com?tools=automation-lab/indonesian-news-aggregation-feed"
Claude Desktop, Cursor, and VS Code setup
Use the equivalent HTTP MCP configuration in Claude Desktop, Cursor, or VS Code:
{"mcpServers": {"apify": {"url": "https://mcp.apify.com?tools=automation-lab/indonesian-news-aggregation-feed"}}}
Example prompts:
- “Run Indonesian News Scraper for
kecerdasan buatanacross Kompas and Detik, with at most 30 articles.” - “Collect 15 Detik records about
ekonomi Indonesiaand summarize the recurring themes.” - “Build a two-topic feed for
kebijakan pemerintahandenergi terbarukan, then group links by publisher.”
Limits and freshness
- Coverage is limited to public Kompas and Detik topic-search results.
- Search ranking, retention, and available pagination are controlled by each publisher.
- Kompas may expose a date without an exact time; such dates are normalized to midnight UTC and the original text is preserved.
- Detik usually exposes a machine-readable timestamp, but fields can still be absent.
summaryis a search excerpt, not the full article body.- The Actor stops at 1000 accepted records per run.
- Public result pages can change. An unrecognized page shape fails visibly instead of silently returning a misleading empty dataset.
- The Actor retries transient network, HTTP 429, and server failures up to a bounded limit. It does not blindly retry deterministic client errors.
Responsible use and legality
This Actor collects publicly displayed news metadata. Users remain responsible for complying with publisher terms, copyright rules, robots guidance, privacy law, and the requirements of their jurisdiction.
Do not republish copyrighted article text or images without permission. Prefer linking to the canonical article, retaining publisher attribution, and using excerpts only for legitimate monitoring, research, or analysis. Avoid collecting or using personal data in ways that create harm.
Troubleshooting
The dataset is empty
Verify the spelling and specificity of the topic. A valid topic can naturally have no current indexed results. Try a broader Indonesian phrase and confirm at least one publisher is selected.
The run reports an unrecognized result page
A publisher may have changed its HTML or returned a challenge page. Retry later once. If the same error persists, keep the run ID and logs when reporting the problem; repeated identical retries will not repair a changed parser.
Fewer records were returned than maxItems
maxItems is a ceiling, not padding. The selected publishers may expose fewer unique matching URLs, and duplicates are removed.
Why are some publication times midnight UTC?
Kompas search results sometimes expose only an Indonesian calendar date. publishedAt uses midnight UTC as a machine-readable date while publishedText preserves the exact displayed value.
Does the Actor use a proxy?
No proxy is enabled by default because both selected public search routes work over direct HTTP. This keeps runs lightweight and avoids unmeasured proxy costs.
Related Automation Lab actors
- Naver News Search Scraper for Korean news discovery by keyword.
- Brave News Search Scraper for broader web news-search workflows.
- Seeking Alpha Analysis Feed Scraper for market-analysis feed collection.
Use those Actors when the buyer job or source differs. They are not silent fallbacks for Kompas or Detik.
FAQ
Does it scrape full articles?
No. It returns public search metadata and available excerpts, not complete article bodies.
Can I select only Kompas or Detik?
Yes. Set publishers to ['kompas'] or ['detik'].
Can I monitor more than one topic?
Yes. Supply up to 20 topics. The Actor distributes available output capacity across them.
Are results translated into English?
No. Headlines and summaries remain in the language exposed by the publisher.
Does it perform sentiment analysis?
No. The normalized records are suitable input for your own sentiment, clustering, or language-model workflow, but this Actor does not claim those enrichments.
How should I deduplicate scheduled runs?
Use articleUrl as the primary key. Keep query and fetchedAt to retain matching and collection provenance.