Google News and Publisher Feed Scraper avatar

Google News and Publisher Feed Scraper

Pricing

from $0.76 / 1,000 article delivereds

Go to Apify Store
Google News and Publisher Feed Scraper

Google News and Publisher Feed Scraper

Google News plus RSS, Atom and JSON feeds, with word filters and recurring monitoring. Base rate: $0.95 per 1,000 articles, usage included. Publisher-link lookup is best effort. Maintained by CleanScrape.

Pricing

from $0.76 / 1,000 article delivereds

Rating

0.0

(0)

Developer

CleanScrape

CleanScrape

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

4 days ago

Last modified

Categories

Share

Collect Google News searches, topic headlines and public RSS, Atom or JSON feeds in one consistent article table. Choose a country and time period from the form, filter by words or publisher domains, and export the results as JSON, CSV or Excel. The price is $0.95 per 1,000 delivered articles, with platform usage included and no startup fee.

Use it for company and industry research, recurring news monitoring or collecting inputs for your own reporting workflow. You do not need a Google login, publisher account or third-party API key. Publisher-link lookup is included on a best-effort basis: when a direct link cannot be resolved, the article keeps its original Google News URL and a clear link status.

Start with a topic

  1. In Search topics, enter a topic or company name, such as electric vehicles.
  2. Choose Google News edition and Published within using the dropdowns. Country codes and manually formatted dates are not needed in the form.
  3. Leave Maximum articles in total at 20 for your first run. Keep Look up publisher article links enabled if you want direct links where available.
  4. Start the Actor. When it finishes, open the output dataset and select Articles to review headlines, publishers, dates and links.
  5. Use Apify's export controls to download JSON, CSV or Excel. Open Run report in the run's output and select RUN_REPORT to check source errors, filtering and coverage limits.

For API users, the equivalent input is:

{
"queries": ["electric vehicles"],
"edition": "US:en",
"timeRange": "week",
"maxArticles": 20,
"resolveUrls": true
}

The total limit is an upper bound, not a promised article count. Start small, inspect the report, then adjust the inputs to suit your workflow.

Choose your sources

SourceHow to use it
Search topicsEnter one Google News search topic or expression per line.
News sectionsSelect sections such as Business, Technology or Science. These are separate feeds, not refinements of your search topics.
Include top headlinesAdd the selected edition's top-headlines feed.
Publisher feed URLsSupply public RSS, Atom or JSON Feed URLs from publishers you want to follow. Use feed URLs, not ordinary article pages or homepages.

Supported Google News editions are United States (English), United Kingdom (English), Germany (German), Finland (Finnish) and France (French). The edition does not translate your query or guarantee that every result has that language or comes from that country.

You can combine sources in one run. For publisher feeds only, clear the prefilled search topic and add your feed URLs. Google search terms do not filter publisher feeds, section feeds or top headlines; use the word filters below to filter every source.

RSS, Atom and JSON Feed are supported feed formats, not three additional news indexes. The Actor reads the public feed entries supplied by each source. It does not crawl a publisher's complete website or bypass paywalls.

Keep the articles that matter

Use Must contain these words or phrases and Must not contain these words or phrases to filter results without writing search operators. These filters apply to every selected source before publisher-link lookup and delivery.

For example, collect news about electric vehicles, keep articles mentioning recall or battery fire, and exclude opinion:

{
"queries": ["electric vehicles"],
"includeTerms": ["recall", "battery fire"],
"excludeTerms": ["opinion"],
"includeMode": "any",
"filterScope": "headline",
"timeRange": "week",
"maxArticles": 20
}

Choose At least one word or phrase or Every word or phrase for included terms. Exclusions always win. Match against the headline alone or the headline and available feed summary. This does not search full article bodies.

Matching ignores case and normalizes Unicode width and whitespace. Terms are literal whole words or phrases, not regular expressions: AI does not match retail. There is no stemming or translation; accents and punctuation remain significant. A phrase must occur within one field. Languages without word separators may need exact longer phrases.

To narrow publishers, enter domains such as bbc.com in Only these publisher domains or Exclude publisher domains. These filters apply to observed entries; they do not make Google discover every article from a publisher.

Filtered-out entries do not use your output allowance or generate article charges. A filter can legitimately return no results within the available feed entries. Loosen the filter or widen the time period before assuming the source is broken.

Dates without formatting guesswork

Choose Past 24 hours, Past 7 days or Past 30 days for a relative period. For a fixed period, select both custom dates using the calendar controls. Custom dates override the preset, use UTC, include the end date and allow a maximum span of 31 days.

Dates are checked against the source's publication timestamp, not the time the Actor collected the article. Entries without a usable timezone-aware publication date are excluded and counted in the report.

Search smaller date windows makes additional Google search requests to improve observed coverage. It does not guarantee complete historical coverage or remove Google's per-feed limits.

What the output contains

The Articles view puts the fields needed for a first review together. Use Sources and link diagnostics when you need provenance or want to understand a link lookup issue.

FieldMeaning
titleHeadline supplied by the source.
publisher, publisherUrlPublisher name and homepage where supplied. Missing values remain null.
publishedAt, publishedRawPublication time normalized to UTC and the source's original date value.
summaryAvailable feed snippet, not a full article or an AI-written summary. It may repeat the headline.
urlAccepted publisher link where resolved; otherwise the original source link. Check urlStatus.
originalUrlOriginal feed link, retained for traceability.
urlStatus, resolutionErrorLink outcome and available diagnostic information.
articleId, sourceKeyIdentifiers used for article identity and source tracking.
collectedAtWhen the Actor observed the article, distinct from publication time.
sourcesSource provenance for the article, including duplicates merged before the output cap.

Missing fields are not invented. A null publisher or missing snippet can be legitimate feed data, not necessarily an extraction error. The Actor does not extract full article text, sentiment or images.

Deduplication uses normalized article URLs and monitoring identities, not similar headlines. Different publishers covering the same story remain separate articles. Rows follow source/feed order, not a globally sorted relevance or recency ranking.

urlStatusWhat it means
decodedA publisher link was returned by Google's lookup and accepted by the checks.
decoded_cachedA previously accepted link was reused from cache.
feed_linkThe publisher feed supplied the link directly; the destination page was not independently verified.
unresolvedLookup failed or was paused. The original Google News link is retained.
not_requestedPublisher-link lookup was switched off.

Lookup checks publisher hostnames where available and rejects homepages and known Google interstitials. It does not confirm every destination page's contents, continued availability or accessibility. Accepted mappings can be cached for seven days.

Pricing and spending limits

Store discounts

Apify applies your subscription tier automatically: Bronze 10% off, Silver 15% off and Gold 20% off the base event prices. Platinum and Diamond use the Gold rates. These discounts apply to every billable event listed below.

Billable unitBase / Free tierBronze (-10%)Silver (-15%)Gold (-20%)
Article delivered (1,000 events)$0.95$0.855$0.8075$0.76

Base-price examples in this README are before discounts. Discounts do not change what is billable or any startup memory multiplier. See the Pricing tab for the rate applicable to your account.

All prices and examples below use the base rate before subscription discounts.

The price is $0.95 per 1,000 delivered articles ($0.00095 each). For example, 20 articles cost $0.019 and 100 cost $0.095. Platform usage and managed datacenter connections are included. There is no startup fee or separate word-filter or link-lookup fee.

An article is billable once saved after filtering and deduplication, including an article whose link remains an unresolved Google News URL. Filtered-out entries, skipped duplicates, previously delivered watch entries and run reports do not generate article events. Empty output has no article charge.

Set both Maximum articles in total and an Apify spending limit. The Actor reduces its collection allowance to the number of article events the spending limit permits. New runs without monitoring can return and charge for the same articles again. Developer-owned runs can still consume platform credits; this is separate from customer article pricing.

Set up recurring monitoring

  1. Choose your sources, filters and article limit.
  2. Enable Return only previously undelivered articles.
  3. Enter a Watch name, such as Industry news.
  4. Run once, then reuse the same watch name, sources and filters for later runs.
  5. If you want automatic runs, save the configuration as a task and schedule it through Apify. Enabling monitoring alone does not create a schedule or send alerts.
{
"queries": ["electric vehicles"],
"edition": "US:en",
"timeRange": "week",
"maxArticles": 20,
"newOnly": true,
"watchName": "Industry news"
}

The first run exports eligible articles. Later runs skip identities already delivered to that customer and watch. Delivery history is retained for up to 90 days, subject to storage availability. Use a new watch name when changing sources or filters. A watch is not a full historical archive and does not track edits to article text.

You can download each dataset or connect it to your existing API, webhook or n8n workflow. Email alerts, summaries and downstream analysis must be configured separately.

Interrupted runs

The Actor records pending deliveries and reconciles them with the output dataset when a run resumes. If it cannot determine whether a delivery was saved, it stops and retains the pending record rather than risk a duplicate billable delivery. Contact support with the run ID if this happens. This is not an exactly-once guarantee across all platform failures.

Already saved rows remain in the dataset. Do not write to the Actor's output dataset from another process while it is running. After an interruption, check the actual dataset and run status; a saved report can lag behind the final stored rows.

Read the run report

A green run status does not establish complete news coverage. Open Run report in the run's output and select RUN_REPORT. You can also find the same RUN_REPORT record in its default key-value store under Storage. Review it alongside the Articles dataset.

The report distinguishes failed sources, empty feeds, invalid entries, unknown dates, date exclusions, word-filter exclusions, previously seen articles and sources skipped at the output cap. Resolver diagnostics include failed lookups, session recoveries, transport failures and paused lookup status.

What you observeWhat to check next
Fewer articles than requestedThe article limit is a maximum. Check filters, available source entries, spending allowance and collection limits.
No articles on a repeated watchEligible articles may already have been delivered. Check the previously seen counts.
Google News links rather than publisher linksCheck urlStatus and resolutionError. Lookup may be disabled, blocked or paused.
A publisher feed returns no rowsCheck that the URL is a public feed, that its entries have usable dates, and whether the report records a request or parsing failure.
Results do not match your intended topicSearch words affect Google searches only. Use word filters to narrow other feeds too.
Some sources contributed nothingEarlier sources may have filled the total cap. Use separate runs when every source needs a chance to contribute.

Coverage and connection limits

  • Up to 20 queries and 20 supplied feed URLs, with a total output cap from 1 to 1,000 articles.
  • Sources are processed in this order: queries, supplied feeds, topic sections, then top headlines. Later sources can be skipped when the cap is reached.
  • Up to 120 planned feed requests, within an overall budget of 650 HTTP connection attempts and approximately four minutes of collection time. Redirects and article lookups count toward the network budget. These limits can stop a run below its requested article count.
  • Only public HTTP(S) feed URLs on standard web ports are accepted. Embedded credentials, private destinations and unsafe redirects are rejected. Each response is bounded to 4 MiB after decompression; DTD/entity declarations are refused.
  • Source availability and feed history are controlled by their publishers. Increasing a limit does not establish exhaustive coverage.

Managed datacenter connection is the default. Direct connection is an advanced alternative. Publisher feeds use direct connections in either mode. Residential bandwidth is not used or offered in this release.

Google requests are paced within each run. A failed link lookup can retry once on a replacement session, with at most two session replacements per run. Server-requested waits are respected. Repeated rate limits or connection failures pause further lookups rather than retrying indefinitely. Cached accepted links can remain available during a pause.

These safeguards reduce repeated failures and resource use; they do not eliminate Google's rate limits. Direct links are best effort, and large exports may contain unresolved links or partial coverage. Disable link lookup if Google News links are sufficient for your workflow.

Maintenance and ongoing use

CleanScrape maintains this Actor for one-off research and recurring workflows. We investigate reported errors, update supported integrations when sources change, and test changes with existing inputs and output formats in mind. Source access and uninterrupted availability cannot be guaranteed.

If something stops working, open an issue or email contact.cleanscrape@gmail.com with the Actor name, run ID, public source URL or search input, and expected result. Never include API tokens, passwords or private customer data.

If the export was useful, an honest review helps other users decide whether it fits their workflow. Critical feedback is welcome too.

Disclaimer

This Actor is an independent tool developed by CleanScrape. It is not affiliated with, endorsed by or sponsored by Google or any publisher referenced in its examples or output. All trademarks belong to their respective owners. Brand names identify sources only. Feed availability does not grant unrestricted republication rights; respect source terms and content rights.