Music Streaming Metrics Scraper avatar

Music Streaming Metrics Scraper

Under maintenance

Pricing

Pay per usage

Go to Apify Store
Music Streaming Metrics Scraper

Music Streaming Metrics Scraper

Under maintenance

Get music performance metrics from YouTube Music, Spotify, JioSaavn, and Gaana in one place. Provide track URLs and retrieve available play, view, engagement, and track metrics in a structured dataset ready for analysis, Google Sheets, APIs, or automation workflows.

Pricing

Pay per usage

Rating

4.0

(1)

Developer

Kamal Pardeshi

Kamal Pardeshi

Maintained by Community

Actor stats

0

Bookmarked

1

Total users

0

Monthly active users

4 days ago

Last modified

Categories

Share

What does Music Streaming Metrics do?

This V1 Actor accepts public track URLs from Spotify, JioSaavn, YouTube Music, and Gaana. It returns identity-checked metrics and metadata. It is currently an audit build, not a production-validated Store product.

Why use it?

The Actor keeps each input URL independent, preserves duplicate input entries, and reports a failure without discarding the rest of the batch. It does not search for songs or match recordings between services. Outputs can be consumed through the Apify dataset API.

What data can it extract?

FieldMeaning
Spotify streamCountExact playcount/playCount from a track entity matching the requested ID
JioSaavn playCountAPI play_count for a song whose canonical URL token matches the request
YouTube viewCountPlayer count after videoId validation
YouTube likeCount / commentCountExact integers from the main video action bar / explicit comments header, when exposed
Gaana favoriteCountTrack-level total_favourite_count, separate from artist favourites and likes
Gaana playCountDisplayThe original play_ct display label, for example 35M+ or <100K; exact playCount remains null
trackIdentityVerified platform ID and available title, artists, album, ISRC and duration
metadata.metricStatusExplains metrics absent from the inspected response
diagnosticsPer-tier runtime and observed response bytes; billing fields remain null

Null is not zero. Rounded counts are not converted into invented exact values. A YouTube uploader is recorded as channelName, not assumed to be the recording artist. ISRC requires an explicit ISRC field and a valid ISRC format; arbitrary platform IDs are never reused as ISRC. Metadata not established from the matched entity remains null. Spotify title and artists are bound to the matched entity; no external recording catalog validation is performed.

How to run

  1. Install dependencies with npm ci and Chromium with npx playwright install chromium.
  2. Build with npm run build.
  3. Save input to storage/key_value_stores/default/INPUT.json.
  4. Run apify run --user-agent apify-codex-plugin/apify-actor-development.
  5. Inspect the dataset and the AUDIT_SUMMARY key-value record.

Apify Cloud uses the included Dockerfile and Actor schema. An authenticated account with the selected proxy entitlement is needed only when a proxy tier is attempted.

Input

See the input tab for configuration options.

In the Console form, use + Add URL to enter each song URL in a text field. The URL-list editor stores entries as { "url": "https://..." } objects. Update saved JSON inputs to this object format: Apify's URL editor validation expects it. The extraction code still understands legacy string entries when invoked without that form validation. Enter individual song URLs; remote URL-list files are not supported.

{
"urls": [
{ "url": "https://open.spotify.com/track/7qiZfU4dY1lWllzX7mPBI3" },
{ "url": "https://www.jiosaavn.com/song/tum-hi-ho/EToxUyFpcwQ" },
{ "url": "https://music.youtube.com/watch?v=dQw4w9WgXcQ" },
{ "url": "https://gaana.com/song/manjha" }
],
"network": {
"spotify": ["direct", "datacenter", "residential"],
"jiosaavn": ["direct", "datacenter", "datacenter-in", "residential"],
"youtube": ["direct", "datacenter", "residential"],
"gaana": ["direct", "datacenter", "residential"]
},
"spotifyMode": "E"
}

Residential uses India. Indian datacenter availability depends on the account and proxy pool; configuration failure is not proof a site requires residential access. Tiers are lazy and platform-specific. Only recognized blocking or transient transport errors advance tiers. Identity errors and parser failures do not trigger costly residential retries. Specify a single tier to benchmark it in isolation. Direct success never creates a proxy.

Spotify benchmark modes

A uses an unblocked browser and a fresh session per track. It is a repaired control, not the original unbuildable implementation. B blocks images, fonts, media, prefetch and service workers with a fresh session per track. C additionally blocks selected analytics/advertising endpoints. D reuses the browser and context for the batch. E also attempts HTTP replay of an observed single-track GraphQL operation using the context's cookies and captured headers, refreshing through the browser if replay fails. E is the local default after the 100-request fixture test; broader catalog validation remains necessary. Tokens are kept in memory and never saved to the dataset.

By default Spotify first inspects the public HTML initialState for a matching track count. It starts a browser only when that count is absent. Required application scripts are retained. Broad script blocking can prevent token bootstrapping and is intentionally avoided. The browser stops after a verified count arrives, without waiting for Spotify's entire UI to finish rendering.

Output

You can download the dataset in various formats such as JSON, HTML, CSV, or Excel. Each input produces one row, with streamCount, playCount, viewCount, likeCount, commentCount, favoriteCount and playCountDisplay at the top level, preserving the latest website edits. status is success when the platform's primary count is verified (track favourites for Gaana), partial when identity is verified but that count is missing, or error when extraction/identity verification fails. A success does not promise every optional metric is available.

Gaana is fetched over HTTP without starting a browser. Track identity is bound to the requested song slug and a consistent numeric track ID. The Actor reports the site's display label without converting rounded values into exact plays. The undocumented popularity field is not used. Missing metrics remain null, with reasons in metadata.metricStatus. Gaana support has only a small live-song smoke test, not broad catalogue validation.

How much will it cost?

Use completed, platform-isolated cloud runs for billing. responseBodyBytes measures decoded HTTP response payloads; browserEncodedDownloadBytes measures completed browser downloads. Neither is residential billed bandwidth. Failed/unfinished transfers and protocol overhead can be absent. Cloud startup, memory allocation, storage, and transfer charges must be included before setting Store prices.

The audit report records actual results and limitations. Do not infer production economics from a two-track local sample or extrapolate a blended platform cost.

Tests and troubleshooting

Run npm test for identity, numeric parsing, URL validation, and fallback regression tests. On a sandbox that forbids test subprocesses, Node 22+ can use node --test --test-isolation=none test/*.test.cjs.

Private endpoints and frontend schemas may change. A missing count is not evidence of zero usage. YouTube comments may require an additional continuation request; this version reports their absence from the initial page explicitly rather than declaring comments disabled. Failed proxy setup is reported as a request failure, not a successful empty result. Use the Issues tab for reproducible inputs and the API tab for programmatic access.