Mudah Cars Scraper avatar

Mudah Cars Scraper

Pricing

from $1.26 / 1,000 results

Go to Apify Store
Mudah Cars Scraper

Mudah Cars Scraper

Car listings from Mudah.my, Malaysia's largest classifieds: price in RM, make, model, year, mileage band, engine, transmission, fuel, body type, condition, state, seller type, verification badges, images, description. Filter by make/model, state, price, year, mileage, fuel, seller type.

Pricing

from $1.26 / 1,000 results

Rating

0.0

(0)

Developer

Ibnu Adzim

Ibnu Adzim

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Categories

Share

Car listings from Mudah.my, Malaysia's largest classifieds site (~88,000 live cars): price in RM, make, model, year, mileage band, engine capacity, transmission, fuel, body type, condition, state and area, dealer or private seller, verification badges, loan eligibility, every image, and the full description.

HTTP only. It reads the structured listing cache the site streams with each page, not the rendered cards, so every field is the site's own value rather than text parsed off a tile. No login, no key, no browser.

What it is for

  • Used-car price monitoring by make/model/state/year band, on a schedule.
  • Dealer inventory tracking — filter to dealers, group by sellerName.
  • Market research — 39 rows per request, ~10,000 per target.

Input

fieldwhat it does
makestoyota, perodua/myvi, mercedes-benz — the slugs in Mudah's URLs.
regionsState slugs (selangor, kuala-lumpur, ...) or malaysia (all, default).
searchTermsFree text. Loose — see below.
condition, priceMin/Max, yearMin/Max, mileageMax, fuelType, transmission, sellerType, mudahCertifiedOnly, sortByFilters, every one verified against the rows it returns.
includeFeaturedAdsOff by default — see below.
maxItems, maxConcurrency, minRequestInterval, proxyConfigurationLimits.

Targets are the product of regions × makes × searchTerms; each gets its own SEARCH_SUMMARY row.

Four things about this site worth knowing before you trust a run

1. A typo in a make or model returns the whole site, silently

/malaysia/cars-for-sale/notamake is HTTP 200 with all 87,695 cars. /toyota/notamodel is all 19,335 Toyotas. No error, no redirect. A scraper that trusts the URL returns 10,000 cars for "toyata" and every one of them looks right, because each one is a real car.

This Actor checks page 1: every row's make must match the slug you gave (and its model must start with the model slug). If it does not, the target stops with zero rows and stoppedReason: make_not_recognised / model_not_recognised. Regions are validated locally against the 16 states.

2. Each target stops at about 9,984 rows, and the site pretends that is zero

Pages are 39 cars; page 256 is the last one served. Page 257 answers with an empty list and totalResults: 0 for a query that has 87,695 — exactly what a search with no matches looks like. The Actor never requests it, reports wallReached: true and stoppedReason: wall, and keeps the real total from page 1. Narrow by state, make, model or price band to get every row.

3. searchTerms is a loose text match

myvi → 3,730 results, every one a Myvi. zzqqxxnotacar → 58 results: an Audi TT, a Mazda CX-5, a Perodua Alza. That is a fuzzy matcher finding partial hits in descriptions, not a fallback catalogue, and there is no honest way to tell "few real matches" from "noise" per row. So nothing is dropped; the summary's keywordHitShare (share of rows whose title/make/model contains a query word) tells you what you got: 1.0 on myvi, 0.0 on nonsense. For exact filtering use makes and regions, which combine with searchTerms.

A browse page carries 6 paid "featured" ads beside its 39 results, and on a make/model/state page they follow the path (a Myvi page's featured ads are Myvis). A ?q= search page carries 39 of them — and a search with zero results still ships all 39, because they ignore the search text. Off by default; with includeFeaturedAds on, each carries isFeatured: true and is never counted in carsReturned or against totalResults.

Other things measured

  • Mileage is a band, not a reading:
    mileageMin: 100000, mileageMax: 109999
    , shown on the site as "100k-110k". There is no exact odometer value.
  • Cloudflare rate-limits at ~20 requests a minute per IP (HTTP 429, clears in ~10 s). Requests are paced (minRequestInterval, default 3 s), a 429 is retried with a long backoff on a fresh IP, and the proxy default is Apify's free datacenter pool so concurrent targets do not share one IP's budget. rateLimitHits in the summary counts what was absorbed.
  • sort=, car_type=, search_only= and page= are accepted by the site and ignored. Only the parameters measured to work are exposed.
  • Dealers are ~78% of the listing; sellerType, dealerVerified and mudahCertified are on every row.

Output

  • CAR — listId, url, title, price (RM), make, model, year, yearVerified, mileageMin/mileageMax/mileageLabel, transmission, fuelType, engineCapacityCc, bodyType, condition, region, subarea, sellerName, sellerType, dealerVerified, mudahCertified, badges, carLoanEligible, carLoanTenureYears, imageUrl, imageUrls, description, postedAt, updatedAt, expiresAt, isFeatured, query, resultPosition, pageFound.
  • SEARCH_SUMMARY — one per target: totalResults (the site's count), carsReturned, featuredReturned, pagesFetched, stoppedReason, wallReached, makeFilterVerified, modelFilterVerified, regionFilterVerified, keywordHitShare, rateLimitHits, minPrice, maxPrice, and the exact filters sent.
  • ERROR — invalid_input, rate_limited, page_shape_changed, and anything else that went wrong, with detail.

Known limits

  • Cars for sale only. Mudah's other verticals (property, phones, jobs) use a different page contract and are not covered; nor are cars for rent.
  • ~9,984 rows per target (see above).
  • The listing page carries no phone number, no dealer address and no exact mileage; those live on the ad page, which this Actor does not fetch.