Global Industrial Scraper avatar

Global Industrial Scraper

Pricing

from $3.00 / 1,000 results

Go to Apify Store
Global Industrial Scraper

Global Industrial Scraper

Scrape Global Industrial (globalindustrial.com) - US industrial & B2B supplies marketplace with 1M+ products. Search by keyword, browse by category, fetch by product ID or URL - prices, brands, availability, ratings, images, specs.

Pricing

from $3.00 / 1,000 results

Rating

0.0

(0)

Developer

Crawler Bros

Crawler Bros

Maintained by Community

Actor stats

0

Bookmarked

4

Total users

3

Monthly active users

7 days ago

Last modified

Share

Scrape Global Industrial (globalindustrial.com) — one of America's largest industrial & B2B supplies marketplaces (1M+ SKUs across tools, material handling, HVAC, safety, packaging, electrical, foodservice and more). Search by keyword, browse by department/category, fetch by item key or direct URL, and get product prices, brands, availability, ratings, images and technical specs. No login, no API key, no residential proxy — runs on the free Apify AUTO proxy tier with US egress.

What this actor does

  • Four modes: search, byCategory, byProductIds, byUrls
  • Full product data: price + price breaks, MSR/sale/overstock prices, brand, SKU/MPN, availability & lead time, ratings & review counts, images, dimensions, specs, flags (best seller, premium, hazardous…)
  • Category browse: every department (/t/…) and category (/c/…) with optional subcategory tree records
  • Filters: price range, brand, keyword-in-title, minimum rating, in-stock only
  • Empty fields are omitted

Data source

Domainwww.globalindustrial.com
TechNext.js server-side rendering — all data comes from the __NEXT_DATA__ JSON embedded in each page (category products also under categoryDetails.productByGroups), plus the same-origin catalogApis/catalog/product lookup for item keys
Searchhttps://www.globalindustrial.com/searchResult?q=<keyword> (SSR; page 2+ adds list=true&cp=<offset>)
Category browsehttps://www.globalindustrial.com/c/<path> and /t/<department> (SSR)
Product pageshttps://www.globalindustrial.com/p/<slug> (SSR)
Imagesimages.globalindustrial.com CDN (/images/{size}/{image}.webp) — accessible from anywhere
EgressThe origin is behind CloudFront with a US-only geo-restriction: every request from a non-US IP gets 403 Request blocked. The actor therefore defaults to Apify AUTO proxy with country US (free on all plans, no residential cost). AUTO rotates IPs automatically, and the actor retries with a fresh proxy exit on 403/429/5xx.

Replacement history (slot lineage). This actor's marketplace slot was originally assigned to Costco (costco.com), then Zoro (zoro.com) — both were proven hard-blocked from Apify cloud egress: direct requests and the AUTO, RESIDENTIAL and SHADER proxy tiers all returned 403 for every page type. Per the project's replacement policy, the slot was reassigned to a same-category platform that is verifiably reachable from Apify's US egress: Global Industrial, a major US wholesale/industrial supplies marketplace (same vertical as Costco Business and Zoro: MRO, warehouse, tools, janitorial, packaging, electrical, HVAC). Cloud egress probe: /c/industrial-supplies and /c/tools return 200 with SSR product data, and all scrape paths below were reverse-engineered from the live site's own client code and archived page captures.

Output per product (recordType = product)

  • itemKey — Global Industrial item ID
  • title, productUrl, brand, sku, legacyNumber, manufacturerPartNumber, unspscCode
  • price, priceOriginal, priceCatalog, priceMsr, priceOverstock, priceSale, priceClearance, minSalePrice, savings, priceType, autoReOrderDiscount
  • priceBreaks[] — quantity-tier pricing (qtyStart, qtyEnd, price, priceOriginal, priceType)
  • availabilityStatus (InStock / …), shipTime, leadTime, freeShipping, inStock
  • avgRating, numReviews
  • imageUrl (177×177), imageUrlLarge (500×500), productImages[] (product pages only)
  • categoryKey, categoryUrl, categoryPath, breadcrumbs[] (url, label)
  • Flags: isBestSeller, isPremium, isNewArrival, isPrivateLabel, isHazardous, isNonReturnable, isNonCancellable, isCallForPrice, isMapPrice, isCustomizable, isFreightFixed, isNoAirShip, isLtl, isReorder, isPromotional
  • dimensions (length, height, width, unit), qty, attributes[] (group, name, value), description, itemLaunchDate, timeStamp
  • recordType: "product", sourceUrl, scrapedAt

Output per category (recordType = category)

  • categoryKey, categoryCode, title, categoryPath, leafUrl, categoryLevel, parentCategoryKey
  • totalItems, sort, showMore, startIndex, categoryUrl, heading, titleTag
  • subcategories[] — child categories with categoryKey, title, url, totalItems (when includeSubcategories is on)

Output per failed lookup (recordType = error)

  • itemKey and/or sourceUrl, error (short reason), recordType: "error", scrapedAt

Input

FieldTypeDefaultDescription
modestringsearchsearch / byCategory / byProductIds / byUrls
searchQuerystringadjustable workbenchFree-text keyword (mode=search)
sortBystringrelevancerelevance / price_low / price_high / best_selling
categoryPresetstringPopular departments & categories (mode=byCategory)
categoryPathstringAny /c/… or /t/… path or URL; overrides preset (mode=byCategory)
includeSubcategoriesbooleanfalseEmit category tree records while browsing
productIdsarrayNumeric item keys, or /p/<slug>-<key> URLs (mode=byProductIds)
startUrlsarrayProduct / category / search URLs (mode=byUrls)
minPrice / maxPricenumberListing-price bounds (USD)
minRatingintegerMinimum average rating (1–5)
brandstringBrand substring filter
containsKeywordstringTitle substring filter
inStockOnlybooleanfalseKeep only InStock products
maxItemsinteger50Hard cap on product records (1–1000)
proxyConfigurationobjectAUTO + USPrefilled: {"useApifyProxy": true, "apifyProxyCountry": "US"}

Example: keyword search with filters

{
"mode": "search",
"searchQuery": "pallet jack",
"sortBy": "price_low",
"minPrice": 100,
"maxPrice": 500,
"inStockOnly": true,
"maxItems": 100
}

Example: browse a department, include subcategories

{
"mode": "byCategory",
"categoryPreset": "Power Tools",
"includeSubcategories": true,
"maxItems": 200
}

Example: exact category path with price filter

{
"mode": "byCategory",
"categoryPath": "/c/tools/abrasives/sanding_discs",
"brand": "3M",
"maxItems": 50
}

Example: lookup by item key, or scrape by URL

{
"mode": "byProductIds",
"productIds": ["33082504", "33201668"]
}
{
"mode": "byUrls",
"startUrls": [
{ "url": "https://www.globalindustrial.com/p/variable-speed-control-switch-with-6-ft-plug" },
{ "url": "https://www.globalindustrial.com/c/tools/sockets_bits" },
{ "url": "https://www.globalindustrial.com/searchResult?q=shelving" }
]
}

Use cases

  • MRO & procurement teams — pull current pricing, availability and specs for industrial supplies
  • Resellers & distributors — monitor competitor assortment and price breaks by brand or category
  • Category intelligence — track SKU counts, best-seller flags and new-arrival flags across departments
  • E-commerce benchmarking — compare list vs MSR vs clearance pricing across 1M+ SKUs
  • Supply-chain research — lead-time and free-shipping signals for sourcing decisions

FAQ

What is the data source? Global Industrial's own website — server-rendered Next.js pages (__NEXT_DATA__) and the site's public catalogApis product lookup. No third-party data.

Why does it need a US proxy if no login is required? The site's CloudFront distribution rejects every non-US IP with HTTP 403 ("Request blocked"). Apify's AUTO proxy with country US is included free on every Apify plan, so there is no extra cost. The actor auto-retries on 403/429 with a fresh proxy exit.

Are prices real-time? Yes — each run scrapes the live listing pages, including current sale/clearance/MAP flags and quantity price breaks.

Why is byCategory sometimes limited? Category pages render products server-side (flat list and/or productByGroups). If a category returns zero embedded products, the actor still emits the category + subcategory tree and reports a status message; this happens only for CMS-only landing pages without product grids. The number of products the site embeds per SSR page varies by page configuration (typically 18–54). Some deep categories (4+ path segments, e.g. /c/tools/tool_storage/chests_roller_cabinets) expose a __1-suffixed URL variant that renders the full product grid, so those browse fully. Other deep paths no longer exist as separate SSR pages — e.g. /c/tools/power_tools/drills and /c/tools/abrasives/sanding_discs now redirect to their parent category (or, for the __1 variant, to /t/tools), so the actor emits what the parent page serves and reports the shortfall in its status message. Shallow categories (3 segments, e.g. /c/tools/power_tools) serve only their first SSR grid, and the actor reports the shortfall in its status message.

How does pagination work? Search and category SSRs only advance when the site's quirks are satisfied: search page 2+ needs list=true&cp=<offset> (cp is an item offset, not a page number) and deep category pages need sort=most_relevant&cp=<offset> on the plain URL. The actor appends these automatically and walks until maxItems, the site's last page, or the point where the SSR stops advancing (reported in the status message).

Why do some searches return 0 records? Queries that exactly match a brand (e.g. dewalt) get redirected by the site to a /shopbybranditems brand landing page whose products load entirely client-side — nothing is in the server-rendered HTML, so nothing can be scraped without a browser. Add a second keyword (e.g. dewalt drill) and the search stays on server-rendered results. The actor reports this case in its status message.

Why do some byProductIds lookups fail? The site's legacy catalogApis product endpoint now returns HTTP 403 for every egress IP, so the actor falls back to the site's own search index, which indexes most — but not all — item keys. Keys that resolve are emitted as full products; keys the index doesn't know produce a typed error record (visible in the dataset). For guaranteed lookups, pass a full product page URL (/p/<slug>) in productIds — those are fetched directly and always resolve.

Why are some products missing images? Products without an image key on the source simply omit the image fields — no placeholders are ever fabricated.

Does this actor require affiliation with Global Industrial? No. This is an independent third-party actor using the public website.

What are recordType: "error" records? Lookups that fail (invalid ID, not found, blocked page) are emitted as typed error records with sourceUrl and a short error reason so you can audit partial runs.

Is this actor affiliated with Apify? It is built on the Apify platform; the data source is a third-party website.

Are there rate limits? The site has no documented API limits. The actor uses polite delays and retries, and rotates proxy exits automatically.