🏭 Industry Classifier - NAICS & SIC Code Assignment avatar

🏭 Industry Classifier - NAICS & SIC Code Assignment

Pricing

from $2.00 / 1,000 classification runs

Go to Apify Store
🏭 Industry Classifier - NAICS & SIC Code Assignment

🏭 Industry Classifier - NAICS & SIC Code Assignment

⚑ Assign a company a NAICS and/or SIC industry code with confidence and evidence, no LLM. βœ… Walks an embedded sector-then-detail code tree, scores with TF-IDF + phrase hits, and uses negative keywords so "SaaS for dentists" lands in software, not dentistry.

Pricing

from $2.00 / 1,000 classification runs

Rating

0.0

(0)

Developer

mohamed alaya

mohamed alaya

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

a month ago

Last modified

Categories

Share

Industry Classifier

Assign a company a NAICS and/or SIC industry code β€” with confidence and the exact evidence that decided it β€” deterministically, with no LLM and no embeddings model. Industry classification is effectively unserved on Apify: most "classify this company" tools either call an LLM per row (slow, non-deterministic, costs a token every time) or don't exist at all. This actor is the same "hierarchy done properly" approach as product-taxonomy-classifier, applied to NAICS/SIC.

What it actually does

1. Hierarchy, done properly. Level 1 picks a NAICS 2-digit sector (24 candidates β€” every official 2017/2022 sector; Manufacturing's 31/32/33 split is represented as three rows). Level 2 is scored only against that sector's own detail codes β€” a company can never land on "Offices of Dentists" unless "Health Care and Social Assistance" was actually the level-1 pick. Either level can stop the walk and route to review instead of forcing a guess.

2. A real embedded code table, not a toy list — see src/codes.js: all 24 NAICS sectors plus 125 hand-picked 4-6 digit NAICS detail codes spanning software, professional services, healthcare, retail, food service, finance, manufacturing, construction, transportation and more, each with a rich keyword/synonym set, plus an approximate NAICS→SIC crosswalk for every code included.

3. Negative keywords β€” the real accuracy lever. "SaaS for dentists" and "a dental clinic" both say "dentist", but only one of them is a dental practice. Codes prone to this confusion (Offices of Dentists, Offices of Physicians, Offices of Lawyers, Real Estate Agents, Full-Service Restaurants, Fitness Centers, Hotels, Beauty Salons, Travel Agencies, Colleges, Veterinary Services, CPA firms, Electronics Stores, Grocery Stores, Electronic Shopping...) carry negative keywords like "software"/"saas"/"platform"/"api". When one fires, that code's score is crushed (see negativePenalty) β€” and crucially this propagates up to the sector level too, so a software company doesn't even win the Health Care sector just because "dental clinic" appears verbatim in its "SaaS for dentists" pitch. Software Publishers / Custom Computer Programming carry no such penalty and win on their own "software"/"saas"/"api" keywords instead.

4. Four-signal lexical scoring per level, computed against the current candidate set only (see src/classify.js): TF-IDF-ish cosine similarity (IDF built from the sibling codes competing at that level), token overlap, character-trigram Dice similarity (typo/word-order tolerance), and exact multi-word phrase hits ("point of sale", "custom software development") β€” the strongest single signal, and exactly what's returned as evidence.

5. Evidence, not a black box. Every classified row carries evidence.matchedPhrases β€” the literal phrases that fired β€” and evidence.suppressed, showing which negative keyword knocked out a competing code, so you can see why, not just what.

6. Confidence + review routing, same shape as the taxonomy classifier: confidence blends how strong the winning score is with how far ahead it is of the runner-up. Anything below reviewBelow at either level stops there and goes to the review queue with its top candidates, instead of being force-fitted.

7. Optional site fetch. Pass a url per company and set fetchSite: true to fetch the page (via the shared fetchPage/extract helpers) and fold its meta description + visible text into the description before classifying β€” otherwise classification runs purely on the text you supply.

Input

{
"companies": [
{ "name": "Bright Smile Dental Clinic", "description": "A friendly dental clinic offering checkups and cleanings." },
{ "name": "DentaFlow", "description": "SaaS practice management software platform for dentists, a cloud app with an API." }
],
"fetchSite": false,
"reviewBelow": 55,
"negativePenalty": 80,
"tfidfWeight": 40, "overlapWeight": 25, "trigramWeight": 15, "phraseWeight": 20
}

reviewBelow, negativePenalty and the four weights are 0-100 integers (Apify input schemas have no float type); weights are normalized against each other, so only their ratio matters.

Output

type: "classified" rows carry naicsSector ({code, title}), naicsDetail ({code, title}), sicCodes (crosswalked), per-level levelConfidences, an overall confidence (the weaker of the two levels), and evidence. type: "review" rows carry the reason and best-guess candidates. type: "sectorCount" rows summarize where companies landed.

Honest limitations

  • Lexical, not semantic. There is no embeddings model β€” matching is TF-IDF/token/trigram/phrase based. Give unusual businesses explicit tags to bridge vocabulary the algorithm can't infer ("Footwear" and "Shoes" share almost no characters, for example).
  • Coverage is the included 125-code subset, not the full official lists. The real NAICS 2022 has ~1,057 six-digit codes and SIC has ~10,000 codes; this actor covers the ~125 detail codes most likely to come up classifying real companies. An unusual or hyper-specific business may correctly land in review rather than being force-fitted to the nearest included code.
  • The SIC crosswalk is an approximate "closest analogue" pick per included NAICS code, not the official many-to-many Census Bureau concordance. Treat it as a helpful cross-reference, not a compliance-grade mapping.
  • Negative keywords are hand-picked for the confusions we know are common (industry vs. the software/app that serves that industry). Confusions outside that pattern aren't specifically guarded against.
  • Very short/sparse company text has little for any lexical method to work with and will often land in review rather than being force-fitted.
  • Confidence is a relative signal (how well the winner beat the field), not a calibrated probability β€” tune reviewBelow against your own precision/recall needs.
  • Capped at 50,000 companies per run.