# Changelog of Company Socials (Crosswalk) (`publicrecords/company-social-profile-finder`) Actor

- **URL**: https://apify.com/publicrecords/company-social-profile-finder/changelog.md
- **Full Actor documentation**: https://apify.com/publicrecords/company-social-profile-finder.md

## Changelog

### 1.2.2 — 2026-09-17 (Gate B name-lookup: Mark fix)

- **names.js auto-resolve:** exact-match override never applies to single-word queries (gap rule only); exact-match override never applies when another candidate with confidence ≥ 0.8 comes from SEC or the user graph — return `candidates` instead (YETI → candidates).
- **evaluate.mjs / domainTruth:** if the truth domain redirects to our chosen domain, or the entity lists the truth domain as an alternate, count as correct and log `truth stale` (Statista .org→.com Wikidata error; Payoff→Happy Money rebrand). BASE confidences, prices, person gate, schema unchanged.
- Tests: YETI / Statista / Payoff in `test/acquisition.test.js`.

### 1.2.1 — 2026-09-17 (launch-day fixes from Apify Master's NO-GO)

- **Public-company index now sourced from Wikidata** (P5531 SEC CIK → P856 website, P249 ticker) with SEC `company_tickers.json` filling tickers/names: SEC's own submissions JSON leaves `website` empty for most registrants (verified: even Apple is blank), so the previous build produced zero domains. Live: 3,431 companies with websites, ~3,300 tickers; `cik:` seeds resolve via `byCik`. Two requests instead of ~10,000; runs in seconds.
- **Cache hits are never charged**: Apify rejects a $0 event price, so `entity-cache-hit` is dropped from Monetization and skipped in code (still counted in RUN\_SUMMARY). PRICING.md updated.

### 1.2.0 — 2026-09-17 (discovery and SEO)

- Actor renamed to the keyword-rich slug `company-social-profile-finder` (Crosswalk stays the brand); all references updated.
- README restructured around search intent (what people use it for, works with, FAQ); LISTING.md with exact seoTitle/seoDescription/categories/screenshots and the MCP description.
- Discovery site generator `site/build.mjs`: 7 pages (one per high-intent query) with canonical URLs, Open Graph, JSON-LD (SoftwareApplication + FAQPage + BreadcrumbList), an embedded live checker, sitemap.xml, robots.txt, llms.txt and og.png.
- Marketing kit (`marketing/`): three articles with disclosure, a cross-promotion template; GROWTH.md gains the Search Console cadence.

### 1.1.1 — 2026-09-17 (discovery optimization)

- Discovery site generator (`site/build.mjs`): 7 query-targeted pages with unique titles/descriptions, canonical URLs, Open Graph, SoftwareApplication + FAQPage + BreadcrumbList JSON-LD, internal linking, embedded live checker, sitemap.xml and robots.txt.
- README: FAQ section written in the phrasing people search (LinkedIn from website, Instagram from website, name to website, enrich Maps leads, canonicalId, MCP).
- `actor.json`: categories (Lead generation, Social media, AI, Developer tools, Business); description rewritten with the task in the first sentence for MCP search. `package.json`: keywords, homepage, repository.
- LISTING.md: exact seoTitle (63 chars), seoDescription (150 chars), screenshots, target-query list, off-Store surfaces. EXECUTION\_BRIEF: Search Console/Bing submission, n8n template submission, partner README links, monthly discovery metrics.

### 1.1.0 — 2026-09-17 (autonomous launch; measured against real truth)

- **Wikidata as a second declared source.** Human-curated profile properties for the organization whose official website is the domain are merged as declared links (0.85, still platform-verified); when a site is blocked, robots-denied, JS-only or down, Wikidata becomes the hub (`website_unreadable_used_wikidata`). Toggle: `wikidataDeclared`.
- **Browser-UA retry on block** for organization sites (robots.txt still honored, no challenge solving), recorded as `siteSignals.uaFallback`.
- **Organization-evidence gate softened:** refuses only when person signals exist; sites with no signals proceed and are stored only with a confident link.
- **Truth without a human:** `scripts/truth-from-wikidata.mjs` builds an independent truth set from Wikidata's social properties; `evaluate.mjs` supports `--no-wikidata` (independent measurement), set-valued truth cells, id-form comparability (channel id vs handle), platform-confirmed stale exclusions, and `--mode auto` with Actor probes via `APIFY_TOKEN`.
- **First real measurements (40 Wikidata organizations, direct mode, container without proxy):** independent precision 0.86 with every residual mismatch on declared-only Instagram/Facebook (brand-vs-corporate or rebrands, e.g. payoff.com → happymoney); direct-probe platforms 1.00; with the Wikidata source enabled, resolved 95%, coverage 88%, precision 0.975. Gate A on the platform runs in `auto` mode so Instagram/Facebook are confirmed.
- EXECUTION\_BRIEF.md: fully autonomous launch and operation plan. Tests: 40.

### 1.0.2 — 2026-09-17 (second adversarial review; see docs/REVIEW\_v1.0.1.md)

- **Shared public graph.** Users read through the publisher's pre-seeded public store (`sharedGraphStoreId` / `CROSSWALK_SHARED_STORE_ID`); writes stay private. Without this, seeding never reached users and the free-first-touch model did not exist.
- Index hygiene: only domains ≥ 0.8 indexed; stale identifier/registry index entries removed on save; `forget` clears the name index; batched writes.
- Dead-organization detection (`site_unreachable` after 3 consecutive failed re-verifications); `registry_dropped` event; compact output includes confident `alternates`; `reverify` reports `stale` separately.
- `undici` pinned; `budget_exhausted` count fixed; evaluation harness reads `declaredOnly`; `scripts/seed.mjs` uses `mode: seed`; dataset view shows compact `domain`; docs synced with code. Tests: 39.

### 1.0.1 — 2026-09-17 (adversarial review; see docs/REVIEW\_v1.0.0.md)

Severity 1 (promise-breaking): people leaked through name lookup (Wikidata humans with websites — fixed with P31 organization classes, humans excluded), through personal websites (fixed: person-site detection + organization-evidence gate; nothing stored without organization signals and a confident link), through `li:in/…` prefixes (fixed), and Wikipedia/Threads/Pinterest/Amazon URLs became domain seeds (fixed: recognized, declared-tier, refused as seeds).
Severity 2: `acme.com.au` no longer confirms `acme.com`; X `expanded_url` list mismatches are `unverified`, not contradictions; render fallback de-duplicates; **maintenance modes default to `direct`** (an `auto` seed of 100K would have billed ~$1,400 of sub-runs to the publisher); Actor probe runs must be SUCCEEDED.
Severity 3: mailto/javascript/IP/localhost/trailing-dot/IDN/over-long inputs handled; `Acme, Inc., Toronto` parsed correctly; full Google Maps URLs (hex CID) and short links supported; CIK/ticker normalized; `inSlice` guard; TikTok URLs without `@`.
Also: `dataset_unreadable` diagnostic; landing page reports browser CORS blocks; structural organization signals (legal pages, commercial nav, corporate copyright) so minimalist company sites still resolve. Live regression set: shopify, apify, stripe, basecamp, fastmail, ycombinator resolve; paulgraham.com refused; "Taylor Swift" returns no candidates. Tests: 37.

### 1.0.0 — 2026-09-17 (final, optimized for acquisition; see GROWTH.md)

- **Dataset on-ramp:** `inputDatasetId` + auto-detected `domainField`; rows pass through with the result under `crosswalk`. Live-tested on a Maps-shaped dataset.
- **Company names, without guessing:** `Acme Plumbing, Toronto` resolves through ranked candidates from the user's graph, Wikidata, the SEC index and Google Maps (auto mode). Auto-resolves only an unambiguous match (exact name with official website, or a ≥ 0.25 gap); otherwise a free `candidates` row with evidence and `resolveWith`. Live: "Shopify" and "Cloudflare, San Francisco" auto-resolve; hosted-page domains (github.io etc.) excluded.
- **Free first touch:** `mode: seed` (list / SEC / Wikidata SPARQL, resumable cursors) and cache hits priced at $0.
- **Agent surface:** compact output by default for a single `query`; `nameMatch` in compact output; MCP description and LangChain/CrewAI tools in `integrations/`.
- **Off-Store:** `site/index.html` free checker (visitor's own token), Google Sheets `=CROSSWALK()`, n8n workflow, Zapier/Make recipe.
- **Listing** renamed for the search bar; GROWTH.md with the bot's weekly/monthly growth cadence and honest expectations.
- Graph: name index (`idx--name--<normalized>`). Tests: 35.

### 1.0.0-rc.1 — 2026-09-17 (strip-down to the resale product)

- **Auto verification is the default.** Instagram and Facebook are now confirmed through the official Store Actors (`apify/instagram-profile-scraper`: `externalUrl`/`externalUrls`/`biography`; `apify/facebook-pages-scraper`: `websites`/`website`/`intro`) under the user's account, with per-entity (2) and per-run (200) budgets and typed errors. Field names verified against the Actors' published output on 2026-09-17.
- **Removed:** free-text name seeds (a guess); Wikipedia, Pinterest, Threads, Amazon-store namespaces (unverifiable, low value).
- **Tiered output:** `identifiers` contains only verifiable namespaces; `declaredOnly` carries Crunchbase/Glassdoor/Indeed/app-store links the site publishes, never counted or billed. `schemaVersion` 3.
- **Docs:** RELIABILITY.md (the contract), PRICING.md, TERMS.md, LAUNCH\_CHECKLIST.md; README rewritten around the contract. Reviews moved to `docs/`.
- Tests: 31.

### 0.3.0 — 2026-09-17 (reliability and ease of use; see REVIEW\_v0.2.md)

- **Stable ids:** `canonicalId` persists across domain moves; new domains index to the existing id; `domain_changed` events still fire.
- **In-run de-duplication:** seeds that resolve to the same organization resolve once and bill once; the second is a cache hit with `dedupedWith`.
- **Probes:** run in parallel (bounded, `probeConcurrency`), ~3× faster on cache misses (6 queries in 4s live). Instagram/Facebook direct attempts removed (`not_attempted`, zero requests). Dedicated TikTok (`bioLink` field) and X (website field / bio) readers.
- **Privacy:** only role email addresses are stored.
- **Ease of use:** `query` single-identifier input; `outputFormat: compact`; `crosswalkVersion` on every record; QUICKSTART.md with curl/Node/Python/MCP.
- **Truth set tooling:** `scripts/truth-assist.mjs` drafts rows with `?` prefixes; `evaluate.mjs` refuses unreviewed rows.
- **Stats:** per-run counters written once (no races).
- **X confirmation restored:** X serves full profile pages only to browser user-agents; public profile reads now use a browser UA (documented in DATA\_POLICY.md), read the embedded `expanded_url` website field, and resolve up to three t.co bio links. Canary caught the regression before release.
- **Bare platform hosts** (`github.com`, `linkedin.com`) resolve as the platform company's own site instead of being refused.
- Tests: 29. Live check 2026-09-17: shopify.com, hubspot.com, stripe.com resolved with LinkedIn/YouTube/Wikidata platform-confirmed; LinkedIn URL seed de-duplicated against domain seed; openai.com blocked from the test container (Cloudflare) and diagnosed honestly.

### 0.2.0 — 2026-09-17 (hardening for commercial use)

Driven by REVIEW\_v0.1.md. Summary of changes:

- **Precision:** platform confirmation now distinguishes `link`/`field` (0.98) from `mention` (0.90); confirmation requires an outbound anchor, a structured website field, or a bio mention — never an unrelated occurrence of the domain. Platform redirect wrappers (LinkedIn `redir`, Instagram `l.instagram.com`, Yelp `biz_redir`) are unwrapped; links that stay on a platform host never contradict. Contradictions are softened to "alternate domain" when the site declared the link strongly and the platform's stated name agrees (shopify.com ↔ shopify.engineering). Name agreement/disagreement feeds the score; the platform's stated name outranks handle derivation. Ambiguous platforms (several candidates, none via sameAs/rel=me/meta) capped at 0.70; Wikipedia/Crunchbase/app-store/Amazon/Pinterest capped at 0.70 without confirmation.
- **Robustness:** bounded retries with backoff (5xx/429/timeouts only, honors Retry-After); per-host token-bucket politeness; robots.txt respected for organization sites (documented public endpoints exempt); login-wall redirects classified as `blocked`; JS-only shells and parked domains diagnosed (`website_js_only`, `website_parked`) with optional render fallback Actor; per-query timeout (`seed_timeout`); structural template substitution for sub-Actor inputs.
- **Billing:** honors the user's max charge (`eventChargeLimitReached` stops new work, emits `budget_exhausted`); `verificationLevel` (`confirmed` / `declared` / `none`) and `confirmedLinks` on every record; output contract validated before storage/billing (`internal_contract_violation` is never charged).
- **Operability:** new input `mode`: `healthcheck` (fails the run on degradation, writes `HEALTH`), `reverify` (deterministic graph slices with change events), `export`, `forget` (full purge), `buildSecIndex` (SEC ticker/CIK ↔ domain index, resumable). `RUN_SUMMARY` includes billing counts, http stats and graph stats. Dataset views for entities, changes, diagnostics.
- **Registry:** SEC index enrichment attaches `sec_cik`/`ticker` to public companies by domain; clearer SEC fair-access guidance.
- **Governance:** DATA\_POLICY.md, SLO.md, updated MAINTENANCE.md and HANDOFF.md.
- **Schema:** `schemaVersion` 2 (adds `verificationLevel`, `confirmedLinks`, `verificationStrength`, `siteSignals`, `probeStats`; no removals).
- Tests: 26 (unit, resolver e2e, hardening incl. transport, scoring guards, contract, modes). Live check 2026-09-17: apify.com, shopify.com, hubspot.com, openai.com resolve with LinkedIn/YouTube/Wikidata confirmed from the platform side; all five modes exercised locally.

### 0.1.0 — 2026-09-17

Initial build (see git history / REVIEW\_v0.1.md for what it lacked).
