# Hacked Website & SEO Spam Checker (casino/pharma injection) (`weio/hacked-website-seo-spam-checker`) Actor

Bulk check sites for public signs of an SEO spam hack, the most common WordPress website malware: casino/pharma spam, CSS-hidden links, Japanese keyword hack sitemap URLs, WordPress spam posts, the ndsw loader. Strong vs weak evidence, so casinos or pharmacies are not called hacked. Read-only.

- **URL**: https://apify.com/weio/hacked-website-seo-spam-checker.md
- **Developed by:** [Weio, Inc.](https://apify.com/weio) (community)
- **Categories:** SEO tools, Developer tools, Lead generation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.50 / 1,000 site checkeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Hacked Website & SEO Spam Checker

Check a list of websites for the most common kind of hack: **search spam injected into a legitimate site**. Attackers add casino, slot, pharma (viagra/cialis) or essay-mill text and links to small-business sites, often hidden from visitors and shown only to Google. The owner rarely notices until rankings drop or Google flags the site.

This actor reads only public pages (homepage, robots.txt, sitemaps and, on WordPress, the public post search). No login, no vulnerability probing, nothing is changed on the site.

### What it checks

- **Visible spam phrases** on the homepage (online casino, slot gacor, situs slot, free spins, replica watches, "buy viagra online" and more, plus Thai, Chinese, Japanese, Korean and Vietnamese gambling terms)
- **Spam links** to casino, slot, betting or pharma sites, and spam doorway subdomains such as `slot88.yoursite.com`
- **Hidden spam**: text placed in `display:none`, zero-size or off-screen blocks
- **Sitemap injection**: casino/pharma/slot or Japanese spam URLs listed in the site's sitemaps (up to 15 sitemap files per site)
- **Spam posts on WordPress**: casino or pharma posts the attacker published, found through WordPress's public post search
- **Known malicious loader**: the `ndsw` script injection used to redirect visitors

### Two levels, so legitimate sites are not called hacked

Words like *casino*, *poker* or *viagra* are normal on many real businesses: a casino-night rental, a poker run, a pharmacy, a news story. So the actor separates:

- `likely_hacked: true` — **strong evidence**: unmistakable spam vocabulary, spam hidden with CSS, links to several unrelated casino sites, many spam sitemap URLs or several spam posts. Each reason is in `issues`.
- `suspicious: true` — anything worth a human look, with the reasons in `warnings` (for example "one spam-like phrase on the homepage" or "3 sitemap URLs mention gambling words; usually legitimate"). Every likely-hacked site is also suspicious.

Before release we ran it on real sites, reading them exactly as it does for you (as WeioBot, within robots.txt). Of 66 small-business sites our team had confirmed hacked by hand, 61 could be checked and it marked **59 of those 61** as likely hacked; the other 5 refused automated visitors or their robots.txt asked crawlers to stay away, so they came back as free "not checked" rows. Of 434 random small-business websites it could check, it marked **6** as likely hacked, and on review all 6 really were serving gambling spam. Treat a result as a lead to confirm, not a verdict.

### Who uses it

- Web and SEO agencies screening prospects or monitoring client sites
- Security and hosting teams triaging lists of customer domains
- Anyone auditing a portfolio of sites for compromise

### Input

```json
{ "websites": ["example.com", "https://www.example.org"], "checkSitemaps": true, "checkWordPress": true, "onlyLikelyHacked": false }
```

Duplicates (`example.com` and `https://example.com/`) are checked and charged once.

### Output (one row per site)

| field | meaning |
|---|---|
| `domain`, `final_url`, `status`, `tls`, `platform` | where the homepage ended up, whether its certificate is valid, and the site builder (WordPress, Wix, Shopify, Squarespace...) when recognisable |
| `likely_hacked` | true when there is strong evidence; the reasons are in `issues` |
| `suspicious` | true when anything is worth a look; weaker reasons are in `warnings` |
| `spam_score` | rough count of spam signals (0 = clean) |
| `visible_spam_terms`, `spam_links`, `hidden_spam_text`, `weak_links`, `comment_spam_links` | the matched evidence (capped) |
| `sitemap_urls_checked`, `sitemap_spam_count`, `sitemap_spam_examples`, `sitemap_weak_examples` | sitemap evidence |
| `wordpress_spam_posts`, `wordpress_weak_posts` | title and link of matching public posts |
| `parked`, `redirects_to`, `checks_skipped` | parked or for-sale domains, sites that send every visitor elsewhere, and checks a bot wall prevented |

Example (a real small-business site): `likely_hacked: true`, `issues: ["spam phrases visible on the homepage: online casinos, free spins", "7 spam link(s) on the homepage"]`.

### Pricing

Pay per checked site. Invalid, unreachable and duplicate sites, sites that answer with an error page or a bot check, and sites whose robots.txt does not allow WeioBot are not charged; with `onlyLikelyHacked` clean and suspicious-only sites are skipped and not charged.

### Limits

A pattern scanner, not a malware scanner: it finds search-spam injection, the most common small-business hack, and will not see server-side backdoors that leave no public trace. A clean result means no public spam signal was found. It identifies itself as WeioBot and follows each site's robots.txt and crawl delay, so a site that blocks crawlers comes back as a free "not checked" row, and some optional checks may be skipped (listed in `checks_skipped`). Built by Weio, Inc.

# Actor input Schema

## `websites` (type: `array`):

Domains or URLs, one per line (max 5,000 per run). Only the homepage, robots.txt and sitemaps are read.

## `checkSitemaps` (type: `boolean`):

Also read robots.txt and up to 15 sitemap files per site and count spam URLs (casino, pharma, slot pages) injected into them.

## `checkWordPress` (type: `boolean`):

On WordPress sites, search the public post list (no login) for spam posts such as casino or pharma articles the attacker published. Up to 5 small requests per WordPress site.

## `onlyLikelyHacked` (type: `boolean`):

Output only sites with strong evidence (likely\_hacked). Clean and suspicious-only sites are skipped and not charged.

## `concurrency` (type: `integer`):

How many sites are checked at the same time (1-32).

## Actor input object example

```json
{
  "websites": [
    "weio.ai"
  ],
  "checkSitemaps": true,
  "checkWordPress": true,
  "onlyLikelyHacked": false,
  "concurrency": 8
}
```

# Actor output Schema

## `siteRows` (type: `string`):

The default dataset with one row per checked website.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "websites": [
        "weio.ai"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("weio/hacked-website-seo-spam-checker").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "websites": ["weio.ai"] }

# Run the Actor and wait for it to finish
run = client.actor("weio/hacked-website-seo-spam-checker").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "websites": [
    "weio.ai"
  ]
}' |
apify call weio/hacked-website-seo-spam-checker --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,weio/hacked-website-seo-spam-checker"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/ZuD3CXPcYYEUO2bmY/builds/LhKKvBjXhdwO1LCVO/openapi.json
