# Mojibake Text Repair (`automation-lab/mojibake-text-repair`) Actor

Repair garbled Unicode strings or selected text fields in supplied JSON records. Export corrected records, preserved identities, changed flags and bounded repair summaries with ftfy; no scraping or remote file access.

- **URL**: https://apify.com/automation-lab/mojibake-text-repair.md
- **Developed by:** [Automation Lab](https://apify.com/automation-lab) (community)
- **Categories:** Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.96 / 1,000 item processeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Mojibake Text Repair

Use mojibake text repair to fix garbled Unicode strings from scraped datasets and legacy exports before analysis or database ingestion. Supply text strings or JSON records and receive corrected text, original identities, changed flags and bounded repair explanations. Processing uses **ftfy 6.3.1** offline: no scraping, translation, text generation, URLs or remote file downloads.

### Who is this for?

Data engineers cleaning recurring ingestion batches, analysts reviewing legacy exports, and scraping teams repairing encoding mistakes without losing unrelated fields or record identity.

### Why use this Actor?

- Keep a before/after result for every selected field, including unchanged text.
- Repair repeated mojibake such as `cafÃƒÂ©` → `café`.
- Preserve unselected record properties and original IDs.
- Choose conservative repair or additional cleanup explicitly.
- Export results using Apify's standard dataset formats.

This is not a general HTML sanitizer or a language model. It does not strip HTML tags, infer missing words or reconstruct bytes already replaced with `�`.

### Getting started

1. Open Input and use the supplied `texts` example.
2. Start the Actor.
3. Open the Repair results dataset and inspect `changed` and `repairs`.
4. Download JSON for nested records or integrate the dataset into your pipeline.

```json
{"texts":["FranÃ§ais","Already clean 日本語"]}
```

### Repair JSON records

```json
{
  "records":[{"id":"sample-1","title":"cafÃ©","count":3}],
  "fields":["title"],
  "idField":"id"
}
```

Supply exactly one of `texts` and `records`. `fields` is required for records and rejected for strings. Selection uses exact, case-sensitive **top-level keys**. Dots are literal, not nested paths. Selected values must be strings; missing, null, array and numeric values fail validation. The entire supplied batch is validated before repair, billing or output.

### Input parameters

| Parameter | Behavior |
|---|---|
| `texts` | 1–1000 strings; empty strings are permitted |
| `records` | 1–1000 JSON objects, mutually exclusive with texts |
| `fields` | 1–20 distinct selected top-level string keys for records |
| `idField` | Optional original string/number identity; otherwise zero-based index |
| `cleanText` | Defaults false; enables extra cleanup when true |
| `normalization` | NFC default, NFKC compatibility folding, or none |
| `maxItems` | Global output-record limit 1–1000, default 1000 |

Unknown options are rejected. Duplicate records and identities are preserved in input order. One output row corresponds to one supplied entry, even with several selected fields. The first `maxItems` entries are processed; SUMMARY reports omitted entries. There is no unlimited mode.

### Repair controls

Encoding repair and surrogate repair are always enabled. With `cleanText:false`, quote straightening, ligature expansion, width folding, line-break cleanup, terminal-escape removal and control-character removal are disabled. NFC normalization still applies unless you select `none`.

With `cleanText:true`, those additional ftfy transformations are enabled. NFKC can alter intentional compatibility characters, so inspect changes before updating your source data. HTML entities are never unescaped. No whitespace trimming, translation or HTML tag removal is performed.

### Output fields

| Field | Meaning |
|---|---|
| `index` | Zero-based supplied array position |
| `recordId` | Original identity or index |
| `changed` | At least one selected string differs |
| `changedFieldCount` | Number of selected changed fields |
| `record` | Corrected object with unselected values preserved; strings use `text` |
| `repairs` | Before/after values, exact changed flag, explanation labels |

Each repair includes `field`, `originalText`, `correctedText`, `changed`, `repairSteps` and `summaryTruncated`. Steps are action/parameter labels, not confidence scores. At most 20 explanation steps are exported per field; corrected strings are not truncated.

Example from the local prefilled-input run:

```json
{
  "index":0,
  "recordId":0,
  "changed":true,
  "changedFieldCount":1,
  "record":{"text":"Français"},
  "repairs":[{
    "field":"text",
    "originalText":"FranÃ§ais",
    "correctedText":"Français",
    "changed":true,
    "repairSteps":["encode:latin-1","decode:utf-8"],
    "summaryTruncated":false
  }]
}
```

The SUMMARY key-value record reports supplied, processed and omitted counts, selected-field count, engine version and active controls.

### How much does it cost to repair mojibake text?

Pricing is per completed output record, including unchanged records, plus a one-time start fee. Multiple selected fields within a record have no separate event charge. The start fee is **$0.0005 per run**. Per-record prices are FREE $0.00184, BRONZE $0.0016, SILVER $0.001248, and GOLD/PLATINUM/DIAMOND $0.00096. Spend tiers depend on qualifying aggregate monthly Apify Store spend, not this Actor's batch size.

BRONZE examples: 1 record ≈ $0.0021; 10 records ≈ $0.0165; 100 records ≈ $0.1605. These are estimated event totals, not guaranteed invoices; applicable platform limits and refunds may change the final bill.

### Limits and failure behavior

The whole serialized JSON input must be at most 1 MiB (UTF-8), including unselected fields. Every selected string must contain at most 100000 JavaScript UTF-16 code units; the sum across **all supplied entries**, including entries beyond `maxItems`, must not exceed 200000. Up to 1000 records and 20 selected fields are supported. The repair worker has a 30-second deadline and fails closed on worker errors. Unknown/malformed inputs produce a failed run, not empty success.

If the configured run charge limit is exhausted, the run fails explicitly and any dataset rows are partial; increase the charge limit before retrying.

ftfy applies conservative heuristics, not certainty about author intent. Already clean multilingual text normally remains unchanged, but normalization may legitimately change its representation. Replacement characters and deleted bytes cannot be recovered. Validate samples before overwriting valuable originals. A storage/billing/platform error after output begins can leave partial rows; a failed run is not a complete batch.

### Integrations

Use JSON exports for nested `record` and `repairs` objects. Join by `recordId` or `index`, review changed rows, and write only approved corrected fields back to your database. For recurring ingestion, provide each fresh batch through the API; this Actor does not discover new records, monitor a website or maintain cross-run state. Retain original data outside the Actor for audit and rollback.

### API usage

Use your Apify token securely; do not put it in supplied text.

```bash
curl -X POST 'https://api.apify.com/v2/acts/automation-lab~mojibake-text-repair/run-sync-get-dataset-items' \
  -H "Authorization: Bearer $APIFY_TOKEN" \
  -H 'Content-Type: application/json' \
  -d '{"texts":["cafÃ©"]}'
```

JavaScript with `apify-client`:

```javascript
import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('automation-lab/mojibake-text-repair')
  .call({ texts: ['cafÃ©'] });
const { items } = await client.dataset(run.defaultDatasetId).listItems({ limit: 20 });
```

Python with `apify-client`:

```python
import os
from apify_client import ApifyClient
client = ApifyClient(os.environ['APIFY_TOKEN'])
run = client.actor('automation-lab/mojibake-text-repair').call(
    run_input={'texts': ['cafÃ©']})
items = client.dataset(run['defaultDatasetId']).list_items(limit=20).items
```

### MCP usage

```bash
claude mcp add --transport http apify \
  'https://mcp.apify.com?tools=automation-lab/mojibake-text-repair'
```

Claude Desktop, Cursor and VS Code configuration (for clients supporting HTTP MCP):

```json
{"mcpServers":{"apify":{"url":"https://mcp.apify.com?tools=automation-lab/mojibake-text-repair"}}}
```

Discover actual tool names and input schemas with scoped `tools/list`. Actor selection also exposes run/data helpers, not exactly one tool. Example prompt: “Repair these supplied strings without additional cleanup, then show changed rows.” Start once and retain the run ID. If pending, check that same run with bounded backoff (2, 4, 8, capped at 10 seconds) for at most 120 seconds; on timeout report pending rather than restart. After success, read pages of 20 untransformed dataset rows, at most 100 rows and 64 KiB of serialized UTF-8 content admitted to model context, whichever comes first. Enforce bytes host-side, disclose truncation and continuation offset, and keep full exports outside model context. Clients unable to intercept oversized results cannot guarantee that byte cap. Discovery is read-only, not permission to run or abort.

### Legality, privacy and responsible use

Only process records you are authorized to handle. Inputs and datasets are stored by Apify under your run's access settings; outputs include original text and unselected record fields. Do not supply secrets or unnecessary personal data. Review export permissions and retention before sharing. Apify storage persists until your account's retention policy or explicit deletion removes it; this Actor does not automatically delete inputs or outputs. Delete runs and their storage through Apify when no longer needed. Report issues through the Actor's Apify Store Issues tab.

Failed operations send sanitized diagnostic input, exceptions and actor/build/run IDs to our private GlitchTip service for repair. Secret fields, email-like addresses and URL queries are removed; reports are retained for 30 days. Sanitization does not remove all ordinary supplied text or personal data. No text is sent to a generative model or external conversion service.

### FAQ and troubleshooting

**Why was my text unchanged?** It may already be valid Unicode, or its corruption is not recoverable by ftfy. `changed:false` is a useful successful result, not a failure.

**Can this decode arbitrary files?** No. Supply Unicode strings in JSON, not paths, remote URLs, byte buffers or base64 files.

**Why did a record fail?** Every selected key must exist and contain a string in every supplied record, even beyond the output cap. Convert non-text fields upstream rather than silently coercing them.

**Does this fix Japanese text?** It preserves valid Japanese text and can repair detectable encoding mistakes. It does not translate or promise recovery of all Japanese encoding damage.

**Can I use nested fields?** No. Flatten them upstream or select a literal top-level key containing dots.

### Related workflows

Use [HTML Readability Markdown Converter](https://apify.com/automation-lab/html-readability-markdown-converter) for complementary readable-document conversion. Mojibake repair is a separate supplied-data cleanup step, not a substitute for extraction or document conversion.

# Changelog

This Actor's version history is a separate document: https://apify.com/automation-lab/mojibake-text-repair/changelog.md

# Actor input Schema

## `texts` (type: `array`):

1–1000 supplied strings, including empty strings. Mutually exclusive with records. Each string is at most 100000 JavaScript UTF-16 code units; total selected text across all supplied entries is at most 200000.

## `records` (type: `array`):

1–1000 JSON objects. Mutually exclusive with texts. Requires fields. Unselected properties are preserved; nested paths are not supported.

## `fields` (type: `array`):

Required with records; rejected with texts. 1–20 distinct case-sensitive top-level keys, each present with a string value in every record. Dots are literal key characters, not path separators. Missing or non-string values fail the whole run.

## `idField` (type: `string`):

Optional with records only. Exact case-sensitive top-level key containing a string or number in every record. Identity is taken from the original record. Otherwise recordId is the zero-based input index. Duplicate identities are preserved, not deduplicated.

## `cleanText` (type: `boolean`):

Both modes: when true, also straighten quotes, expand Latin ligatures, normalize character width and line breaks, and remove terminal escapes and control characters. When false these transformations are disabled. HTML entities are never unescaped. Unicode normalization is controlled separately.

## `normalization` (type: `string`):

Both modes: NFC composes canonical characters; NFKC also folds compatibility characters and may alter intentional formatting. none disables normalization. This does not disable mojibake repair or surrogate repair.

## `maxItems` (type: `integer`):

Global output-record cap, not text-field count. Outputs the first N entries in supplied order without deduplication. All supplied entries are validated first, including omitted entries. Zero/unlimited is unsupported. SUMMARY reports omitted entries.

## Actor input object example

```json
{
  "texts": [
    "FranÃ§ais",
    "Already clean 日本語"
  ],
  "cleanText": false,
  "normalization": "NFC",
  "maxItems": 20
}
```

# Actor output Schema

## `dataset` (type: `string`):

Corrected records with identity, before/after text and change flags.

## `summary` (type: `string`):

Supplied, processed and omitted counts with active repair controls.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "texts": [
        "FranÃ§ais",
        "Already clean 日本語"
    ],
    "maxItems": 20
};

// Run the Actor and wait for it to finish
const run = await client.actor("automation-lab/mojibake-text-repair").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "texts": [
        "FranÃ§ais",
        "Already clean 日本語",
    ],
    "maxItems": 20,
}

# Run the Actor and wait for it to finish
run = client.actor("automation-lab/mojibake-text-repair").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "texts": [
    "FranÃ§ais",
    "Already clean 日本語"
  ],
  "maxItems": 20
}' |
apify call automation-lab/mojibake-text-repair --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,automation-lab/mojibake-text-repair"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/CqrkK1eMlka6ED63J/builds/DG43Zgai6HV14xnEs/openapi.json
