Mojibake Text Repair
Pricing
from $0.96 / 1,000 item processeds
Mojibake Text Repair
Repair garbled Unicode strings or selected text fields in supplied JSON records. Export corrected records, preserved identities, changed flags and bounded repair summaries with ftfy; no scraping or remote file access.
Pricing
from $0.96 / 1,000 item processeds
Rating
0.0
(0)
Developer
Automation Lab
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
a day ago
Last modified
Categories
Share
Use mojibake text repair to fix garbled Unicode strings from scraped datasets and legacy exports before analysis or database ingestion. Supply text strings or JSON records and receive corrected text, original identities, changed flags and bounded repair explanations. Processing uses ftfy 6.3.1 offline: no scraping, translation, text generation, URLs or remote file downloads.
Who is this for?
Data engineers cleaning recurring ingestion batches, analysts reviewing legacy exports, and scraping teams repairing encoding mistakes without losing unrelated fields or record identity.
Why use this Actor?
- Keep a before/after result for every selected field, including unchanged text.
- Repair repeated mojibake such as
café→café. - Preserve unselected record properties and original IDs.
- Choose conservative repair or additional cleanup explicitly.
- Export results using Apify's standard dataset formats.
This is not a general HTML sanitizer or a language model. It does not strip HTML tags, infer missing words or reconstruct bytes already replaced with �.
Getting started
- Open Input and use the supplied
textsexample. - Start the Actor.
- Open the Repair results dataset and inspect
changedandrepairs. - Download JSON for nested records or integrate the dataset into your pipeline.
{"texts":["Français","Already clean 日本語"]}
Repair JSON records
{"records":[{"id":"sample-1","title":"café","count":3}],"fields":["title"],"idField":"id"}
Supply exactly one of texts and records. fields is required for records and rejected for strings. Selection uses exact, case-sensitive top-level keys. Dots are literal, not nested paths. Selected values must be strings; missing, null, array and numeric values fail validation. The entire supplied batch is validated before repair, billing or output.
Input parameters
| Parameter | Behavior |
|---|---|
texts | 1–1000 strings; empty strings are permitted |
records | 1–1000 JSON objects, mutually exclusive with texts |
fields | 1–20 distinct selected top-level string keys for records |
idField | Optional original string/number identity; otherwise zero-based index |
cleanText | Defaults false; enables extra cleanup when true |
normalization | NFC default, NFKC compatibility folding, or none |
maxItems | Global output-record limit 1–1000, default 1000 |
Unknown options are rejected. Duplicate records and identities are preserved in input order. One output row corresponds to one supplied entry, even with several selected fields. The first maxItems entries are processed; SUMMARY reports omitted entries. There is no unlimited mode.
Repair controls
Encoding repair and surrogate repair are always enabled. With cleanText:false, quote straightening, ligature expansion, width folding, line-break cleanup, terminal-escape removal and control-character removal are disabled. NFC normalization still applies unless you select none.
With cleanText:true, those additional ftfy transformations are enabled. NFKC can alter intentional compatibility characters, so inspect changes before updating your source data. HTML entities are never unescaped. No whitespace trimming, translation or HTML tag removal is performed.
Output fields
| Field | Meaning |
|---|---|
index | Zero-based supplied array position |
recordId | Original identity or index |
changed | At least one selected string differs |
changedFieldCount | Number of selected changed fields |
record | Corrected object with unselected values preserved; strings use text |
repairs | Before/after values, exact changed flag, explanation labels |
Each repair includes field, originalText, correctedText, changed, repairSteps and summaryTruncated. Steps are action/parameter labels, not confidence scores. At most 20 explanation steps are exported per field; corrected strings are not truncated.
Example from the local prefilled-input run:
{"index":0,"recordId":0,"changed":true,"changedFieldCount":1,"record":{"text":"Français"},"repairs":[{"field":"text","originalText":"Français","correctedText":"Français","changed":true,"repairSteps":["encode:latin-1","decode:utf-8"],"summaryTruncated":false}]}
The SUMMARY key-value record reports supplied, processed and omitted counts, selected-field count, engine version and active controls.
How much does it cost to repair mojibake text?
Pricing is per completed output record, including unchanged records, plus a one-time start fee. Multiple selected fields within a record have no separate event charge. The start fee is $0.0005 per run. Per-record prices are FREE $0.00184, BRONZE $0.0016, SILVER $0.001248, and GOLD/PLATINUM/DIAMOND $0.00096. Spend tiers depend on qualifying aggregate monthly Apify Store spend, not this Actor's batch size.
BRONZE examples: 1 record ≈ $0.0021; 10 records ≈ $0.0165; 100 records ≈ $0.1605. These are estimated event totals, not guaranteed invoices; applicable platform limits and refunds may change the final bill.
Limits and failure behavior
The whole serialized JSON input must be at most 1 MiB (UTF-8), including unselected fields. Every selected string must contain at most 100000 JavaScript UTF-16 code units; the sum across all supplied entries, including entries beyond maxItems, must not exceed 200000. Up to 1000 records and 20 selected fields are supported. The repair worker has a 30-second deadline and fails closed on worker errors. Unknown/malformed inputs produce a failed run, not empty success.
If the configured run charge limit is exhausted, the run fails explicitly and any dataset rows are partial; increase the charge limit before retrying.
ftfy applies conservative heuristics, not certainty about author intent. Already clean multilingual text normally remains unchanged, but normalization may legitimately change its representation. Replacement characters and deleted bytes cannot be recovered. Validate samples before overwriting valuable originals. A storage/billing/platform error after output begins can leave partial rows; a failed run is not a complete batch.
Integrations
Use JSON exports for nested record and repairs objects. Join by recordId or index, review changed rows, and write only approved corrected fields back to your database. For recurring ingestion, provide each fresh batch through the API; this Actor does not discover new records, monitor a website or maintain cross-run state. Retain original data outside the Actor for audit and rollback.
API usage
Use your Apify token securely; do not put it in supplied text.
curl -X POST 'https://api.apify.com/v2/acts/automation-lab~mojibake-text-repair/run-sync-get-dataset-items' \-H "Authorization: Bearer $APIFY_TOKEN" \-H 'Content-Type: application/json' \-d '{"texts":["café"]}'
JavaScript with apify-client:
import { ApifyClient } from 'apify-client';const client = new ApifyClient({ token: process.env.APIFY_TOKEN });const run = await client.actor('automation-lab/mojibake-text-repair').call({ texts: ['café'] });const { items } = await client.dataset(run.defaultDatasetId).listItems({ limit: 20 });
Python with apify-client:
import osfrom apify_client import ApifyClientclient = ApifyClient(os.environ['APIFY_TOKEN'])run = client.actor('automation-lab/mojibake-text-repair').call(run_input={'texts': ['café']})items = client.dataset(run['defaultDatasetId']).list_items(limit=20).items
MCP usage
claude mcp add --transport http apify \'https://mcp.apify.com?tools=automation-lab/mojibake-text-repair'
Claude Desktop, Cursor and VS Code configuration (for clients supporting HTTP MCP):
{"mcpServers":{"apify":{"url":"https://mcp.apify.com?tools=automation-lab/mojibake-text-repair"}}}
Discover actual tool names and input schemas with scoped tools/list. Actor selection also exposes run/data helpers, not exactly one tool. Example prompt: “Repair these supplied strings without additional cleanup, then show changed rows.” Start once and retain the run ID. If pending, check that same run with bounded backoff (2, 4, 8, capped at 10 seconds) for at most 120 seconds; on timeout report pending rather than restart. After success, read pages of 20 untransformed dataset rows, at most 100 rows and 64 KiB of serialized UTF-8 content admitted to model context, whichever comes first. Enforce bytes host-side, disclose truncation and continuation offset, and keep full exports outside model context. Clients unable to intercept oversized results cannot guarantee that byte cap. Discovery is read-only, not permission to run or abort.
Legality, privacy and responsible use
Only process records you are authorized to handle. Inputs and datasets are stored by Apify under your run's access settings; outputs include original text and unselected record fields. Do not supply secrets or unnecessary personal data. Review export permissions and retention before sharing. Apify storage persists until your account's retention policy or explicit deletion removes it; this Actor does not automatically delete inputs or outputs. Delete runs and their storage through Apify when no longer needed. Report issues through the Actor's Apify Store Issues tab.
Failed operations send sanitized diagnostic input, exceptions and actor/build/run IDs to our private GlitchTip service for repair. Secret fields, email-like addresses and URL queries are removed; reports are retained for 30 days. Sanitization does not remove all ordinary supplied text or personal data. No text is sent to a generative model or external conversion service.
FAQ and troubleshooting
Why was my text unchanged? It may already be valid Unicode, or its corruption is not recoverable by ftfy. changed:false is a useful successful result, not a failure.
Can this decode arbitrary files? No. Supply Unicode strings in JSON, not paths, remote URLs, byte buffers or base64 files.
Why did a record fail? Every selected key must exist and contain a string in every supplied record, even beyond the output cap. Convert non-text fields upstream rather than silently coercing them.
Does this fix Japanese text? It preserves valid Japanese text and can repair detectable encoding mistakes. It does not translate or promise recovery of all Japanese encoding damage.
Can I use nested fields? No. Flatten them upstream or select a literal top-level key containing dots.
Related workflows
Use HTML Readability Markdown Converter for complementary readable-document conversion. Mojibake repair is a separate supplied-data cleanup step, not a substitute for extraction or document conversion.