MCP Agentic Data Transformer avatar

MCP Agentic Data Transformer

Pricing

from $1.00 / 1,000 results

Go to Apify Store
MCP Agentic Data Transformer

MCP Agentic Data Transformer

Pricing

from $1.00 / 1,000 results

Rating

0.0

(0)

Developer

Martin B.

Martin B.

Maintained by Community

Actor stats

0

Bookmarked

1

Total users

0

Monthly active users

11 hours ago

Last modified

Categories

Share

An Apify Python actor that transforms raw HTML or text with a multi-LLM routing architecture. It either cleans the source into Markdown or extracts data using a schema you provide.

Features

  • Native routing for OpenAI GPT-4o-mini
  • Native routing for Anthropic Claude 3.5 Sonnet
  • Native routing for Google Gemini 1.5 Flash
  • Native routing to Cerebras hardware using Qwen
  • Markdown cleanup and schema-based structured extraction

Requirements

  • Python 3.11 (the Docker image is based on apify/actor-python:3.11)
  • A provider API key, or TRIAL_OPENAI_API_KEY for Playground runs
  • Apify CLI for running locally or publishing the actor

Run locally

Install the Apify CLI, then install the Python dependencies:

npm install -g apify-cli
python -m pip install -r requirements.txt

Create storage/key_value_stores/default/INPUT.json with actor input (see examples below), then run:

apify run

The result is stored in the actor's default dataset. You can also run the actor on the Apify platform by pushing it with apify push and supplying the input in the Console or API.

Input

raw_text_html is the preferred source property. raw_html_or_text is also accepted as a backward-compatible fallback. The following provider API keys are optional masked inputs; provide the key for each provider you want the actor to route requests to:

  • openai_api_key
  • anthropic_api_key
  • gemini_api_key
  • cerebras_api_key

The optional model input selects the OpenAI model when the OpenAI route is used. It defaults to gpt-4o-mini; Playground runs always use that model.

Clean to Markdown

Omit extraction_schema to clean and format the input as Markdown:

{
"raw_text_html": "<html><body><h1>Example</h1><p>Source text.</p></body></html>",
"openai_api_key": "YOUR_OPENAI_API_KEY"
}

The dataset item has this shape:

{
"markdown": "# Example\n\nSource text."
}

Structured extraction

Provide extraction_schema to request a JSON object matching the schema. For example:

{
"raw_text_html": "Acme invoice INV-42 dated 2026-09-25, total $125.00.",
"anthropic_api_key": "YOUR_ANTHROPIC_API_KEY",
"extraction_schema": {
"type": "object",
"properties": {
"invoice_number": { "type": "string" },
"date": { "type": "string" },
"total": { "type": "number" }
},
"required": ["invoice_number", "date", "total"],
"additionalProperties": false
}
}

The actor validates the returned object against common JSON Schema constraints, including types, required properties, enums, array items, and additional properties. It also accepts a simple key/type template, for example { "name": "string", "age": "integer" }.

LLM configuration

The actor routes requests natively to the configured provider. Set one or more provider API keys in the input using the optional masked parameters listed above. When multiple providers are configured, the actor selects the appropriate native route for the requested model or provider.

Environment variableDefaultDescription
OPENAI_API_KEYNoneOpenAI API key.
ANTHROPIC_API_KEYNoneAnthropic API key.
GEMINI_API_KEYNoneGoogle Gemini API key.
CEREBRAS_API_KEYNoneCerebras API key.
TRIAL_OPENAI_API_KEYNoneOpenAI key used for Playground runs when no provider key is configured.
LLM_MODELqwen-3.8-27bCerebras model used when the Cerebras route is selected.

Provider-specific environment variables can be used when keys are not supplied in the input. If multiple provider keys are configured, routing priority is OpenAI, Anthropic, Gemini, then Cerebras. Anthropic and Gemini use fixed models; Cerebras uses LLM_MODEL.

When no provider key is configured, the actor uses TRIAL_OPENAI_API_KEY for a Playground run. These runs are limited to 4,000 input characters and always use gpt-4o-mini.

The actor expects each provider's native response to contain the requested Markdown or JSON object.

Publish

Authenticate with Apify, then push the actor from this directory:

apify login
apify push

Output

Each successful run stores one result object in the default dataset. The actor output schema links to the dataset items endpoint. Errors such as missing input, invalid model output, or an unreachable LLM endpoint fail the actor run and appear in its logs.