MCP Agentic Data Transformer
Pricing
from $1.00 / 1,000 results
MCP Agentic Data Transformer
Pricing
from $1.00 / 1,000 results
Rating
0.0
(0)
Developer
Martin B.
Maintained by CommunityActor stats
0
Bookmarked
1
Total users
0
Monthly active users
11 hours ago
Last modified
Categories
Share
An Apify Python actor that transforms raw HTML or text with a multi-LLM routing architecture. It either cleans the source into Markdown or extracts data using a schema you provide.
Features
- Native routing for OpenAI GPT-4o-mini
- Native routing for Anthropic Claude 3.5 Sonnet
- Native routing for Google Gemini 1.5 Flash
- Native routing to Cerebras hardware using Qwen
- Markdown cleanup and schema-based structured extraction
Requirements
- Python 3.11 (the Docker image is based on
apify/actor-python:3.11) - A provider API key, or
TRIAL_OPENAI_API_KEYfor Playground runs - Apify CLI for running locally or publishing the actor
Run locally
Install the Apify CLI, then install the Python dependencies:
npm install -g apify-clipython -m pip install -r requirements.txt
Create storage/key_value_stores/default/INPUT.json with actor input (see examples below), then run:
apify run
The result is stored in the actor's default dataset. You can also run the actor on the Apify platform by pushing it with apify push and supplying the input in the Console or API.
Input
raw_text_html is the preferred source property. raw_html_or_text is also accepted as a backward-compatible fallback. The following provider API keys are optional masked inputs; provide the key for each provider you want the actor to route requests to:
openai_api_keyanthropic_api_keygemini_api_keycerebras_api_key
The optional model input selects the OpenAI model when the OpenAI route is used. It defaults to gpt-4o-mini; Playground runs always use that model.
Clean to Markdown
Omit extraction_schema to clean and format the input as Markdown:
{"raw_text_html": "<html><body><h1>Example</h1><p>Source text.</p></body></html>","openai_api_key": "YOUR_OPENAI_API_KEY"}
The dataset item has this shape:
{"markdown": "# Example\n\nSource text."}
Structured extraction
Provide extraction_schema to request a JSON object matching the schema. For example:
{"raw_text_html": "Acme invoice INV-42 dated 2026-09-25, total $125.00.","anthropic_api_key": "YOUR_ANTHROPIC_API_KEY","extraction_schema": {"type": "object","properties": {"invoice_number": { "type": "string" },"date": { "type": "string" },"total": { "type": "number" }},"required": ["invoice_number", "date", "total"],"additionalProperties": false}}
The actor validates the returned object against common JSON Schema constraints, including types, required properties, enums, array items, and additional properties. It also accepts a simple key/type template, for example { "name": "string", "age": "integer" }.
LLM configuration
The actor routes requests natively to the configured provider. Set one or more provider API keys in the input using the optional masked parameters listed above. When multiple providers are configured, the actor selects the appropriate native route for the requested model or provider.
| Environment variable | Default | Description |
|---|---|---|
OPENAI_API_KEY | None | OpenAI API key. |
ANTHROPIC_API_KEY | None | Anthropic API key. |
GEMINI_API_KEY | None | Google Gemini API key. |
CEREBRAS_API_KEY | None | Cerebras API key. |
TRIAL_OPENAI_API_KEY | None | OpenAI key used for Playground runs when no provider key is configured. |
LLM_MODEL | qwen-3.8-27b | Cerebras model used when the Cerebras route is selected. |
Provider-specific environment variables can be used when keys are not supplied in the input. If multiple provider keys are configured, routing priority is OpenAI, Anthropic, Gemini, then Cerebras. Anthropic and Gemini use fixed models; Cerebras uses LLM_MODEL.
When no provider key is configured, the actor uses TRIAL_OPENAI_API_KEY for a Playground run. These runs are limited to 4,000 input characters and always use gpt-4o-mini.
The actor expects each provider's native response to contain the requested Markdown or JSON object.
Publish
Authenticate with Apify, then push the actor from this directory:
apify loginapify push
Output
Each successful run stores one result object in the default dataset. The actor output schema links to the dataset items endpoint. Errors such as missing input, invalid model output, or an unreachable LLM endpoint fail the actor run and appear in its logs.