# API Docs Scraper & Postman Collection Generator (`express_kingfisher/api-docs-scraper-generator`) Actor

Scrape API documentation pages, extract endpoints, parameters, and generate Postman collections or OpenAPI specs.

- **URL**: https://apify.com/express\_kingfisher/api-docs-scraper-generator.md
- **Developed by:** [Prince Raj](https://apify.com/express_kingfisher) (community)
- **Categories:** Developer tools, Automation, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, NaN bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per usage

This Actor is paid per platform usage. The Actor is free to use, and you only pay for the Apify platform usage, which gets cheaper the higher subscription plan you have.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-usage

## What's an Apify Actor?

Actors are a software tools running on the Apify platform, for all kinds of web data extraction and automation use cases.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

In JavaScript/TypeScript projects, use official [JavaScript/TypeScript client](https://docs.apify.com/api/client/js.md):

```bash
npm install apify-client
```

In Python projects, use official [Python client library](https://docs.apify.com/api/client/python.md):

```bash
pip install apify-client
```

In shell scripts, use [Apify CLI](https://docs.apify.com/cli/docs.md):

````bash
# MacOS / Linux
curl -fsSL https://apify.com/install-cli.sh | bash
# Windows
irm https://apify.com/install-cli.ps1 | iex
```bash

In AI frameworks, you might use the [Apify MCP server](https://docs.apify.com/platform/integrations/mcp.md).

If your project is in a different language, use the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).


# README

## API Docs Scraper & Postman Generator

Scrape API documentation pages and generate Postman collections or OpenAPI specs. Built for developers, QA teams, and technical writers who need to work with API documentation.

### What It Extracts

- **Endpoints**: API paths, HTTP methods, descriptions
- **Parameters**: Query params, path params, headers, request bodies
- **Examples**: Request/response examples from documentation
- **Generated output**: Markdown docs, Postman collections, or OpenAPI specs
- **Auth signals**: Authentication requirements and notes

### Why It's Better

Not just scraping - generates ready-to-use Postman collections and OpenAPI specs from documentation pages. Handles multiple documentation formats (REST, GraphQL, code examples).

### Input

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `apiDocUrls` | string[] | required | URLs of API documentation pages |
| `outputFormat` | string | "markdown" | Output format: markdown, postman, openapi-lite |
| `includeExamples` | boolean | true | Include request/response examples |

### Output Example

```json
{
  "sourceUrl": "https://docs.example.com/api",
  "endpoints": [
    {
      "method": "GET",
      "path": "/api/v1/users",
      "description": "List all users",
      "parameters": [
        { "name": "page", "type": "integer", "required": false, "description": "Page number" }
      ],
      "authNote": "Requires Bearer token"
    }
  ],
  "generatedDocs": "## API Documentation\n\n### GET /api/v1/users\nList all users...",
  "endpointCount": 15,
  "pagesScraped": 8,
  "scrapedAt": "2025-01-15T10:30:00Z"
}
````

### Use Cases

- **API integration**: Quickly understand an API before coding
- **Postman setup**: Generate Postman collections from docs
- **API documentation**: Create standardized docs from existing APIs
- **QA testing**: Generate test cases from API documentation
- **Technical writing**: Extract API specs for documentation projects

### PPE Pricing

| Event | Description | Suggested Price |
|-------|-------------|-----------------|
| `api-doc-processed` | One documentation page processed | $0.005 |
| `endpoint-extracted` | Per endpoint extracted | $0.001 |

### Limitations

- Endpoint detection is pattern-based, may miss unconventional formats
- Generated OpenAPI specs are "lite" - not full OpenAPI 3.0 compliance
- Cannot access authenticated documentation pages
- Rate limited to respect documentation sites

### Legal/Ethical Use

This actor only accesses publicly available API documentation. It does not access private APIs, bypass authentication, or scrape rate-limited endpoints. Users are responsible for compliance with documentation site terms of service.

### Local Run

```bash
cd actors/api-docs-scraper-generator
apify run --input-file .actor/sample_input.json
```

### Deploy

```bash
cd actors/api-docs-scraper-generator
apify push
```

### FAQ

**Q: What documentation formats does it support?**
A: It handles common patterns: headings with method+path, code blocks with HTTP methods, table-based endpoint listings.

**Q: Can it generate a complete OpenAPI spec?**
A: The "openapi-lite" output generates a basic spec. Full OpenAPI 3.0 compliance requires manual refinement.

**Q: Does it work with GraphQL documentation?**
A: Limited support. It can extract query/mutation names but not full GraphQL schemas.

### Tags

API documentation scraper, Postman collection generator, OpenAPI generator, API docs, REST API, developer tools, API integration, technical documentation

# Actor input Schema

## `apiDocUrls` (type: `array`):

URLs of API documentation pages to scrape. Supports OpenAPI/Swagger, REST API docs, and developer portals.

## `outputFormat` (type: `string`):

Format for the generated API specification.

## `includeExamples` (type: `boolean`):

Extract and include example requests and responses where available.

## Actor input object example

```json
{
  "apiDocUrls": [
    "https://docs.apify.com/api"
  ],
  "outputFormat": "postman",
  "includeExamples": true
}
```

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "apiDocUrls": [
        "https://docs.apify.com/api"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("express_kingfisher/api-docs-scraper-generator").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "apiDocUrls": ["https://docs.apify.com/api"] }

# Run the Actor and wait for it to finish
run = client.actor("express_kingfisher/api-docs-scraper-generator").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "apiDocUrls": [
    "https://docs.apify.com/api"
  ]
}' |
apify call express_kingfisher/api-docs-scraper-generator --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=express_kingfisher/api-docs-scraper-generator",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

```json
{
    "openapi": "3.0.1",
    "info": {
        "title": "API Docs Scraper & Postman Collection Generator",
        "description": "Scrape API documentation pages, extract endpoints, parameters, and generate Postman collections or OpenAPI specs.",
        "version": "1.0",
        "x-build-id": "NFxtRWGbGjEYpa4g9"
    },
    "servers": [
        {
            "url": "https://api.apify.com/v2"
        }
    ],
    "paths": {
        "/acts/express_kingfisher~api-docs-scraper-generator/run-sync-get-dataset-items": {
            "post": {
                "operationId": "run-sync-get-dataset-items-express_kingfisher-api-docs-scraper-generator",
                "x-openai-isConsequential": false,
                "summary": "Executes an Actor, waits for its completion, and returns Actor's dataset items in response.",
                "tags": [
                    "Run Actor"
                ],
                "requestBody": {
                    "required": true,
                    "content": {
                        "application/json": {
                            "schema": {
                                "$ref": "#/components/schemas/inputSchema"
                            }
                        }
                    }
                },
                "parameters": [
                    {
                        "name": "token",
                        "in": "query",
                        "required": true,
                        "schema": {
                            "type": "string"
                        },
                        "description": "Enter your Apify token here"
                    }
                ],
                "responses": {
                    "200": {
                        "description": "OK"
                    }
                }
            }
        },
        "/acts/express_kingfisher~api-docs-scraper-generator/runs": {
            "post": {
                "operationId": "runs-sync-express_kingfisher-api-docs-scraper-generator",
                "x-openai-isConsequential": false,
                "summary": "Executes an Actor and returns information about the initiated run in response.",
                "tags": [
                    "Run Actor"
                ],
                "requestBody": {
                    "required": true,
                    "content": {
                        "application/json": {
                            "schema": {
                                "$ref": "#/components/schemas/inputSchema"
                            }
                        }
                    }
                },
                "parameters": [
                    {
                        "name": "token",
                        "in": "query",
                        "required": true,
                        "schema": {
                            "type": "string"
                        },
                        "description": "Enter your Apify token here"
                    }
                ],
                "responses": {
                    "200": {
                        "description": "OK",
                        "content": {
                            "application/json": {
                                "schema": {
                                    "$ref": "#/components/schemas/runsResponseSchema"
                                }
                            }
                        }
                    }
                }
            }
        },
        "/acts/express_kingfisher~api-docs-scraper-generator/run-sync": {
            "post": {
                "operationId": "run-sync-express_kingfisher-api-docs-scraper-generator",
                "x-openai-isConsequential": false,
                "summary": "Executes an Actor, waits for completion, and returns the OUTPUT from Key-value store in response.",
                "tags": [
                    "Run Actor"
                ],
                "requestBody": {
                    "required": true,
                    "content": {
                        "application/json": {
                            "schema": {
                                "$ref": "#/components/schemas/inputSchema"
                            }
                        }
                    }
                },
                "parameters": [
                    {
                        "name": "token",
                        "in": "query",
                        "required": true,
                        "schema": {
                            "type": "string"
                        },
                        "description": "Enter your Apify token here"
                    }
                ],
                "responses": {
                    "200": {
                        "description": "OK"
                    }
                }
            }
        }
    },
    "components": {
        "schemas": {
            "inputSchema": {
                "type": "object",
                "required": [
                    "apiDocUrls"
                ],
                "properties": {
                    "apiDocUrls": {
                        "title": "API Documentation URLs",
                        "maxItems": 10,
                        "type": "array",
                        "description": "URLs of API documentation pages to scrape. Supports OpenAPI/Swagger, REST API docs, and developer portals.",
                        "items": {
                            "type": "string"
                        }
                    },
                    "outputFormat": {
                        "title": "Output format",
                        "enum": [
                            "postman",
                            "openapi",
                            "markdown"
                        ],
                        "type": "string",
                        "description": "Format for the generated API specification.",
                        "default": "postman"
                    },
                    "includeExamples": {
                        "title": "Include request/response examples",
                        "type": "boolean",
                        "description": "Extract and include example requests and responses where available.",
                        "default": true
                    }
                }
            },
            "runsResponseSchema": {
                "type": "object",
                "properties": {
                    "data": {
                        "type": "object",
                        "properties": {
                            "id": {
                                "type": "string"
                            },
                            "actId": {
                                "type": "string"
                            },
                            "userId": {
                                "type": "string"
                            },
                            "startedAt": {
                                "type": "string",
                                "format": "date-time",
                                "example": "2025-01-08T00:00:00.000Z"
                            },
                            "finishedAt": {
                                "type": "string",
                                "format": "date-time",
                                "example": "2025-01-08T00:00:00.000Z"
                            },
                            "status": {
                                "type": "string",
                                "example": "READY"
                            },
                            "meta": {
                                "type": "object",
                                "properties": {
                                    "origin": {
                                        "type": "string",
                                        "example": "API"
                                    },
                                    "userAgent": {
                                        "type": "string"
                                    }
                                }
                            },
                            "stats": {
                                "type": "object",
                                "properties": {
                                    "inputBodyLen": {
                                        "type": "integer",
                                        "example": 2000
                                    },
                                    "rebootCount": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "restartCount": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "resurrectCount": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "computeUnits": {
                                        "type": "integer",
                                        "example": 0
                                    }
                                }
                            },
                            "options": {
                                "type": "object",
                                "properties": {
                                    "build": {
                                        "type": "string",
                                        "example": "latest"
                                    },
                                    "timeoutSecs": {
                                        "type": "integer",
                                        "example": 300
                                    },
                                    "memoryMbytes": {
                                        "type": "integer",
                                        "example": 1024
                                    },
                                    "diskMbytes": {
                                        "type": "integer",
                                        "example": 2048
                                    }
                                }
                            },
                            "buildId": {
                                "type": "string"
                            },
                            "defaultKeyValueStoreId": {
                                "type": "string"
                            },
                            "defaultDatasetId": {
                                "type": "string"
                            },
                            "defaultRequestQueueId": {
                                "type": "string"
                            },
                            "buildNumber": {
                                "type": "string",
                                "example": "1.0.0"
                            },
                            "containerUrl": {
                                "type": "string"
                            },
                            "usage": {
                                "type": "object",
                                "properties": {
                                    "ACTOR_COMPUTE_UNITS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATASET_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATASET_WRITES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "KEY_VALUE_STORE_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "KEY_VALUE_STORE_WRITES": {
                                        "type": "integer",
                                        "example": 1
                                    },
                                    "KEY_VALUE_STORE_LISTS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "REQUEST_QUEUE_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "REQUEST_QUEUE_WRITES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATA_TRANSFER_INTERNAL_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATA_TRANSFER_EXTERNAL_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "PROXY_RESIDENTIAL_TRANSFER_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "PROXY_SERPS": {
                                        "type": "integer",
                                        "example": 0
                                    }
                                }
                            },
                            "usageTotalUsd": {
                                "type": "number",
                                "example": 0.00005
                            },
                            "usageUsd": {
                                "type": "object",
                                "properties": {
                                    "ACTOR_COMPUTE_UNITS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATASET_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATASET_WRITES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "KEY_VALUE_STORE_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "KEY_VALUE_STORE_WRITES": {
                                        "type": "number",
                                        "example": 0.00005
                                    },
                                    "KEY_VALUE_STORE_LISTS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "REQUEST_QUEUE_READS": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "REQUEST_QUEUE_WRITES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATA_TRANSFER_INTERNAL_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "DATA_TRANSFER_EXTERNAL_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "PROXY_RESIDENTIAL_TRANSFER_GBYTES": {
                                        "type": "integer",
                                        "example": 0
                                    },
                                    "PROXY_SERPS": {
                                        "type": "integer",
                                        "example": 0
                                    }
                                }
                            }
                        }
                    }
                }
            }
        }
    }
}
```
