# Changelog of Website Content Crawler (`apify/website-content-crawler`) Actor

- **URL**: https://apify.com/apify/website-content-crawler/changelog.md
- **Full Actor documentation**: https://apify.com/apify/website-content-crawler.md

## Changelog

All notable changes to this project will be documented in this file.

### 0.3.97 (2026-09-08)

### 0.3.96 (2026-09-07)

#### 🐛 Bug Fixes

- Yield between clicks  to avoid site deadlocks (#905)
- Retry XML navigation on load after a domcontentloaded timeout (#900)
- Discover blocked sitemaps (#903)
- Deduplicate pages with uppercase ETags (#920)

### 0.3.95 (2026-08-25)

#### 🐛 Bug Fixes

- Keep the requested page when it navigates away (#884)

### 0.3.94 (2026-08-10)

#### 🚀 Features

- Log how to fix certificate and robots.txt failures (#863)

#### 🐛 Bug Fixes

- Synthetic llms.txt pages do not count as start URLs (#857)

### 0.3.93 (2026-07-27)

#### 🚀 Features

- Add optional AI page summarization (#814)

#### 🐛 Bug Fixes

- Decrease amount of attempted detection on longer runs (#809)

### 0.3.92 (2026-07-15)

#### 🐛 Bug Fixes

- Upgrade crawlee for request-queue write dedup (#826)

### 0.3.91 (2026-07-13)

#### 🐛 Bug Fixes

- Preserve valid special characters (: +) in URL paths (#813)
- Dedup cache-busted URL variants and stabilize final run status (#820)

### 0.3.90 (2026-07-03)

### 0.3.89 (2026-07-02)

#### 🐛 Bug Fixes

- **adaptive:** Don't fallback to browser on session errors (#808)
- Retry on AWS challenge (#806)

### 0.3.88 (2026-06-19)

#### 🚀 Features

- Retry blocked content through the UNBLOCKER proxy group (#784)

#### 🐛 Bug Fixes

- Stop enqueueing after crawl limits (#763)
- Avoid performance enhance when start url is a sitemap (#776)
- Playwright download capture (#724)

#### ⚡ Performance

- Enforce crawl budget in SitemapCrawler (#778)

### 0.3.87 (2026-06-04)

#### 🐛 Bug Fixes

- Detect cloudflare challenge header (#771)

### 0.3.86 (2026-06-01)

#### 🐛 Bug Fixes

- Allow non-object items in dataset jsonLd schema (#767)

### 0.3.84 (2026-05-26)

#### 🚀 Features

- Add dataset schema (#761)

#### ⚡ Performance

- **docker:** Use npm ci, node 24 and consistent firefox base image (#754)

### 0.3.83 (2026-05-13)

#### 🚀 Features

- Force reuse detection on 1 result runs (#683)

#### 🐛 Bug Fixes

- Pass signed request flag to adaptive crawler (#726)
- Enable `impit` fallbacks on TLS errors (#739)
- Recycle memory from workers on long runs (#727)

#### ⚡ Performance

- Process file downloads in adaptive crawler without fallback (#732)

### 0.3.82 (2026-04-29)

#### ⚡ Performance

- Add missing indexes (#718)

### 0.3.81 (2026-04-21)

#### 🚀 Features

- Downscale the crawler on recurring network errors (#677)

#### 🐛 Bug Fixes

- Expand the guessTxtSitemap to accept both text/plain and \*/xml (#682)
- Try-catch page object to avoid throwing (#684)
- Tighten output contract typing (#689)
- Polling race condition. Replace .running with .hasFinishedBefore (#710)
- Avoid indefinite selector waits when dynamic content timeout is zero (#712)

#### ⚡ Performance

- Initiate CountMinSketch only if used (#676)
- Fix polling interval sleeping unnecessarily for short runs (#678)
- Only create 1 worker thread for 1 request crawl (#673)
- Ensure worker threads don't pull unrelated dependencies (#674)
- Don't persist sessionPool and stats on 1 request runs (#701)

### 0.3.80 (2026-04-02)

#### 🚀 Features

- Add contentLengthBytes for every page scraped in cheerio runs (#605)

#### 🐛 Bug Fixes

- Pass `user-agent` correctly to `web-bot-auth`-signed requests (#600)
- Migration safe trimmedStartUrls (#608)
- Guessing if a text file is a sitemap on keywords (#661)

#### ⚡ Performance

- Replace 'crawlee' with '@crawlee/' subpackages (-400ms startup) (#636)
- Node dist/main.js instead of npm run start:prod (#646)
- Skip checking proxy access and initialize state concurrently (#647)
- Import playwright and adaptive db storage packages dynamically, only when needed (#648)
- Use memory queue for single request inputs (#660)
- Properly dynamically import packages for adaptive crawling (#670)

### 0.3.79 (2026-02-13)

#### 🐛 Bug Fixes

- Add timeout to sitemap discovery to avoid startup hang (#599)

#### ⚡ Performance

- Faster check for onSkippedRequest (#597)

### 0.3.78 (2026-02-12)

#### 🚀 Features

- Deduplicate with etag (#579)
- Save errors to dataset with extended reporting & user-friendly advice (#571)
- Introduce multiple datasets (#592)

#### 🐛 Bug Fixes

- Use array of regex as scope constraint (#591)
- Fallbacks for title element extraction (#584)

### 0.3.77 (2026-02-05)

#### 🚀 Features

- Utilize stored rendering types (#552)

#### 🐛 Bug Fixes

- **adaptive-crawler:** Allow overriding sendRequest and ensure proxy info for HTTP requests (#567)
- **file-downloader:** Correctly skip downloads (#568)
- Use the custom user-agent to be allowed by the robots file (#577)

### 0.3.76 (2026-01-21)

#### 🚀 Features

- Prefix KVS record keys with type (#543)
- Use crawlee's built-in request depth tracking (#469)
- Introduce `saveContentTypes` input option, allowing to download images (#563)

#### 🐛 Bug Fixes

- Log a warning on failed screenshot (#553)
- Decouple `maxCrawlPages` from the result count (#564)

### 0.3.75 (2025-12-02)

#### 🚀 Features

- Support llms.txt top files (#476)
- Better README (#527)

### 0.3.74 (2025-11-14)

#### 🐛 Bug Fixes

- Reinstate explicit crawler termination checks (#528)

### 0.3.73 (2025-11-13)

#### 🐛 Bug Fixes

- Don't stall on low-concurrency `maxResults`-limited runs (#526)

### 0.3.72 (2025-11-12)

#### 🚀 Features

- Track rendering type detection results across all runs (#514)
- Enable signing requests with the HTTP message signatures (#486)

#### ⚡ Performance

- Speed-up main crawler termination (#522)

### 0.3.71 (2025-11-04)

#### 🐛 Bug Fixes

- Stop waiting for start URLs in case of empty `RequestQueue` (#517)

### 0.3.70 (2025-11-03)

#### 🚀 Features

- Deprecate JSDOM crawler mode (#510)

#### 🐛 Bug Fixes

- Don't stall on special characters in start URLs (#495)
- Prevent the crawler from getting stuck on finished runs (#498)
- Render absolute links in markdown exports (#513)
- Don't terminate main crawler too soon after migration (#515)

#### ⚡ Performance

- Make auxiliary crawlers lazy-initialized (#505)

### 0.3.68 (2025-11-03)

#### 🚀 Features

- Deprecate JSDOM crawler mode (#510)

#### 🐛 Bug Fixes

- Don't stall on special characters in start URLs (#495)
- Prevent the crawler from getting stuck on finished runs (#498)
- Render absolute links in markdown exports (#513)

#### ⚡ Performance

- Make auxiliary crawlers lazy-initialized (#505)

### 0.3.68 (2025-10-14)

#### 🚀 Features

- Log skipped urls with reasons into a KVS object (#438)
- Add `customHttpHeaders` option for passing custom HTTP header pairs (#434)
- Enable encryption for the `initialCookies` input option (#448)
- Add defuddle support (#474)

#### 🐛 Bug Fixes

- Respect globs with sitemap-discovered requests (#432)
- Improve `SitemapCrawler` stability (#440)
- Add grace period before checking whether to stop crawlers (#463)
- Include url fragment in snapshot hash input (#466)
- Valid public KVS url in `screenshotUrl` (#468)
- Fix title extraction (#480)

### 0.3.67 (2025-06-19)

#### 🚀 Features

- Lower default `maxRequestRetries` limit to `3` (#419)
- Add `ignoreHttpsErrors` input option (#422)
- Add a flag to disable image and video loading in headless browser (#423)

#### 🐛 Bug Fixes

- `keepElementsCssSelector` doesn't affect link enqueuing (#413)
- Clean up after `SitemapRequestList` to avoid excessive timeouts (#415)

### 0.3.66 (2025-05-26)

#### 🚀 Features

- Use code from the Ghostery adblocker to remove cookie modals (#403)

#### 🐛 Bug Fixes

- Only use titles from the document head for the metadata section (#401)
- Store downloaded HTML explicitly as UTF-8 (#410)

#### ⚡ Performance

- Add more "unparsable" file types (`jpg`, `png`, `gif`...) (#402)

### 0.3.65 (2025-04-17)

#### 🚀 Features

- Allow toggling the crawlee respectRobotsTxtFile via input (#392)
- Sticky containers (#390)

#### 🐛 Bug Fixes

- Fix SitemapRequestList persistence (#385)

### 0.3.64 (2025-03-21)

#### 🚀 Features

- Handle more types of cookie modals (#378)
- Store original file name and size for downloads (#377)

### 0.3.63 (2025-03-17)

#### 🚀 Features

- Use RequestQueueV1 (#376)

### (2025-03-14)

#### 🐛 Bug Fixes

- Automatically close cookie modals by usercentrics.eu (#373)

### 0.3.61 (2025-03-04)

#### 🐛 Bug Fixes

- Generate unique KVS keys for `saveHtmlAsFile` results (#369)
- Correctly validate `waitForSelector` input option (#370)

### 0.3.60 (2025-02-21)

#### 🚀 Features

- Add support for `softWaitForSelector` (#365)

#### 🐛 Bug Fixes

- More stable crawler stopping (#366)

### 0.3.59 (2025-01-22)

#### 🐛 Bug Fixes

- Work around missing `content-type` header in `FileDownload` (#360)
- Extract only top-level `title` (skip iframes) (#362)

### 0.3.58 (2025-01-10)

#### 🚀 Features

- Push `latest` to Apify with GH actions, add auto changelog bumps (#347)
- More lenient behaviour on missing `content-type` response header (#353)
- Improve antiblocking performance (#354)

### 0.3.57 (2024-12-13)

This release only contains documentation changes in readme and input schema.

### 0.3.56 (2024-11-25)

- Input:
  - Empty `includeUrlGlobs` are now filtered out with a warning log message. To enforce the old behavior (i.e. matching everything), use `**` instead.
- Behaviour:
  - The Actor now automatically drops the request queue associated with the file download.
  - Only the first `<title>` element on the page is prepended to the exported content.
  - The crawler now uses the correct scope with all `startUrls` being sitemaps.
  - Sitemaps are now processed in a separate thread. The 30 seconds per sitemap request limit is now strongly enforced.

### 0.3.55 (2024-11-11)

- Behaviour:
  - The sitemap timeout warning is only logged the first time.

### 0.3.54 (2024-11-07)

- Input:
  - The default `removeElementsCssSelector` now removes `img` elements with `data:` URLs to prevent cluttering the text output.

- Behaviour:
  - `expandIframes` now skips broken `iframe` elements instead of failing the whole request.
  - Actor now parses formatted (indented or newline-separated) sitemaps correctly.
  - The sitemap discovery process is now parallelized. Logging is improved to show the progress of the sitemap discovery.
  - Sitemap processing now has stricter per-URL time limits to prevent indefinite hangs.

### 0.3.53 (2024-10-22)

- Input:
  - New `keepElementsCssSelector` accepts a CSS selector targeting elements to extract from the page to the output.
- Behaviour:
  - Actor optimizes RQ writes by following the `maxCrawlPages` limits better.

### 0.3.52 (2024-10-10)

- Behavior:
  - Handle sitemap-based requests correctly. Solves the indefinite `RequestQueue` hanging read issue.

### 0.3.51 (2024-10-10)

- Behavior:
  - Revert internal library update to mitigate the indefinite `RequestQueue` hanging read issue.

### 0.3.50 (2024-10-07)

- Behavior:
  - Actor terminates sitemap loading prematurely in case of a timeout.
  - Sitemap loading now respects `maxRequestRetries` option.

### 0.3.49 (2024-09-23)

- Behavior:
  - Use the correct proxy settings when loading sitemap files.
  - Mitigate sitemap persistence issues with premature stopping on `maxRequests` (`ERR_STREAM_PUSH_AFTER_EOF`).

### 0.3.48 (2024-09-10)

- Input:
  - `useSitemaps` option is now pre-filled to `true` to automatically enable it for new users and in API examples.

### 0.3.47 (2024-09-04)

- Behaviour:
  - Use crawlee 3.11.3 which should help with the crawler being stuck because of some stale requests in the queue.

### 0.3.46 (2024-08-30)

- Behaviour:
  - Process markdown in a separate worker thread so it won't block the main process on too large pages.
  - Sitemap loading with `useSitemaps` doesn't block indefinitely on large sitemaps anymore.

### 0.3.45 (2024-08-20)

- Input:
  - New `keepFragmentUrls` (URL `#fragments` identify unique pages) input option to consider fragment URLs as separate pages.
- Behaviour:
  - Ensure canonical URL is only taken from the main page and not embedded pages from `<iframe>`.
  - Too large dataset items used to fail, now we retry with a trimmed payload (`html`, `text` and `markdown` fields are trimmed to the first three million characters each).

### 0.3.44 (2024-07-30)

- Behaviour:
  - `waitForSelector` option allows users to specify a CSS selector to wait for before extracting the page content. This is useful for pages that load content dynamically and break the automatic waiting mechanism.

### 0.3.43 (2024-07-24)

- Behaviour:
  - Change the shadow DOM expansion logic to handle edge cases better.

### 0.3.42 (2024-07-12)

- Behaviour:
  - Fix edge cases with the improved `startUrls` sanitization.

### 0.3.41 (2024-07-11)

- Behaviour:
  - Better input URLs sanitization to prevent issues with the `startUrls` input.

### 0.3.40 (2024-07-10)

- Input:
  - New `expandIframes` (Expand iframe elements) option for extracting content from on-page `iframe` elements. Available only in `playwright:firefox`.

### 0.3.39 (2024-06-28)

- Behaviour:
  - Mitigating excessive Request Queue writes in some cases.

### 0.3.38 (2024-06-25)

- Behaviour:
  - The `saveScreenshots` option now correctly prints warnings with crawler types that don't support screenshots.
  - The screenshot KVS key now contains the website hostname and a hash of the original URL to avoid collisions.

### 0.3.37 (2024-06-17)

- Behaviour:
  - The Actor now respects the advanced request configuration passed through the `Start URLs` input.

### 0.3.36 (2024-06-10)

- Input:
  - New `saveHtmlAsFile` option is available which enables storing the HTML into a key-value store, replacing their values in the dataset with links to make the dataset value smaller, since there is a hard limit for its size.
  - Deprecated `saveHtml` in favor of `saveHtmlAsFile`.
- Output:
  - The new `saveHtmlAsFile` option saves the URL under a new `htmlUrl` key in the dataset.
- Behaviour:
  - HTML processors don't block the main thread and can safely time out.
  - Better fallback logic for the HTML processing pipeline.
  - When pushing to dataset, we now detect too large payload and skip retrying (while suggesting to use the new `saveHtmlAsFile` option to get around this problem).

### 0.3.35 (2024-05-23)

- Behaviour:
  - `RequestQueue` race condition fixed.
- Output:
  - The Readable text extractor now correctly handles the article titles.

### 0.3.34 (2024-05-17)

- Behaviour:
  - If any of the Start URLs lead to a sitemap file, it is processed and the links are enqueued.
  - Performance / QoL improvements (see [Crawlee 3.10.0 changelog](https://github.com/apify/crawlee/releases/tag/v3.10.0) for more details)

### 0.3.33 (2024-04-22)

- Input:
  - `AdaptiveCrawler` is the new default (prefill) crawler type.
- Behaviour:
  - The use of Chrome + Chromium browsers was deprecated. The Actor now uses only Firefox internally for the browser-based crawlers.
  - Reimplementation of the file download feature for better stability and performance.
  - On smaller websites, the `AdaptiveCrawler` skips the adaptive scanning to speed up the crawl.

### 0.3.32 (2024-03-28)

- Output:
  - The Actor now stores `metadata.headers` with the HTTP response headers of the crawled page.

### 0.3.31 (2024-03-14)

- Input:
  - Invalid `startUrls` are now filtered out and don't cause the actor to fail.

### 0.3.30 (2024-02-24)

- Behavior:
  - File download now respects `excludeGlobs` and `includeGlobs` input options, stores filenames, and understands `Content-Disposition` HTTP headers (i.e. "forced download").
- Output:
  - Better `og:` metatags coverage (`article:`, `movie:` etc.).
  - Stores JSON-LD metatags in the `metadata` field.

### 0.3.29 (2024-02-05)

- Input:
  - `useSitemaps` toggle for sitemap discovery - leads to more consistent results, scrapes also unreachable webpages.
  - Do not fail on empty globs in input (ignore them instead).
  - Experimental `playwright:adaptive` crawling mode
- Output:
  - Add `metadata.openGraph` output field for contents of `og:*` metatags.

### 0.3.27 (2024-01-25)

- Input:
  - `maxRequestRetries` input option for limiting the number of request retries on network, server or parsing errors.
- Behavior:
  - Allow large lists of start URLs with deep crawling (`maxCrawlDepth > 0`), as the memory overflow issue from `0.3.18` is now fixed.

### 0.3.26 (2024-01-02)

- Output:
  - The Actor now stores `metadata.mimeType` for downloaded files (only applicable when `saveFiles` is enabled).

### 0.3.25 (2023-12-21)

- Input:
  - Add `maxSessionRotations` input option for limiting the number of session rotations when recognized as a bot.
- Behavior:
  - Fail on `401`, `403` and `429` HTTP status codes.

### 0.3.24 (2023-12-07)

- Behavior:
  - Fix a bug within the `expandClickableElements` utility function.

### 0.3.23 (2023-12-06)

- Behavior:
  - Respect empty `<body>` tag when extracting text from HTML.
  - Fix a bug with the `simplifiedBody === null` exception.

### 0.3.22 (2023-12-04)

- Input:
  - `ignoreCanonicalUrl` toggle to deduplicate pages based on their actual URLs (useful when two different pages share the same canonical URL).
- Behavior:
  - Improve the large content detection - this fixes a regression from `0.3.21`.

### 0.3.21 (2023-11-29)

- Output:
  - The `debug` mode now stores the results of all the extractors (+ raw HTML) as Key-Value Store objects.
  - New extractor "Readable text with fallback" checks the results of the "Readable text" extractor and checks the content integrity on the fly.
- Behavior:
  - Skip text-cleaning and markdown processing step on large responses to avoid indefinite hangs.

### 0.3.20 (2023-11-08)

- Output:
  - The `debug` mode now stores the raw page HTML (without the page preprocessing) under the `rawHtml` key.

### 0.3.19 (2023-10-18)

- Input:
  - Add a default for `proxyConfiguration` option (which is now required since 0.3.18). This fixes the actor usage via API, falling back to the default proxy settings if they are not explicitly provided.

### 0.3.18 (2023-10-17)

- Input:
  - Adds `includeUrlGlobs` option to allow explicit control over enqueuing logic (overrides the default scoping logic).
  - Adds `requestTimeoutSecs` option to allow overriding the default request processing timeout.
- Behavior:
  - Disallow using large list of start URLs (more than 100) with deep crawling (`maxCrawlDepth > 0`) as it can lead to memory overflow.

### 0.3.17 (2023-10-05)

- Input:
  - Adds `debugLog` option to enable debug logging.

### 0.3.16 (2023-09-06)

- Behavior:
  - Raw HTTP Client (Cheerio) now works correctly with proxies again.

### 0.3.15 (2023-08-30)

- Input:
  - `startUrls` is now a required input field.
  - Input tooltips now provide more detailed description of crawler types and other input options.

### 0.3.14 (2023-07-19)

- Behavior:
  - When using the Cheerio based crawlers, the actor processes links from removed elements correctly now.
  - Crawlers now follow links in `<link>` tags (`rel=next, prev, help, search`).
  - Relative canonical URLs are now getting correctly resolved during the deduplication phase.
  - Actor now automatically recognizes blocked websites and retries the crawl with a new proxy / fingerprint combination.

### 0.3.13 (2023-06-30)

- Input:
  - The Actor now uses new defaults for input settings for better user experience:
    - default `crawlerType` is now `playwright:firefox`
    - `saveMarkdown` is `true`
    - `removeCookieWarnings` is `true`

### 0.3.12 (2023-06-14)

- Input:
  - add `excludeUrlGlobs` for skipping certain URLs when enqueueing
  - add `maxScrollHeightPixels` for scrolling down the page (useful for dynamically loaded pages)
- Behavior:
  - by default, the crawler scrolls down on every page to trigger dynamic loading (disable by setting `maxScrollHeightPixels` to 0)
  - the crawler now handles HTML processing errors gracefully
  - the actor now stays alive and restarts the crawl on certain known errors (Playwright Assertion Error).
- Output:
  - `crawl.httpStatusCode` now contains HTTP response status code.

### 0.3.10 (2023-06-05)

- Input:
  - move `processedHtml` under `debug` object
  - add `removeCookieWarnings` option to automatically remove cookie consent modals (via [I don't care about cookies](https://addons.mozilla.org/en-US/firefox/addon/i-dont-care-about-cookies/) browser extension)
- Behavior:
  - consider redirected start URL for prefix matching
  - make URL deduplication case-insensitive
  - wait at least 3s in playwright to let dynamic content to load
  - retry the start URLs 10 times and regular URLs 5 times (to get around issues with retries on burnt proxies)
  - ignore links not starting with `http`
  - skip parsing non-html files
  - support `startUrls` as text-file

### 0.3.9 (2023-05-18)

- Input:
  - Updated README and input hints.

### 0.3.8 (2023-05-17)

- Input:
  - `initialCookies` option for passing cookies to the crawler. Provide the cookies as a JSON array of objects with `"name"` and `"value"` keys. Example: `[{ "name": "token", "value": "123456" }]`.
- Behavior:
  - `textExtractor` option is now removed in favour of `htmlTransformer`
  - `unfluff` extractor has been completely removed
  - HTML is always simplified by removing some elements from it, those are configurable via `removeElementsCssSelector` option, which now defaults to a larger set of elements, including `<nav>`, `<footer>`, `<svg>`, and elements with `role` attribute set to one of `alert`, `banner`, `dialog`, `alertdialog`.
  - New `htmlTransformer` option has been introduced which allows to configure how the simplified HTML is further processed. The output of this is still an HTML, which can be later used to generate markdown or plain text from it.
  - The crawler will now try to expand collapsible sections automatically. This works by clicking on elements with `aria-expanded="false"` attribute. You can configure this selector via `clickElementsCssSelector` option.
  - When using Playwright based crawlers, the dynamic content is awaited for based on network activity, rather than webpage changes. This should improve the reliability of the crawler.
  - Firefox `SEC_ERROR_UNKNOWN_ISSUER` has been solved by preloading the recognized intermediate TLS certificates into the Docker image.
  - Crawled URLs are retried in case their processing timeouts.

### 0.3.7 (2023-05-10)

- Behavior:
  - URLs with redirects are now enqueued based on the original (unredirected) URL. This should prevent the actor from skipping relevant pages which are hidden behind redirects.
  - The actor now considers all start URLs when enqueueing new links. This way, the user can specify multiple start URLs as a workaround for the actor skipping some relevant pages on the website.
  - Error with not enqueueing URLs with certain query parameters is now fixed.
- Output:
  - The `.url` field now contains the main resource URL without the fragment (`#`) part.

### 0.3.6 (2023-05-04)

- Input:
  - Made the `initialConcurrency` option visible in the input editor.
  - Added `aggressivePruning` option. With this option set to `true`, the crawler will try to deduplicate the scraped content. This can be useful when the crawler is scraping a website with a lot of duplicate content (header menus, footers, etc.)
- Behavior:
  - The actor now stays alive and restarts the crawl on certain known errors (Playwright Assertion Error).

### 0.3.4 (2023-05-04)

- Input:
  - Added a new hidden option `initialConcurrency`. This option sets the initial number of web browsers or HTTP clients running in parallel during the actor run. Increasing this number can speed up the crawling process. Bear in mind this option is hidden and can be changed only by editing the actor input using the JSON editor.

### 0.3.3 (2023-04-28)

- Input:
  - Added a new option `maxResults` to limit the total number of results. If used with `maxCrawlPages`, the crawler will stop when either of the limits is reached.

### 0.3.1 (2023-04-24)

- Input:
  - Added an option to download linked document files from the page - `saveFiles`. This is useful for downloading pdf, docx, xslx... files from the crawled pages. The files are saved to the default key-value store of the run and the links to the files are added to the dataset.
  - Added a new crawler - Stealthy web browser - that uses a Firefox browser with a stealthy profile. It is useful for crawling websites that block scraping.

### 0.0.13 (2023-04-18)

- Input:
  - Added new `textExtractor` option `readableText`. It is generally very accurate and has a good ratio of coverage to noise. It extracts only the main article body (similar to `unfluff`) but can work for more complex pages.
  - Added `readableTextCharThreshold` option. This only applies to `readableText` extractor. It allows fine-tuning which part of the text should be focused on. That only matters for very complex pages where it is not obvious what should be extracted.
- Output:
  - Added simplified output view `Overview` that has only `url` and `text` for quick output check
- Behavior:
  - Domains starting with `www.` are now considered equal to ones without it. This means that the start URL `https://apify.com` can enqueue  `https://www.apify.com` and vice versa.

### 0.0.10 (2023-04-05)

- Input:
  - Added new `crawlerType` option `jsdom` for processing with JSDOM. It allows client-side script processing, trying to mimic the browser behavior in Node.js but with much better performance. This is still experimental and may crash on some particular pages.
  - Added `dynamicContentWaitSecs` option (defaults to 10s), which is the maximum waiting time for dynamic waiting.
- Output (BREAKING CHANGE):
  - Renamed `crawl.date` to `crawl.loadedTime`
  - Moved `crawl.screenshotUrl` to top-level object
  - The `markdown` field was made visible
  - Renamed `metadata.language` to `metadata.languageCode`
  - Removed `metadata.createdAt` (for now)
  - Added `metadata.keywords`
- Behavior:
  - Added waiting for dynamically rendered content (supported in Headless browser and JSDOM crawlers). The crawler checks every half a second for content changes. When there are no changes for 2 seconds, the crawler proceeds to extraction.

### 0.0.7 (2023-03-30)

- Input:
  - BREAKING CHANGE: Added `textExtractor` input option to choose how strictly to parse the content. Swapped the previous `unfluff` for `CrawleeHtmlToText` as default which in general will extract more text. We chose to output more text rather than less by default.
  - Added `removeElementsCssSelector` which allows passing extra CSS selectors to further strip down the HTML before it is converted to text. This can help fine-tuning. By default, the actor removes the page navigation bar, header, and footer.
- Output:
  - Added markdown to output if `saveMarkdown` option is chosen
  - All extractor outputs + HTML as a link can be obtained if `debugMode` is set.
  - Added `pageType` to the output (only as `debug` for now), it will be fine-tuned in the future.
- Behavior:
  - Added deduplication by `canonicalUrl`. E.g. if more different URLs point to the same canonical URL, they are skipped
  - Skip pages that redirect outside the original start URLs domain.
  - Only run a single text extractor unless in debug mode. This improves performance.
