# Changelog of Blog Scraper (`naive_zing/blog-scraper`) Actor

- **URL**: https://apify.com/naive\_zing/blog-scraper/changelog.md
- **Full Actor documentation**: https://apify.com/naive\_zing/blog-scraper.md

## Changelog

### \[3.1.1] - 2025-12-31

#### Fixed

- **Blog Detection**: Fixed an issue where company homepages were incorrectly identified as blog posts, preventing the crawler from finding the actual blog section.
- **Content Detection**: Added `og:type` meta tag check for more accurate blog post detection.

### \[3.1.0] - 2025-12-31

#### Changed

- **Performance Optimization**: Switched from Playwright (Headless Chrome) to BeautifulSoup (HTTP requests) for significantly faster scraping and lower resource usage.
- **Resource Usage**: Reduced memory requirement from 2048 MB to 512 MB.
- **Input Schema**: Simplified `company_urls` input to strictly accept a list of strings (URLs), removing the object format support.
- **Docker Image**: Optimized Dockerfile by removing browser dependencies, resulting in a much smaller image size.
- **Dependencies**: Replaced `crawlee[playwright]` with `crawlee[beautifulsoup]`.

### \[3.0.0] - 2025-12-29

#### Changed

- **Complete Rewrite**: Transformed from Numbeo Cost of Living Scraper to Company Blog Post Scraper
- **Core Functionality**:
  - Now scrapes blog posts from company domains instead of cost-of-living data
  - Automatic blog section discovery on company websites
  - Intelligent blog post detection using URL patterns and content analysis
- **Input Schema**:
  - Changed from `start_urls` to `company_urls` (list of company domains)
  - Added `number_of_blog_posts_to_fetch` parameter (1-50, default: 10)
  - Added `max_concurrency` parameter for performance tuning
- **Data Extraction**:
  - Extracts comprehensive blog post metadata: title, author, published date, content, excerpt, tags, category
  - Removed cost-of-living specific extraction logic
- **Actor Metadata**:
  - Updated actor name from `cost-of-living-scraper` to `blog-scraper`
  - Updated version to 1.0.0 for new functionality
- **Documentation**: Completely rewrote README with blog scraper features and use cases

#### Added

- Smart blog post detection with multiple URL pattern matching
- Content extraction using multiple fallback strategies for different blog platforms
- Metadata extraction (author, date, tags, categories)
- Per-domain post limit enforcement
- Automatic URL normalization (handles URLs with or without https://)

#### Removed

- Numbeo-specific crawling logic (country/city navigation)
- Cost-of-living data extraction
- Price table parsing

### \[2.0.0] - 2025-12-29

#### Changed

- **Complete Rewrite**: Transformed the actor from an SEO Schema Validator to a Numbeo Cost of Living Scraper.
- **Core Logic**: Replaced generic sitemap crawling with targeted Numbeo directory traversal (Main -> Country -> City).
- **Extraction**: Added specialized parsing for Numbeo's cost tables.
- **Dependencies**: Removed `src.utils` and `src.validator` as they are no longer needed.
- **Input**: Simplified input to accept optional `start_urls`, defaulting to Numbeo's main page.

### \[1.1.0] - 2025-12-17

#### Changed

- **Validation Logic**: Added recursive object validation for deeply nested schemas.
- **Validation Logic**: Added specific checks for `Article` (headline, datePublished) and `Product` (aggregateRating/offers/review).
- **Output Schema**: Corrected field types in `output_schema.json` to match actual data types (integer, array, object).
- **Metadata**: Updated actor title and description for better SEO.
- **Documentation**: Revamped README with "Problem-Solution-Benefit" structure and keywords.

### \[1.0.0] - 2025-12-16

#### Added

- Initial release of SEO Schema & Meta Validator.
