Python news agent (Smolagents)
An AI agent that finds and summarizes news on topics you choose.
my_actor/main.py
my_actor/__main__.py
1"""Module defines the main entry point for the Apify Actor.2
3Feel free to modify this file to suit your specific needs.4
5To build Apify Actors, utilize the Apify SDK toolkit, read more at the official documentation:6https://docs.apify.com/sdk/python7"""8
9from __future__ import annotations10
11import os12import sys13from io import TextIOWrapper14
15import requests16from apify import Actor17from smolagents import CodeAgent, OpenAIServerModel, WebSearchTool18
19# Configure stdout to use UTF-8 encoding for proper unicode support20if hasattr(sys.stdout, 'reconfigure'):21 sys.stdout.reconfigure(encoding='utf-8') # ty: ignore[call-non-callable]22else:23 # Fall back to TextIOWrapper for environments where reconfigure is unavailable24 sys.stdout = TextIOWrapper(sys.stdout.buffer, encoding='utf-8')25
26OPENAI_API_KEY = os.environ.get('OPENAI_API_KEY')27OPENAI_API_BASE = 'https://api.openai.com/v1'28
29# Bounds that keep a run short when the search engine is slow or blocks the client.30SEARCH_TIMEOUT_SECS = 1531MAX_AGENT_STEPS = 632
33
34class BoundedWebSearchTool(WebSearchTool):35 """DuckDuckGo search with a request timeout.36
37 The stock tool sends its request without a timeout, so a stalled connection blocks the agent step for minutes.38 """39
40 def search_duckduckgo(self, query: str) -> list:41 """Search DuckDuckGo Lite and parse its result rows."""42 try:43 response = requests.get(44 'https://lite.duckduckgo.com/lite/',45 params={'q': query},46 headers={'User-Agent': 'Mozilla/5.0'},47 timeout=SEARCH_TIMEOUT_SECS,48 )49 response.raise_for_status()50 except requests.RequestException as exc:51 msg = f'Web search failed: {exc}. Try again later or answer from the results you already have.'52 raise RuntimeError(msg) from exc53 parser = self._create_duckduckgo_parser()54 parser.feed(response.text)55 return parser.results56
57
58async def main() -> None:59 """Define a main entry point for the Apify Actor.60
61 This coroutine is executed using `asyncio.run()`, so it must remain an asynchronous function for proper execution.62 Asynchronous execution is required for communication with Apify platform, and it also enhances performance in63 the field of web scraping significantly.64 """65 async with Actor:66 # Retrieve input parameters from the Apify Actor configuration67 actor_input = await Actor.get_input() or {}68
69 model = actor_input.get('model')70 if not model:71 raise ValueError('Missing "model" attribute in Actor input!')72
73 user_interests = actor_input.get('interests')74 if not user_interests:75 raise ValueError('Missing "interests" attribute in Actor input!')76
77 if not OPENAI_API_KEY:78 raise ValueError('Missing OPENAI_API_KEY environment variable!')79
80 # Initialize the OpenAI model for text processing81 model = OpenAIServerModel(82 model_id=model,83 api_base=OPENAI_API_BASE,84 api_key=OPENAI_API_KEY,85 )86
87 # Create the search tool and AI agent. The step cap bounds the run even when every search fails.88 search_tool = BoundedWebSearchTool()89 agent = CodeAgent(tools=[search_tool], model=model, max_steps=MAX_AGENT_STEPS)90
91 # Search and summarize in one run, so the agent doesn't search again for the summary.92 query = (93 f'Find the latest news on {", ".join(user_interests)} and return a concise summary of the most important '94 'stories as plain text. If a search fails, summarize what you found so far.'95 )96 summary = agent.run(query)97 Actor.log.info('News search and summarization completed successfully.')98
99 # Push the results to the dataset by wrapping it in an object.100 Actor.log.info('The results will be stored in the dataset.')101 await Actor.push_data({'summary': str(summary)})An AI news aggregator that fetches and summarizes the latest news based on user-defined interests using DuckDuckGo search and OpenAI models, built with Python Smolagents.
This Actor works as an AI-powered news aggregator:
- The user provides a list of topics they are interested in.
- The Actor searches for relevant news articles using DuckDuckGo.
- The retrieved articles are processed and summarized using an OpenAI model, all in one agent run.
- Each search times out after 15 seconds and the agent stops after 6 steps, so a slow or blocked search can't stall the run.
- The final summarized news output is stored in a dataset.
- Provide input: Define your topics of interest by setting the
interestsfield in the Actor input. - Choose an OpenAI model: Specify the OpenAI model to use in the
modelfield. - Run the Actor: Execute the Actor on the Apify platform or locally.
- Retrieve results: The summarized news articles will be available in the default dataset.
- You can modify the
my_actor/main.pyfile to adjust the query structure or change how the results are summarized. - If needed, you can replace the
DuckDuckGosearch tool with another search API. - Update the prompt used for summarization to fine-tune the output.
- Apify SDK for Python - a toolkit for building Apify Actors and scrapers in Python
- Input schema - define and easily validate a schema for your Actor's input
- Dataset - store structured data where each object stored has the same attributes
- Smolagents - lightweight AI agent framework
Python site crawler (Crawlee + BeautifulSoup)
A crawler that follows links and gets data from static pages, with Crawlee handling retries and request queues. Uses BeautifulSoup, Python's most popular HTML parser. Can't run client-side JavaScript.
Empty Python Actor
An Actor with the Apify SDK set up, so you can build any tool you need.
Python one-page scraper (BeautifulSoup)
A scraper that gets data from one web page with BeautifulSoup. The simplest way to start scraping.
Empty Python Actor (uv)
An Actor with the Apify SDK set up and dependencies managed by the uv package manager, so you can build any tool you need.
Python site crawler (BeautifulSoup)
A lightweight crawler that follows links and gets data from static pages with BeautifulSoup. Good for blogs, news, or product listings, but it can't run client-side JavaScript.
Python browser crawler (Playwright)
A crawler that uses a real browser through Playwright, so it gets data HTTP crawlers miss. Good for social feeds, dashboards, or single-page apps.