Paraguay DNCP Procurement Processes - Per Process
Pricing
from $33.50 / 1,000 dncp process records
Paraguay DNCP Procurement Processes - Per Process
Paraguay DNCP procurement processes (contrataciones.gov.py OCDS) as clean per-record data by year - process id, tender title, status, buyer, awarded organisation supplier and value. Year page-walk; natural-person (non-RUC-80) suppliers dropped. CC BY 4.0. $0.05 per record.
Pricing
from $33.50 / 1,000 dncp process records
Rating
0.0
(0)
Developer
NexGen Signal
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
11 hours ago
Last modified
Categories
Share
Paraguay's public-procurement processes from the DNCP (Direccion Nacional de Contrataciones Publicas) OCDS API as clean, per-process records - for the year you choose. Process id, tender object, status, buyer, and the awarded organisation supplier. Natural-person suppliers are never delivered.
What one record represents
The source is the DNCP open-data OCDS API at contrataciones.gov.py. Each record is one procurement process
(OCDS ocid, compiled-release grain): the process id and release date, the tender id, title, status detail,
main category and method, the tender period, the buyer (contracting entity) id and name, and - where the process
has been awarded to a legal entity - the awarded supplier's name and RUC and the award value and currency. It is
the process-registry view of Paraguayan public procurement, at the grain of one process.
Coverage, the 10,000 ceiling, and the year walk
The DNCP search API reports a nominal total_items of 10,000 for every query window - even a single day -
which is a fixed ceiling, not a real count. Retrieval, however, is not capped there: paging past the nominal
last page keeps returning fresh, distinct processes. The Actor therefore takes a required year input
(prefilled to the current year) and walks every page within that year at 1,000 records a page, deduplicating
on ocid, until a short page marks the true end. That recovers the complete year rather than the first ten
thousand rows. You raise Maximum records to pull the whole year or lower it to sample.
Person data: the RUC-80 supplier gate
Paraguayan supplier identifiers in OCDS take the form PY-RUC-<number>-<check>. Legal entities carry an
eight-digit RUC beginning 80; natural persons use their cedula, a lower number that does not. This cell
emits a process only when it has an award supplier whose RUC begins 80 - a legal entity. Processes with no
award supplier, and processes awarded to a natural person, are dropped and never delivered; the run receipt
reports how many were removed. In practice this removes a large share of raw processes (many are at tender stage
with no award yet, or are awarded to individuals), and it means the delivered award_supplier_name is always a
company name, never a person. Contact blocks (contact points, emails, phones) are never read from the source at
all. The result is a clean, person-free record of which organisations won which public processes in Paraguay.
Licence
The licence is read verbatim from the OCDS package's own license field on every run and echoed on every
record: Creative Commons Attribution 4.0 International (CC BY 4.0), as declared verbatim in the OCDS package license field (https://creativecommons.org/licenses/by/4.0/): free to share and adapt for any purpose, including commercially, with attribution to the DNCP. The attribution to the DNCP rides on every record.
Who buys this, and how they use it
This cell is bought by teams that need the source's full published set as a typed, keyed table they can hold and refresh, rather than a page they scrape. Market-intelligence and lead-generation teams use it to size a market and track who is active in it; analysts and journalists use it to build a longitudinal series that the source's own portal does not expose; data engineers use it as a clean upstream feed into a warehouse, keyed so it upserts without duplication. The common thread is that the record grain and the stable key are chosen so the output is a building block, not a one-off export: you run it on a schedule, keep the delta, and join it to your other sources on the identifiers it preserves verbatim.
Field-by-field, and why the grain is what it is
Every field in the record is either a source-native identifier, a source-native attribute, or one of the six
provenance fields (source, source_dataset, licence, attribution, caveat, observed_at) the fleet
attaches to every record. Nothing is derived or inferred beyond the small, documented transforms noted above,
and nothing is dropped silently: the person-handling section spells out exactly which fields are excluded and
why. The grain - one record per the natural unit of the source - is deliberate: it keeps each row independently
meaningful, keeps the key stable across runs so re-running is a cheap upsert rather than a re-import, and lets
you aggregate up to whatever unit you need without having to unpick a pre-joined table. If you need a different
grain, you compose it downstream from these rows; the cell's job is to deliver the atomic, person-safe,
licence-clean records that everything else is built from.
Reconciling counts honestly
Where the live count differs from any previously published figure, the live measure is the honest one and is what this listing quotes; sources re-issue and consolidate their data over time, so a figure drifts. The run receipt always states what was actually delivered and charged and confirms the two agree, so every run is auditable against itself regardless of what any external index expected.
Sibling Actors
Provenance and compliance
Every run reads the door host's robots.txt at runtime; the gate result (URL, status, byte length and, where a
policy is served, its SHA-256) is written to the run's RUN_RECEIPT. Where the host serves no applicable
robots rule, or redirects its policy to another host, the gate records that (flagged) and proceeds on the
licence, which grants re-use. The endpoint is keyless. The Actor never bypasses a block or fetches through a
mirror, and it reads only the public listing endpoint - never a per-record detail page.
Data quality and freshness
Numeric columns are delivered as real numbers and every other column as a string or null, so the dataset loads
without a cleaning pass. Delivery is keyed on a stable source identifier, so the data is safe to diff,
deduplicate or upsert. Every run re-reads the live door, so the data is as fresh as the source publishes, and
each record's observed_at stamp dates the snapshot. The run's RUN_RECEIPT records the source URL and how
many records were delivered and charged, and confirms charge_equals_delivered.
Billing, delivery and joins
Pricing is per record: you are billed only for records the Actor actually delivers, with the charge raised after each record is pushed (push-then-charge), so a failed or empty run costs nothing. The Maximum records cap bounds every run, so you control spend precisely - sample cheaply, then raise it. Every record is a flat, typed object keyed on a stable id, so the data loads without a cleaning pass, diffs cleanly between runs, and upserts into a table you maintain over time; re-running keeps that table current without re-paying for rows you already hold, and each receipt reconciles delivered against charged. Because the source's own identifiers are preserved verbatim, the dataset joins cleanly onto other sources keyed on the same identifier.
Scaling and scheduling
Set Maximum records low to sample the shape of the data cheaply, then raise it once the cell fits your use.
The Actor delivers incrementally and streams its source, so memory stays flat regardless of how many records you
request, and you are billed only for what is delivered. Because the source republishes on its own cadence, a
scheduled run keeps a downstream table current: new and changed records upsert over the old ones on the stable
key, and the observed_at stamp on every record tells you when each was last seen live. There is no
subscription and no minimum - the per-record price and the record cap together mean the spend on any run is
known in advance and matched exactly to the data you receive.