Colombia Tourism Registry — Lodging Supply, Per Record avatar

Colombia Tourism Registry — Lodging Supply, Per Record

Pricing

from $33.50 / 1,000 tourism registry records

Go to Apify Store
Colombia Tourism Registry — Lodging Supply, Per Record

Colombia Tourism Registry — Lodging Supply, Per Record

Colombia National Tourism Registry (RNT, thwd-ivmp) as clean, per-record lodging-capacity data — registry ID, status, department/municipality, category and declared rooms/beds. Proprietor name and tax-ID excluded. CC BY-SA 4.0, $0.05 per record.

Pricing

from $33.50 / 1,000 tourism registry records

Rating

0.0

(0)

Developer

NexGen Signal

NexGen Signal

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

4 days ago

Last modified

Share

Turn Colombia's National Tourism Registry (RNT) into clean, per-record lodging-capacity data — one row per registered tourism establishment, with registry ID, status, locality, category and declared rooms/beds, ready to qualify legal lodging supply across Colombia.

Each source row becomes one clean, flat record with numeric fields coerced to real numbers, a stable source-native record_id, and provenance stamped on every row: source, resource id, the licence notice, the required attribution, a UTC retrieval timestamp and an interpretation caveat.

What one record represents

The source is thwd-ivmpRegistro Nacional de Turismo - RNT on the Colombian national open-data portal (datos.gov.co). Each record is one registered tourism establishment in the RNT: its registry number and status, its department and municipality, its provider category and sub-category, and — for lodging providers — the declared number of rooms and beds, plus a declared employee count and the registry year.

For each record you get a composite record_id built from the source-native key, the analytic columns listed below (reproduced verbatim, numbers as numbers), and the provenance block. Columns include codigo_rnt (registry number), estado_rnt (status), departamento/municipio with their codes, categoria/sub_categoria, habitaciones (rooms), camas (beds), num_emp1 (employees) and ano (year).

Coverage and volume

The live registry holds 679,548 establishment records across Colombia's departments and municipalities and multiple tourism-provider categories — that is the record capacity of a full pull.

Live count: 679,548 records — matches the Wave-3 index figure exactly.

The Actor pages the source with keyless SODA $query requests ordered by the source-native key for a stable total order, and stops as soon as your Maximum records cap is met. Only the registry, locality, category and capacity columns are selected; the proprietor legal-name (razon_social) and tax-ID (nit) columns are never selected.

Licence and attribution

This dataset is published by Colombia's Ministry of Commerce, Industry and Tourism under Creative Commons Attribution-ShareAlike 4.0 (CC BY-SA 4.0). It is redistributed unchanged with attribution; the ShareAlike term applies to any database you derive and redistribute. The full notice travels on every record:

Colombia National Tourism Registry (MinCIT), CC BY-SA 4.0. Reproduced unchanged with attribution; ShareAlike applies to any redistributed database. No proprietor/owner name (razon_social) or tax-ID (NIT) column is read or delivered.

The required attribution — Ministerio de Comercio, Industria y Turismo - MinCIT, Bogota D.C. — travels on every record.

Interpretation caveat

A registry snapshot: estado_rnt reflects the registration status at capture; habitaciones/camas are declared lodging capacity, not verified or occupied capacity. Categories cover all tourism providers, not only lodging, so rooms/beds are zero for non-lodging categories.

Values are reproduced verbatim: the Actor never rescales, re-derives or editorialises a number.

Person-data policy

The registry can carry natural-person proprietors, so the proprietor legal-name column (razon_social_establecimiento) and the tax-ID column (nit) are structurally excluded — never selected, never delivered, enforced by a never-in-SELECT assertion and a planted-name test. Registry ID, status, locality, category and rooms/beds remain. This is a supply-capacity product about establishments and places, not about people.

Data quality and freshness

Numeric fields are coerced from the source's string encoding into real numbers (integers where whole, floats otherwise); genuinely missing cells are delivered as null, never as zero. Text is passed through verbatim. Every run re-reads the live source, so the data is as fresh as the portal itself, and each record's observed_at stamp records exactly when the row was retrieved. Delivery order is fixed by the source-native key, so a capped sample and a later full pull agree on their overlap and a repeated run returns rows in the same order. The RUN_RECEIPT reports source rows scanned and records delivered and charged for a per-run reconciliation.

Provenance, licensing and compliance

Every run begins with a live source-preflight: the Actor reads the exact host's robots.txt at runtime and refuses to proceed if the crawl policy disallows the data path. The gate result — URL, HTTP status, byte length and a SHA-256 of the policy — is written to the run's RUN_RECEIPT, so each run carries its own audit trail. The Actor identifies itself with a transparent, non-impersonating User-Agent and never bypasses a block, solves a challenge, or fetches through a cache or mirror. When the door is genuinely unavailable the run fails loudly and bills nothing.

Inputs

  • Maximum records (maxRecords) — hard cap on records delivered and billed. Raise it to pull the full set; lower it to sample cheaply. Records arrive in a stable, source-native order.

Output

Records land in the Actor's default dataset and export as JSON, CSV, Excel or via the Apify API. A tabular overview view surfaces the most useful columns for quick inspection while the full record retains every selected field and provenance stamp.

Fields in detail

The record leads with record_id — a stable composite key drawn from the source's own grain — followed by the analytic columns described above and closed by a provenance block: source, source_dataset (the Socrata resource id), licence, attribution, caveat and observed_at. Every one of those provenance fields is present on every record, so a single row is self-describing: hand it to a colleague or a downstream system and it carries its own origin, licence and retrieval time without reference back to this page. Because delivery is ordered by the source-native key, the same record always carries the same record_id across runs, which makes the dataset safe to diff, deduplicate, or upsert into a warehouse. Nothing in the record is computed or inferred beyond the explicit count where one is stated — every other value is the source's own, reproduced byte-for-byte.

Sibling Actors

This Actor sits alongside the fleet's eu-hotel-occupancy-benchmark-records cell — but that measures tourism demand (nights spent) in Europe, while this one qualifies legal lodging supply in Colombia: distinct jobs, distinct geographies. This Actor also shares its engineering — the runtime robots gate, push-then-charge billing and verbatim-value discipline — with the fleet's other public-data records Actors.

Pricing

This Actor uses Apify's pay-per-event model: a flat $0.05 per record actually delivered to the dataset, and nothing else — no monthly rental, no per-run base fee, no compute charge. Deliver 40 records and you pay $2.00; deliver 10,000 and you pay $500.00. Billing is wired after delivery — each record is pushed first and only then does the per-record event fire — so a mid-run failure can only ever under-charge you, never over-charge. Use Maximum records to cap spend precisely.

Scaling and limits

Set Maximum records low to sample the leading slice cheaply, or high to pull the full set. The Actor paginates server-side and delivers incrementally, so memory stays flat regardless of how many records you request, and you are billed only for what is actually delivered. Because the source is a live public API, extremely deep pagination is ultimately bounded by the source's own paging behaviour; for the vast majority of uses — sampling, a full refresh, or a scheduled top-up — the default paging is more than sufficient. Schedule the Actor on Apify to keep a downstream table current: each run re-reads the live source and re-stamps observed_at, so a nightly or weekly run gives you a dated, reproducible snapshot.

Typical uses

Qualify legal lodging supply before onboarding partners; size room/bed capacity by municipality; screen establishments by registry status; map tourism-provider density across departments; or feed a travel-supply or market-entry model with clean registry records.

What this Actor does not do

It does not forecast or model, does not merge multiple source tables into one record, and it does not include the proprietor name or tax-ID columns some source views expose. It gives you faithful, analysis-ready records — with a provenance trail you can audit on every run.