How to use an open-data portal's API: CKAN, DCAT and OpenDataSoft

Most open-data portals run one of three software stacks, and each one answers a different URL. How to tell which one you are looking at, how to search and download from it without clicking, and what to do when the portal has no API at all.

2026-09-25

A portal’s download button gives you one file. Its API gives you every file, every update, and the metadata that says when the file last changed. The good news is that you almost never meet a bespoke system: a few thousand public portals worldwide run a handful of platforms, and once you recognise the platform you already know its URLs.

First: find out what you are looking at

Ask the portal one question before you write any code:

https://portal.example/api/3/action/package_search?rows=0

If the answer is JSON with "success": true, it is a CKAN, and everything in the next section applies. That single request is how we classify every portal in our own register — it costs one HTTP call and it is the difference between writing a scraper and writing three lines.

If it answers something else, look for these in order:

  • /data.json at the root — a DCAT / Project Open Data catalogue, the standard every US federal agency publishes and the one the European hub aggregates.
  • /api/explore/v2.1/catalog/datasets — an OpenDataSoft portal.
  • /resource/abcd-1234.json in the links on a dataset page — Socrata.

If nothing answers, that is an answer too: a surprising share of the portals listed in public directories are dead or have moved, and no API will fix a domain that no longer resolves.

CKAN: search the catalogue, then the rows

CKAN separates the catalogue (which datasets exist) from the datastore (the rows inside a tabular one). Both are read-only over HTTP, no key needed.

Search the catalogue:

/api/3/action/package_search?q=air+quality&rows=50
/api/3/action/package_search?fq=organization:city-of-x&rows=50

The response holds a results array; each item is a dataset with a resources list, and each resource has a url — the actual CSV, the actual GeoJSON — plus format, size and last_modified. Fetching those URLs is the whole download pipeline.

Read one dataset you already know by name or id:

/api/3/action/package_show?id=air-quality-hourly

And when the resource was pushed into CKAN’s datastore, you can query rows instead of downloading the file:

/api/3/action/datastore_search?resource_id=<id>&limit=100
/api/3/action/datastore_search?resource_id=<id>&q=Madrid

Two practical notes. package_search paginates with start and rows, and most portals cap rows at 1000 — loop, do not ask for everything. And the datastore is optional: plenty of CKANs only ever host the file, so check for datastore_active: true on the resource before you rely on it.

DCAT: one file that describes everything

A DCAT catalogue is not a query API. It is a single document — /data.json, or an RDF serialisation at catalog.rdf or catalog.ttl — listing every dataset with its title, publisher, licence, update frequency and distribution URLs. You download it once, filter it locally, and you are done.

That makes it the least fashionable and the most reliable option: no rate limit, no pagination, no API version to track. It is also how catalogues federate, which is why a dataset published by a Spanish province shows up on data.europa.eu without anyone copying the file.

The trade-off is freshness and depth. The document describes the datasets; it says nothing about the rows inside them, and portals regenerate it on their own schedule.

OpenDataSoft and Socrata: the rows are the API

These two platforms are built the other way round: every dataset is queryable, with filters, aggregations and exports as URL parameters.

OpenDataSoft, current API version:

/api/explore/v2.1/catalog/datasets                       list the datasets
/api/explore/v2.1/catalog/datasets/<id>/records?limit=20
/api/explore/v2.1/catalog/datasets/<id>/records?where=year>2020&group_by=region
/api/explore/v2.1/catalog/datasets/<id>/exports/csv      the whole thing

Socrata uses SoQL, which reads like SQL pressed into a query string:

/resource/abcd-1234.json?$where=year>2020&$select=region,count(*)&$group=region
/resource/abcd-1234.csv?$limit=50000&$offset=0

Both throttle anonymous traffic and both offer a free application token that raises the limit. If you are pulling tens of thousands of rows, ask for the export URL instead of paginating — one large file is cheaper for you and for them.

When there is no API

Some of the most valuable datasets sit behind an HTML table, an FTP directory, or a “request by email” form. In that order of preference: look for the underlying file that the page’s JavaScript fetches (the browser’s network tab shows it in seconds), check whether a statistics office publishes the same series through a real API, and only then consider asking for it — our guide on how to file a FOIA request covers the case where the data exists but is not published.

Before you republish what you pulled

An open API does not mean an open licence, and the two are recorded in different places. CKAN keeps license_id and license_title on the dataset; DCAT keeps a license on each distribution; OpenDataSoft and Socrata keep it in the dataset metadata. Read that field before the data leaves your laptop: attribution is usually the only condition, it is trivial to comply with and embarrassing to miss, and what each licence actually requires is a guide of its own.

The Spanish version of this guide is at Cómo usar la API de un portal de datos abiertos.