Skip to content

feat(api): metadata mapping and export (DCAT, Croissant, Dublin Core) - #211

Draft
anantjain341 wants to merge 7 commits into
CivicDataLab:devfrom
anantjain341:feat/metadata-mapping
Draft

anantjain341 wants to merge 7 commits into
CivicDataLab:devfrom
anantjain341:feat/metadata-mapping

Conversation

@anantjain341

Copy link
Copy Markdown

Built on top of #204. The last two commits are this PR; the rest belong to #204.

What it does

Any published dataset can now be downloaded as a metadata file in a standard
format: DCAT, Croissant or Dublin Core, as JSON-LD, Turtle, RDF/XML or
N-Triples (Croissant is JSON-LD only). The file is built when requested from
the data we already have. Nothing new is stored.

GET /api/datasets/<id>/export/?standard=dcat&format=turtle
GET /api/metadata/export-options/      (lists standards and formats)

Owners can export their drafts; everyone else can export published datasets.
Adding report=1 returns the file together with a short report of what could
not be expressed in that standard.

How the mapping works

The rules for which of our fields becomes which property live in one JSON
file, api/services/metadata_export/contracts/crosswalk.json, with three
CSVs of allowed licences, sectors and geographies next to it. The code reads
the file; it has no standard-specific logic. To change a mapping, edit the
file in a PR. These files belong to this repo and do not depend on anything
outside it.

Admin-defined metadata fields

Extra dataset fields created in Django admin (for example ds:source_website
and ds:createdOn on dev) now work with import and export. The platform
import fills them in where it can, and the export puts their values under the
right property. A field the mapping does not recognise is left alone and
listed in the report.

Checked

Exports of real datasets from Hugging Face, Kaggle, GitHub and an uploaded
CSV were checked against the DCAT-AP and Croissant requirements and parsed
as valid RDF.

Not in this PR

  • Croissant wants a sha256 hash on each file. We do not store one yet.
    Follow-up PR on the upload path.
  • Unknown licences from imports still fall back to CC BY 4.0. Needs an OTHER
    option; product decision.

Deploy notes

  • New dependency: rdflib.
  • No migration.

Let a publisher add a dataset that already lives on a third-party platform
by giving its identifier (or page URL). Only metadata is fetched; files are
never copied or listed, and downloads redirect to the platform.

- previewPlatformDataset query: fetch title, description, license, tags,
  author and last-updated from the platform, with no side effects.
- importPlatformDataset mutation: create a DRAFT dataset prefilled from the
  platform, attach sectors/geographies whose names match the platform's
  tags, fill matching dataset metadata fields, add one EXTERNAL resource
  linking to the dataset page, record provenance, and grant the owner role,
  all in one transaction. Optional title override. Duplicate imports within
  the same organisation/user are rejected.
- DatasetSource model (one-to-one with Dataset) for provenance; exposed as
  TypeDataset.source. TypeResource.url is now exposed.
- Importers for Hugging Face, GitHub and Kaggle behind a registry; all work
  without API keys for public datasets. HF_TOKEN, GITHUB_TOKEN and
  KAGGLE_USERNAME/KAGGLE_KEY are optional.
- Download view redirects EXTERNAL resources to their URL and returns 404
  instead of raising when a resource has no file.
- Dataset search document gains source_platform; /api/search/dataset/
  returns it, aggregates on it and filters by it (NATIVE = not imported).
  formats indexing skips resources without file details.

After deploy, run `manage.py search_index --rebuild` once so Elasticsearch
maps source_platform as a keyword before the first import is indexed.

Refs CivicDataLab/DataSpace#174, CivicDataLab#190, CivicDataLab#191, CivicDataLab#192
The Hub's dataset endpoint returns `siblings`, one entry per file, by
default. For large repos that key dominates the response: 9.6 MB for an
85k-file repo against ~4 KB for everything else. The importer stored the
whole response in DatasetSource.raw_metadata and fetched it on every
preview and import, although files are never listed.

Request fields by name with `expand[]` (every default field except
`siblings`, plus `citation`), and drop `siblings` defensively before the
payload is kept.
…load

DatasetSource no longer keeps the platform's raw JSON. Every field we
fetch now lands in a typed column, chosen by one rule: a metadata
standard reads it on export (DCAT / Croissant / Dublin Core) or the
platform itself reads it (attribution, duplicate check, licence review).

New columns: revision (commit hash or version), source_created_at,
source_readme (full card; Dataset.description keeps a 1,000-char cut),
citation, languages, source_homepage, is_archived. Column definitions a
platform declares (Hugging Face dataset_info) become ResourceSchema rows
on the link resource, so the columns list works without fetching data.

Importers fill what each platform provides: Hugging Face all of the
above; GitHub adds one small call for the branch head SHA and reads
homepage/archived; Kaggle uses the version number and the earliest
version date. Fields nothing reads (downloads, likes, stars, size
categories, task taxonomy) are no longer fetched.
Found by importing many real datasets and fuzzing identifiers:

- Hugging Face: keep the platform's canonical id ("imdb" is really
  "stanfordnlp/imdb"), so the same dataset cannot be imported twice under
  a legacy name. Duplicate check re-runs under the canonical id.
- GitHub: SPDX "NOASSERTION"/"other" means no detectable licence; treat as
  empty instead of a licence called NOASSERTION. Branch names and folder
  paths are validated (no "..", no whitespace, plain segments only).
- Kaggle: a 403 covers "does not exist" as well as private, so say both.
  The created date is taken from the earliest version only when the view
  lists every version (it lists one of 2,324 for kaggle/meta-kaggle).
- All: identifiers over 200 characters are rejected; README kept to
  200 KB; citation to 20 KB.
- Service: platform calls now happen before the transaction is opened, so
  a slow platform never holds a database connection.
GET /api/datasets/<id>/export/?standard=dcat|croissant|dublin_core
    &format=jsonld|turtle|rdfxml|ntriples   (Croissant: jsonld only)
GET /api/metadata/export-options/

Generated on request from the platform's own model; nothing is stored.
Published datasets are public; owners can preview drafts. `report=1`
returns the document together with the gap report (values the standard
wanted as URIs but got as names, fields the standard cannot carry,
mandatory properties we had nothing for).

The mapping is data, not code. `contracts/crosswalk.json` binds each
concept to a record field and, per standard, to a property with its
value type and obligation; `contracts/*.csv` hold the allowed licences,
sectors and geographies with URIs. Both are owned here and edited like
any other file (see contracts/README.md). `crosswalk.py` is the engine
that applies the file and knows no standard by name; `adapter.py` is the
only module that reads Django models and resolves names to URIs.

Serialisation the mapping cannot express lives in exporter.py: dates as
xsd:date literals, media types as IANA IRIs, the licence repeated on each
distribution, and the @id/contentUrl/encodingFormat Croissant requires on
every FileObject. Imported datasets export with provenance to the source
page and the platform author as creator; the link resource becomes a
distribution with accessURL, never downloadURL.

Not covered yet: Croissant's sha256 on file objects (no file hash is
stored; separate change on the upload path).

New dependency: rdflib (pure Python) for the non-JSON serialisations.
Metadata definitions (label, URN, type) created in Django admin were
opaque strings: the platform import could not fill them and the export
could not place them. metadata_mapping.py gives each definition a
crosswalk concept, matched by URN (ds:createdOn -> created), then by a
standard's own property name (dcterms:issued -> issued), then by label.

Import: each enabled dataset definition is prefilled with the platform
value for its concept (source page, creator, dates, licence, version,
homepage, citation, languages); the definition's validators still apply.

Export: definition values are read by concept. Core columns always win;
a definition only supplies what the model has no column for. ds:createdOn
becomes dcterms:created (the data's origin) while our created column
stays dcterms:issued (when the record appeared). Definitions the mapping
cannot place are listed under `dropped` in the report instead of vanishing.
Use the official Croissant @context (every cr: term, @language, dct)
instead of three bare prefixes, so validators recognise the document;
type the dataset as sc:Dataset; declare conformsTo 1.1; and emit the
source platform identifier as alternateName for imported datasets, as
Hugging Face's own Croissant does. No change to the properties' meaning.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant