feat(api): metadata mapping and export (DCAT, Croissant, Dublin Core) - #211
Draft
anantjain341 wants to merge 7 commits into
Draft
anantjain341 wants to merge 7 commits into
anantjain341 wants to merge 7 commits into
Conversation
Let a publisher add a dataset that already lives on a third-party platform by giving its identifier (or page URL). Only metadata is fetched; files are never copied or listed, and downloads redirect to the platform. - previewPlatformDataset query: fetch title, description, license, tags, author and last-updated from the platform, with no side effects. - importPlatformDataset mutation: create a DRAFT dataset prefilled from the platform, attach sectors/geographies whose names match the platform's tags, fill matching dataset metadata fields, add one EXTERNAL resource linking to the dataset page, record provenance, and grant the owner role, all in one transaction. Optional title override. Duplicate imports within the same organisation/user are rejected. - DatasetSource model (one-to-one with Dataset) for provenance; exposed as TypeDataset.source. TypeResource.url is now exposed. - Importers for Hugging Face, GitHub and Kaggle behind a registry; all work without API keys for public datasets. HF_TOKEN, GITHUB_TOKEN and KAGGLE_USERNAME/KAGGLE_KEY are optional. - Download view redirects EXTERNAL resources to their URL and returns 404 instead of raising when a resource has no file. - Dataset search document gains source_platform; /api/search/dataset/ returns it, aggregates on it and filters by it (NATIVE = not imported). formats indexing skips resources without file details. After deploy, run `manage.py search_index --rebuild` once so Elasticsearch maps source_platform as a keyword before the first import is indexed. Refs CivicDataLab/DataSpace#174, CivicDataLab#190, CivicDataLab#191, CivicDataLab#192
The Hub's dataset endpoint returns `siblings`, one entry per file, by default. For large repos that key dominates the response: 9.6 MB for an 85k-file repo against ~4 KB for everything else. The importer stored the whole response in DatasetSource.raw_metadata and fetched it on every preview and import, although files are never listed. Request fields by name with `expand[]` (every default field except `siblings`, plus `citation`), and drop `siblings` defensively before the payload is kept.
…load DatasetSource no longer keeps the platform's raw JSON. Every field we fetch now lands in a typed column, chosen by one rule: a metadata standard reads it on export (DCAT / Croissant / Dublin Core) or the platform itself reads it (attribution, duplicate check, licence review). New columns: revision (commit hash or version), source_created_at, source_readme (full card; Dataset.description keeps a 1,000-char cut), citation, languages, source_homepage, is_archived. Column definitions a platform declares (Hugging Face dataset_info) become ResourceSchema rows on the link resource, so the columns list works without fetching data. Importers fill what each platform provides: Hugging Face all of the above; GitHub adds one small call for the branch head SHA and reads homepage/archived; Kaggle uses the version number and the earliest version date. Fields nothing reads (downloads, likes, stars, size categories, task taxonomy) are no longer fetched.
Found by importing many real datasets and fuzzing identifiers:
- Hugging Face: keep the platform's canonical id ("imdb" is really
"stanfordnlp/imdb"), so the same dataset cannot be imported twice under
a legacy name. Duplicate check re-runs under the canonical id.
- GitHub: SPDX "NOASSERTION"/"other" means no detectable licence; treat as
empty instead of a licence called NOASSERTION. Branch names and folder
paths are validated (no "..", no whitespace, plain segments only).
- Kaggle: a 403 covers "does not exist" as well as private, so say both.
The created date is taken from the earliest version only when the view
lists every version (it lists one of 2,324 for kaggle/meta-kaggle).
- All: identifiers over 200 characters are rejected; README kept to
200 KB; citation to 20 KB.
- Service: platform calls now happen before the transaction is opened, so
a slow platform never holds a database connection.
GET /api/datasets/<id>/export/?standard=dcat|croissant|dublin_core
&format=jsonld|turtle|rdfxml|ntriples (Croissant: jsonld only)
GET /api/metadata/export-options/
Generated on request from the platform's own model; nothing is stored.
Published datasets are public; owners can preview drafts. `report=1`
returns the document together with the gap report (values the standard
wanted as URIs but got as names, fields the standard cannot carry,
mandatory properties we had nothing for).
The mapping is data, not code. `contracts/crosswalk.json` binds each
concept to a record field and, per standard, to a property with its
value type and obligation; `contracts/*.csv` hold the allowed licences,
sectors and geographies with URIs. Both are owned here and edited like
any other file (see contracts/README.md). `crosswalk.py` is the engine
that applies the file and knows no standard by name; `adapter.py` is the
only module that reads Django models and resolves names to URIs.
Serialisation the mapping cannot express lives in exporter.py: dates as
xsd:date literals, media types as IANA IRIs, the licence repeated on each
distribution, and the @id/contentUrl/encodingFormat Croissant requires on
every FileObject. Imported datasets export with provenance to the source
page and the platform author as creator; the link resource becomes a
distribution with accessURL, never downloadURL.
Not covered yet: Croissant's sha256 on file objects (no file hash is
stored; separate change on the upload path).
New dependency: rdflib (pure Python) for the non-JSON serialisations.
Metadata definitions (label, URN, type) created in Django admin were opaque strings: the platform import could not fill them and the export could not place them. metadata_mapping.py gives each definition a crosswalk concept, matched by URN (ds:createdOn -> created), then by a standard's own property name (dcterms:issued -> issued), then by label. Import: each enabled dataset definition is prefilled with the platform value for its concept (source page, creator, dates, licence, version, homepage, citation, languages); the definition's validators still apply. Export: definition values are read by concept. Core columns always win; a definition only supplies what the model has no column for. ds:createdOn becomes dcterms:created (the data's origin) while our created column stays dcterms:issued (when the record appeared). Definitions the mapping cannot place are listed under `dropped` in the report instead of vanishing.
Use the official Croissant @context (every cr: term, @language, dct) instead of three bare prefixes, so validators recognise the document; type the dataset as sc:Dataset; declare conformsTo 1.1; and emit the source platform identifier as alternateName for imported datasets, as Hugging Face's own Croissant does. No change to the properties' meaning.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Built on top of #204. The last two commits are this PR; the rest belong to #204.
What it does
Any published dataset can now be downloaded as a metadata file in a standard
format: DCAT, Croissant or Dublin Core, as JSON-LD, Turtle, RDF/XML or
N-Triples (Croissant is JSON-LD only). The file is built when requested from
the data we already have. Nothing new is stored.
Owners can export their drafts; everyone else can export published datasets.
Adding
report=1returns the file together with a short report of what couldnot be expressed in that standard.
How the mapping works
The rules for which of our fields becomes which property live in one JSON
file,
api/services/metadata_export/contracts/crosswalk.json, with threeCSVs of allowed licences, sectors and geographies next to it. The code reads
the file; it has no standard-specific logic. To change a mapping, edit the
file in a PR. These files belong to this repo and do not depend on anything
outside it.
Admin-defined metadata fields
Extra dataset fields created in Django admin (for example
ds:source_websiteand
ds:createdOnon dev) now work with import and export. The platformimport fills them in where it can, and the export puts their values under the
right property. A field the mapping does not recognise is left alone and
listed in the report.
Checked
Exports of real datasets from Hugging Face, Kaggle, GitHub and an uploaded
CSV were checked against the DCAT-AP and Croissant requirements and parsed
as valid RDF.
Not in this PR
Follow-up PR on the upload path.
option; product decision.
Deploy notes