feat(api): store a SHA-256 for uploaded resource files - #212
Draft
anantjain341 wants to merge 8 commits into
Draft
anantjain341 wants to merge 8 commits into
anantjain341 wants to merge 8 commits into
Conversation
Let a publisher add a dataset that already lives on a third-party platform by giving its identifier (or page URL). Only metadata is fetched; files are never copied or listed, and downloads redirect to the platform. - previewPlatformDataset query: fetch title, description, license, tags, author and last-updated from the platform, with no side effects. - importPlatformDataset mutation: create a DRAFT dataset prefilled from the platform, attach sectors/geographies whose names match the platform's tags, fill matching dataset metadata fields, add one EXTERNAL resource linking to the dataset page, record provenance, and grant the owner role, all in one transaction. Optional title override. Duplicate imports within the same organisation/user are rejected. - DatasetSource model (one-to-one with Dataset) for provenance; exposed as TypeDataset.source. TypeResource.url is now exposed. - Importers for Hugging Face, GitHub and Kaggle behind a registry; all work without API keys for public datasets. HF_TOKEN, GITHUB_TOKEN and KAGGLE_USERNAME/KAGGLE_KEY are optional. - Download view redirects EXTERNAL resources to their URL and returns 404 instead of raising when a resource has no file. - Dataset search document gains source_platform; /api/search/dataset/ returns it, aggregates on it and filters by it (NATIVE = not imported). formats indexing skips resources without file details. After deploy, run `manage.py search_index --rebuild` once so Elasticsearch maps source_platform as a keyword before the first import is indexed. Refs CivicDataLab/DataSpace#174, CivicDataLab#190, CivicDataLab#191, CivicDataLab#192
The Hub's dataset endpoint returns `siblings`, one entry per file, by default. For large repos that key dominates the response: 9.6 MB for an 85k-file repo against ~4 KB for everything else. The importer stored the whole response in DatasetSource.raw_metadata and fetched it on every preview and import, although files are never listed. Request fields by name with `expand[]` (every default field except `siblings`, plus `citation`), and drop `siblings` defensively before the payload is kept.
…load DatasetSource no longer keeps the platform's raw JSON. Every field we fetch now lands in a typed column, chosen by one rule: a metadata standard reads it on export (DCAT / Croissant / Dublin Core) or the platform itself reads it (attribution, duplicate check, licence review). New columns: revision (commit hash or version), source_created_at, source_readme (full card; Dataset.description keeps a 1,000-char cut), citation, languages, source_homepage, is_archived. Column definitions a platform declares (Hugging Face dataset_info) become ResourceSchema rows on the link resource, so the columns list works without fetching data. Importers fill what each platform provides: Hugging Face all of the above; GitHub adds one small call for the branch head SHA and reads homepage/archived; Kaggle uses the version number and the earliest version date. Fields nothing reads (downloads, likes, stars, size categories, task taxonomy) are no longer fetched.
Found by importing many real datasets and fuzzing identifiers:
- Hugging Face: keep the platform's canonical id ("imdb" is really
"stanfordnlp/imdb"), so the same dataset cannot be imported twice under
a legacy name. Duplicate check re-runs under the canonical id.
- GitHub: SPDX "NOASSERTION"/"other" means no detectable licence; treat as
empty instead of a licence called NOASSERTION. Branch names and folder
paths are validated (no "..", no whitespace, plain segments only).
- Kaggle: a 403 covers "does not exist" as well as private, so say both.
The created date is taken from the earliest version only when the view
lists every version (it lists one of 2,324 for kaggle/meta-kaggle).
- All: identifiers over 200 characters are rejected; README kept to
200 KB; citation to 20 KB.
- Service: platform calls now happen before the transaction is opened, so
a slow platform never holds a database connection.
GET /api/datasets/<id>/export/?standard=dcat|croissant|dublin_core
&format=jsonld|turtle|rdfxml|ntriples (Croissant: jsonld only)
GET /api/metadata/export-options/
Generated on request from the platform's own model; nothing is stored.
Published datasets are public; owners can preview drafts. `report=1`
returns the document together with the gap report (values the standard
wanted as URIs but got as names, fields the standard cannot carry,
mandatory properties we had nothing for).
The mapping is data, not code. `contracts/crosswalk.json` binds each
concept to a record field and, per standard, to a property with its
value type and obligation; `contracts/*.csv` hold the allowed licences,
sectors and geographies with URIs. Both are owned here and edited like
any other file (see contracts/README.md). `crosswalk.py` is the engine
that applies the file and knows no standard by name; `adapter.py` is the
only module that reads Django models and resolves names to URIs.
Serialisation the mapping cannot express lives in exporter.py: dates as
xsd:date literals, media types as IANA IRIs, the licence repeated on each
distribution, and the @id/contentUrl/encodingFormat Croissant requires on
every FileObject. Imported datasets export with provenance to the source
page and the platform author as creator; the link resource becomes a
distribution with accessURL, never downloadURL.
Not covered yet: Croissant's sha256 on file objects (no file hash is
stored; separate change on the upload path).
New dependency: rdflib (pure Python) for the non-JSON serialisations.
Metadata definitions (label, URN, type) created in Django admin were opaque strings: the platform import could not fill them and the export could not place them. metadata_mapping.py gives each definition a crosswalk concept, matched by URN (ds:createdOn -> created), then by a standard's own property name (dcterms:issued -> issued), then by label. Import: each enabled dataset definition is prefilled with the platform value for its concept (source page, creator, dates, licence, version, homepage, citation, languages); the definition's validators still apply. Export: definition values are read by concept. Core columns always win; a definition only supplies what the model has no column for. ds:createdOn becomes dcterms:created (the data's origin) while our created column stays dcterms:issued (when the record appeared). Definitions the mapping cannot place are listed under `dropped` in the report instead of vanishing.
Use the official Croissant @context (every cr: term, @language, dct) instead of three bare prefixes, so validators recognise the document; type the dataset as sc:Dataset; declare conformsTo 1.1; and emit the source platform identifier as alternateName for imported datasets, as Hugging Face's own Croissant does. No change to the properties' meaning.
Croissant requires a sha256 on every FileObject and DCAT-AP has a spdx:checksum slot for the same fact; we stored no file hash, so the Croissant export of an uploaded dataset could not pass a validator. - ResourceFileDetails.sha256 (nullable, 64 chars), migration 0049: one ALTER TABLE, no data rewrite. - A pre_save hook computes the hash when a file is saved, streaming it in chunks. It is skipped when the stored file is unchanged, and an unreadable file logs a warning instead of blocking the save. - `manage.py backfill_file_hashes [--dry-run]` fills the hash for files uploaded before this change. Writes through a queryset update so the DVC versioning signal does not run. - GraphQL fileDetails exposes sha256. - Exports: Croissant FileObject.sha256; DCAT distribution spdx:checksum as an spdx:Checksum node (algorithm sha256). Link-only imported datasets are unaffected: there is no file to hash and the property is omitted. Deploy: run migrations, then `backfill_file_hashes` once.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Built on top of #211. Only the last commit is this PR.
Croissant needs a sha256 hash on every file it describes. We never stored
one, so the Croissant export of an uploaded dataset would fail a validator.
This adds it.
What changes:
sha256column on file details. Migration 0049 just adds the column.not changed, nothing is recomputed. If the file cannot be read, we log a
warning and the save still goes through.
python manage.py backfill_file_hashesto fill the hash forfiles uploaded before this change.
--dry-runonly counts them. It doesnot trigger DVC versioning.
sha256is available onfileDetailsin GraphQL.Imported datasets are not affected. We do not hold their files, so there is
nothing to hash.
Cost: one extra read of the file at upload time. DVC already reads it once.
After deploy: run migrations, then run
backfill_file_hashesonce.