Skip to content

feat(api): store a SHA-256 for uploaded resource files - #212

Draft
anantjain341 wants to merge 8 commits into
CivicDataLab:devfrom
anantjain341:feat/resource-file-hash
Draft

anantjain341 wants to merge 8 commits into
CivicDataLab:devfrom
anantjain341:feat/resource-file-hash

Conversation

@anantjain341

Copy link
Copy Markdown

Built on top of #211. Only the last commit is this PR.

Croissant needs a sha256 hash on every file it describes. We never stored
one, so the Croissant export of an uploaded dataset would fail a validator.
This adds it.

What changes:

  • New sha256 column on file details. Migration 0049 just adds the column.
  • When a file is saved, we read it once and store its hash. If the file has
    not changed, nothing is recomputed. If the file cannot be read, we log a
    warning and the save still goes through.
  • New command python manage.py backfill_file_hashes to fill the hash for
    files uploaded before this change. --dry-run only counts them. It does
    not trigger DVC versioning.
  • sha256 is available on fileDetails in GraphQL.
  • The Croissant and DCAT exports include the hash for uploaded files.

Imported datasets are not affected. We do not hold their files, so there is
nothing to hash.

Cost: one extra read of the file at upload time. DVC already reads it once.

After deploy: run migrations, then run backfill_file_hashes once.

Let a publisher add a dataset that already lives on a third-party platform
by giving its identifier (or page URL). Only metadata is fetched; files are
never copied or listed, and downloads redirect to the platform.

- previewPlatformDataset query: fetch title, description, license, tags,
  author and last-updated from the platform, with no side effects.
- importPlatformDataset mutation: create a DRAFT dataset prefilled from the
  platform, attach sectors/geographies whose names match the platform's
  tags, fill matching dataset metadata fields, add one EXTERNAL resource
  linking to the dataset page, record provenance, and grant the owner role,
  all in one transaction. Optional title override. Duplicate imports within
  the same organisation/user are rejected.
- DatasetSource model (one-to-one with Dataset) for provenance; exposed as
  TypeDataset.source. TypeResource.url is now exposed.
- Importers for Hugging Face, GitHub and Kaggle behind a registry; all work
  without API keys for public datasets. HF_TOKEN, GITHUB_TOKEN and
  KAGGLE_USERNAME/KAGGLE_KEY are optional.
- Download view redirects EXTERNAL resources to their URL and returns 404
  instead of raising when a resource has no file.
- Dataset search document gains source_platform; /api/search/dataset/
  returns it, aggregates on it and filters by it (NATIVE = not imported).
  formats indexing skips resources without file details.

After deploy, run `manage.py search_index --rebuild` once so Elasticsearch
maps source_platform as a keyword before the first import is indexed.

Refs CivicDataLab/DataSpace#174, CivicDataLab#190, CivicDataLab#191, CivicDataLab#192
The Hub's dataset endpoint returns `siblings`, one entry per file, by
default. For large repos that key dominates the response: 9.6 MB for an
85k-file repo against ~4 KB for everything else. The importer stored the
whole response in DatasetSource.raw_metadata and fetched it on every
preview and import, although files are never listed.

Request fields by name with `expand[]` (every default field except
`siblings`, plus `citation`), and drop `siblings` defensively before the
payload is kept.
…load

DatasetSource no longer keeps the platform's raw JSON. Every field we
fetch now lands in a typed column, chosen by one rule: a metadata
standard reads it on export (DCAT / Croissant / Dublin Core) or the
platform itself reads it (attribution, duplicate check, licence review).

New columns: revision (commit hash or version), source_created_at,
source_readme (full card; Dataset.description keeps a 1,000-char cut),
citation, languages, source_homepage, is_archived. Column definitions a
platform declares (Hugging Face dataset_info) become ResourceSchema rows
on the link resource, so the columns list works without fetching data.

Importers fill what each platform provides: Hugging Face all of the
above; GitHub adds one small call for the branch head SHA and reads
homepage/archived; Kaggle uses the version number and the earliest
version date. Fields nothing reads (downloads, likes, stars, size
categories, task taxonomy) are no longer fetched.
Found by importing many real datasets and fuzzing identifiers:

- Hugging Face: keep the platform's canonical id ("imdb" is really
  "stanfordnlp/imdb"), so the same dataset cannot be imported twice under
  a legacy name. Duplicate check re-runs under the canonical id.
- GitHub: SPDX "NOASSERTION"/"other" means no detectable licence; treat as
  empty instead of a licence called NOASSERTION. Branch names and folder
  paths are validated (no "..", no whitespace, plain segments only).
- Kaggle: a 403 covers "does not exist" as well as private, so say both.
  The created date is taken from the earliest version only when the view
  lists every version (it lists one of 2,324 for kaggle/meta-kaggle).
- All: identifiers over 200 characters are rejected; README kept to
  200 KB; citation to 20 KB.
- Service: platform calls now happen before the transaction is opened, so
  a slow platform never holds a database connection.
GET /api/datasets/<id>/export/?standard=dcat|croissant|dublin_core
    &format=jsonld|turtle|rdfxml|ntriples   (Croissant: jsonld only)
GET /api/metadata/export-options/

Generated on request from the platform's own model; nothing is stored.
Published datasets are public; owners can preview drafts. `report=1`
returns the document together with the gap report (values the standard
wanted as URIs but got as names, fields the standard cannot carry,
mandatory properties we had nothing for).

The mapping is data, not code. `contracts/crosswalk.json` binds each
concept to a record field and, per standard, to a property with its
value type and obligation; `contracts/*.csv` hold the allowed licences,
sectors and geographies with URIs. Both are owned here and edited like
any other file (see contracts/README.md). `crosswalk.py` is the engine
that applies the file and knows no standard by name; `adapter.py` is the
only module that reads Django models and resolves names to URIs.

Serialisation the mapping cannot express lives in exporter.py: dates as
xsd:date literals, media types as IANA IRIs, the licence repeated on each
distribution, and the @id/contentUrl/encodingFormat Croissant requires on
every FileObject. Imported datasets export with provenance to the source
page and the platform author as creator; the link resource becomes a
distribution with accessURL, never downloadURL.

Not covered yet: Croissant's sha256 on file objects (no file hash is
stored; separate change on the upload path).

New dependency: rdflib (pure Python) for the non-JSON serialisations.
Metadata definitions (label, URN, type) created in Django admin were
opaque strings: the platform import could not fill them and the export
could not place them. metadata_mapping.py gives each definition a
crosswalk concept, matched by URN (ds:createdOn -> created), then by a
standard's own property name (dcterms:issued -> issued), then by label.

Import: each enabled dataset definition is prefilled with the platform
value for its concept (source page, creator, dates, licence, version,
homepage, citation, languages); the definition's validators still apply.

Export: definition values are read by concept. Core columns always win;
a definition only supplies what the model has no column for. ds:createdOn
becomes dcterms:created (the data's origin) while our created column
stays dcterms:issued (when the record appeared). Definitions the mapping
cannot place are listed under `dropped` in the report instead of vanishing.
Use the official Croissant @context (every cr: term, @language, dct)
instead of three bare prefixes, so validators recognise the document;
type the dataset as sc:Dataset; declare conformsTo 1.1; and emit the
source platform identifier as alternateName for imported datasets, as
Hugging Face's own Croissant does. No change to the properties' meaning.
Croissant requires a sha256 on every FileObject and DCAT-AP has a
spdx:checksum slot for the same fact; we stored no file hash, so the
Croissant export of an uploaded dataset could not pass a validator.

- ResourceFileDetails.sha256 (nullable, 64 chars), migration 0049: one
  ALTER TABLE, no data rewrite.
- A pre_save hook computes the hash when a file is saved, streaming it
  in chunks. It is skipped when the stored file is unchanged, and an
  unreadable file logs a warning instead of blocking the save.
- `manage.py backfill_file_hashes [--dry-run]` fills the hash for files
  uploaded before this change. Writes through a queryset update so the
  DVC versioning signal does not run.
- GraphQL fileDetails exposes sha256.
- Exports: Croissant FileObject.sha256; DCAT distribution spdx:checksum
  as an spdx:Checksum node (algorithm sha256).

Link-only imported datasets are unaffected: there is no file to hash and
the property is omitted.

Deploy: run migrations, then `backfill_file_hashes` once.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant