Skip to content

feat(api): link-only dataset import from Hugging Face, GitHub and Kaggle - #204

Draft
anantjain341 wants to merge 4 commits into
devfrom
feat/platform-based-import
Draft

anantjain341 wants to merge 4 commits into
devfrom
feat/platform-based-import

Conversation

@anantjain341

Copy link
Copy Markdown

What

Backend for platform-based dataset import (CivicDataLab/DataSpace#174, sub-issues #190 Hugging Face, #191 Kaggle, #192 GitHub).

A publisher gives a platform and a dataset identifier or URL. We fetch the dataset's metadata and create a DRAFT dataset prefilled from it, with one link back to the source. Files are never copied, listed, or counted, and downloads redirect to the platform. No API key is needed for public datasets on any of the three platforms.

API

  • previewPlatformDataset(platform, identifier): fetches title, description, author, license, mapped license, tags, last updated. Writes nothing. Requires a logged-in user.
  • importPlatformDataset(importInput: { platform, identifier, title? }): creates the draft, attaches sectors and geographies whose names match the platform's tags, fills matching dataset metadata fields, adds one EXTERNAL resource pointing at the dataset page, records provenance, grants the owner role. One transaction. Same permission and organization / dataspace headers as addDataset. Duplicate imports in the same organisation or by the same user are rejected.
  • TypeDataset.source (null for native datasets) and TypeResource.url are new read fields.

Also changed

  • New DatasetSource model, migration 0048_platform_import.
  • Download view redirects EXTERNAL resources to their URL, and returns 404 instead of raising when a resource has no file.
  • Dataset search document gains source_platform. /api/search/dataset/ returns it, aggregates on it, and filters with ?source_platform=HUGGINGFACE,GITHUB,KAGGLE or NATIVE.
  • formats indexing skips resources without file details.

Deploy note

Run manage.py search_index --rebuild once after deploy, so Elasticsearch maps source_platform as a keyword before the first imported dataset is indexed. Until a dataset is imported, nothing changes for existing behaviour.

Status

Draft. Backend is functionally complete and additive. Frontend is tracked separately in CivicDataLab/DataSpace#251. Not yet built: re-sync with the source platform.

Flow and API write-up: https://docs.google.com/document/d/1OJgSkZcmVvKajx4e8rVcOn1qmG5Rnw9LQ2J5795Mqio/edit?usp=sharing

Let a publisher add a dataset that already lives on a third-party platform
by giving its identifier (or page URL). Only metadata is fetched; files are
never copied or listed, and downloads redirect to the platform.

- previewPlatformDataset query: fetch title, description, license, tags,
  author and last-updated from the platform, with no side effects.
- importPlatformDataset mutation: create a DRAFT dataset prefilled from the
  platform, attach sectors/geographies whose names match the platform's
  tags, fill matching dataset metadata fields, add one EXTERNAL resource
  linking to the dataset page, record provenance, and grant the owner role,
  all in one transaction. Optional title override. Duplicate imports within
  the same organisation/user are rejected.
- DatasetSource model (one-to-one with Dataset) for provenance; exposed as
  TypeDataset.source. TypeResource.url is now exposed.
- Importers for Hugging Face, GitHub and Kaggle behind a registry; all work
  without API keys for public datasets. HF_TOKEN, GITHUB_TOKEN and
  KAGGLE_USERNAME/KAGGLE_KEY are optional.
- Download view redirects EXTERNAL resources to their URL and returns 404
  instead of raising when a resource has no file.
- Dataset search document gains source_platform; /api/search/dataset/
  returns it, aggregates on it and filters by it (NATIVE = not imported).
  formats indexing skips resources without file details.

After deploy, run `manage.py search_index --rebuild` once so Elasticsearch
maps source_platform as a keyword before the first import is indexed.

Refs CivicDataLab/DataSpace#174, #190, #191, #192
The Hub's dataset endpoint returns `siblings`, one entry per file, by
default. For large repos that key dominates the response: 9.6 MB for an
85k-file repo against ~4 KB for everything else. The importer stored the
whole response in DatasetSource.raw_metadata and fetched it on every
preview and import, although files are never listed.

Request fields by name with `expand[]` (every default field except
`siblings`, plus `citation`), and drop `siblings` defensively before the
payload is kept.
…load

DatasetSource no longer keeps the platform's raw JSON. Every field we
fetch now lands in a typed column, chosen by one rule: a metadata
standard reads it on export (DCAT / Croissant / Dublin Core) or the
platform itself reads it (attribution, duplicate check, licence review).

New columns: revision (commit hash or version), source_created_at,
source_readme (full card; Dataset.description keeps a 1,000-char cut),
citation, languages, source_homepage, is_archived. Column definitions a
platform declares (Hugging Face dataset_info) become ResourceSchema rows
on the link resource, so the columns list works without fetching data.

Importers fill what each platform provides: Hugging Face all of the
above; GitHub adds one small call for the branch head SHA and reads
homepage/archived; Kaggle uses the version number and the earliest
version date. Fields nothing reads (downloads, likes, stars, size
categories, task taxonomy) are no longer fetched.
Found by importing many real datasets and fuzzing identifiers:

- Hugging Face: keep the platform's canonical id ("imdb" is really
  "stanfordnlp/imdb"), so the same dataset cannot be imported twice under
  a legacy name. Duplicate check re-runs under the canonical id.
- GitHub: SPDX "NOASSERTION"/"other" means no detectable licence; treat as
  empty instead of a licence called NOASSERTION. Branch names and folder
  paths are validated (no "..", no whitespace, plain segments only).
- Kaggle: a 403 covers "does not exist" as well as private, so say both.
  The created date is taken from the earliest version only when the view
  lists every version (it lists one of 2,324 for kaggle/meta-kaggle).
- All: identifiers over 200 characters are rejected; README kept to
  200 KB; citation to 20 KB.
- Service: platform calls now happen before the transaction is opened, so
  a slow platform never holds a database connection.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant