feat(api): link-only dataset import from Hugging Face, GitHub and Kaggle - #204
Draft
anantjain341 wants to merge 4 commits into
Draft
anantjain341 wants to merge 4 commits into
anantjain341 wants to merge 4 commits into
Conversation
Let a publisher add a dataset that already lives on a third-party platform by giving its identifier (or page URL). Only metadata is fetched; files are never copied or listed, and downloads redirect to the platform. - previewPlatformDataset query: fetch title, description, license, tags, author and last-updated from the platform, with no side effects. - importPlatformDataset mutation: create a DRAFT dataset prefilled from the platform, attach sectors/geographies whose names match the platform's tags, fill matching dataset metadata fields, add one EXTERNAL resource linking to the dataset page, record provenance, and grant the owner role, all in one transaction. Optional title override. Duplicate imports within the same organisation/user are rejected. - DatasetSource model (one-to-one with Dataset) for provenance; exposed as TypeDataset.source. TypeResource.url is now exposed. - Importers for Hugging Face, GitHub and Kaggle behind a registry; all work without API keys for public datasets. HF_TOKEN, GITHUB_TOKEN and KAGGLE_USERNAME/KAGGLE_KEY are optional. - Download view redirects EXTERNAL resources to their URL and returns 404 instead of raising when a resource has no file. - Dataset search document gains source_platform; /api/search/dataset/ returns it, aggregates on it and filters by it (NATIVE = not imported). formats indexing skips resources without file details. After deploy, run `manage.py search_index --rebuild` once so Elasticsearch maps source_platform as a keyword before the first import is indexed. Refs CivicDataLab/DataSpace#174, #190, #191, #192
The Hub's dataset endpoint returns `siblings`, one entry per file, by default. For large repos that key dominates the response: 9.6 MB for an 85k-file repo against ~4 KB for everything else. The importer stored the whole response in DatasetSource.raw_metadata and fetched it on every preview and import, although files are never listed. Request fields by name with `expand[]` (every default field except `siblings`, plus `citation`), and drop `siblings` defensively before the payload is kept.
…load DatasetSource no longer keeps the platform's raw JSON. Every field we fetch now lands in a typed column, chosen by one rule: a metadata standard reads it on export (DCAT / Croissant / Dublin Core) or the platform itself reads it (attribution, duplicate check, licence review). New columns: revision (commit hash or version), source_created_at, source_readme (full card; Dataset.description keeps a 1,000-char cut), citation, languages, source_homepage, is_archived. Column definitions a platform declares (Hugging Face dataset_info) become ResourceSchema rows on the link resource, so the columns list works without fetching data. Importers fill what each platform provides: Hugging Face all of the above; GitHub adds one small call for the branch head SHA and reads homepage/archived; Kaggle uses the version number and the earliest version date. Fields nothing reads (downloads, likes, stars, size categories, task taxonomy) are no longer fetched.
Found by importing many real datasets and fuzzing identifiers:
- Hugging Face: keep the platform's canonical id ("imdb" is really
"stanfordnlp/imdb"), so the same dataset cannot be imported twice under
a legacy name. Duplicate check re-runs under the canonical id.
- GitHub: SPDX "NOASSERTION"/"other" means no detectable licence; treat as
empty instead of a licence called NOASSERTION. Branch names and folder
paths are validated (no "..", no whitespace, plain segments only).
- Kaggle: a 403 covers "does not exist" as well as private, so say both.
The created date is taken from the earliest version only when the view
lists every version (it lists one of 2,324 for kaggle/meta-kaggle).
- All: identifiers over 200 characters are rejected; README kept to
200 KB; citation to 20 KB.
- Service: platform calls now happen before the transaction is opened, so
a slow platform never holds a database connection.
anantjain341
force-pushed
the
feat/platform-based-import
branch
from
September 24, 2026 13:38
9f46470 to
f7fe195
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Backend for platform-based dataset import (CivicDataLab/DataSpace#174, sub-issues #190 Hugging Face, #191 Kaggle, #192 GitHub).
A publisher gives a platform and a dataset identifier or URL. We fetch the dataset's metadata and create a DRAFT dataset prefilled from it, with one link back to the source. Files are never copied, listed, or counted, and downloads redirect to the platform. No API key is needed for public datasets on any of the three platforms.
API
previewPlatformDataset(platform, identifier): fetches title, description, author, license, mapped license, tags, last updated. Writes nothing. Requires a logged-in user.importPlatformDataset(importInput: { platform, identifier, title? }): creates the draft, attaches sectors and geographies whose names match the platform's tags, fills matching dataset metadata fields, adds oneEXTERNALresource pointing at the dataset page, records provenance, grants the owner role. One transaction. Same permission andorganization/dataspaceheaders asaddDataset. Duplicate imports in the same organisation or by the same user are rejected.TypeDataset.source(null for native datasets) andTypeResource.urlare new read fields.Also changed
DatasetSourcemodel, migration0048_platform_import.EXTERNALresources to their URL, and returns 404 instead of raising when a resource has no file.source_platform./api/search/dataset/returns it, aggregates on it, and filters with?source_platform=HUGGINGFACE,GITHUB,KAGGLEorNATIVE.formatsindexing skips resources without file details.Deploy note
Run
manage.py search_index --rebuildonce after deploy, so Elasticsearch mapssource_platformas a keyword before the first imported dataset is indexed. Until a dataset is imported, nothing changes for existing behaviour.Status
Draft. Backend is functionally complete and additive. Frontend is tracked separately in CivicDataLab/DataSpace#251. Not yet built: re-sync with the source platform.
Flow and API write-up: https://docs.google.com/document/d/1OJgSkZcmVvKajx4e8rVcOn1qmG5Rnw9LQ2J5795Mqio/edit?usp=sharing