4 min read
Kaggle and Hugging Face
Search, preview and import datasets from Kaggle and Hugging Face, understand how credentials are stored, and choose between importing Parquet splits and linking them live.
Kaggle
Connect your account
Open Import from Kaggle (from the Data Hub toolbar or the Add data wizard). The first time, choose Connect your Kaggle account and paste a Kaggle API token. Older username-and-key credentials still work: enter the username and key in the two fields. DLWAY tests the credential, shows Connected as and keeps the result of the last check. The credential is encrypted on the server and is never sent back to the browser; you must be signed in to DLWAY to use it.
Search, preview, attach
Search for datasets by keyword, or paste a kaggle.com/datasets URL or an owner/slug. Pick a result to see its files, licence and size before attaching. Choose Attach from Kaggle to download the files into your project.
If the download is a ZIP, its contents are extracted. Every file inside becomes its own dataset, and each one keeps its Kaggle provenance. The confirmation step names the children it will create.
Keeping up to date
Datasets remember where they came from. A pipeline can follow a Kaggle dataset with Refresh from Kaggle, which downloads only when Kaggle has a newer version.
Hugging Face
Find a dataset
Choose Add from Hugging Face, search for a dataset (press Enter), or paste a huggingface.co/datasets URL or owner/name. Pick a dataset to see its splits before adding it. Public repositories need no token. For private or gated datasets, enter a token that starts with hf_.
Where the token lives: it is sent only to huggingface.co as an authorisation header, never to the DLWAY server, and never stored in job records, bundles or logs. Remember on this device (the default) keeps it in the browser's storage so it survives a reload; untick it to keep the token in memory for the session only.
Import or link
You can add a Parquet split two ways:
- Import: stream the file into the project's browser storage. It is never fully buffered in memory, so large splits are fine.
- Link (remote): attach the split as a live reference. Queries read only the byte ranges they need over HTTP, so nothing large is downloaded up front.
If the Hub has converted only part of a dataset to Parquet, DLWAY says the rows available are a subset. For gated datasets you must first accept the terms on huggingface.co.
Imported and linked datasets show their Hugging Face provenance, and an up-to-date check compares them with the upstream commit.
Models
Hugging Face also supplies pretrained models for Deployment (run in the browser through transformers.js) and for Fine-Tuning. Searching models works without signing in.
Privacy
Searching and downloading contacts Kaggle or Hugging Face; that is the only time these imports touch the network. See Privacy and storage.