Import Datasets from Kaggle and Hugging Face
Connect a Kaggle account, attach a dataset from a ZIP, then add a Hugging Face Parquet split both as an import and as a live remote link.
What you will do
Bring one dataset from each hub into your project, and learn when to import and when to link.
Before you start
- Signed in to DLWAY (Kaggle credentials are stored encrypted on your account)
- A Kaggle account and an API token from your Kaggle settings
- Optional: a Hugging Face token (
hf_...), only needed for private or gated datasets
Part 1: Kaggle
Step 1: Connect
In the Data Hub choose Import from Kaggle. The first time, choose Connect your Kaggle account, paste your token and save. DLWAY tests it and shows Connected as with your username.
Step 2: Find a dataset
Search for iris, or paste a kaggle.com/datasets address, or type uciml/iris. Pick the result and read its file list, licence and size before downloading anything.
Step 3: Attach it
Choose Attach from Kaggle. If the download is a ZIP, DLWAY extracts it and names every dataset it will create before you confirm. Confirm, and each file appears as its own dataset with its Kaggle provenance recorded.
Part 2: Hugging Face
Step 4: Find a dataset
Choose Add from Hugging Face. Search for scikit-learn/iris and press Enter, or paste a huggingface.co/datasets address. Pick it to see its splits.
Step 5: Import or link
For a Parquet split you have two choices:
- Import streams the file into the project's storage. Use this for data you will transform, or that you need offline.
- Link (remote) adds a live reference. Queries read just the byte ranges they need over HTTP, so a very large split costs almost nothing to attach.
Add the split both ways to compare them. Linked datasets keep their Hugging Face provenance, and an up-to-date check compares them with the upstream commit.
Step 6: Private and gated data
For a private or gated dataset, enter your token when asked. It is sent only to huggingface.co and is never stored in the DLWAY server. Leave Remember on this device on if you want it to survive a reload. For a gated dataset, accept its terms on huggingface.co first.
Check your work
Open each dataset. Profile one and look at its Versions tab: new imports start at v1. Open Lineage to see the source recorded for each.
Keep it fresh
A pipeline can follow either source. Add Refresh from Kaggle or Refresh from Hugging Face to download only when the upstream has a newer version. See Schedule an Incremental Data Refresh.