6 min read
Importing Data
Every way to bring data into the Data Hub: files, folders, URLs and pasted text, the supported formats, parse options, duplicate handling and the size limits to know.
The Add data wizard
Choose Add data in the Data Hub to open a full-screen wizard with four steps: Source (where the data lives), Files & parse (pick what to attach and review parse options), Structure (only when a folder with a complex layout is detected) and Confirm (name, collection and visibility). The wizard can also open straight on a source's file step from the library's Kaggle and From URL buttons.
You can also drop files or a folder onto the library, use Add files, Add folder, Import from URL, or Connect database from the toolbar.
Sources
- Files: one or many files from disk.
- A folder: loaded as one structured dataset. DLWAY looks at the paths and suggests a layout. See Folders and structure.
- A URL: a public address to a file. Invalid URLs, unsupported protocols and HTTP errors such as 404 show a visible error and keep the confirm button disabled.
- Kaggle and Hugging Face: see Kaggle and Hugging Face.
- A live database: see Database connections.
Supported formats
The format is detected from the file's extension and its content.
- Tables: CSV, TSV, pipe-separated (PSV), fixed-width text, JSON, JSON Lines, Excel (.xlsx, .xls, .xlsm, .xlsb), Parquet, Arrow and Feather, Avro, Apache ORC, SQLite, Stata, SAS (.sas7bdat) and MATLAB files up to v7.2
- Scientific and array data: HDF5, NetCDF, NumPy (.npy, .npz), SafeTensors, ONNX models, TFRecord, NIfTI and DICOM
- Geospatial: GeoJSON, TopoJSON, KML, KMZ and GeoTIFF
- Documents: PDF, Word (.docx), Rich Text, Markdown, HTML, EPUB, XML, plain text, and subtitle files (SubRip and WebVTT)
- Images, audio and video: common image formats including HEIC, TIFF and SVG; WAV, MP3, FLAC, OGG, Opus, M4A and AAC; MP4, WebM, MKV, MOV and AVI
- 3D: PLY, OBJ, STL, glTF and GLB
- Annotations: COCO, YOLO, Pascal VOC and LabelMe
- Archives: ZIP, TAR and gzip, whose contents become datasets of their own
Documents, media, images and meshes are imported as a table of metadata, one row per file. For example, an MP3 and an MP4 each become a one-row table with duration, sample rate and channel information, and a PDF becomes a row with its extracted text and an OCR flag. Video support reads container metadata; there is no playback or frame extraction.
Parse options
For delimited and similar text files you can adjust how the file is read before confirming: the delimiter, whether the first row is a header, the text encoding (including UTF-16) and whether to read every column as text instead of guessing types. Your choices are saved with the dataset and carried into profiles, bundles and re-imports.
Large delimited files (over about 2 MB) are read with a streaming parser. It requires UTF-8, a header in the first row and double quotes, and it infers column types; the delimiter remains editable, and the import screen tells you when these rules apply.
Duplicates
Each file is hashed before it is stored and compared with what the project already holds. If it is a duplicate, you choose:
- Keep copy: import it as a new dataset.
- Link it: create the dataset over the existing stored bytes, copying nothing.
Previews come first
Importing opens the dataset's preview. Full profiling is optional: choose Profile dataset when you want statistics. Small CSV, TSV and PSV tables use a bounded in-memory path; larger and native formats are read by DuckDB. The bundled Iris sample is available for trying things out.
Stored but not read
Python pickle files (.pkl, .pickle, .pt, .pth, .joblib) are stored but never opened, because loading a pickle can run arbitrary code. Convert the contents to SafeTensors or NumPy first. ONNX models are previewed (their structure is shown) but not loaded as a table.
Limits to know
- Formats DuckDB cannot read natively are parsed in memory with a bounded path: 16 MB or 200,000 rows. For larger data, convert to Parquet first.
- Some Transform steps work only on the bounded preview. SQL and native formats support much larger tables.
- A random split of a folder dataset treats each file as a sample. Use grouping or folder rules to keep an image with its label.
- Cloud object-store imports (S3, GCS, Azure) and shared dataset visibility are not available.
After you import
Open the dataset to see its profile, versions and lineage, clean it in the Transform window, or send it to PrepFlow.