All documentation pages

    4 min read

    Data Hub

    Load files and folders, connect live databases, and import datasets from Kaggle and Hugging Face into one versioned library.

    Connecting data

    Open individual files, an entire folder, or connect a live database (Postgres, MySQL, MongoDB, Snowflake, BigQuery, and more). Dozens of formats are supported, including CSV, Parquet, Arrow, COCO/YOLO annotations, and DICOM.

    Dataset hubs

    Search Kaggle and Hugging Face from inside the Data Hub and pull a dataset straight into your project, without downloading and re-uploading files by hand.

    Preview and filter

    Datasets render in a responsive, virtualized grid. Parsing happens in a dedicated Web Worker so large files don't block the interface, and SQL runs locally on DuckDB compiled to WebAssembly.

    Profiles, versions and lineage

    Each dataset keeps a profile of its columns, a history of versions, and a lineage graph showing what it was derived from, so a change can be traced or undone.

    Using data elsewhere

    Once a dataset is loaded, it becomes available to PrepFlow, to the Model Builder for training, and to the Dashboard for chart building, with no re-uploading.

    Go deeper