All tutorials
    PrepFlow
    Intermediate
    30 min

    Prepare Training Data with PrepFlow

    Build a visual preparation graph for a messy table: drop useless columns, impute, encode and scale, mark the target, split, preview every step, then materialize the result and generate the equivalent Python.

    What you will build

    A PrepFlow graph that turns the Titanic passenger table into model-ready numeric data. The same pattern works for any tabular classification task.

    Before you start

    Step 1: Open PrepFlow with your data

    In the sidebar open the Data Hub, then PrepFlow. Attach the dataset from the Datasets panel (you can load more from the Data Hub there). A Source node appears on the canvas.

    Step 2: Drop columns you will not use

    Search the palette for Keep / Remove Columns and drag it onto the canvas. Connect the Source to it and remove Name, Ticket and Cabin. Click the node: the live preview shows the table after this step. These columns are free text or mostly empty, so a model cannot use them directly.

    Step 3: Fill missing values

    Add Missing Value Imputation (Statistic) and fill Age with its median. Add Missing Value Imputation (Constant) for Embarked if it has gaps. Watch the preview: the missing counts should drop to zero for those columns.

    Step 4: Encode categories

    Add One-Hot Encoding for Sex and Embarked. Each category becomes its own 0/1 column. Alternatives in the palette include Ordinal / Label Encoding and Target / Mean Encoding; pick one-hot here because the categories are few and unordered.

    Step 5: Scale numbers

    Add Z-Score Standardization to Age and Fare. Scaling puts numeric columns on a comparable range, which helps neural networks train.

    Step 6: Mark the target

    Mark Survived as the prediction target. The mark feeds the steps that need a label and becomes the Output node in the Model Builder.

    Step 7: Split

    Add Dataset Split (Train / Val / Test), set the proportions and a seed, and stratify by the target so each partition keeps the same survival rate. The seed makes the split repeatable.

    Step 8: Check each step

    Click through the nodes from the Source to the end. Any step with a problem lists it, for example a column that no longer exists. Use undo and redo from the toolbar, or open the outline view to see the graph as a list.

    Step 9: Materialize

    Press Materialize. The log shows what ran, and the result is saved as a new dataset in the Data Hub, ready to train on. If a Jupyter or Kaggle kernel is attached, the button offers to run on it instead.

    Step 10: Read and keep the code

    Press Python to generate inspectable pandas and scikit-learn code with a requirements file. Read it to confirm it does what you drew. Press Save to keep the chain as a preprocessing artifact the Model Builder can use at prediction time.

    A note on leakage

    Scalers and imputers learn from data. If you fit them on the whole table before splitting, information from the test rows leaks into training. In a pipeline, use Fit preprocessing, which fits on training rows only and replays on the rest. See Build Your First Pipeline.

    PrepFlow and the transform reference.

    Try it in DLWΛY

    Open the Studio and follow along in a real project. There is nothing to install.

    Open Studio