9 min read
Pipeline Step Reference
Every step you can put in a pipeline, grouped by job: reading data, preparing it, training and evaluating models, checking quality, refreshing and writing to outside sources, control flow and delivery.
How steps connect
A pipeline is a graph of steps joined by wires. Each step has typed ports, so the editor only lets you connect an output to an input that accepts it; a step with several inputs labels them (data, model, baseline and so on) so two wires cannot be confused. You can add steps from the Steps tab, drag a project asset (a dataset, a saved model, a PrepFlow workflow) from the Assets tab, or build from a template. See Running pipelines for how a pipeline is saved, published and run.
Steps that are deterministic can reuse their result when neither their settings nor their inputs have changed. Steps that run your code, train a model or talk to the outside world are never reused from a cache.
Reading data
- Read Dataset: brings a dataset from the Data Hub into the pipeline. Choose which version to read: the latest ready version, or a pinned one.
- Apply Recipe: applies a saved Data Hub recipe (from the Transform window) to a dataset.
- Run PrepFlow: runs a whole PrepFlow workflow.
- Use Model: brings a trained model into the pipeline.
- Use Fitted Preprocessing: brings in preprocessing that an earlier step fitted.
Getting data from outside
- Fetch from URL: downloads a file from a public address and only publishes a new version when its content changed.
- Refresh from Kaggle: follows a Kaggle dataset and downloads only when it has a newer version.
- Refresh from Hugging Face: follows a public Hugging Face dataset and downloads only when it has a newer commit.
- Refresh from Database: reads a table of a saved connection, only when its rows changed. It can read only the rows past a growing column, such as an id or timestamp, since the last successful run.
Keeping a table current
- New Rows: passes on only the rows that arrived since the last successful run. With nothing new the step finishes without output and what depends on it is skipped, not failed. Deletions are not passed on.
- Upsert Rows: merges rows into a table by key. A row whose key exists replaces it, a new key is added, everything else stays. Merging the same rows twice gives the same table, so a retry cannot duplicate anything, and a change that adds nothing creates no new version. Large merges run inside the database engine and are written as Parquet, so a table can grow past what a tab holds as rows.
- Replace Partitions: replaces whole partitions. Rows of a partition that arrives again replace its old rows, and partitions that do not arrive stay. Unlike Upsert, a row that vanished from the source also vanishes from its partition, so re-running a time window gives the same table.
- Run SQL: queries the connected datasets and publishes the result as a new version. The upstream data is the table named
data.
Preparing data for ML
- Split Dataset: assigns every row to train, validation or test with a seed, adding a
splitcolumn. - Fit Preprocessing: fits a PrepFlow workflow on the training rows only, then applies it to validation and test rows. Held-out rows never influence the statistics.
- Apply Preprocessing: replays fitted preprocessing on another dataset without refitting.
Training and evaluating
- Train Model: trains a neural network drawn in the Model Builder, in the browser, on the partitions a split produced. The model is stored in the browser and registered with a training record. It supports tabular and image data, classification and regression, and one or several outputs.
- Train Classical Model: trains one of the Model Builder's classical algorithms (linear and logistic regression, decision tree, random forest, AdaBoost, k-nearest neighbours, naive Bayes, SVM) and reports held-out metrics. These produce metrics only, not a model that can be deployed.
- Evaluate Model: scores a model on the test or validation partition and records classification or regression metrics.
- Batch Predict: adds a model's predictions to every row of a dataset. Every input row keeps its place and a
row_id; rows with missing or non-numeric features are kept, left without a prediction and marked, unless you choose to fail the run. - Fine-tune Model: continues training a trained model on new train rows. See Fine-Tuning.
- Export Fine-tune Script: writes a Hugging Face LoRA script and its data as an export bundle. DLWAY does not run it.
Checking quality
- Validate Data: checks a dataset against the rules saved on it in the schema editor (range, unique, not-null), plus required columns, a minimum row count and extra rules. A broken rule fails the step, so nothing downstream runs on bad data, or, when set to warn, the run finishes "with warnings". The dataset passes through unchanged.
- Quality Gate: compares an evaluation's metrics with conditions you set, optionally against a baseline evaluation, and records approve or reject. A missing metric fails its condition. A rejection is not a failure: the step succeeds, and whatever depends on its decision, such as a package, is withheld while the run ends "gated". A baseline comparison is only made between comparable evaluations.
Delivering results
- Refresh Dashboard: gives the Dashboard new data after checking that its fields still match.
- Build Package: packages an approved model as a browser demo page and/or a model file with a hashed manifest, after a one-input smoke check. See Packaging and export.
- Save Files: writes a dataset (as stored or as CSV), export files, a model or metrics into a folder inside the connected project folder. It needs the folder connected and permitted, never writes outside the chosen folder, never overwrites unless told to, and reports a file saved only after reading its size back.
- Write to Database: puts rows into a table of a saved connection. Writing must be allowed on that connection, and the connection's mode decides whether it asks, types, dry-runs or writes automatically. See Database connections.
Your own code
Run Code runs JavaScript or Python over one or more datasets and publishes the rows it returns as new dataset versions. The step keeps its own copy of the script and only runs code whose exact content you have trusted. It runs in a worker created for the step, which separates it from the page but is not a security boundary. Code steps are never cached and never retried; cancelling terminates the worker. Python runs in the browser, or on an attached kernel.
Control flow
- Branch: sends a result down one of two paths by a condition, looking at the artifact itself or at another one wired to its test port (typically an evaluation). The path not taken produces nothing, and steps that depend only on it are skipped as "branch not taken".
- Merge: brings the paths back together. It runs when at least one input was produced.
- Stop: ends the run there, either successfully ("nothing new to process") or as a failure with your message. Put it on a branch so it only fires when that path is taken.
- Run Pipeline: runs another pipeline in the same project (its published version) and waits for it. Its parameters take the values given here, and one of its steps can be handed on as the result. A pipeline cannot run itself directly or through others, and nesting stops at three levels.
Typical recipes
- Incremental load: Refresh from Database → Upsert Rows → Validate Data → Refresh Dashboard.
- Retrain with a gate: Read Dataset → Split → Train Model → Evaluate → Quality Gate → Build Package.
- Keep a table current from a file: Fetch from URL → Upsert Rows → Validate Data.
The Templates dialog builds these shapes for you from the assets you pick.