4 min read
The Data Model
Draw relationships between datasets on a canvas, validate them for orphan rows, define reusable measures such as SUM(orders.amount), and export the model as JSON, SQL DDL, Pandera or Pydantic.
Model relationships across datasets
Choose Model relationships across your datasets in the Data Hub to open the data model. Each dataset is drawn as a table node listing its columns. Select the datasets you want on the canvas, then connect them:
- Drag between two fields to draw a relationship, or click an existing line to inspect it.
- Set the cardinality of the relationship: one-to-one, one-to-many or many-to-many.
- The canvas uses the same pointer tools and clipboard shortcuts as the other canvases. See Keyboard shortcuts.
The Studio also suggests relationships by comparing column names and values. Suggestions are drawn dashed and labelled suggested until you accept them.
Validate a relationship
A relationship can be validated against the data. The check measures referential integrity in both directions: how many rows on each side have no partner (orphans) and whether the key is really unique where it should be. For example, linking two Iris tables by Id reports a one-to-one relationship with zero orphans both ways.
Measures
A measure is a named formula such as SUM(orders.amount) that you define once. Use New measure, give it a name and a formula over fields of the modelled datasets. Formulas are checked as you type, and table names that need quoting are quoted for you. A measure can be previewed: it is computed against the live tables, so you see a real number before you rely on it. Measures are saved with the model and persist after reload.
Export the model
Export model writes the relationships and column information in the form you need:
- JSON: the model itself
- SQL DDL: tables, keys and foreign keys
- Pandera: a Python schema for validating DataFrames
- Pydantic: typed models for records
Where it fits
The model documents how your datasets join, and gives the exports a single source of truth for the shapes of your data. The joins you actually run are still done in the Transform window or a pipeline. Roles and constraints on individual columns are set in the schema editor; see Profiles, versions and lineage.