Skip to content

smelt explain

Inspect the logical and physical execution plan for a project, or the derived maintenance plan for a single incremental model.

smelt explain [MODEL_NAME] [--json] [--select <selector>] [--project-dir <path>] [--show-sql] [--period <start>..<end>] [--technique <name>]
smelt explain --diff [<ref>] [--json] [--markdown] [--fail-on <downgrade|any>] [--select <selector>] [--project-dir <path>]

Options

Flag Description
MODEL_NAME Optional. Name of a single model to print the maintenance-plan report for instead of the whole-project graph.
--json Output as JSON instead of human-readable text. With MODEL_NAME, emits the per-model maintenance-plan report as JSON — with or without --show-sql, and with the same schema either way.
--select Select models to include (repeatable). Ignored when MODEL_NAME is given.
--project-dir Path to the smelt project root. Defaults to the current directory.
--show-sql With MODEL_NAME, also print the maintenance statements each cell executes. Never connects to a backend.
--period With --show-sql, real literal date bounds (<start>..<end>, end exclusive) for the printed statements' region. Without it, the symbolic placeholders {{window_start}}/{{window_end}} stand in.
--technique Requires --show-sql. Render a named technique's own preview statements instead of the admitted one's, per cell — including a NotApplicable reason where that technique doesn't apply to a given cell. Accepts delete_insert, keyed_fold, column_scoped_merge, in_place_update, per_group_recompute, recompute.
--diff [<ref>] Diff the project's property profile between a git baseline (default: merge-base with main) and the working tree, instead of printing a graph or a report. See Property diff below. Exclusive with MODEL_NAME, --show-sql, --period, and --technique — combining them is a usage error (exit 2).
--fail-on Only with --diff. downgrade exits 1 when any downgrade is present; any exits 1 when any model shifted at all. Without it, --diff exits 0 whenever the diff was computed.
--markdown Only with --diff. Emit the diff as a single GitHub-flavoured Markdown comment body instead of text. Exclusive with --json. See Property diff below.

Human-readable output

Without --json or a MODEL_NAME, smelt prints a summary of the logical graph (models, dependencies, materialization, incremental config) followed by the physical graph (planner optimisations, execution order, per-model strategy).

Per-model maintenance plan

smelt explain <model> prints that model's derived maintenance plan instead of the whole-project graph: every cell (trigger, corner, technique), the ledger_catch_up flag (whether the cell routes through the reconciliation ledger, which lives in the target backend, not .smelt/), the derived per-source scan clamps, each source's partition-locality verdict, any admission refusals, and the model's inbound propagation edges. A cell whose ideal technique needed a state structure unavailable on the target (no ledger builder yet, or state.warehouse_tables: none) prints its recorded MaintenanceStateDowngraded downgrade alongside the technique it actually runs — visible here whether or not --json is given. This only applies to refresh: incremental models with a grain: declared — other models print a one-line notice instead.

Add --show-sql to also print, after each cell's block, the maintenance statements that cell executes — the output of the same pure emitters a run executes. Each cell's SELECT body is compiled through the real discovered project's ephemeral resolver and upstream column types, the same way a run compiles it, so the printed SQL matches what a run would compile (referenced ephemeral models are CTE-inlined, and ref-column aggregates cast to their real type). A suppressible cell's ColumnScopedMerge/KeyedFold matched arm prints in whichever variant that cell's own write-suppression resolution actually resolves — the change-suppressed IS DISTINCT FROM-guarded arm or the plain unconditional one, matching the report's own "write variant: …" line rather than always printing the unconditional shape. A technique: suppress pin whose proof refused prints no statements for that cell. A transactional group (e.g. a paired region DELETE+INSERT) is bracketed by BEGIN/COMMIT lines. --show-sql never connects to a backend. Combine with --period <start>..<end> for the real literal window bounds, or omit it to see the symbolic {{window_start}}/{{window_end}} placeholders instead. Combined with --json, the report is emitted as JSON with a statements array per cell.

Add --technique <name> alongside --show-sql to preview a different technique than the one the plan admitted — every technique smelt has an emitter for, rendered against each cell's own contract/identity/column data and labelled with its admissibility: Admitted (the plan's own choice), InterchangeableAlternative (proven sound here, but not the one picked — region recompute always qualifies when it isn't itself admitted), or NotApplicable with a reason (the technique's preconditions aren't met for this cell — reported, never omitted). Accepts delete_insert, keyed_fold, column_scoped_merge, in_place_update, per_group_recompute, recompute (recompute and delete_insert are the same technique). The preview always uses the symbolic {{window_start}}/{{window_end}} placeholders regardless of --period — it's a display-only illustration, not a windowed dry run. --json output always carries every technique's preview per cell (a technique_previews array) plus the model's full derived property set, whether or not --technique is given.

One narrow gap: a column aggregated directly off an ephemeral ref (rather than a materialized upstream model) still casts to the BIGINT default — a compile-order limitation shared identically by a real run, not an explain-specific divergence. See docs/specs/cli.md Known Divergences.

Execution postures

For a grain: key model, smelt explain <model> derives and prints three model-level properties — never declared, folded from the model's column families: re-run tolerance (may a merged window be blindly re-merged over unchanged input?), order-independence (may windows apply out of order or in parallel?), and reprocessing refusal (a window whose input changed since merging must not be re-merged — unconditional across every family). Each verdict comes with a reason naming the deciding column, alongside the derived run shape (window-forward or snapshot-reconcile) the postures qualify:

Execution postures:
  run shape: window-forward
  re-run tolerance: no (column `total_amount` folds through an additive combiner (`Sum`) — a re-merged window double-counts or cancels)
  order-independence: yes (every combiner is order-independent)
  reprocessing: refused (a window whose input changed since merging must not be re-merged, for every family)

A model that never classifies as grain: key prints no execution-postures section. Order-independence is a derived verdict only — a windowed-keyed run still applies windows sequentially even where the verdict holds, forgoing the parallel/out-of-order application as an optimisation. With --json, the same information appears as a top-level execution_postures object: {"run_shape": "window-forward", "rerun_tolerant": {"holds": false, "reason": "..."}, "order_independent": {"holds": true, "reason": "..."}, "reprocessing_refused": {"holds": true, "reason": "..."}}.

Internal state columns

Some presented columns (AVG, the STDDEV_*/VAR_* family, MAX_BY/MIN_BY, and the fallback-bearing or multi-candidate once-write spellings) don't fold their presented value directly — they fold hidden state columns instead, and recompute the presented value from that state on every read. smelt explain <model> lists these state columns as internal state, distinct from the model's public schema:

State columns:
  - avg_amount (presented) folds through: avg_amount__sum, avg_amount__count
      presentation: avg_amount__sum / avg_amount__count

A model with no decomposed-state columns prints no state section. With --json, the same information appears as a top-level state_columns array: [{"presented_column": "avg_amount", "state_columns": ["avg_amount__sum", "avg_amount__count"], "presentation_expr": "avg_amount__sum / avg_amount__count"}].

Refusals

smelt explain <model> also prints any maintenance admission refusals — the cases where no technique could be admitted for a cell, or the plan admitted the model with a caveat:

Refusals (1):
  - ScanUnbounded { source: "raw.orders", why: "no partition_column declared" }

A model whose plan admitted every cell prints Refusals: (none). With --json, the same set appears as a top-level refusals array: [{"code": "MaintenanceScanUnbounded", "text": "ScanUnbounded { source: \"raw.orders\", why: \"no partition_column declared\" }"}] — code names the diagnostic code a refusal of this shape raises, absent for a refusal that raises no diagnostic today; text is the report's own rendering of the refusal, verbatim. Empty when the model's plan admitted every cell.

Retention reach

For a model referencing a source that declares a rolling retention: bound, smelt explain <model> prints a Retention: section: one row per source, naming the retained bound and the model's required reach into it in seconds, and the verdict — within bound, exceeding it (SourceRetentionExceeded), or unprovable and recorded as a downgrade (SourceRetentionDowngraded):

Retention:
  - events: retained 3888000s, required reach 604800s — within bound
  - events: retained 3888000s, required reach 8640000s — exceeds bound (SourceRetentionExceeded)
  - events: retained 3888000s, reach unprovable — downgraded (SourceRetentionDowngraded): the model's reach into it is unbounded

A model referencing no retention:-bearing source prints no Retention: section at all — never an empty one. With --json, the same rows appear as a top-level retention array, omitted entirely when empty: [{"source": "events", "verdict": "within"|"exceeds"|"unprovable", "retained_secs": 3888000, "required_lookback_secs": 604800, "reason": "..."}] — required_lookback_secs is present only for the bounded verdicts (within/exceeds), and reason only for unprovable.

Probes

A declared world-fact (functional_dependencies:, bounded_domain:, assert_monotonic, mutation_profile: {kind: append_only}, referential_integrity:, key_recurrence) licenses a maintenance technique only because a cheap runtime probe checks it before every consuming write — see probes: in the smelt.yml reference. smelt explain <model> prints the model's declared-fact probe set: per probe, the declared fact, the named diagnostic it raises, the maintenance cell it licenses, the project's dispatch cadence, and its cost:

Probes (1):
  cadence: per_run
  - fact: functional_dependencies:
      probe: DeclaredFunctionalDependencyViolated
      licensed cell: main.subscriptions (declared)
      cost: +1 query per consuming run

A model declaring no probe-backed fact prints Probes (0):, not a missing section. The probe set stays offline: probe SQL is built to confirm a declaration is probe-backed, never executed, so cost is always the static "+1 query per consuming run" statement — smelt explain makes no backend connection. With --json, the same set appears as a top-level probes array: [{"fact": "functional_dependencies:", "probe": "DeclaredFunctionalDependencyViolated", "cell": "main.subscriptions (declared)", "cadence": "per_run", "cost": "+1 query per consuming run"}]. See the run manifest for where a live run records each probe's actual dispatched/skipped outcome.

See smelt explain in the CLI reference for the full flag list and a sample maintenance-plan report. The web UI's model diagnostics page renders the same technique previews and admissibility verdicts interactively, alongside the model's full derived property set.

JSON output schema

With --json, smelt prints a single JSON object with the following top-level fields:

{
  "models": { "<model-name>": { ... }, ... },
  "execution_order": ["<model-name>", ...],
  "physical": { ... }
}

models — per-model entry

Each entry in models has:

Field Type Description
dependencies string[] Upstream model names.
materialization string Resolved storage materialization ("view", "table", "ephemeral", "materialized_view").
refresh string Resolved refresh strategy: "incremental" or "materialized_view". Omitted when "full".
incremental object Present only when the model is incremental. See below.
tags string[] Model tags from frontmatter or smelt.yml. Omitted when empty.
owner string Model owner from frontmatter owner: key. Omitted when absent.
origin object Present only for generator-emitted models. Contains type, generator_file, generator_name.

incremental object

Present on a model when refresh: incremental and grain: are set and a timeseries: block is declared.

Field Type Description
granularity string Partition granularity ("day", "week", etc.).
partition_column string The output column used as the partition key.
event_time_column string The source timestamp column.
unique_key string[] Columns for MERGE-strategy deduplication. Omitted when empty.
batch_safety string Batch-safety classification. One of "fully_batch_safe", "bounded_safe(chunk=Nd,context=Nd)", "per_partition_only".
source_bounds object Per-source bound map derived from the model's SQL. See below. Omitted when there are no timeseries upstream references.

source_bounds field

The source_bounds object maps each timeseries upstream reference name to its derived bound. Lookup sources (those without a timeseries: declaration) do not appear.

Each entry is a tagged object with "type":

"bounded" — the planner derived an explicit lookback/lookahead:

{
  "events_parsed": {
    "type": "bounded",
    "partition_col": "event_date",
    "before": "PT30M",
    "after": "PT0S"
  }
}
Field Type Description
partition_col string The source's declared timeseries.partition_column.
before string ISO-8601 duration — how far before the run window to read the source. "PT0S" means partition-local (no extra lookback).
after string ISO-8601 duration — how far after the run window to read the source.
scan_start string Present only when --period <start>..<end> is supplied: the resolved run-relative scan window start, rendered in the model's own axis domain (YYYY-MM-DD on the calendar axis, a bare integer on the integer axis) — the same window the run's pushdown filter reads.
scan_end string Present alongside scan_start: the resolved scan window end.
scan_unresolved string Present instead of scan_start/scan_end when --period was supplied but the margin could not be resolved to a fixed value (e.g. a non-uniform month/year offset) — names the reason.

Without --period, scan_start/scan_end/scan_unresolved are all omitted:

smelt explain --json --period 2026-01-01..2026-01-08 | jq '.models.sessions.incremental.source_bounds'
{
  "events_parsed": {
    "type": "bounded",
    "partition_col": "event_date",
    "before": "PT30M",
    "after": "PT0S",
    "scan_start": "2026-01-01",
    "scan_end": "2026-01-08"
  }
}

"unbounded" — the source requires reading unbounded history (cumulative aggregation, UNBOUNDED PRECEDING):

{ "type": "unbounded" }

"not_derivable" — the planner could not determine the bound (bare LAG/LEAD without a RANGE clause, or a computed-expression join without an explicit interval filter). A model with any not_derivable bound is refused at planning time.

{ "type": "not_derivable" }

Duration format

All before and after values use ISO-8601 duration strings:

Value Duration
"PT0S" Zero — partition-local read.
"PT30M" 30 minutes.
"PT2H" 2 hours.
"P1D" 1 day.
"P7D" 7 days.

physical object

Field Type Description
execution_order string[] Order the physical nodes are executed.
nodes object Per-node metadata.
ephemerals string[] Logical models inlined as CTEs.
transformations string[] Human-readable planner optimisation descriptions.

Each node in nodes has strategy, materialization, target, and logical_origins.

Property diff

smelt explain --diff [<ref>] derives every model's property profile at two versions of the project — a git baseline (<ref>, defaulting to the merge-base with main) and the working tree, uncommitted edits included — and reports every model whose profile shifted: grain, bound/reach, per-cell technique, refusals, contract point, and probes. A model whose own file was edited is reported (edited); one that shifted only because an upstream model or source it depends on changed is reported (downstream of <model>[, <model>…]), naming the nearest edited ancestors — this is how a downgrade that silently propagates into a downstream mart gets caught.

The baseline never touches a warehouse, a deployed snapshot, or the maintenance ledger — it is a pure comparison of two source trees.

A new maintenance cell is graded on how it reads its data, not on the fact that it exists: a cell whose maintenance is not partition-local — it has no time column to scan, so every run reads it in full and any change to it rebuilds the whole model — is a downgrade when it's added to a model that was already maintained, because the model just gained a dependency that can silently force a full rebuild. A partition-local new cell is neutral instead; cell_added never grades upgrade on its own — a model going from unmaintained to maintained is reported once, as maintenance_gained. Symmetrically, a grain that widens (its keys now cover a strict superset of the old column set, a weaker uniqueness claim — proving one row per (date, user) implies one row per (date, user, name), never the reverse) is a downgrade; a grain that narrows (a strict subset — a stronger claim) is an upgrade.

Text (default):

$ smelt explain --diff

property diff vs merge-base(main) = 3e9c1a4a (1 file(s) changed, 2 model(s) shifted)

  user_daily_spend  (edited)
    [cost] Costlier maintenance: Changes from raw.transactions are applied to {total_amount} by DeleteInsert instead of KeyedFold.
    verdicts:
      ▼ cell_technique {total_amount}@NewData { source: "raw.transactions" }: KeyedFold → DeleteInsert

  user_spend_running_total  (downstream of user_daily_spend)
    [cost] Maintenance route lost: {running_total} no longer has a maintenance route for changes from user_daily_spend.
    verdicts:
      ▼ cell_removed {running_total}@NewData { source: "user_daily_spend" }: {...} → null

2 model(s) shifted · 2 read more per run

When nothing shifted, the whole output is one line: property diff vs <ref>: no models shifted.

JSON (--diff --json):

{
  "baseline": { "ref": "<as given>", "commit": "<sha>", "resolved_as": "merge_base" | "explicit" },
  "edited_files": ["<project-relative path>", ...],
  "summary": { "downgrades": 1, "upgrades": 0, "neutral": 0, "shifted_models": 2 },
  "headline": "2 model(s) shifted · 2 read more per run",
  "models": [
    {
      "model": "<name>",
      "cause": { "kind": "edited" | "added" | "removed" | "downstream", "of": ["<model>", ...] },
      "stories": [
        {
          "kind": "technique",
          "severity": "cost",
          "subject": "{total_amount}@NewData { source: \"raw.transactions\" }",
          "lead": "Costlier maintenance",
          "detail": "Changes from raw.transactions are applied to {total_amount} by DeleteInsert instead of KeyedFold.",
          "changes": [0]
        }
      ],
      "changes": [
        { "dimension": "cell_technique", "subject": "...", "direction": "downgrade", "old": "KeyedFold", "new": "DeleteInsert" }
      ]
    }
  ]
}

The full schema — every dimension value, the per-dimension direction rules, the story-folding rules, and the attribution algorithm — is normative in docs/specs/property_diff.md.

Stories

Above the raw per-dimension changes, each shifted model carries stories: the short list of sentences a reviewer reads first. A story folds one or more changes into a lead and a detail sentence, and every change belongs to exactly one story — nothing is ever dropped, even a change no dedicated rule recognizes (it lands in a catch-all other story).

Each story carries a severity, derived from the directions of the changes it folds rather than assigned by hand:

  • risk — the story folds a downgrade to one of the model's correctness guarantees (its grain, row identity, a maintained cell disappearing entirely, a relaxed contract, a column losing determinism or comparability, and similar). This is what a careful reviewer needs to see before approving.
  • cost — the story folds a downgrade that makes the model more expensive to maintain (a wider read window, a costlier maintenance technique, a new dependency read in full on every run) without weakening a correctness guarantee.
  • improvement — the story folds an upgrade and no downgrade at all.
  • info — everything else: a schema change, a cleared refusal, or any other change with no downgrade or upgrade behind it.

Because severity is derived this way, --fail-on downgrade, the editor's PropertyDowngrade warnings, and the risk/cost stories can never disagree with one another.

The story kinds are: maintenance_lost/maintenance_gained (a model lost or gained incremental maintenance entirely), refusal (a maintenance admission refusal appeared or cleared), rows_may_duplicate (a join or row-identity change means rows are no longer identifiable by their old key), row_key (the proven row key was lost, gained, widened, or narrowed), reads (a source's read window widened or narrowed), dependency (a new or dropped upstream dependency, or a maintenance route lost for one that survives), technique (a cell's maintenance technique got cheaper or costlier), contract (a contract-lattice relaxation tightened or loosened), probe (a runtime check on a declared fact appeared or disappeared), column_semantics (a column's determinism or comparability changed), schema (columns were added or removed), and other (anything the rules above don't recognize).

The headline is the report's one-line summary, joining the number of shifted models with a clause for each thing worth calling out — models that lost incremental maintenance, models with correctness risks, models that read more per run, and models that only improved — ending in "no downgrades" when the diff is entirely clean. The editor's code lens shows the same idea per model, as <n> risk(s), <n> costlier vs <ref> (or changed vs <ref> when the model has neither).

Markdown (--diff --markdown): a single comment body suitable for gh pr comment --body-file, leading with the headline and one bullet per story, each model's raw verdicts collapsed under a Verdict table — never open by default — and a trailing <!-- smelt-property-diff --> marker a CI job uses to update its previous comment instead of stacking a new one on every push. The marker is emitted even when nothing shifted, so a stale downgrade comment can be cleared once the regression is fixed. See Continuous integration for the documented GitHub Actions job.

$ smelt explain --diff --markdown

### property diff vs merge-base(main) @ 3e9c1a4a

**2 model(s) shifted · 2 read more per run**

**user_daily_spend** (edited)

- ⚠️ **Costlier maintenance.** Changes from raw.transactions are applied to {total_amount} by DeleteInsert instead of KeyedFold.

<details>
<summary>Verdict table</summary>

| dimension | subject | direction | old | new | reason |
|---|---|---|---|---|---|
| cell_technique | {total_amount}@NewData { source: "raw.transactions" } | downgrade | KeyedFold | DeleteInsert |  |

</details>

**user_spend_running_total** (downstream of user_daily_spend)

- ⚠️ **Maintenance route lost.** {running_total} no longer has a maintenance route for changes from user_daily_spend.

<details>
<summary>Verdict table</summary>

| dimension | subject | direction | old | new | reason |
|---|---|---|---|---|---|
| cell_removed | {running_total}@NewData { source: "user_daily_spend" } | downgrade | {...} | null |  |

</details>

<!-- smelt-property-diff -->

--select <selector> narrows the reported set only; every model is still compared at both versions so a downstream model's cause stays correct even when its edited ancestor is filtered out of the printed set.

Exit codes: 0 whenever the diff was computed (whether or not anything shifted); 1 only under --fail-on; 2 for an unresolvable baseline (not a git work tree, an unknown ref, or a ref with no project at that path) or combining --diff with MODEL_NAME, --show-sql, --period, or --technique.

Example

# Human-readable
smelt explain

# JSON, piped to jq for the sessions model's source bounds
smelt explain --json | jq '.models.sessions.incremental.source_bounds'

# Property diff against the default baseline (merge-base with main)
smelt explain --diff

# Property diff against an explicit ref, machine-readable, gating CI on any downgrade
smelt explain --diff origin/main --json --fail-on downgrade