smelt explain¶
Inspect the logical and physical execution plan for a project, or the derived maintenance plan for a single incremental model.
smelt explain [MODEL_NAME] [--json] [--select <selector>] [--project-dir <path>] [--show-sql] [--period <start>..<end>] [--technique <name>]
smelt explain --diff [<ref>] [--json] [--markdown] [--fail-on <downgrade|any>] [--select <selector>] [--project-dir <path>]
Options¶
| Flag | Description |
|---|---|
MODEL_NAME |
Optional. Name of a single model to print the maintenance-plan report for instead of the whole-project graph. |
--json |
Output as JSON instead of human-readable text. With MODEL_NAME, emits the per-model maintenance-plan report as JSON — with or without --show-sql, and with the same schema either way. |
--select |
Select models to include (repeatable). Ignored when MODEL_NAME is given. |
--project-dir |
Path to the smelt project root. Defaults to the current directory. |
--show-sql |
With MODEL_NAME, also print the maintenance statements each cell executes. Never connects to a backend. |
--period |
With --show-sql, real literal date bounds (<start>..<end>, end exclusive) for the printed statements' region. Without it, the symbolic placeholders {{window_start}}/{{window_end}} stand in. |
--technique |
Requires --show-sql. Render a named technique's own preview statements instead of the admitted one's, per cell — including a NotApplicable reason where that technique doesn't apply to a given cell. Accepts delete_insert, keyed_fold, column_scoped_merge, in_place_update, per_group_recompute, recompute. |
--diff [<ref>] |
Diff the project's property profile between a git baseline (default: merge-base with main) and the working tree, instead of printing a graph or a report. See Property diff below. Exclusive with MODEL_NAME, --show-sql, --period, and --technique — combining them is a usage error (exit 2). |
--fail-on |
Only with --diff. downgrade exits 1 when any downgrade is present; any exits 1 when any model shifted at all. Without it, --diff exits 0 whenever the diff was computed. |
--markdown |
Only with --diff. Emit the diff as a single GitHub-flavoured Markdown comment body instead of text. Exclusive with --json. See Property diff below. |
Human-readable output¶
Without --json or a MODEL_NAME, smelt prints a summary of the logical graph (models, dependencies, materialization, incremental config) followed by the physical graph (planner optimisations, execution order, per-model strategy).
Per-model maintenance plan¶
smelt explain <model> prints that model's derived maintenance plan instead of the whole-project graph: every cell (trigger, corner, technique), the ledger_catch_up flag (whether the cell routes through the reconciliation ledger, which lives in the target backend, not .smelt/), the derived per-source scan clamps, each source's partition-locality verdict, any admission refusals, and the model's inbound propagation edges. A cell whose ideal technique needed a state structure unavailable on the target (no ledger builder yet, or state.warehouse_tables: none) prints its recorded MaintenanceStateDowngraded downgrade alongside the technique it actually runs — visible here whether or not --json is given. This only applies to refresh: incremental models with a grain: declared — other models print a one-line notice instead.
Add --show-sql to also print, after each cell's block, the maintenance statements that cell executes — the output of the same pure emitters a run executes. Each cell's SELECT body is compiled through the real discovered project's ephemeral resolver and upstream column types, the same way a run compiles it, so the printed SQL matches what a run would compile (referenced ephemeral models are CTE-inlined, and ref-column aggregates cast to their real type). A suppressible cell's ColumnScopedMerge/KeyedFold matched arm prints in whichever variant that cell's own write-suppression resolution actually resolves — the change-suppressed IS DISTINCT FROM-guarded arm or the plain unconditional one, matching the report's own "write variant: …" line rather than always printing the unconditional shape. A technique: suppress pin whose proof refused prints no statements for that cell. A transactional group (e.g. a paired region DELETE+INSERT) is bracketed by BEGIN/COMMIT lines. --show-sql never connects to a backend. Combine with --period <start>..<end> for the real literal window bounds, or omit it to see the symbolic {{window_start}}/{{window_end}} placeholders instead. Combined with --json, the report is emitted as JSON with a statements array per cell.
Add --technique <name> alongside --show-sql to preview a different technique than the one the plan admitted — every technique smelt has an emitter for, rendered against each cell's own contract/identity/column data and labelled with its admissibility: Admitted (the plan's own choice), InterchangeableAlternative (proven sound here, but not the one picked — region recompute always qualifies when it isn't itself admitted), or NotApplicable with a reason (the technique's preconditions aren't met for this cell — reported, never omitted). Accepts delete_insert, keyed_fold, column_scoped_merge, in_place_update, per_group_recompute, recompute (recompute and delete_insert are the same technique). The preview always uses the symbolic {{window_start}}/{{window_end}} placeholders regardless of --period — it's a display-only illustration, not a windowed dry run. --json output always carries every technique's preview per cell (a technique_previews array) plus the model's full derived property set, whether or not --technique is given.
One narrow gap: a column aggregated directly off an ephemeral ref (rather than a materialized upstream model) still casts to the BIGINT default — a compile-order limitation shared identically by a real run, not an explain-specific divergence. See docs/specs/cli.md Known Divergences.
Execution postures¶
For a grain: key model, smelt explain <model> derives and prints three model-level properties — never declared, folded from the model's column families: re-run tolerance (may a merged window be blindly re-merged over unchanged input?), order-independence (may windows apply out of order or in parallel?), and reprocessing refusal (a window whose input changed since merging must not be re-merged — unconditional across every family). Each verdict comes with a reason naming the deciding column, alongside the derived run shape (window-forward or snapshot-reconcile) the postures qualify:
Execution postures:
run shape: window-forward
re-run tolerance: no (column `total_amount` folds through an additive combiner (`Sum`) — a re-merged window double-counts or cancels)
order-independence: yes (every combiner is order-independent)
reprocessing: refused (a window whose input changed since merging must not be re-merged, for every family)
A model that never classifies as grain: key prints no execution-postures section. Order-independence is a derived verdict only — a windowed-keyed run still applies windows sequentially even where the verdict holds, forgoing the parallel/out-of-order application as an optimisation. With --json, the same information appears as a top-level execution_postures object: {"run_shape": "window-forward", "rerun_tolerant": {"holds": false, "reason": "..."}, "order_independent": {"holds": true, "reason": "..."}, "reprocessing_refused": {"holds": true, "reason": "..."}}.
Internal state columns¶
Some presented columns (AVG, the STDDEV_*/VAR_* family, MAX_BY/MIN_BY, and the fallback-bearing or multi-candidate once-write spellings) don't fold their presented value directly — they fold hidden state columns instead, and recompute the presented value from that state on every read. smelt explain <model> lists these state columns as internal state, distinct from the model's public schema:
State columns:
- avg_amount (presented) folds through: avg_amount__sum, avg_amount__count
presentation: avg_amount__sum / avg_amount__count
A model with no decomposed-state columns prints no state section. With --json, the same information appears as a top-level state_columns array: [{"presented_column": "avg_amount", "state_columns": ["avg_amount__sum", "avg_amount__count"], "presentation_expr": "avg_amount__sum / avg_amount__count"}].
Refusals¶
smelt explain <model> also prints any maintenance admission refusals — the cases where no
technique could be admitted for a cell, or the plan admitted the model with a caveat:
A model whose plan admitted every cell prints Refusals: (none). With --json, the same set
appears as a top-level refusals array: [{"code": "MaintenanceScanUnbounded", "text":
"ScanUnbounded { source: \"raw.orders\", why: \"no partition_column declared\" }"}] — code
names the diagnostic code a refusal of this shape raises, absent for a refusal that raises no
diagnostic today; text is the report's own rendering of the refusal, verbatim. Empty when the
model's plan admitted every cell.
Retention reach¶
For a model referencing a source that declares a rolling retention:
bound, smelt explain <model> prints a Retention: section: one row per
source, naming the retained bound and the model's required reach into it in seconds, and the
verdict — within bound, exceeding it (SourceRetentionExceeded), or unprovable and recorded as
a downgrade (SourceRetentionDowngraded):
Retention:
- events: retained 3888000s, required reach 604800s — within bound
- events: retained 3888000s, required reach 8640000s — exceeds bound (SourceRetentionExceeded)
- events: retained 3888000s, reach unprovable — downgraded (SourceRetentionDowngraded): the model's reach into it is unbounded
A model referencing no retention:-bearing source prints no Retention: section at all — never
an empty one. With --json, the same rows appear as a top-level retention array, omitted
entirely when empty: [{"source": "events", "verdict": "within"|"exceeds"|"unprovable",
"retained_secs": 3888000, "required_lookback_secs": 604800, "reason": "..."}] —
required_lookback_secs is present only for the bounded verdicts (within/exceeds), and
reason only for unprovable.
Probes¶
A declared world-fact (functional_dependencies:, bounded_domain:, assert_monotonic,
mutation_profile: {kind: append_only}, referential_integrity:, key_recurrence) licenses a
maintenance technique only because a cheap runtime probe checks it before every consuming write —
see probes: in the smelt.yml reference. smelt explain
<model> prints the model's declared-fact probe set: per probe, the declared fact, the named
diagnostic it raises, the maintenance cell it licenses, the project's dispatch cadence, and its
cost:
Probes (1):
cadence: per_run
- fact: functional_dependencies:
probe: DeclaredFunctionalDependencyViolated
licensed cell: main.subscriptions (declared)
cost: +1 query per consuming run
A model declaring no probe-backed fact prints Probes (0):, not a missing section. The probe set
stays offline: probe SQL is built to confirm a declaration is probe-backed, never executed, so
cost is always the static "+1 query per consuming run" statement — smelt explain makes no
backend connection. With --json, the same set appears as a top-level probes array:
[{"fact": "functional_dependencies:", "probe": "DeclaredFunctionalDependencyViolated", "cell":
"main.subscriptions (declared)", "cadence": "per_run", "cost": "+1 query per consuming run"}].
See the run manifest for where a live run records each probe's actual
dispatched/skipped outcome.
See smelt explain in the CLI reference for the full flag list and a sample maintenance-plan report. The web UI's model diagnostics page renders the same technique previews and admissibility verdicts interactively, alongside the model's full derived property set.
JSON output schema¶
With --json, smelt prints a single JSON object with the following top-level fields:
{
"models": { "<model-name>": { ... }, ... },
"execution_order": ["<model-name>", ...],
"physical": { ... }
}
models — per-model entry¶
Each entry in models has:
| Field | Type | Description |
|---|---|---|
dependencies |
string[] |
Upstream model names. |
materialization |
string |
Resolved storage materialization ("view", "table", "ephemeral", "materialized_view"). |
refresh |
string |
Resolved refresh strategy: "incremental" or "materialized_view". Omitted when "full". |
incremental |
object | Present only when the model is incremental. See below. |
tags |
string[] |
Model tags from frontmatter or smelt.yml. Omitted when empty. |
owner |
string |
Model owner from frontmatter owner: key. Omitted when absent. |
origin |
object | Present only for generator-emitted models. Contains type, generator_file, generator_name. |
incremental object¶
Present on a model when refresh: incremental and grain: are set and a timeseries: block is declared.
| Field | Type | Description |
|---|---|---|
granularity |
string |
Partition granularity ("day", "week", etc.). |
partition_column |
string |
The output column used as the partition key. |
event_time_column |
string |
The source timestamp column. |
unique_key |
string[] |
Columns for MERGE-strategy deduplication. Omitted when empty. |
batch_safety |
string |
Batch-safety classification. One of "fully_batch_safe", "bounded_safe(chunk=Nd,context=Nd)", "per_partition_only". |
source_bounds |
object | Per-source bound map derived from the model's SQL. See below. Omitted when there are no timeseries upstream references. |
source_bounds field¶
The source_bounds object maps each timeseries upstream reference name to its derived bound. Lookup sources (those without a timeseries: declaration) do not appear.
Each entry is a tagged object with "type":
"bounded" — the planner derived an explicit lookback/lookahead:
{
"events_parsed": {
"type": "bounded",
"partition_col": "event_date",
"before": "PT30M",
"after": "PT0S"
}
}
| Field | Type | Description |
|---|---|---|
partition_col |
string |
The source's declared timeseries.partition_column. |
before |
string |
ISO-8601 duration — how far before the run window to read the source. "PT0S" means partition-local (no extra lookback). |
after |
string |
ISO-8601 duration — how far after the run window to read the source. |
scan_start |
string |
Present only when --period <start>..<end> is supplied: the resolved run-relative scan window start, rendered in the model's own axis domain (YYYY-MM-DD on the calendar axis, a bare integer on the integer axis) — the same window the run's pushdown filter reads. |
scan_end |
string |
Present alongside scan_start: the resolved scan window end. |
scan_unresolved |
string |
Present instead of scan_start/scan_end when --period was supplied but the margin could not be resolved to a fixed value (e.g. a non-uniform month/year offset) — names the reason. |
Without --period, scan_start/scan_end/scan_unresolved are all omitted:
smelt explain --json --period 2026-01-01..2026-01-08 | jq '.models.sessions.incremental.source_bounds'
{
"events_parsed": {
"type": "bounded",
"partition_col": "event_date",
"before": "PT30M",
"after": "PT0S",
"scan_start": "2026-01-01",
"scan_end": "2026-01-08"
}
}
"unbounded" — the source requires reading unbounded history (cumulative aggregation, UNBOUNDED PRECEDING):
"not_derivable" — the planner could not determine the bound (bare LAG/LEAD without a RANGE clause, or a computed-expression join without an explicit interval filter). A model with any not_derivable bound is refused at planning time.
Duration format¶
All before and after values use ISO-8601 duration strings:
| Value | Duration |
|---|---|
"PT0S" |
Zero — partition-local read. |
"PT30M" |
30 minutes. |
"PT2H" |
2 hours. |
"P1D" |
1 day. |
"P7D" |
7 days. |
physical object¶
| Field | Type | Description |
|---|---|---|
execution_order |
string[] |
Order the physical nodes are executed. |
nodes |
object | Per-node metadata. |
ephemerals |
string[] |
Logical models inlined as CTEs. |
transformations |
string[] |
Human-readable planner optimisation descriptions. |
Each node in nodes has strategy, materialization, target, and logical_origins.
Property diff¶
smelt explain --diff [<ref>] derives every model's property profile at two versions of the
project — a git baseline (<ref>, defaulting to the merge-base with main) and the working
tree, uncommitted edits included — and reports every model whose profile shifted: grain,
bound/reach, per-cell technique, refusals, contract point, and probes. A model whose own file
was edited is reported (edited); one that shifted only because an upstream model or source it
depends on changed is reported (downstream of <model>[, <model>…]), naming the nearest edited
ancestors — this is how a downgrade that silently propagates into a downstream mart gets caught.
The baseline never touches a warehouse, a deployed snapshot, or the maintenance ledger — it is a pure comparison of two source trees.
A new maintenance cell is graded on how it reads its data, not on the fact that it exists: a
cell whose maintenance is not partition-local — it has no time column to scan, so every run reads
it in full and any change to it rebuilds the whole model — is a downgrade when it's added to a
model that was already maintained, because the model just gained a dependency that can silently
force a full rebuild. A partition-local new cell is neutral instead; cell_added never grades
upgrade on its own — a model going from unmaintained to maintained is reported once, as
maintenance_gained. Symmetrically, a grain that widens (its keys now cover a strict
superset of the old column set, a weaker uniqueness claim — proving one row per (date, user)
implies one row per (date, user, name), never the reverse) is a downgrade; a grain that
narrows (a strict subset — a stronger claim) is an upgrade.
Text (default):
$ smelt explain --diff
property diff vs merge-base(main) = 3e9c1a4a (1 file(s) changed, 2 model(s) shifted)
user_daily_spend (edited)
[cost] Costlier maintenance: Changes from raw.transactions are applied to {total_amount} by DeleteInsert instead of KeyedFold.
verdicts:
▼ cell_technique {total_amount}@NewData { source: "raw.transactions" }: KeyedFold → DeleteInsert
user_spend_running_total (downstream of user_daily_spend)
[cost] Maintenance route lost: {running_total} no longer has a maintenance route for changes from user_daily_spend.
verdicts:
▼ cell_removed {running_total}@NewData { source: "user_daily_spend" }: {...} → null
2 model(s) shifted · 2 read more per run
When nothing shifted, the whole output is one line: property diff vs <ref>: no models shifted.
JSON (--diff --json):
{
"baseline": { "ref": "<as given>", "commit": "<sha>", "resolved_as": "merge_base" | "explicit" },
"edited_files": ["<project-relative path>", ...],
"summary": { "downgrades": 1, "upgrades": 0, "neutral": 0, "shifted_models": 2 },
"headline": "2 model(s) shifted · 2 read more per run",
"models": [
{
"model": "<name>",
"cause": { "kind": "edited" | "added" | "removed" | "downstream", "of": ["<model>", ...] },
"stories": [
{
"kind": "technique",
"severity": "cost",
"subject": "{total_amount}@NewData { source: \"raw.transactions\" }",
"lead": "Costlier maintenance",
"detail": "Changes from raw.transactions are applied to {total_amount} by DeleteInsert instead of KeyedFold.",
"changes": [0]
}
],
"changes": [
{ "dimension": "cell_technique", "subject": "...", "direction": "downgrade", "old": "KeyedFold", "new": "DeleteInsert" }
]
}
]
}
The full schema — every dimension value, the per-dimension direction rules, the story-folding
rules, and the attribution algorithm — is normative in docs/specs/property_diff.md.
Stories¶
Above the raw per-dimension changes, each shifted model carries stories: the short list of
sentences a reviewer reads first. A story folds one or more changes into a lead and a detail
sentence, and every change belongs to exactly one story — nothing is ever dropped, even a change
no dedicated rule recognizes (it lands in a catch-all other story).
Each story carries a severity, derived from the directions of the changes it folds rather than assigned by hand:
risk— the story folds a downgrade to one of the model's correctness guarantees (its grain, row identity, a maintained cell disappearing entirely, a relaxed contract, a column losing determinism or comparability, and similar). This is what a careful reviewer needs to see before approving.cost— the story folds a downgrade that makes the model more expensive to maintain (a wider read window, a costlier maintenance technique, a new dependency read in full on every run) without weakening a correctness guarantee.improvement— the story folds an upgrade and no downgrade at all.info— everything else: a schema change, a cleared refusal, or any other change with no downgrade or upgrade behind it.
Because severity is derived this way, --fail-on downgrade, the editor's PropertyDowngrade
warnings, and the risk/cost stories can never disagree with one another.
The story kinds are: maintenance_lost/maintenance_gained (a model lost or gained
incremental maintenance entirely), refusal (a maintenance admission refusal appeared or
cleared), rows_may_duplicate (a join or row-identity change means rows are no longer
identifiable by their old key), row_key (the proven row key was lost, gained, widened, or
narrowed), reads (a source's read window widened or narrowed), dependency (a new or dropped
upstream dependency, or a maintenance route lost for one that survives), technique (a cell's
maintenance technique got cheaper or costlier), contract (a contract-lattice relaxation
tightened or loosened), probe (a runtime check on a declared fact appeared or disappeared),
column_semantics (a column's determinism or comparability changed), schema (columns were
added or removed), and other (anything the rules above don't recognize).
The headline is the report's one-line summary, joining the number of shifted models with a
clause for each thing worth calling out — models that lost incremental maintenance, models with
correctness risks, models that read more per run, and models that only improved — ending in
"no downgrades" when the diff is entirely clean. The editor's code lens shows the same idea per
model, as <n> risk(s), <n> costlier vs <ref> (or changed vs <ref> when the model has neither).
Markdown (--diff --markdown): a single comment body suitable for gh pr comment
--body-file, leading with the headline and one bullet per story, each model's raw verdicts
collapsed under a Verdict table — never open by default — and a trailing
<!-- smelt-property-diff --> marker a CI job uses to update its previous comment instead of
stacking a new one on every push. The marker is emitted even when nothing shifted, so a stale
downgrade comment can be cleared once the regression is fixed. See
Continuous integration for the documented GitHub Actions job.
$ smelt explain --diff --markdown
### property diff vs merge-base(main) @ 3e9c1a4a
**2 model(s) shifted · 2 read more per run**
**user_daily_spend** (edited)
- ⚠️ **Costlier maintenance.** Changes from raw.transactions are applied to {total_amount} by DeleteInsert instead of KeyedFold.
<details>
<summary>Verdict table</summary>
| dimension | subject | direction | old | new | reason |
|---|---|---|---|---|---|
| cell_technique | {total_amount}@NewData { source: "raw.transactions" } | downgrade | KeyedFold | DeleteInsert | |
</details>
**user_spend_running_total** (downstream of user_daily_spend)
- ⚠️ **Maintenance route lost.** {running_total} no longer has a maintenance route for changes from user_daily_spend.
<details>
<summary>Verdict table</summary>
| dimension | subject | direction | old | new | reason |
|---|---|---|---|---|---|
| cell_removed | {running_total}@NewData { source: "user_daily_spend" } | downgrade | {...} | null | |
</details>
<!-- smelt-property-diff -->
--select <selector> narrows the reported set only; every model is still compared at both
versions so a downstream model's cause stays correct even when its edited ancestor is filtered
out of the printed set.
Exit codes: 0 whenever the diff was computed (whether or not anything shifted); 1 only
under --fail-on; 2 for an unresolvable baseline (not a git work tree, an unknown ref, or a
ref with no project at that path) or combining --diff with MODEL_NAME, --show-sql,
--period, or --technique.
Example¶
# Human-readable
smelt explain
# JSON, piped to jq for the sessions model's source bounds
smelt explain --json | jq '.models.sessions.incremental.source_bounds'
# Property diff against the default baseline (merge-base with main)
smelt explain --diff
# Property diff against an explicit ref, machine-readable, gating CI on any downgrade
smelt explain --diff origin/main --json --fail-on downgrade