Skip to content

Project Configuration (smelt.yml)

The smelt.yml file is the main configuration file for a smelt project. It must be located at the root of your project directory.

Top-Level Fields

Field Type Required Default Description
name string yes Project name
version integer no 1 Configuration version (currently 1). Defaults to 1 when omitted.
paths string[] no ["models"] Workspace-relative directories scanned for project files (.sql, .py, .csv, .yml). Kind is determined by file format/content, not by which directory the file lives in.
targets map yes Named execution environments (see Targets)
default_materialization string no "view" Default materialization for all models
models map no {} Per-model configuration overrides (see Model Configuration)
python string no Path to Python interpreter. Can also be set via the SMELT_PYTHON environment variable, which takes precedence over this field.
maintenance object no Project-level maintenance-plan baseline (today only scan_bounds); a per-model maintenance: block in SQL frontmatter refines it (see Maintenance Configuration)
probes object no {cadence: per_run} Project-wide dispatch cadence for declared-fact probes (see Probes Configuration)
state object no {mode: stateless, warehouse_tables: allowed} Project state posture: mode (stateless | intervals | environments) and warehouse_tables (allowed | none) (see State Configuration)

Targets

Targets define execution environments. Each target has a name (the map key) and specifies a backend type with its connection details.

targets:
  <target_name>:
    type: <backend_type>
    # backend-specific fields...

You can define multiple targets and select one at runtime with the --target CLI flag (default: dev).

State isolation per target. Run state (interval coverage, deployed-schema snapshots, run history) is stored under .smelt/targets/<target_name>/, so each target has its own closed, disjoint state store — a dev run can never mask a coverage gap in prod, and vice versa. The reconciliation ledger is separately isolated per target by living in that target's own backend schema rather than under .smelt/. See docs/reference/cli.md §"State isolation per target" and docs/specs/run_state.md §".smelt/ directory layout" for the full on-disk shape.

DuckDB Target

Field Type Required Description
type string yes Must be duckdb
database string yes Path to the DuckDB database file (relative to project root)
schema string yes Database schema to use
settings map no Connection-time DuckDB settings applied as SET key = value on open (e.g. memory_limit, threads, temp_directory). Unknown keys are rejected with an error.
targets:
  dev:
    type: duckdb
    database: target/dev.duckdb
    schema: main
    settings:
      memory_limit: "16GB"          # cap DuckDB's buffer pool
      temp_directory: target/spill  # where to spill when over the limit

Memory limits and spilling

Left to itself, DuckDB sizes its buffer pool at ~80% of total host RAM and only spills to disk once it reaches that ceiling. On a machine shared with other work (your editor, a language server, a second build, an agent), one heavy model scanning a large source can grow toward that ceiling and push the whole host into memory pressure.

To keep smelt a good citizen by default, when a DuckDB target's settings: omits these keys smelt fills them in:

  • memory_limit — defaulted to roughly min(50% of RAM, RAM − 20 GB) (floored at 40% of RAM on small hosts), leaving generous headroom for the OS and other processes. It is set conservatively because DuckDB's limit caps its buffer pool, not total process memory — actual RSS runs a few GB higher.
  • temp_directory — defaulted to .smelt-duckdb-tmp next to your .duckdb file, so queries that exceed the limit spill to disk instead of failing or growing unbounded.

Any value you set yourself is used exactly as written and never overridden — set memory_limit: "48GB" on a dedicated box to opt back into a larger budget, or point temp_directory at a faster disk. threads is left at DuckDB's default.

Spark Target

Field Type Required Description
type string yes Must be spark
connect_url string yes Spark Connect URL (e.g., sc://localhost:15002)
catalog string no Spark catalog name
schema string yes Database schema to use
format string no Table format: delta (default) or parquet. Affects schema evolution capabilities. See Schema Evolution.
targets:
  spark_prod:
    type: spark
    connect_url: sc://localhost:15002
    catalog: spark_catalog
    schema: main
    format: delta  # default; can also be "parquet"

BigQuery Target

Field Type Required Description
type string yes Must be bigquery
project string yes GCP project the jobs are billed to
dataset string no Dataset holding created tables and views (BigQuery's analogue of a schema). Defaults to schema.
location string no Dataset location (e.g. US, europe-west2). Must match at query time.
schema string yes Database schema to use
targets:
  bigquery_prod:
    type: bigquery
    project: my-gcp-project
    dataset: analytics
    location: US
    schema: analytics

Credentials come from SMELT_BQ_ACCESS_TOKEN; smelt never falls back to Google application-default credentials. See Targets.

Environment interpolation

Any string value anywhere in smelt.yml may reference an environment variable with ${VAR_NAME}. The reference is resolved once, at config load, before validation — this is how secrets (a Spark connect_url, a warehouse credential) stay out of the checked-in file. Write $$ for a literal $ that must not trigger a lookup.

targets:
  spark_prod:
    type: spark
    connect_url: ${SPARK_CONNECT_URL}   # e.g. sc://spark.internal:15002 from CI secrets
    catalog: spark_catalog
    schema: main

If SPARK_CONNECT_URL is unset, loading the config fails immediately with an error naming both the variable and the key path (targets.spark_prod.connect_url) — it never silently resolves to an empty string. When a config references more than one missing variable, every one of them is reported together, not just the first.


Materialization Types

The default_materialization field and per-model materialization field accept these values:

Value Description
table Persisted as a physical table. Required for incremental models.
view Created as a database view. Re-computed on each query.
ephemeral Not materialized at all. Inlined as a CTE into downstream models. Cannot have incremental config or target overrides.
materialized_view Backend-managed persistent view (e.g., PostgreSQL, Databricks). Refreshed atomically.

Test files are identified by a smelt.test declaration in the SQL file, not by a materialization value. See Testing for details.

A key-grain running-state table is opted in with materialization: table + refresh: incremental + grain: key — see Materializations for details.

Precedence for materialization resolution:

  1. SQL file frontmatter (materialization: in the model file)
  2. smelt.yml per-model config (models.<name>.materialization)
  3. smelt.yml top-level default_materialization
  4. Built-in default: view

When materialization is omitted at every level, a model is materialized as a view. See the Materializations guide for when to override this.


Model Configuration

Per-model configuration is specified under the models key, using the model name (filename without extension) as the key.

models:
  <model_name>:
    materialization: <type>
    tags: [<tag>, ...]
    target: <target_name>
    refresh: incremental
    unique_key: [<column>, ...]
    merge_key: [<column>, ...]  # MERGE-dedup-only write key — see the callout below
    safety_overrides:
      # safety-override fields...

refresh: incremental is admitted on the two shape-defining facts alone — a timeseries: block (the clock) and/or a top-level unique_key: (the identity). Declaring one or both is enough; declaring neither is a hard error naming what's missing. grain: itself is never required to admit a model — it is an optional, check-only partition / key label the two facts derive; write it only when you want the friendly name in frontmatter, and it errors if it disagrees with what the facts derive. key_per_partition is a third derived label (a clock plus an identity whose partition_column is a member) with no writable spelling — declaring grain: key_per_partition is a hard error.

Model Fields

Field Type Required Default Description
materialization string no (project default) Materialization type for this model
tags string[] no [] Tags for model selection (used with --select tag:X)
target string no (CLI default) Override which target to execute this model on
timeseries object no Time-dimension declaration — the clock shape-defining fact (see Timeseries Configuration)
refresh string no full Refresh axis: full, incremental, or materialized_view
unique_key string | string[] no The identity shape-defining fact — the output's row identity. A single string is sugar for a one-element list. Together with timeseries:, this is what admits refresh: incremental; frontmatter wins over this smelt.yml override when both set it. Distinct from merge_key below (a MERGE-dedup write key, never identity).
grain string no Optional check-only assertion — partition or key — validated against the label timeseries:/unique_key: derive; never a driver. key_per_partition is derived-only; declaring it is a hard error.
merge_key string | string[] no The write/dedup key a column-scoped MERGE technique writes on — never identity-conferring, never a driver of grain. A single string is sugar for a one-element list. Admitted on any refresh: incremental output; frontmatter wins over this smelt.yml override when both set it (see the callout in Incremental Configuration).
safety_overrides object no (all false) Named escape hatches for the partition-grain safety checks (see Safety Overrides). Admitted only on a partition-shaped output (a declared clock, no declared identity); same frontmatter-wins precedence as unique_key:.

Target precedence: SQL file frontmatter > smelt.yml model config > CLI --target flag.

Tags from smelt.yml and SQL frontmatter are merged (union, deduplicated).

Layer split: smelt.yml vs SQL frontmatter

The fields listed above are the complete set of per-model configuration accepted in smelt.yml. Other per-model settings — schema_evolution (schema-change strategy), columns (per-column defaults and backfills), and per-model table format — are declared in the model's SQL frontmatter, not in smelt.yml. Placing these keys under models.<name>: in smelt.yml has no effect. See the SQL Models guide and the Schema Evolution guide for how to declare them.

Timeseries Configuration

Models that process time-partitioned data must declare a timeseries: block. This is required for refresh: incremental + grain: partition models. A grain: key model may also declare timeseries: to time-partition its keyed output — the key and clock axes are independent, not alternatives — but only when key temporal locality is established (a proof, or a checked declaration, that every duplicate delivery of one key stays within a bounded window of itself on the event axis; see the composed shape). A grain: key model whose timeseries: block satisfies none of the three locality routes is refused (KeyedForbidsTimeseries, naming the missing route). The timeseries: and merge_key: keys are siblings, not nested.

models:
  daily_revenue:
    materialization: table
    refresh: incremental
    grain: partition
    timeseries:
      event_time_column: transaction_timestamp  # column in SOURCE data (WHERE filter)
      partition_column: revenue_date             # column in OUTPUT (DELETE target)
      granularity: day
    merge_key:
      - transaction_id

Timeseries Fields

Field Type Required Default Description
event_time_column string yes Column in source data to filter on (used in the injected WHERE clause). Must be a timestamp or date.
partition_column string yes Column in the output table to delete by (for DELETE+INSERT strategy).
granularity string yes Partition granularity: hour, day, week, month, quarter, or year.
week_start string no Start day for weekly partitions. Required when granularity: week. One of: monday, tuesday, wednesday, thursday, friday, saturday, sunday.

Example with weekly granularity:

models:
  weekly_rollup:
    materialization: table
    refresh: incremental
    grain: partition
    timeseries:
      event_time_column: event_ts
      partition_column: week_start_date
      granularity: week
      week_start: monday

Incremental Configuration

refresh: incremental + grain: partition processes only new or changed data instead of rebuilding the entire table. This implies a stored table. A timeseries: block must also be present (see above); merge_key: itself is optional.

models:
  daily_revenue:
    materialization: table
    refresh: incremental
    grain: partition
    timeseries:
      event_time_column: transaction_timestamp
      partition_column: revenue_date
      granularity: day
    safety_overrides:
      allow_window_functions: false
    merge_key:
      - transaction_id

batched: is retired

The batched: sub-block is retired outright on both surfaces — declaring it in a model file or a smelt.yml model entry is a hard parse-time error naming the top-level replacement for each key you wrote (unique_key → top-level merge_key:, safety_overrides → top-level safety_overrides:, nondeterministic_columns → columns.<c>.contract: plausible). merge_key: never confers identity or reshapes the output — it names only the row key a column-scoped MERGE technique writes on, for a partition-shaped model whose column-scoped MERGE needs a row key without becoming key-addressed (declaring the identity fact instead means unique_key:, which makes the output key-shaped and requires an aggregated GROUP BY body — see Incremental models). merge_key: is declarable in both .sql frontmatter and as a smelt.yml model override (frontmatter wins). nondeterministic_columns has no merge_key:-adjacent replacement: columns.<c>.contract: plausible is SQL-frontmatter-only, with no smelt.yml spelling.

Incremental Fields

Field Type Required Default Description
merge_key string[] no [] A MERGE-dedup write key for a partition-shaped output — never the identity-conferring fact unique_key: is. When present, the backend may choose a column-scoped MERGE strategy instead of DELETE+INSERT for a mutable-dimension cell. Declarable in .sql frontmatter and as a smelt.yml model override (frontmatter wins).

safety_overrides is a top-level model key (see Model Fields and Safety Overrides below), not part of merge_key:.

Safety Overrides

Smelt validates incremental models to ensure they produce the same results whether run on the full dataset or on individual partitions. Certain SQL patterns can violate this guarantee. Safety overrides let you acknowledge and allow these patterns when you know they are safe for your use case.

Field Type Default Description
allow_window_functions bool false Allow window functions (e.g., ROW_NUMBER(), LAG()) which may produce different results on partial data
allow_having bool false Allow HAVING clauses which filter on aggregates that may differ per-partition
allow_limit bool false Allow LIMIT which produces non-deterministic results on partial data
allow_subqueries bool false Allow subqueries which may reference data outside the current partition
allow_nondeterministic bool false Allow nondeterministic functions (e.g., RANDOM(), NOW())
allow_distinct bool false Allow DISTINCT which may produce different results when data is split across partitions

Maintenance Configuration

maintenance: is a sibling of timeseries: in SQL frontmatter (it is not a smelt.yml per-model field). It constrains the derived maintenance plan — the per-(column-group × trigger) cell matrix smelt explain <model> prints — without ever choosing a strategy the derivation didn't already admit.

maintenance:
  scan_bounds:
    per_source:
      raw.users:
        allow_full_scan: true

The most common use is scan_bounds.per_source.<source>.allow_full_scan: true, which names your acceptance of a full read of an unclocked source (one with no partition_column to bound a scan by). Some maintenance cells — e.g. a column-scoped MERGE driven by a mutable dimension's UpstreamMutation trigger — can only be admitted by reading that dimension in full; without this acceptance, the plan refuses the cell and smelt run falls back to region-recompute (DELETE+INSERT) instead. See Enrichment joins and dimension updates for the full example.

Maintenance Fields

Field Type Required Default Description
scan_bounds.require string no partition_local The partition-locality guardrail: partition_local or none.
scan_bounds.on_violation string no error What to do when the derived plan exceeds the stated expectation: error or warn.
scan_bounds.per_source.<address>.allow_full_scan bool no false Named acceptance of a full (unbounded) read of the source at <address>.

A project-level maintenance.scan_bounds block in smelt.yml's top level sets the baseline; a per-model maintenance: block in the SQL frontmatter refines it (narrower wins). on_violation: warn admits the derived plan for the source in place of refusing it, reporting the violation as a Warning-severity diagnostic instead of an Error-severity one; the guardrail itself never widens what the plan admits — error vs. warn only changes whether the outcome is a refusal or a Warning.

maintenance.defaults.prefer, maintenance.cells[].prefer, and maintenance.cells[].technique primarily choose among the techniques a cell's derived plan admits (fold vs. region recompute vs. rederiving columns). This family choice is live: a technique: pin is a hard override that a run honours directly (bypassing the cost-model default, never bypassing admission — an inadmissible pin fails the run loudly, naming the cell), and prefer: nudges the same default without ever refusing. There is no separate config surface for this — the same keys drive both a live run's resolution and smelt bakeoff's offline measurement of what each admissible technique costs, uniformly across every dispatch route including the ordinary windowed/partition-grain region path.

The same keys also carry a suppress/unconditional value that steers the orthogonal conditional-write dimension: whether a ColumnScopedMerge/KeyedFold cell's matched arm is suppressed for unchanged rows. By default this follows a structural rule (a steady-state trigger prefers suppression; a first-build/backfill trigger prefers the plain matched arm), never bypassing the underlying row-identity/column-comparability proof — prefer: suppress/prefer: unconditional nudge the default without ever refusing, and technique: suppress/technique: unconditional force it, refusing loudly if the proof itself never admitted suppression. This suppression ladder only drives the live run path for ColumnScopedMerge cells today; a KeyedFold cell's refresh: keyed executor still always honours a proven Suppressed verdict unconditionally, regardless of trigger or override — smelt explain resolves and prints the ladder's answer for it, but that answer doesn't yet reach the live keyed-fold write.

Pinning a measured technique

smelt bakeoff <model> --pin measures every admissible technique for a cell against replayed windows of real data and prints the cheapest one as a ready-to-paste cells[] entry — the same shape as the block above, with technique: set to the winner. The command never edits the .sql file itself; paste the printed block into the model's frontmatter yourself:

maintenance:
  cells:
    - columns: [user_name]
      on: users
      technique: fold

Once pasted, the pin is an ordinary override, re-validated through admission on every compile — if the plan later stops admitting that technique for the cell (e.g. a schema change removes the proof it relied on), the next compile fails loud rather than silently reverting to the default.

cells[].write — the physical addressing pin

maintenance.cells[].write is a separate axis from prefer/technique: it pins how a cell physically locates the rows it writes (region DELETE+INSERT, keyed MERGE, column-scoped MERGE, in-place UPDATE, full rebuild, or a backend-contributed pattern), not which technique family runs.

maintenance:
  cells:
    - columns: [amount]
      on: backfill
      technique: recompute
      write: region

write: is an open name, not a sealed keyword set — it resolves against a registry of currently-known write patterns, and the set grows as backends contribute new ones. A pin is validated, never silently honoured or downgraded:

  • An unrecognised pattern name, or one the model's target backend cannot execute (e.g. write: column on a backend with no column-scoped MERGE), fails the build with MaintenanceWritePatternUnavailable, naming the pattern and the backend.
  • A recognised, backend-capable pattern that this cell's own declared facts cannot support (e.g. write: keyed on an output with no unique_key) fails with MaintenanceWriteAddressingRefused, naming the cell and the pattern.
  • A refused pin never falls back to a different addressing — fix the pin or the model's declared facts.

smelt explain <model> prints each cell's admissible pattern set and its active pin (if any), so you can see what a pin would resolve against before setting one.

Within a keyed-fold cell, write: keyed (or its explicit alias write: keyed_conditional) and write: staged_candidate are not two different techniques — both keep the fold, they just pin a different mechanism for it. keyed/keyed_conditional pin the ordinary MERGE; unavailable on a backend without MERGE fails MaintenanceWritePatternUnavailable rather than silently downgrading. staged_candidate pins the merge-less staged conditional DELETE+INSERT — even on a backend that can run MERGE, since an explicit pin is a deliberate choice, not a downgrade to second-guess. A staged_candidate pin over a cell whose write-suppression proof resolved to the unconditional (always-rewrite) form fails MaintenanceWriteAddressingRefused: the staged-candidate mechanism has no unconditional shape to fall back to.


State Configuration

state: controls how much observability bookkeeping a run persists under .smelt/, and whether smelt may author correctness-bearing tables of its own in the target backend:

state:
  mode: stateless             # default — stateless | intervals | environments
  warehouse_tables: allowed   # default — allowed | none
Field Type Required Default Description
mode string no stateless stateless writes nothing under .smelt/. intervals additionally writes run manifests, reports, the interval ledger, landed deltas, deployed-schema snapshots, source postures, and probe/migration baselines. environments writes everything intervals does plus the snapshot/environment store. Correctness structures (the reconciliation ledger, the transactional merge ledger) are unaffected by this key — they live in the target backend under every posture. See State — state.mode and what is written.
warehouse_tables string no allowed allowed lets smelt create the engine-resident correctness tables (e.g. the reconciliation ledger) a maintenance technique needs. none forbids smelt from authoring any table of its own in the target backend: every cell whose technique needs one downgrades to its recompute-family equivalent, recorded and printed as MaintenanceStateDowngraded, and a declaration whose semantics require one (e.g. contract.deferral) refuses with DeclaredContractRequiresState. Project-wide and binary — there is no per-table or per-model opt-out.

MaintenanceStateDowngraded and DeclaredContractRequiresState are warehouse_tables: none's only two consequences: a downgrade never changes what a maintained table equals, only how it's computed, while a declaration whose guarantee cannot be verified without the missing structure refuses outright rather than silently going unchecked. See State — The reconciliation ledger for the structure these keys govern.


Probes Configuration

A declared world-fact (functional_dependencies:, bounded_domain:, assert_monotonic, mutation_profile: {kind: append_only}, referential_integrity:, key_recurrence) is checked at run time by a cheap runtime probe before the maintenance write it licenses — "declared" means "checked", not "trusted forever". probes: controls how often those checks run project-wide:

probes:
  cadence: per_run       # default
Field Type Required Default Description
cadence string or object no per_run per_run — dispatch every probe on every consuming run. off — never dispatch; every declaration is trusted and recorded unverified on the run manifest. {periodic: {every_n_runs: N}} — dispatch once every N runs (a model's first run always dispatches).
# Dispatch every 10th run instead of every run
probes:
  cadence:
    periodic:
      every_n_runs: 10

A firing probe fails the run with a named diagnostic before any write — never a silent continue or a downgraded warning. smelt explain <model> lists the model's declared-fact probe set (the fact, the diagnostic it raises, the licensed maintenance cell, and its per-run cost) so you can see what a run will check before running it; see smelt explain and the run manifest's probes field for where a probe's outcome is recorded.


Validation Rules

Smelt validates model configurations and reports errors or warnings:

Errors (block execution):

  • Ephemeral models cannot declare refresh: incremental / grain: / merge_key: / safety_overrides:
  • Ephemeral models cannot have a target override

Warnings (printed to stderr):

  • View models with refresh: set (the refresh axis only applies to stored tables)
  • refresh: materialized_view models declaring merge_key: / safety_overrides: (materialized views are refreshed atomically by the engine, not smelt's incremental loop)

Complete Example

The following is a fully annotated smelt.yml based on the timeseries example project:

# Project identity
name: smelt_examples
version: 1

# Workspace-relative directories scanned for project files (.sql, .py, .csv, .yml).
# Default: ["models"]. Kind is determined by file format/content, not by directory.
paths:
  - models
  - seeds

# Execution environments
targets:
  # Local development with DuckDB (default target)
  dev:
    type: duckdb
    database: target/dev.duckdb
    schema: main

  # Remote Spark cluster
  spark:
    type: spark
    connect_url: sc://localhost:15002
    catalog: spark_catalog
    schema: main

# Default materialization for models not explicitly configured
default_materialization: view

# Per-model configuration
models:
  # Simple table materialization
  users:
    materialization: table

  events:
    materialization: table

  user_activity:
    materialization: table

  transactions:
    materialization: table

  # Incremental model (grain: partition)
  daily_revenue:
    materialization: table
    refresh: incremental
    grain: partition
    timeseries:
      event_time_column: transaction_timestamp  # column in SOURCE data (WHERE filter)
      partition_column: revenue_date             # column in OUTPUT (DELETE target)
      granularity: day

  cube_metrics:
    materialization: table