Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
23 commits
Select commit Hold shift + click to select a range
92a6876
add redshift support
brendan-cook-87 Nov 24, 2023
8ba69db
Merge branch 'main' into REDSHIFT_SUPPORT
hanslemm Oct 21, 2024
6ad19c4
chore(profiles): clean up
hanslemm Oct 21, 2024
2be379f
chore(type_helpers): redshift SUPER type for JSON and ARRAY
hanslemm Oct 21, 2024
dfe7572
Merge branch 'brooklyn-data:main' into REDSHIFT_SUPPORT
hanslemm Dec 19, 2024
72f5e4a
Update dbt_project.yml
hanslemm Dec 19, 2024
42b11cd
Merge branch 'main' into REDSHIFT_SUPPORT
hanslemm Jun 12, 2025
d6bd6b3
Update README.md
hanslemm Jun 12, 2025
70d0081
Update type_helpers.sql
shiv-io Aug 5, 2025
ab81712
Update macros/database_specific_helpers/type_helpers.sql
hanslemm Sep 23, 2025
c5f3ed1
Update CONTRIBUTING.md
hanslemm Sep 23, 2025
4069ed7
Add Trino array type macro to type_helpers.sql
hanslemm Sep 23, 2025
637586e
Merge branch 'main' into REDSHIFT_SUPPORT
hanslemm Sep 23, 2025
c58f50d
Merge branch 'REDSHIFT_SUPPORT' into REDSHIFT_SUPPORT
hanslemm Aug 12, 2026
2021662
Merge pull request #1 from shiv-io/REDSHIFT_SUPPORT
hanslemm Aug 12, 2026
91ad526
chore: merge upstream 2.11.0 into REDSHIFT_SUPPORT
hanslemm Sep 8, 2026
495aecf
chore: nest generic test arguments under arguments (dbt 1.10+)
hanslemm Sep 8, 2026
b93ed16
feat: dim_dbt__current_relations for deferral state
hanslemm Sep 11, 2026
54bb9cb
feat: export_state run-operation for deferral state
hanslemm Sep 11, 2026
f1b2849
fix: validate and escape export_state's interpolated arguments
hanslemm Sep 11, 2026
a455f14
fix: filter target_name in Jinja and widen the identifier pattern
hanslemm Sep 11, 2026
d152d7a
fix(dim_dbt__current_relations): rank per (node_id, target_name); ord…
hanslemm Sep 11, 2026
2179d12
fix(export_state): fall back to a node's own target; emit ISO 8601 ti…
hanslemm Sep 11, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
44 changes: 44 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,6 +25,7 @@ The package currently supports
- Postgres :white_check_mark:
- SQL Server :white_check_mark:
- Trino :white_check_mark:
- Redshift ✅

Models included:

Expand Down Expand Up @@ -238,6 +239,49 @@ An example operation is as follows:
dbt run-operation migrate_from_v0_to_v1 --args '{old_database: analytics, old_schema: dbt_artifacts, new_database: analytics, new_schema: artifact_sources}'
```

## Using dbt_artifacts as deferral state

`dbt build --defer --state` reads a `manifest.json` from disk, and the only
thing it takes from that manifest for unselected nodes is where each relation
lives. `dim_dbt__current_relations` knows that from your production runs, so a
CI job can parse its own project for a structurally complete manifest and then
correct the relation coordinates from the warehouse, instead of shipping a
manifest between jobs.

```bash
dbt --quiet run-operation dbt_artifacts.export_state \
--args '{schema: dbt_artifacts, target_name: prod}' > export.json
```

The document is versioned; refuse a `dbt_artifacts_state_version` whose major
differs from the one you support.

```json
{
"dbt_artifacts_state_version": 1,
"generated_at": "2026-09-11T09:00:00+00:00",
"source": {"database": "analytics", "schema": "dbt_artifacts", "target_name": "prod"},
"nodes": {
"model.my_project.dim_customer": {
"resource_type": "model",
"name": "dim_customer",
"package_name": "my_project",
"database": "analytics",
"schema": "marts",
"alias": "dim_customer",
"materialization": "table",
"checksum": "9f2c…",
"last_success_at": "2026-09-10T02:14:09+00:00",
"command_invocation_id": "0f0a…"
}
}
}
```

A node appears when production executed it successfully at least once; the row
shown is the most recent success. An empty `nodes` object is a valid document
and means the tables hold no successful executions yet.

## Acknowledgements

Thank you to [Tails.com](https://tails.com/gb/careers/) for initial development and maintenance of this package. On 2021/12/20, the repository was transferred from the Tails.com GitHub organization to Brooklyn Data Co.
Expand Down
4 changes: 4 additions & 0 deletions integration_test_project/example-env.sh
Original file line number Diff line number Diff line change
Expand Up @@ -16,6 +16,10 @@ export DBT_ENV_SECRET_DATABRICKS_TOKEN=
export DBT_ENV_SECRET_GCP_PROJECT=
export DBT_ENV_SPARK_DRIVER_PATH= # /Library/simba/spark/lib/libsparkodbc_sbu.dylib on a Mac
export DBT_ENV_SPARK_ENDPOINT= # The endpoint ID from the Databricks HTTP path
export DBT_ENV_SECRET_REDSHIFT_HOST=
export DBT_ENV_SECRET_REDSHIFT_CLUSTER_ID=
export DBT_ENV_SECRET_REDSHIFT_DB=
export DBT_ENV_SECRET_REDSHIFT_USER=

# dbt environment variables, change these
export DBT_VERSION="1_5_0"
Expand Down
11 changes: 11 additions & 0 deletions integration_test_project/profiles.yml
Original file line number Diff line number Diff line change
Expand Up @@ -66,6 +66,17 @@ dbt_artifacts:
trust_cert: True
Encrypt: False
user: dbt
password: "123"
redshift:
Comment thread
hanslemm marked this conversation as resolved.
type: redshift
method: iam
threads: 8
host: "{{ env_var('DBT_ENV_SECRET_REDSHIFT_HOST') }}"
port: 5439
dbname: "{{ env_var('DBT_ENV_SECRET_REDSHIFT_DB') }}"
user: "{{ env_var('DBT_ENV_SECRET_REDSHIFT_USER') }}"
schema: dbt_artifacts_test_commit_{{ env_var('DBT_VERSION', '') }}_{{ env_var('GITHUB_SHA_OVERRIDE', '') if env_var('GITHUB_SHA_OVERRIDE', '') else env_var('GITHUB_SHA') }}
cluster_id: "{{ env_var('DBT_ENV_SECRET_REDSHIFT_CLUSTER_ID') }}"
password: "123Administrator"
trino:
type: trino
Expand Down
25 changes: 25 additions & 0 deletions integration_test_project/tests/assert_current_relations_unique.sql
Original file line number Diff line number Diff line change
@@ -0,0 +1,25 @@
{{ config(enabled = target.type in ["postgres", "redshift"]) }}
-- dim_dbt__current_relations must hold at most one row per (node_id,
-- target_name) and never a null relation coordinate: a consumer patches a
-- manifest with these values, so a duplicate silently picks a winner and a
-- null produces an unusable relation name. Returns rows (fails) on either
-- condition.
with duplicates as (
select node_id
from {{ ref("dim_dbt__current_relations") }}
group by node_id, target_name
having count(*) > 1
),

nulls as (
select node_id
from {{ ref("dim_dbt__current_relations") }}
where database is null
or schema is null
or alias is null
or resource_type is null
)

select node_id from duplicates
union all
select node_id from nulls
32 changes: 32 additions & 0 deletions macros/_macros.yml
Original file line number Diff line number Diff line change
Expand Up @@ -57,6 +57,10 @@ macros:
description: |
Dependent on the adapter type, returns the native type for storing JSON.

- name: type_string
description: |
Dependent on the adapter type, returns the native type for storing a string.

## MIGRATION ##
- name: migrate_from_v0_to_v1
description: |
Expand Down Expand Up @@ -285,3 +289,31 @@ macros:
type: list[any]
description: |
The results object from dbt.

## DEFERRAL STATE ##
- name: export_state
description: |
Prints dim_dbt__current_relations as a versioned JSON document for use as
dbt deferral state. See "Using dbt_artifacts as deferral state" in the
README for the contract.
arguments:
- name: resource_types
type: list[string]
description: |
Which resource types to export. Defaults to model, seed and snapshot;
deferral resolves refable nodes only.
- name: database
type: string
description: |
Overrides the database of dim_dbt__current_relations. Needed when the
session runs on a target whose ref() would resolve elsewhere.
- name: schema
type: string
description: |
Overrides the schema of dim_dbt__current_relations.
- name: target_name
type: string
description: |
When set, exports each node's latest success on this target only.
Use it when several targets write to the same artifacts tables.
When unset, exports each node's latest success across all targets.
22 changes: 21 additions & 1 deletion macros/database_specific_helpers/type_helpers.sql
Original file line number Diff line number Diff line change
Expand Up @@ -26,10 +26,14 @@
json
{% endmacro %}

{% macro redshift__type_json() %}
super
{% endmacro %}

{#- ARRAY -#}

{% macro type_array() %}
{{ return(adapter.dispatch('type_array', 'dbt_artifacts')()) }}
{{ return(adapter.dispatch('type_array', 'dbt_artifacts')()) }}
{% endmacro %}

{% macro default__type_array() %}
Expand All @@ -44,6 +48,22 @@
array<string>
{% endmacro %}

{% macro redshift__type_array() %}
super
{% endmacro %}

{% macro type_string() %}
{{ return(adapter.dispatch('type_string', 'dbt_artifacts')()) }}
{% endmacro %}

{% macro default__type_string() %}
{{ return(api.Column.translate_type("string")) }}
{% endmacro %}

{% macro redshift__type_string() %}
varchar(max)
{% endmacro %}

{% macro trino__type_array() %}
array(varchar)
{% endmacro %}
Expand Down
204 changes: 204 additions & 0 deletions macros/export_state.sql
Original file line number Diff line number Diff line change
@@ -0,0 +1,204 @@
{#
Prints dim_dbt__current_relations as a versioned JSON document, for use as
dbt deferral state. Emits with print() so `dbt --quiet run-operation` gives
clean JSON on stdout. Never raises for a missing relation, an empty
resource_types list, or a zero-row result: each yields an empty `nodes`
object and a warning, leaving the caller to decide. Raises only on
invalid arguments (an unknown resource_type, or a database/schema value
that isn't a safe identifier).

Emits one node per node_id. dim_dbt__current_relations is grained on
(node_id, target_name), since several targets can write to the same
artifacts tables; when target_name is given, each node's row for that
target is used, and when it isn't, each node's most recently completed
success across all targets is used, matching the single-target-project
behaviour where there is only ever one target to pick from.

Usage:
dbt --quiet run-operation dbt_artifacts.export_state \
--args '{schema: dbt_artifacts, target_name: betterdata_prod}' > export.json
#}
{% macro export_state(resource_types=['model', 'seed', 'snapshot'], database=none, schema=none, target_name=none) %}

{% set state_version = 1 %}
{% set relation = ref("dim_dbt__current_relations") %}
{% set source_database = database if database is not none else relation.database %}
{% set source_schema = schema if schema is not none else relation.schema %}

{# `resource_types` is interpolated straight into an IN (...) list below, and
`source_database`/`source_schema` are interpolated indirectly: they become
the FROM-clause relation via api.Relation.create(), whose quoted() wraps
each part in the adapter's quote character but does not escape a quote
character embedded in the value, so an unvalidated value could still break
out of the quoted identifier. Validate both, ahead of any query, so a bad
value fails loudly instead of reaching the database. This is the one place
this macro may raise: validation happens before querying starts, so the
"never raises on a missing relation or zero rows" contract, which is about
behaviour once the query runs, is untouched. `target_name` is not
constrained to a known set (target names are project-specific) and is not
interpolated into SQL at all — see the row loop below — so it needs no
validation here. #}
{% set known_resource_types = ['model', 'seed', 'snapshot'] %}
{% for resource_type in resource_types %}
{% if resource_type not in known_resource_types %}
{% do exceptions.raise_compiler_error(
"export_state: unknown resource_type '" ~ resource_type ~ "'; expected one of " ~ known_resource_types | join(", ")
) %}
{% endif %}
{% endfor %}

{# Not an exhaustive identifier grammar: the point is only to exclude
quotes, semicolons, whitespace and backslashes (the characters that
could break out of a quoted identifier or a SQL string literal),
while still accepting what real warehouses use here in practice —
including hyphens, which are near-universal in BigQuery project IDs
used as `database`. #}
{% set safe_identifier_pattern = '^[A-Za-z0-9_$.\-]+$' %}
{% for value, label in [(source_database, 'database'), (source_schema, 'schema')] %}
{% if not modules.re.fullmatch(safe_identifier_pattern, value) %}
{% do exceptions.raise_compiler_error(
"export_state: unsafe " ~ label ~ " value '" ~ value ~ "'; expected a plain SQL identifier matching " ~ safe_identifier_pattern
) %}
{% endif %}
{% endfor %}

{% set target_relation = api.Relation.create(
database=source_database, schema=source_schema, identifier=relation.identifier
) %}

{% set nodes = {} %}

{% if execute %}
{# load_relation (a dbt-core built-in) checks existence without raising,
unlike selecting from the relation directly, which errors out when
it is absent. #}
{% if resource_types | length == 0 %}
{# An empty IN (...) list is invalid SQL on most adapters; treat
"nothing requested" the same as "nothing found". #}
{% do log("export_state: resource_types is empty; exporting an empty document", info=False) %}
{% elif load_relation(target_relation) is none %}
{% do log("export_state: " ~ target_relation ~ " does not exist; exporting an empty document", info=False) %}
{% else %}
{% set query %}
select
node_id,
resource_type,
name,
package_name,
{% if target.type == "sqlserver" %} "database" {% else %} database {% endif %}, -- noqa
{% if target.type == "sqlserver" %} "schema" {% else %} schema {% endif %}, -- noqa
alias,
materialization,
checksum,
last_success_at,
command_invocation_id,
target_name
from {{ target_relation }}
where resource_type in ({{ "'" ~ resource_types | join("', '") ~ "'" }})
{% endset %}

{% set results = run_query(query) %}

{# dim_dbt__current_relations is grained on (node_id,
target_name): several targets writing to the same artifacts
tables can each have their own latest success for the same
node. Two cases:
- target_name given: at most one row per node_id already
matches it, so filtering is a straight lookup.
- target_name not given: more than one row per node_id can
come back (one per target that has ever built the node).
Collapse to one, keeping the most recently completed
success, so the single-target-project default behaviour
(one row per node) is unchanged and the multi-target case
picks the truly freshest relation instead of an arbitrary
target's.
`last_success_at_sort_keys` holds the raw (pre-formatting)
completion time used only for that comparison; it never
reaches the emitted document. #}
{% set last_success_at_sort_keys = {} %}

{# target_name is project-specific — not a closed set like
resource_types — so it isn't validated or interpolated into
SQL at all. Quote-doubling would not be a safe escape on
every adapter this package supports: Snowflake, BigQuery,
Spark and Databricks treat backslash as an escape character
in string literals, so a value ending in an odd run of
backslashes could desynchronise a doubled quote and reopen
the hole. Filtering here in Jinja, over a result set that is
at most hundreds of rows, avoids the per-adapter escaping
question entirely. #}
{% for row in results.rows %}
{% if target_name is none or row["target_name"] == target_name %}
{% set completed_at = row["last_success_at"] %}
{% set current_best = last_success_at_sort_keys.get(row["node_id"]) %}
{# A specific target_name already guarantees at most one
matching row per node_id, so every match wins
outright. With no target_name, only overwrite the
node's current winner when this row is strictly more
recent; a null completion time never displaces an
existing real one, but still seeds the entry the
first time a node is seen. #}
{% set wins = target_name is not none
or row["node_id"] not in nodes
or (completed_at is not none and (current_best is none or completed_at > current_best)) %}
{% if wins %}
{% set last_success_at = none %}
{% if completed_at is not none %}
{% set last_success_at = completed_at %}
{# Every adapter this package writes to records
query_completed_at from dbt's own UTC
run-results timing. A value that comes back
with no tzinfo (e.g. Postgres/Redshift's
timestamp-without-time-zone) is UTC wall-clock
time, not an unknown offset, so attaching it
explicitly turns the naive value into a real
instant. A value that already carries tzinfo
(e.g. Snowflake's TIMESTAMP_TZ) is left as-is.
isoformat() then gives the `T` separator and
explicit offset the README documents, instead
of the space-separated, offset-less string
`| string` produced. #}
{% if last_success_at.tzinfo is none %}
{% set last_success_at = last_success_at.replace(tzinfo=modules.pytz.utc) %}
{% endif %}
{% set last_success_at = last_success_at.isoformat() %}
{% endif %}
{% do nodes.update({
row["node_id"]: {
"resource_type": row["resource_type"],
"name": row["name"],
"package_name": row["package_name"],
"database": row["database"],
"schema": row["schema"],
"alias": row["alias"],
"materialization": row["materialization"],
"checksum": row["checksum"],
"last_success_at": last_success_at,
"command_invocation_id": row["command_invocation_id"],
}
}) %}
{% do last_success_at_sort_keys.update({row["node_id"]: completed_at}) %}
{% endif %}
{% endif %}
{% endfor %}

{% if nodes | length == 0 %}
{% do log("export_state: no successful executions found in " ~ target_relation ~ "; exporting an empty document", info=False) %}
{% endif %}
{% endif %}
{% endif %}

{% set document = {
"dbt_artifacts_state_version": state_version,
"generated_at": modules.datetime.datetime.now(modules.pytz.utc).isoformat(),
"source": {
"database": source_database,
"schema": source_schema,
"target_name": target_name,
},
"nodes": nodes,
} %}

{% do print(tojson(document)) %}

{% endmacro %}
Loading