-
Notifications
You must be signed in to change notification settings - Fork 183
chore: add Redshift support #453
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Open
hanslemm
wants to merge
23
commits into
brooklyn-data:main
Choose a base branch
from
hanslemm:REDSHIFT_SUPPORT
base: main
Could not load branches
Branch not found: {{ refName }}
Loading
Could not load tags
Nothing to show
Loading
Are you sure you want to change the base?
Some commits from the old base branch may be removed from the timeline,
and old review comments may become outdated.
Open
Changes from all commits
Commits
Show all changes
23 commits
Select commit
Hold shift + click to select a range
92a6876
add redshift support
brendan-cook-87 8ba69db
Merge branch 'main' into REDSHIFT_SUPPORT
hanslemm 6ad19c4
chore(profiles): clean up
hanslemm 2be379f
chore(type_helpers): redshift SUPER type for JSON and ARRAY
hanslemm dfe7572
Merge branch 'brooklyn-data:main' into REDSHIFT_SUPPORT
hanslemm 72f5e4a
Update dbt_project.yml
hanslemm 42b11cd
Merge branch 'main' into REDSHIFT_SUPPORT
hanslemm d6bd6b3
Update README.md
hanslemm 70d0081
Update type_helpers.sql
shiv-io ab81712
Update macros/database_specific_helpers/type_helpers.sql
hanslemm c5f3ed1
Update CONTRIBUTING.md
hanslemm 4069ed7
Add Trino array type macro to type_helpers.sql
hanslemm 637586e
Merge branch 'main' into REDSHIFT_SUPPORT
hanslemm c58f50d
Merge branch 'REDSHIFT_SUPPORT' into REDSHIFT_SUPPORT
hanslemm 2021662
Merge pull request #1 from shiv-io/REDSHIFT_SUPPORT
hanslemm 91ad526
chore: merge upstream 2.11.0 into REDSHIFT_SUPPORT
hanslemm 495aecf
chore: nest generic test arguments under arguments (dbt 1.10+)
hanslemm b93ed16
feat: dim_dbt__current_relations for deferral state
hanslemm 54bb9cb
feat: export_state run-operation for deferral state
hanslemm f1b2849
fix: validate and escape export_state's interpolated arguments
hanslemm a455f14
fix: filter target_name in Jinja and widen the identifier pattern
hanslemm d152d7a
fix(dim_dbt__current_relations): rank per (node_id, target_name); ord…
hanslemm 2179d12
fix(export_state): fall back to a node's own target; emit ISO 8601 ti…
hanslemm File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
25 changes: 25 additions & 0 deletions
25
integration_test_project/tests/assert_current_relations_unique.sql
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,25 @@ | ||
| {{ config(enabled = target.type in ["postgres", "redshift"]) }} | ||
| -- dim_dbt__current_relations must hold at most one row per (node_id, | ||
| -- target_name) and never a null relation coordinate: a consumer patches a | ||
| -- manifest with these values, so a duplicate silently picks a winner and a | ||
| -- null produces an unusable relation name. Returns rows (fails) on either | ||
| -- condition. | ||
| with duplicates as ( | ||
| select node_id | ||
| from {{ ref("dim_dbt__current_relations") }} | ||
| group by node_id, target_name | ||
| having count(*) > 1 | ||
| ), | ||
|
|
||
| nulls as ( | ||
| select node_id | ||
| from {{ ref("dim_dbt__current_relations") }} | ||
| where database is null | ||
| or schema is null | ||
| or alias is null | ||
| or resource_type is null | ||
| ) | ||
|
|
||
| select node_id from duplicates | ||
| union all | ||
| select node_id from nulls |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,204 @@ | ||
| {# | ||
| Prints dim_dbt__current_relations as a versioned JSON document, for use as | ||
| dbt deferral state. Emits with print() so `dbt --quiet run-operation` gives | ||
| clean JSON on stdout. Never raises for a missing relation, an empty | ||
| resource_types list, or a zero-row result: each yields an empty `nodes` | ||
| object and a warning, leaving the caller to decide. Raises only on | ||
| invalid arguments (an unknown resource_type, or a database/schema value | ||
| that isn't a safe identifier). | ||
|
|
||
| Emits one node per node_id. dim_dbt__current_relations is grained on | ||
| (node_id, target_name), since several targets can write to the same | ||
| artifacts tables; when target_name is given, each node's row for that | ||
| target is used, and when it isn't, each node's most recently completed | ||
| success across all targets is used, matching the single-target-project | ||
| behaviour where there is only ever one target to pick from. | ||
|
|
||
| Usage: | ||
| dbt --quiet run-operation dbt_artifacts.export_state \ | ||
| --args '{schema: dbt_artifacts, target_name: betterdata_prod}' > export.json | ||
| #} | ||
| {% macro export_state(resource_types=['model', 'seed', 'snapshot'], database=none, schema=none, target_name=none) %} | ||
|
|
||
| {% set state_version = 1 %} | ||
| {% set relation = ref("dim_dbt__current_relations") %} | ||
| {% set source_database = database if database is not none else relation.database %} | ||
| {% set source_schema = schema if schema is not none else relation.schema %} | ||
|
|
||
| {# `resource_types` is interpolated straight into an IN (...) list below, and | ||
| `source_database`/`source_schema` are interpolated indirectly: they become | ||
| the FROM-clause relation via api.Relation.create(), whose quoted() wraps | ||
| each part in the adapter's quote character but does not escape a quote | ||
| character embedded in the value, so an unvalidated value could still break | ||
| out of the quoted identifier. Validate both, ahead of any query, so a bad | ||
| value fails loudly instead of reaching the database. This is the one place | ||
| this macro may raise: validation happens before querying starts, so the | ||
| "never raises on a missing relation or zero rows" contract, which is about | ||
| behaviour once the query runs, is untouched. `target_name` is not | ||
| constrained to a known set (target names are project-specific) and is not | ||
| interpolated into SQL at all — see the row loop below — so it needs no | ||
| validation here. #} | ||
| {% set known_resource_types = ['model', 'seed', 'snapshot'] %} | ||
| {% for resource_type in resource_types %} | ||
| {% if resource_type not in known_resource_types %} | ||
| {% do exceptions.raise_compiler_error( | ||
| "export_state: unknown resource_type '" ~ resource_type ~ "'; expected one of " ~ known_resource_types | join(", ") | ||
| ) %} | ||
| {% endif %} | ||
| {% endfor %} | ||
|
|
||
| {# Not an exhaustive identifier grammar: the point is only to exclude | ||
| quotes, semicolons, whitespace and backslashes (the characters that | ||
| could break out of a quoted identifier or a SQL string literal), | ||
| while still accepting what real warehouses use here in practice — | ||
| including hyphens, which are near-universal in BigQuery project IDs | ||
| used as `database`. #} | ||
| {% set safe_identifier_pattern = '^[A-Za-z0-9_$.\-]+$' %} | ||
| {% for value, label in [(source_database, 'database'), (source_schema, 'schema')] %} | ||
| {% if not modules.re.fullmatch(safe_identifier_pattern, value) %} | ||
| {% do exceptions.raise_compiler_error( | ||
| "export_state: unsafe " ~ label ~ " value '" ~ value ~ "'; expected a plain SQL identifier matching " ~ safe_identifier_pattern | ||
| ) %} | ||
| {% endif %} | ||
| {% endfor %} | ||
|
|
||
| {% set target_relation = api.Relation.create( | ||
| database=source_database, schema=source_schema, identifier=relation.identifier | ||
| ) %} | ||
|
|
||
| {% set nodes = {} %} | ||
|
|
||
| {% if execute %} | ||
| {# load_relation (a dbt-core built-in) checks existence without raising, | ||
| unlike selecting from the relation directly, which errors out when | ||
| it is absent. #} | ||
| {% if resource_types | length == 0 %} | ||
| {# An empty IN (...) list is invalid SQL on most adapters; treat | ||
| "nothing requested" the same as "nothing found". #} | ||
| {% do log("export_state: resource_types is empty; exporting an empty document", info=False) %} | ||
| {% elif load_relation(target_relation) is none %} | ||
| {% do log("export_state: " ~ target_relation ~ " does not exist; exporting an empty document", info=False) %} | ||
| {% else %} | ||
| {% set query %} | ||
| select | ||
| node_id, | ||
| resource_type, | ||
| name, | ||
| package_name, | ||
| {% if target.type == "sqlserver" %} "database" {% else %} database {% endif %}, -- noqa | ||
| {% if target.type == "sqlserver" %} "schema" {% else %} schema {% endif %}, -- noqa | ||
| alias, | ||
| materialization, | ||
| checksum, | ||
| last_success_at, | ||
| command_invocation_id, | ||
| target_name | ||
| from {{ target_relation }} | ||
| where resource_type in ({{ "'" ~ resource_types | join("', '") ~ "'" }}) | ||
| {% endset %} | ||
|
|
||
| {% set results = run_query(query) %} | ||
|
|
||
| {# dim_dbt__current_relations is grained on (node_id, | ||
| target_name): several targets writing to the same artifacts | ||
| tables can each have their own latest success for the same | ||
| node. Two cases: | ||
| - target_name given: at most one row per node_id already | ||
| matches it, so filtering is a straight lookup. | ||
| - target_name not given: more than one row per node_id can | ||
| come back (one per target that has ever built the node). | ||
| Collapse to one, keeping the most recently completed | ||
| success, so the single-target-project default behaviour | ||
| (one row per node) is unchanged and the multi-target case | ||
| picks the truly freshest relation instead of an arbitrary | ||
| target's. | ||
| `last_success_at_sort_keys` holds the raw (pre-formatting) | ||
| completion time used only for that comparison; it never | ||
| reaches the emitted document. #} | ||
| {% set last_success_at_sort_keys = {} %} | ||
|
|
||
| {# target_name is project-specific — not a closed set like | ||
| resource_types — so it isn't validated or interpolated into | ||
| SQL at all. Quote-doubling would not be a safe escape on | ||
| every adapter this package supports: Snowflake, BigQuery, | ||
| Spark and Databricks treat backslash as an escape character | ||
| in string literals, so a value ending in an odd run of | ||
| backslashes could desynchronise a doubled quote and reopen | ||
| the hole. Filtering here in Jinja, over a result set that is | ||
| at most hundreds of rows, avoids the per-adapter escaping | ||
| question entirely. #} | ||
| {% for row in results.rows %} | ||
| {% if target_name is none or row["target_name"] == target_name %} | ||
| {% set completed_at = row["last_success_at"] %} | ||
| {% set current_best = last_success_at_sort_keys.get(row["node_id"]) %} | ||
| {# A specific target_name already guarantees at most one | ||
| matching row per node_id, so every match wins | ||
| outright. With no target_name, only overwrite the | ||
| node's current winner when this row is strictly more | ||
| recent; a null completion time never displaces an | ||
| existing real one, but still seeds the entry the | ||
| first time a node is seen. #} | ||
| {% set wins = target_name is not none | ||
| or row["node_id"] not in nodes | ||
| or (completed_at is not none and (current_best is none or completed_at > current_best)) %} | ||
| {% if wins %} | ||
| {% set last_success_at = none %} | ||
| {% if completed_at is not none %} | ||
| {% set last_success_at = completed_at %} | ||
| {# Every adapter this package writes to records | ||
| query_completed_at from dbt's own UTC | ||
| run-results timing. A value that comes back | ||
| with no tzinfo (e.g. Postgres/Redshift's | ||
| timestamp-without-time-zone) is UTC wall-clock | ||
| time, not an unknown offset, so attaching it | ||
| explicitly turns the naive value into a real | ||
| instant. A value that already carries tzinfo | ||
| (e.g. Snowflake's TIMESTAMP_TZ) is left as-is. | ||
| isoformat() then gives the `T` separator and | ||
| explicit offset the README documents, instead | ||
| of the space-separated, offset-less string | ||
| `| string` produced. #} | ||
| {% if last_success_at.tzinfo is none %} | ||
| {% set last_success_at = last_success_at.replace(tzinfo=modules.pytz.utc) %} | ||
| {% endif %} | ||
| {% set last_success_at = last_success_at.isoformat() %} | ||
| {% endif %} | ||
| {% do nodes.update({ | ||
| row["node_id"]: { | ||
| "resource_type": row["resource_type"], | ||
| "name": row["name"], | ||
| "package_name": row["package_name"], | ||
| "database": row["database"], | ||
| "schema": row["schema"], | ||
| "alias": row["alias"], | ||
| "materialization": row["materialization"], | ||
| "checksum": row["checksum"], | ||
| "last_success_at": last_success_at, | ||
| "command_invocation_id": row["command_invocation_id"], | ||
| } | ||
| }) %} | ||
| {% do last_success_at_sort_keys.update({row["node_id"]: completed_at}) %} | ||
| {% endif %} | ||
| {% endif %} | ||
| {% endfor %} | ||
|
|
||
| {% if nodes | length == 0 %} | ||
| {% do log("export_state: no successful executions found in " ~ target_relation ~ "; exporting an empty document", info=False) %} | ||
| {% endif %} | ||
| {% endif %} | ||
| {% endif %} | ||
|
|
||
| {% set document = { | ||
| "dbt_artifacts_state_version": state_version, | ||
| "generated_at": modules.datetime.datetime.now(modules.pytz.utc).isoformat(), | ||
| "source": { | ||
| "database": source_database, | ||
| "schema": source_schema, | ||
| "target_name": target_name, | ||
| }, | ||
| "nodes": nodes, | ||
| } %} | ||
|
|
||
| {% do print(tojson(document)) %} | ||
|
|
||
| {% endmacro %} |
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
Uh oh!
There was an error while loading. Please reload this page.