Repository navigation
materialize-s3-json: new connector writing gzipped NDJSON - #5253
Open
SeanWhelan wants to merge 1 commit into
Open
SeanWhelan wants to merge 1 commit into
SeanWhelan wants to merge 1 commit into
Conversation
…JSON Adds a JSON file format alongside the existing CSV and Parquet object-store sinks. Each line of a materialized file is a single JSON document, gzip compressed and written as `.json.gz`, matching the extension already used for staged newline-delimited JSON elsewhere in the repo. The format helper wraps `writer.JsonWriter`, which already satisfies `filesink.StreamWriter` and is in use by the Redshift, Databricks, BigQuery, MotherDuck and Snowflake stagers, so the connector is the same shape as materialize-s3-csv with the writer and file extension swapped. Object and array fields are written as nested JSON rather than encoded strings. The one format option exposed is `skipNulls`, which omits null-valued fields instead of writing them as null. It maps to `writer.WithJsonSkipNulls`, which had no test coverage, so `TestJsonWriterSkipNulls` is added alongside the existing writer tests. No CHANGELOG entry: the connector is new, so there is no existing behaviour to describe a change to.
Contributor
Strix Security ReviewNo security issues found. Updated for Reviewed by Strix |
|
Docs preview: https://connectors-pr-5253-estuary-docs.mike-39e.workers.dev Changed pages:
Rebuilt on every push, and the link stays the same. |
danielnelson
approved these changes
Sep 17, 2026
| @@ -0,0 +1 @@ | |||
| v1 No newline at end of file | |||
Contributor
There was a problem hiding this comment.
Nit: can you add a trailing EOL.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description:
Adds
materialize-s3-json, a third object-store file format alongside the existing CSV and Parquet sinks. Each line of a materialized file is one JSON document, gzip compressed and written as.json.gz. Object and array fields are written as nested JSON rather than encoded strings.The connector is materialize-s3-csv with the writer and file extension swapped. The new
filesink.NewJsonStreamWriterwrapswriter.JsonWriter, which already satisfiesfilesink.StreamWriterand is what the Redshift, Databricks, BigQuery, MotherDuck and Snowflake stagers use, so there is no new serialization code here.Workflow steps:
Same as the CSV and Parquet S3 connectors: configure bucket, region, credentials, optional prefix and upload interval, then one binding per collection with a
path. The only format-specific option isjsonConfig.skipNulls, which omits null-valued fields instead of writing them as null.Notes for reviewers:
The file extension is
.json.gzrather than.jsonl.gz. The contents are newline-delimited, but.jsonis what materialize-motherduck already stages NDJSON as, and it is what Athena, Glue crawlers and Snowflake key off. Easy to change if you'd rather be literal.writer.WithJsonSkipNullshad no test coverage anywhere, soTestJsonWriterSkipNullsis added next to the existing writer tests. Nothing was added tofilesink, which does not test its other format constructors.Requested by a customer who asked for JSON output from the S3 materialization. Still needs a control-plane connectors row, a UI tile, and the
go.estuary.dev/materialize-s3-jsonshort link, none of which live in this repo.Checklist:
CHANGELOG.md: the connector is new, so there is no prior behaviour to describe a change to. materialize-s3-csv likewise has none.amazon-s3-json.mdpage plus the materialization-connectors index entry.