Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
100 changes: 80 additions & 20 deletions docs/manage-data.md
Original file line number Diff line number Diff line change
@@ -1,11 +1,11 @@
---
title: "Data Management"
description: "Loading STAC collections and items into PostgreSQL using pypgstac"
description: "Loading and exporting STAC collections and items in a PgSTAC database"
external_links:
- name: "eoapi-k8s Repository"
url: "https://github.com/developmentseed/eoapi-k8s"
- name: "pypgstac Documentation"
url: "https://github.com/stac-utils/pypgstac"
- name: "PgSTAC Documentation"
url: "https://github.com/stac-utils/pgstac"
- name: "STAC Specification"
url: "https://stacspec.org/"
---
Expand All @@ -14,33 +14,93 @@ external_links:

eoAPI-k8s provides a basic data ingestion process that consist of manual operations on the components of the stack.

# Load data
Data management is built from two layers:

You will have to have STAC records for the collection and items you wish to load (e.g., `collections.json` and `items.json`).
[This repo](https://github.com/vincentsarago/MAXAR_opendata_to_pgstac) contains a few script that may help you to generate sample input data.
- **`scripts/raw/`** — the actual load/export logic. Each script only needs `psql` (and, for
ingest, `jq`) on `PATH` and a reachable Postgres DSN — no `kubectl`, no cluster, nothing
Kubernetes-specific at all. Run these directly against any PgSTAC database: local, external,
port-forwarded, whatever.
- **`eoapi-cli ingest` / `eoapi-cli export`** (`scripts/data-management.sh`) — a thin Kubernetes
convenience wrapper around the scripts above. Pass an explicit DSN and it just passes straight
through, no `kubectl` calls at all; omit it and it auto-discovers and port-forwards the
*current* cluster's own database for you.

## Preshipped bash script
You will need STAC records for the collections and items you wish to load (e.g. `collections.json`
and `items.json`, NDJSON or a plain JSON document either way).
[This repo](https://github.com/vincentsarago/MAXAR_opendata_to_pgstac) contains a few scripts that
may help you generate sample input data.

Execute `make ingest` to load data into the eoAPI service - it expects `collections.json` and `items.json` in the current directory.
# Load data

## Manual steps
`scripts/raw/ingest.sh --dsn DSN [COLLECTIONS_FILE] [ITEMS_FILE]` loads STAC collections/items into
a PgSTAC database, using the same "insert, ignore duplicates" semantics as `pypgstac load --method
insert_ignore` — but talking to pgstac's own SQL functions (`pgstac.upsert_collection`, the
`pgstac.items_staging_ignore` staging table) directly instead of depending on `pypgstac`. Files
default to `./collections.json` / `./items.json`.

In order to add raster data to eoAPI you can load STAC collections and items into the PostgreSQL database using pgSTAC and the tool `pypgstac`.
```bash
scripts/raw/ingest.sh --dsn "postgresql://user:pass@host:5432/postgis" collections.json items.json
```

First, ensure your Kubernetes cluster is running and `kubectl` is configured to access and modify it.
`eoapi-cli ingest [COLLECTIONS_FILE] [ITEMS_FILE]` wraps this for the current cluster: with
`--target-dsn`, it passes straight through (no `kubectl` involved); without it, it auto-discovers
and port-forwards the cluster's own database, the same way `eoapi-cli export` does (see below).

In a second step, you'll have to upload the data into the pod running the raster eoAPI service. You can use the following commands to copy the data:
# Export data

STAC collections and items can be exported from a PgSTAC database to `collections.ndjson` and
`items.ndjson`, loadable straight back in with `eoapi-cli ingest OUTPUT_DIR/collections.ndjson
OUTPUT_DIR/items.ndjson`. This is primarily meant for migrating data between PgSTAC instances:
export from an old instance, then ingest the result into a new one.

```bash
kubectl cp collections.json "$NAMESPACE/$EOAPI_POD_RASTER":/tmp/collections.json
kubectl cp items.json "$NAMESPACE/$EOAPI_POD_RASTER":/tmp/items.json
scripts/raw/export.sh --dsn "postgresql://user:pass@old-pgstac-host:5432/postgis" ./stac-export
```
Then, bash into the pod or server running the raster eoAPI service, you can use the following commands to load the data:

Use `--collection <id>` (repeatable) to export only specific collections instead of the whole
catalog.

`eoapi-cli export [OUTPUT_DIR]` wraps this the same way as ingest: pass `--source-dsn` to export
from an external PgSTAC instance (e.g. the old one you're migrating away from) — this just passes
straight through to `scripts/raw/export.sh`, no `kubectl` calls happen at all:

```bash
#!/bin/bash
apt update -y && apt install python3 python3-pip -y && pip install pypgstac[psycopg]';
pypgstac pgready --dsn $PGADMIN_URI
pypgstac load collections /tmp/collections.json --dsn $PGADMIN_URI --method insert_ignore
pypgstac load items /tmp/items.json --dsn $PGADMIN_URI --method insert_ignore
eoapi-cli export --source-dsn "postgresql://user:pass@old-pgstac-host:5432/postgis" ./stac-export
```

## Auto-discovery (no --source-dsn / --target-dsn)

Without an explicit DSN, both `eoapi-cli` commands auto-discover and port-forward the *current*
cluster's own database (useful for testing/round-tripping against a dev deployment): they read the
`{release}-pguser-postgres` secret's `uri` key, find the CrunchyData PGO primary pod (via its
`postgres-operator.crunchydata.com/role=master` label — the `-primary` Service itself is headless
with no selector, so `kubectl port-forward` can't target it directly), port-forward to that pod,
and call the matching `scripts/raw/` script with the resulting local DSN. This only works when
`postgresql.type` is `postgrescluster` (the chart's default) — for external-database
configurations, pass `--source-dsn`/`--target-dsn` directly instead. Verified against a real k3d
deployment.

## How export hydration works

Items are stored dehydrated in PgSTAC (fields shared with the collection are stripped out to save
space), so a plain `SELECT * FROM items` would not produce valid, complete STAC Items. The script
calls PgSTAC's own `pgstac.content_hydrate(items)` SQL function to reassemble full items, and
streams query output straight to a file via psql's `\o` redirection (with unaligned, tuples-only
output — one JSON object per line, i.e. NDJSON) — all within a single psql session per invocation,
rather than one connection per query, since some port-forward setups only tolerate a small number
of sequential connections before the tunnel drops.

**Caveat:** `content_hydrate()`'s exact signature has changed across PgSTAC releases (and may
change again). The script verifies the function exists on the source database before exporting and
fails with a clear error otherwise. If you hit that error against an unusual/very old or very new
PgSTAC version, fall back to inspecting the source database directly (`\df pgstac.*hydrate*` in
`psql`) and adjust the export query manually — collections are stored whole and can always be
exported directly with `select content from pgstac.collections`.

Ingest has the same kind of version sensitivity in reverse: it relies on `pgstac.upsert_collection`
and the `pgstac.items_staging_ignore` staging table existing, and verifies both before loading.
`scripts/raw/ingest.sh` deliberately avoids `psql \copy` for the bulk item load — its default TEXT
format treats backslash as its own escape character, which can corrupt JSON strings that already
contain backslash sequences. Instead it streams one `insert into pgstac.items_staging_ignore
(content) values ('<escaped>'::jsonb);` statement per record through a single `psql -f -` session,
inside one transaction.
23 changes: 20 additions & 3 deletions eoapi-cli
Original file line number Diff line number Diff line change
Expand Up @@ -18,6 +18,7 @@ readonly COMMANDS=(
"test"
"load"
"ingest"
"export"
"docs"
)

Expand All @@ -41,6 +42,7 @@ COMMANDS:
test Run tests (helm, integration, autoscaling)
load Run load testing scenarios
ingest Load sample data into eoAPI services
export Export STAC collections/items to NDJSON (e.g. for migrating to another PgSTAC)
docs Generate and serve documentation

Use 'eoapi-cli <COMMAND> --help' for more information about a specific command.
Expand All @@ -67,6 +69,9 @@ EXAMPLES:
# Ingest sample data
eoapi-cli ingest sample-data

# Export STAC data to NDJSON (e.g. before migrating to a new PgSTAC)
eoapi-cli export ./stac-export

# Serve documentation locally
eoapi-cli docs serve

Expand Down Expand Up @@ -108,7 +113,10 @@ get_command_script() {
echo "${SCRIPTS_DIR}/load.sh"
;;
ingest)
echo "${SCRIPTS_DIR}/ingest.sh"
echo "${SCRIPTS_DIR}/data-management.sh"
;;
export)
echo "${SCRIPTS_DIR}/data-management.sh"
;;
docs)
echo "${SCRIPTS_DIR}/docs.sh"
Expand Down Expand Up @@ -137,8 +145,17 @@ execute_command() {
chmod +x "$script_path"
fi

# Execute the command script with remaining arguments
exec "$script_path" "$@"
# data-management.sh backs both the 'ingest' and 'export' commands and
# needs to know which mode to run in; every other script is invoked with
# just its own remaining arguments, unchanged.
case "$cmd" in
ingest|export)
exec "$script_path" "$cmd" "$@"
;;
*)
exec "$script_path" "$@"
;;
esac
}

main() {
Expand Down
28 changes: 20 additions & 8 deletions scripts/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,13 +7,16 @@ This directory contains the implementation scripts for the eoAPI CLI.
```
scripts/
├── lib/
│ ├── common.sh # Shared utilities (logging, validation)
│ └── k8s.sh # Kubernetes helper functions
├── cluster.sh # Cluster management (start, stop, clean, status, inspect)
├── deployment.sh # Deployment operations (run, debug)
├── test.sh # Test suites (schema, lint, unit, integration)
├── ingest.sh # Data ingestion
└── docs.sh # Documentation (generate, serve)
│ ├── common.sh # Shared utilities (logging, validation)
│ └── k8s.sh # Kubernetes helper functions
├── raw/
│ ├── ingest.sh # Data ingestion (pure: psql/jq + a DSN, no kubectl)
│ └── export.sh # Data export to NDJSON (pure: psql + a DSN, no kubectl)
├── cluster.sh # Cluster management (start, stop, clean, status, inspect)
├── deployment.sh # Deployment operations (run, debug)
├── test.sh # Test suites (schema, lint, unit, integration)
├── data-management.sh # Kubernetes wrapper for raw/ingest.sh and raw/export.sh
└── docs.sh # Documentation (generate, serve)
```

## Usage
Expand All @@ -28,6 +31,7 @@ All scripts are accessed through the main CLI:
./eoapi-cli deployment run
./eoapi-cli test all
./eoapi-cli ingest collections.json items.json
./eoapi-cli export ./stac-export
./eoapi-cli docs serve
```

Expand Down Expand Up @@ -80,10 +84,18 @@ The eoAPI CLI provides a unified interface for all operations:
./eoapi-cli test integration # Run integration tests
```

### Data Ingestion
### Data Ingestion and Export
```bash
# Ingest sample data
./eoapi-cli ingest <collections-file> <items-file>

# Export STAC data to NDJSON (e.g. before migrating to a new PgSTAC)
./eoapi-cli export ./stac-export

# Both auto-discover/port-forward the current cluster's own database by
# default; pass --target-dsn / --source-dsn to talk to any other PgSTAC
# database instead (see docs/manage-data.md, or scripts/raw/*.sh directly
# for a Kubernetes-free standalone tool).
```

### Documentation
Expand Down
Loading