Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 6 additions & 1 deletion .github/workflows/npd-nvsentinel-object-monitor-e2e.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,11 @@
# The path filter is deliberately narrow. This lane costs a cluster spin-up and
# a full image pull, and pkg/recipe/** or pkg/bundler/** would drag it onto
# most core PRs; the allowlist machinery those packages own is already covered
# by unit tests that run everywhere.
# by unit tests that run everywhere. The localformat template directory is the
# one carve-out: this lane installs through the scripts those templates render.
# Unit tests cover those scripts as rendered text, against golden files and
# stub binaries; what they cannot cover is the script driving a real apiserver,
# which is the part that broke last time.

name: NPD Object Monitor E2E

Expand All @@ -44,6 +48,7 @@ on:
- 'recipes/overlays/kind.yaml'
- 'recipes/overlays/base.yaml'
- 'tests/e2e/npd-nvsentinel-object-monitor/**'
- 'pkg/bundler/deployer/localformat/templates/**'
# run.sh sources tools/common and is invoked through a make target.
- 'tools/common'
- 'Makefile'
Expand Down
7 changes: 6 additions & 1 deletion .github/workflows/nvsentinel-object-monitor-e2e.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -24,7 +24,11 @@
# The path filter is deliberately narrow. This lane costs a cluster spin-up and
# a full image pull, and pkg/recipe/** or pkg/bundler/** would drag it onto
# most core PRs; the allowlist machinery those packages own is already covered
# by unit tests that run everywhere.
# by unit tests that run everywhere. The localformat template directory is the
# one carve-out: this lane installs through the scripts those templates render.
# Unit tests cover those scripts as rendered text, against golden files and
# stub binaries; what they cannot cover is the script driving a real apiserver,
# which is the part that broke last time.

name: NVSentinel Object Monitor E2E

Expand All @@ -41,6 +45,7 @@ on:
- 'recipes/overlays/kind.yaml'
- 'recipes/overlays/base.yaml'
- 'tests/e2e/nvsentinel-object-monitor/**'
- 'pkg/bundler/deployer/localformat/templates/**'
# run.sh sources tools/common and is invoked through a make target.
- 'tools/common'
- 'Makefile'
Expand Down
7 changes: 6 additions & 1 deletion .github/workflows/nvsentinel-preflight-e2e.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -26,7 +26,11 @@
# The path filter is deliberately narrow. This lane costs a cluster spin-up and
# a full image pull, and pkg/recipe/** or pkg/bundler/** would drag it onto
# most core PRs; the allowlist machinery those packages own is already covered
# by unit tests that run everywhere.
# by unit tests that run everywhere. The localformat template directory is the
# one carve-out: this lane installs through the scripts those templates render.
# Unit tests cover those scripts as rendered text, against golden files and
# stub binaries; what they cannot cover is the script driving a real apiserver,
# which is the part that broke last time.

name: NVSentinel Preflight E2E

Expand All @@ -47,6 +51,7 @@ on:
- 'recipes/components/kai-scheduler/**'
- 'recipes/components/cert-manager/**'
- 'tests/e2e/nvsentinel-preflight/**'
- 'pkg/bundler/deployer/localformat/templates/**'
# run.sh sources tools/common and is invoked through a make target.
- 'tools/common'
- 'Makefile'
Expand Down
6 changes: 5 additions & 1 deletion pkg/bundler/deployer/helm/templates/README.md.tmpl
Original file line number Diff line number Diff line change
Expand Up @@ -109,7 +109,11 @@ bash install.sh
> Every helm and kubectl call it makes is bounded, 30s by default and
> overridable with `AICR_CRD_STEP_TIMEOUT`. The bound is per call and
> `deploy.sh` retries a failing component, so the budget it consumes is a
> multiple of that. The step is skipped when `DRY_RUN_FLAG` is set, and when
> multiple of that. `KUBECONFIG_FLAG` reaches this step as well, but only
> `--kube-context` and `--kubeconfig` are understood there: helm's remaining
> connection flags have no `kubectl` spelling, so the step stops with an error
> naming the flag rather than apply CRDs to an unintended cluster. The step is
> skipped when `DRY_RUN_FLAG` is set, and when
> neither the release nor any of the chart's CRDs exist yet, since only then
> does `helm install` create them. A release that was uninstalled leaves its
> CRDs behind, so a reinstall still applies them. Only components audited as
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -93,7 +93,11 @@ bash install.sh
> Every helm and kubectl call it makes is bounded, 30s by default and
> overridable with `AICR_CRD_STEP_TIMEOUT`. The bound is per call and
> `deploy.sh` retries a failing component, so the budget it consumes is a
> multiple of that. The step is skipped when `DRY_RUN_FLAG` is set, and when
> multiple of that. `KUBECONFIG_FLAG` reaches this step as well, but only
> `--kube-context` and `--kubeconfig` are understood there: helm's remaining
> connection flags have no `kubectl` spelling, so the step stops with an error
> naming the flag rather than apply CRDs to an unintended cluster. The step is
> skipped when `DRY_RUN_FLAG` is set, and when
> neither the release nor any of the chart's CRDs exist yet, since only then
> does `helm install` create them. A release that was uninstalled leaves its
> CRDs behind, so a reinstall still applies them. Only components audited as
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -93,7 +93,11 @@ bash install.sh
> Every helm and kubectl call it makes is bounded, 30s by default and
> overridable with `AICR_CRD_STEP_TIMEOUT`. The bound is per call and
> `deploy.sh` retries a failing component, so the budget it consumes is a
> multiple of that. The step is skipped when `DRY_RUN_FLAG` is set, and when
> multiple of that. `KUBECONFIG_FLAG` reaches this step as well, but only
> `--kube-context` and `--kubeconfig` are understood there: helm's remaining
> connection flags have no `kubectl` spelling, so the step stops with an error
> naming the flag rather than apply CRDs to an unintended cluster. The step is
> skipped when `DRY_RUN_FLAG` is set, and when
> neither the release nor any of the chart's CRDs exist yet, since only then
> does `helm install` create them. A release that was uninstalled leaves its
> CRDs behind, so a reinstall still applies them. Only components audited as
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -93,7 +93,11 @@ bash install.sh
> Every helm and kubectl call it makes is bounded, 30s by default and
> overridable with `AICR_CRD_STEP_TIMEOUT`. The bound is per call and
> `deploy.sh` retries a failing component, so the budget it consumes is a
> multiple of that. The step is skipped when `DRY_RUN_FLAG` is set, and when
> multiple of that. `KUBECONFIG_FLAG` reaches this step as well, but only
> `--kube-context` and `--kubeconfig` are understood there: helm's remaining
> connection flags have no `kubectl` spelling, so the step stops with an error
> naming the flag rather than apply CRDs to an unintended cluster. The step is
> skipped when `DRY_RUN_FLAG` is set, and when
> neither the release nor any of the chart's CRDs exist yet, since only then
> does `helm install` create them. A release that was uninstalled leaves its
> CRDs behind, so a reinstall still applies them. Only components audited as
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -93,7 +93,11 @@ bash install.sh
> Every helm and kubectl call it makes is bounded, 30s by default and
> overridable with `AICR_CRD_STEP_TIMEOUT`. The bound is per call and
> `deploy.sh` retries a failing component, so the budget it consumes is a
> multiple of that. The step is skipped when `DRY_RUN_FLAG` is set, and when
> multiple of that. `KUBECONFIG_FLAG` reaches this step as well, but only
> `--kube-context` and `--kubeconfig` are understood there: helm's remaining
> connection flags have no `kubectl` spelling, so the step stops with an error
> naming the flag rather than apply CRDs to an unintended cluster. The step is
> skipped when `DRY_RUN_FLAG` is set, and when
> neither the release nor any of the chart's CRDs exist yet, since only then
> does `helm install` create them. A release that was uninstalled leaves its
> CRDs behind, so a reinstall still applies them. Only components audited as
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -93,7 +93,11 @@ bash install.sh
> Every helm and kubectl call it makes is bounded, 30s by default and
> overridable with `AICR_CRD_STEP_TIMEOUT`. The bound is per call and
> `deploy.sh` retries a failing component, so the budget it consumes is a
> multiple of that. The step is skipped when `DRY_RUN_FLAG` is set, and when
> multiple of that. `KUBECONFIG_FLAG` reaches this step as well, but only
> `--kube-context` and `--kubeconfig` are understood there: helm's remaining
> connection flags have no `kubectl` spelling, so the step stops with an error
> naming the flag rather than apply CRDs to an unintended cluster. The step is
> skipped when `DRY_RUN_FLAG` is set, and when
> neither the release nor any of the chart's CRDs exist yet, since only then
> does `helm install` create them. A release that was uninstalled leaves its
> CRDs behind, so a reinstall still applies them. Only components audited as
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -43,6 +43,64 @@ if ! command -v kubectl >/dev/null 2>&1; then
exit 1
fi

# KUBECONFIG_FLAG holds helm's spelling of the connection options, because every
# other consumer forwards it unexamined into `helm upgrade`.
# kubectl spells one of them differently -- helm's --kube-context is kubectl's
# --context -- so the value is translated here rather than forwarded.
#
# An option with no known kubectl spelling is refused. Forwarding it aborts the
# deploy on an unknown flag, and dropping it is worse: the reads and the CRD
# force-apply below would then land on whatever cluster the ambient context
# names, which is the wrong-cluster write this script must never make.
KUBECTL_CONN=()
if [[ -n "${KUBECONFIG_FLAG:-}" ]]; then
# Deliberate word-split: this slot holds a flag list, not a single word.
# shellcheck disable=SC2206
helm_conn=(${KUBECONFIG_FLAG})
while (( ${#helm_conn[@]} > 0 )); do
case "${helm_conn[0]}" in
--kube-context|--kubeconfig)
# An option-looking value is a malformed list, not a context named
# "--kubeconfig". Accepting it consumes the next real option as this
# one's argument, which silently discards a connection option the
# operator did set.
if (( ${#helm_conn[@]} < 2 )) || [[ "${helm_conn[1]}" == --* ]]; then
echo "ERROR: KUBECONFIG_FLAG gives ${helm_conn[0]} no usable value." >&2
exit 1
fi
if [[ "${helm_conn[0]}" == "--kube-context" ]]; then
KUBECTL_CONN+=(--context "${helm_conn[1]}")
else
KUBECTL_CONN+=(--kubeconfig "${helm_conn[1]}")
fi
helm_conn=("${helm_conn[@]:2}")
;;
# An empty joined value is refused rather than forwarded. Both helm and
# kubectl read an empty --context as "use the current context", so
# passing it through turns a stated target into the ambient one without
# saying so -- the silent retarget this step exists to prevent.
--kube-context=|--kubeconfig=)
echo "ERROR: KUBECONFIG_FLAG gives ${helm_conn[0]%=} an empty value." >&2
exit 1
;;
--kube-context=*)
KUBECTL_CONN+=(--context "${helm_conn[0]#*=}")
helm_conn=("${helm_conn[@]:1}")
;;
--kubeconfig=*)
KUBECTL_CONN+=("${helm_conn[0]}")
helm_conn=("${helm_conn[@]:1}")
;;
*)
echo "ERROR: KUBECONFIG_FLAG carries '${helm_conn[0]}', which has no known" >&2
echo " kubectl spelling. The ${RELEASE} CRD step refuses to guess rather" >&2
echo " than read and apply CRDs against an unintended cluster." >&2
exit 1
;;
esac
done
fi

# Every helm and kubectl call below runs through run_bounded. This script runs
# inside the deploy path, where a command that never returns hangs the whole
# rollout rather than failing it: deploy.sh retries a component that exits
Expand Down Expand Up @@ -273,7 +331,8 @@ fi
# rather than an error. A failure here is indeterminate and fails closed, for
# the same reason the release lookup does.
if [[ "${RELEASE_EXISTS}" == "false" ]]; then
if ! capture_bounded kubectl get -f "${CRD_DIR}" --ignore-not-found -o name ${KUBECONFIG_FLAG:-}; then
if ! capture_bounded kubectl get -f "${CRD_DIR}" --ignore-not-found -o name \
${KUBECTL_CONN[@]+"${KUBECTL_CONN[@]}"}; then
echo "ERROR: cannot determine whether ${RELEASE} CRDs are already present; refusing" >&2
echo " to skip and risk pairing a new controller with a retained schema: $(cat "${BOUNDED_OUT}")" >&2
exit 1
Expand Down Expand Up @@ -314,7 +373,7 @@ while IFS= read -r doc; do
# sorted it can do so before the real CRDs are ever reached.
grep -q '[^[:space:]]' "${doc}" || continue
if ! capture_bounded kubectl apply --server-side --force-conflicts \
--field-manager=helm -f "${doc}" ${KUBECONFIG_FLAG:-}; then
--field-manager=helm -f "${doc}" ${KUBECTL_CONN[@]+"${KUBECTL_CONN[@]}"}; then
echo "ERROR: could not apply a ${RELEASE} CRD: $(cat "${BOUNDED_OUT}")" >&2
exit 1
fi
Expand Down
6 changes: 5 additions & 1 deletion pkg/bundler/deployer/helm/testdata/owns_crds/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -93,7 +93,11 @@ bash install.sh
> Every helm and kubectl call it makes is bounded, 30s by default and
> overridable with `AICR_CRD_STEP_TIMEOUT`. The bound is per call and
> `deploy.sh` retries a failing component, so the budget it consumes is a
> multiple of that. The step is skipped when `DRY_RUN_FLAG` is set, and when
> multiple of that. `KUBECONFIG_FLAG` reaches this step as well, but only
> `--kube-context` and `--kubeconfig` are understood there: helm's remaining
> connection flags have no `kubectl` spelling, so the step stops with an error
> naming the flag rather than apply CRDs to an unintended cluster. The step is
> skipped when `DRY_RUN_FLAG` is set, and when
> neither the release nor any of the chart's CRDs exist yet, since only then
> does `helm install` create them. A release that was uninstalled leaves its
> CRDs behind, so a reinstall still applies them. Only components audited as
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -93,7 +93,11 @@ bash install.sh
> Every helm and kubectl call it makes is bounded, 30s by default and
> overridable with `AICR_CRD_STEP_TIMEOUT`. The bound is per call and
> `deploy.sh` retries a failing component, so the budget it consumes is a
> multiple of that. The step is skipped when `DRY_RUN_FLAG` is set, and when
> multiple of that. `KUBECONFIG_FLAG` reaches this step as well, but only
> `--kube-context` and `--kubeconfig` are understood there: helm's remaining
> connection flags have no `kubectl` spelling, so the step stops with an error
> naming the flag rather than apply CRDs to an unintended cluster. The step is
> skipped when `DRY_RUN_FLAG` is set, and when
> neither the release nor any of the chart's CRDs exist yet, since only then
> does `helm install` create them. A release that was uninstalled leaves its
> CRDs behind, so a reinstall still applies them. Only components audited as
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -93,7 +93,11 @@ bash install.sh
> Every helm and kubectl call it makes is bounded, 30s by default and
> overridable with `AICR_CRD_STEP_TIMEOUT`. The bound is per call and
> `deploy.sh` retries a failing component, so the budget it consumes is a
> multiple of that. The step is skipped when `DRY_RUN_FLAG` is set, and when
> multiple of that. `KUBECONFIG_FLAG` reaches this step as well, but only
> `--kube-context` and `--kubeconfig` are understood there: helm's remaining
> connection flags have no `kubectl` spelling, so the step stops with an error
> naming the flag rather than apply CRDs to an unintended cluster. The step is
> skipped when `DRY_RUN_FLAG` is set, and when
> neither the release nor any of the chart's CRDs exist yet, since only then
> does `helm install` create them. A release that was uninstalled leaves its
> CRDs behind, so a reinstall still applies them. Only components audited as
Expand Down
Loading
Loading