Skip to content

Harden OpenShell eval pipeline: judge-safe model routing, gateway alias, best-effort cleanup - #110

Open
shricharan-ks wants to merge 11 commits into
RHEcosystemAppEng:mainfrom
shricharan-ks:forge-nommen-eval-hardening
Open

shricharan-ks wants to merge 11 commits into
RHEcosystemAppEng:mainfrom
shricharan-ks:forge-nommen-eval-hardening

Conversation

@shricharan-ks

Copy link
Copy Markdown
Collaborator

Why this is needed

The OpenShell OpenClaw profile (eval-engine=aeh_openshell_openclaw) evaluates the Forge "Chief of Staff" agent end-to-end: it stages a submission workspace, boots the agent in an OpenShell sandbox on a KubeVirt VM, drives scene-based cases (morning-briefing, analysis-panel), scores them with deterministic + LLM judges, and publishes results to MLflow. Running it in a per-user namespace (verification was done in forge-nommen) exposed five gaps that made runs either fail outright or silently lose all LLM-judge signal. This PR closes those gaps; none of them are covered by #108/#109.

Problems earlier → how this fixes them

# Problem before Root cause Fix in this PR
1 Every LLM judge errored; mean_reward unusable. The judge client (score.py) only speaks the Anthropic dialect (POST /v1/messages). For an openai/ model, LiteLLM converts that to the upstream's /responses API, which the upstream vLLM gateway rejects with 403 Authentication parameters missing. Model routing, not code: the judge override pointed at an openai/ route. New judge-glm-5-3 LiteLLM route (hosted_vllm/) that keeps judge calls on /chat/completions, plus aeh-judge-model-override=judge-glm-5-3 default.
2 Slow, off-production eval model. Default llm-model was rits/zai-org/glm-5-3 (reasoning model), while the deployed agent runs GLM-5-3-Flash. Evals were slower and didn't measure the model users actually get. Pipeline defaults drifted from the deployed agent config. Defaults switched to rits/zai-org/GLM-5-3-Flash (pipeline + example run + aeh-model-override).
3 Gateway endpoint broke outside its home namespace. The default openshell-gateway-endpoint embedded another namespace's DNS and didn't validate against the gateway certificate, so the same Pipeline failed when created anywhere else. The gateway runs on the openshell-saw-agent VMI in the workspace namespace; there was no stable local name for it. New namespace-local alias Service openshell (port 17670) selecting the VMI; default endpoint https://openshell:17670, valid wherever the alias is installed.
4 The finally task could fail or hang the whole run and leak PVCs. Cleanup used bitnami/kubectl from Docker Hub (rate-limited pulls) with no request timeout — a hung API call held cleanup until the finally budget expired and a failed delete left the PVC behind; a failed finally marks the PipelineRun failed. Cleanup acted like a gate when it is a janitor. In-cluster OpenShift CLI image (image-registry.openshift-image-registry.svc:5000/openshell/cli), --request-timeout=20s, 3 retries, onError: continue.
5 oc apply -n <ns> -f … failed for anyone but the original namespace. Manifests pinned namespace: ab-eval-flow / gz-forge-eval, so applying into a per-user namespace returned "the namespace from the provided object does not match". Hardcoded metadata. Namespace fields removed from all deployable manifests; the example PipelineRun uses short service names (http://litellm:4000, https://openshell:17670) that resolve in whatever namespace the run is created in.

Additionally, failures were slow to diagnose because nothing could tell which dependency (litellm, mlflow, minio, k8s API, github, pypi, agent gateway, postgres) was unreachable — hence the standalone probe task below.

Changes, one by one

  1. feat(forge-saw): add stable namespace-local openshell gateway alias Service — new config/forge-saw/openshell-eval-alias.yaml. A selector Service for the openshell-saw-agent VMI on port 17670. Justification: gives every namespace a stable, certificate-valid https://openshell:17670 endpoint without cross-namespace DNS or cert pinning; installed with the stack, referenced by the Pipeline default.
  2. feat(litellm): add GLM-5-3-Flash routes for agent and judge traffic — config/litellm/configmap.yaml adds rits/zai-org/GLM-5-3-Flash (openai/ lane, agent traffic) and judge-glm-5-3 (hosted_vllm lane, judge traffic) over the same INFERENCE_ENDPOINT_URL/INFERENCE_API_KEY env. Justification: the dialect-safe judge route is the only way the Anthropic-dialect judge client gets served by this OpenAI-dialect upstream (problem 1); the Flash route matches the deployed agent (problem 2).
  3. feat(pipeline): add eval-stack-probe task — new pipeline/tasks/eval-stack-probe.yaml: HTTP probes (mlflow, litellm, minio, k8s API, github, pypi, local wheels) + TCP probes (agent gateway, integration gateway, postgres). Deliberately not labelled part-of so it can run without the pipeline's NetworkPolicy allowances. Justification: one TaskRun replaces reading thousands of lines of evaluate logs to find a dead dependency; it is a triage tool, not a pipeline stage.
  4. fix(cleanup): make PVC cleanup best-effort and use in-cluster CLI image — pipeline/tasks/post/cleanup_pvc.yaml. Justification: problem 4 — cleanup must never fail the run or leak PVCs; the in-cluster registry removes the Docker Hub dependency.
  5. feat(pipeline): default to GLM-5-3-Flash with judge route and gateway alias — pipeline/pipelines/ci-pipeline-openshell.yaml + pipeline/runs/openshell-openclaw-pipelinerun.yaml. Justification: wires problems 1–3 into the Pipeline defaults so a plain oc create -f of the example run works in any namespace; keeps explicit timeout budgets (pipeline 3h / tasks 2h30m / finally 15m) and the part-of label required by the canonical NetworkPolicies.
  6. test: cover Flash defaults, judge route, alias Service, and cleanup hardening — tests/test_openshell_pipeline_profile.py (10 tests). Justification: locks in every default this PR changes — judge route must stay hosted_vllm, example run must stay namespace-agnostic, cleanup must stay best-effort — so a future edit can't silently regress them. Also asserts the evaluate step takes the model key from $(params.llm-api-key) (no namespace-secret dependency, per Fix OpenShell evaluation Secret dependency #108).
  7. chore: drop hardcoded namespaces from deployable manifests — 8 files, one line each. Justification: problem 5; makes every manifest apply cleanly with -n <any-namespace>.

Dependencies

Verification

  • End-to-end: PipelineRun scharan-verify-hardening-2 (namespace forge-nommen, 2026-10-09): all 5 tasks Succeeded; recommendation pass, mean_reward 0.7000 (threshold 0.5); 12 judges × 2 cases with 0 errors (previously all LLM judges errored); scorecard 6/6 gates; MLflow experiment recorded with traces + judge feedback; the hardened finally task deleted the run's PVC.
  • Baseline comparison: same-day runs on the pre-hardening stack (sana-morning-briefing-pr105-single-*, -collector-diag-*) scored 0.00–0.35 and failed the aeh-low-score gate.
  • Tests: tests/test_openshell_pipeline_profile.py — 10/10 pass.
  • Cluster parity: deployed objects in forge-nommen verified identical to this branch (all 44 Pipeline param defaults, every task step script, litellm configmap, alias Service selector).
Ops note (out of scope, for the record)

During verification the run-namespace LiteLLM lost its upstream key (the k8s secret holding it had been deleted; only the old pods still had it in memory). #108 already removed the eval task's dependency on that secret; the key for the namespace's LiteLLM deployment itself was restored from the vault-injected value on the running forge-ai-gateway pod. Keeping that namespace secret in sync with Vault is an ops concern outside this repo.

🤖 Generated with Claude Code

tarun-etikala and others added 11 commits October 8, 2026 13:24
Port the verified OIDC refresh, USER.md fixture, pinned SAW image, and
published_brief gate onto main. Add same-namespace NetworkPolicy templates
that match the canonical Pipeline name or app.kubernetes.io/part-of=abevalflow
so ad-hoc Pipeline copies no longer time out on gateway preflight.

Co-authored-by: Cursor <cursoragent@cursor.com>
Depth-1 --branch fails for bare SHAs; fall back to full clone + checkout
so clean-source pins work the same way as the submission revision.

Co-authored-by: Cursor <cursoragent@cursor.com>
…ervice

Selects the agent VMI directly so Pipeline defaults need no deployment
namespace. The gateway server leaf includes 'openshell' as a DNS SAN,
letting the evaluate task keep OPENSHELL_GATEWAY_INSECURE=false.

Co-Authored-By: Claude Code <noreply@anthropic.com>
- openai/ route for rits/zai-org/GLM-5-3-Flash (agent model)
- judge-glm-5-3 hosted_vllm route: LLM judges call /v1/messages
  (Anthropic dialect); with an openai/ model LiteLLM converts that to
  the upstream /responses API which answers 403, so every judge errors.
  hosted_vllm stays on /chat/completions. Select with
  aeh-judge-model-override=judge-glm-5-3.

Also drop the hardcoded gz-forge-eval namespace so the ConfigMap
applies anywhere.

Co-Authored-By: Claude Code <noreply@anthropic.com>
Probes every dependency an eval pod needs: mlflow, litellm, minio,
k8s api, github, pypi, github releases/raw, and TCP reachability of
the agent/integ gateways and postgres. Deliberately not labelled
part-of=abevalflow so NetworkPolicy sees it as an ordinary Tekton pod
and reports real connectivity.

Co-Authored-By: Claude Code <noreply@anthropic.com>
- image-registry.openshift-image-registry openshift/cli instead of
  registry.redhat.io (avoids pull failures without a redhat registry
  secret on the pipeline SA)
- --request-timeout on oc calls, 3 retries with backoff
- onError: continue so the finally task never fails the PipelineRun

Co-Authored-By: Claude Code <noreply@anthropic.com>
… alias

- llm-model / aeh-model-override default to the Flash model; the
  submission's model entry carries the 128000-token budget it needs
- aeh-judge-model-override defaults to judge-glm-5-3 (hosted_vllm
  route; a plain openai/ judge model 403s on /v1/messages)
- openshell-gateway-endpoint default becomes https://openshell:17670,
  the alias Service from config/forge-saw/openshell-eval-alias.yaml,
  so the Pipeline applies to any namespace
- example PipelineRun: short-name endpoints (namespace-agnostic,
  drops the gz-forge-eval hostAliases pin), main revisions, and
  timeouts sized for 15-minute cases (3h pipeline / 2h30m tasks)

Co-Authored-By: Claude Code <noreply@anthropic.com>
…ardening

Extends the OpenShell profile tests for the carried forge-nommen
changes: Flash model defaults, judge-glm-5-3 override, namespace-
agnostic example run, gateway alias Service selector, LiteLLM
Flash/judge routes, and best-effort PVC cleanup.

Co-Authored-By: Claude Code <noreply@anthropic.com>
Pinned namespaces (ab-eval-flow, gz-forge-eval) made `oc apply -n <ns>`
fail with a namespace mismatch for anyone deploying the OpenShell
profile into their own workspace (e.g. forge-nommen). Remove the
namespace fields so the manifests land in whatever namespace they are
applied to; the example PipelineRun already uses short service names
that resolve in the run namespace.

Co-Authored-By: Claude Code <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants