Skip to content

Latest commit

 

History

History

Folders and files

NameName
Last commit message
Last commit date

parent directory

..
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

README.md

Documentation

Deep dives for each component of inference-reliability-platform — a Brev-launchable, single-GPU vLLM inference platform on Kubernetes engineered for inference reliability.

The root README.md is the fast path (what it is, how to launch, how to send a request). Everything else lives here.

Component docs

Doc What it covers
01 — Architecture Big-picture: namespaces, request path, data flow, why every piece exists
02 — Reliability model The core of the repo. How every layer contributes to inference reliability — SLOs, error budgets, KV-cache pressure, preemption, graceful drain, priority, GPU health
03 — Bootstrap & GitOps (ArgoCD) bootstrap/install.sh, root-app, sync waves, app-of-apps, how to add a new app
04 — Kubernetes cluster Cluster assumptions, storage class, RuntimeClass, priority classes, node layout
05 — GPU Operator & DCGM NVIDIA GPU Operator config, DCGM exporter, XID/ECC alerts, /dev/shm and RuntimeClass gotchas
06 — Inference stack (vLLM Helm chart) charts/llama-8b — every value, probe, PVC, priority class, network policy, rollout gate
07 — Gateway API Inference Extension (EPP) InferencePool, endpoint picker, KV-cache-aware routing, RBAC, gRPC ext_proc wiring
08 — Gateway (Envoy Gateway) GatewayClass, listeners, EnvoyProxy telemetry, PodMonitor, HTTPRoutes, rate-limit policy, ext_proc extension policy
09 — KEDA autoscaling The ScaledObject, Prometheus trigger query, activation, why min/max is 1 today, how to scale on multi-GPU
10 — Secrets (ESO + Vault) ClusterSecretStore, ExternalSecrets, seeding Vault, rotating tokens, moving to prod Vault
11 — Kyverno policies Every ClusterPolicy — what it enforces, why it exists, mutate vs. validate, audit vs. enforce
12 — Observability stack Prometheus, Grafana, Loki, Tempo, OTel collector, Alertmanager — wiring and retention
13 — Dashboards Each Grafana dashboard: what it shows, which metrics, when to use it (vLLM, GPU, gateway, cost, loadtests, model quality)
14 — Alerts Every Prometheus alert with condition, severity, runbook link, and error-budget math
15 — Load testing vllm bench serve wrapper, Argo WorkflowTemplate, suite DAG, nightly CronWorkflow, Pushgateway metrics, how to adapt for your workload
16 — Model quality evals evals/ — prompt set, evaluator, Pushgateway metrics, dashboard, alerts, adding new prompts
17 — CI/CD (GitHub Actions) ci.yml (lint/test/render), e2e.yml (kind cluster), how to add a check
18 — Operations runbook Day-2: rollouts, rollbacks, upgrades, incident triage per alert, GPU quarantine, cost levers
19 — Extending the platform Add another model, another route, another dashboard, another eval, swap Llama for Mixtral, run on 2+ GPUs
20 — Making inference reliable (design notes) Beyond this repo: what else matters — request shaping, canary, saturation, capacity planning, chaos, safety, evaluation cadence

Screenshots

Live screenshots of the running platform live under ../images/screenshots/ and are embedded in the relevant docs:

How to read these docs

  • Read 01-architecture.md and 02-reliability.md first — they frame everything else.
  • Each component doc is standalone: cite the files, explain the config, note the gotchas.
  • Every doc ends with "Extending / Operating" — the levers you'll actually touch.