-
Notifications
You must be signed in to change notification settings - Fork 15
Add agent harness documentation #52
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Merged
kubernetes-prow
merged 6 commits into
kubernetes-sigs:main
from
jessie1111101:add-agents-docs
Aug 21, 2026
Merged
Changes from all commits
Commits
Show all changes
6 commits
Select commit
Hold shift + click to select a range
f9b3d6e
Add agent harness documentation
jessie1111101 27beb6a
Address review on the agent harness docs
jessie1111101 2a29ff4
Drop the Codex column and the notes from the capability map
jessie1111101 a0a1c59
Fix the harness signature in the places people copy, and note two ant…
jessie1111101 66d4667
Import Path in the harness skeleton
jessie1111101 1985f3d
Use vendor-neutral examples in the agent harness docs
jessie1111101 File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,26 @@ | ||
| # Harness capabilities map | ||
|
|
||
| The eval skills and references are **agent-agnostic**: they describe what they | ||
| need as generic capabilities ("spawn a sub-agent", "run a command detached", | ||
| "schedule a wakeup", "keep durable state", "use an isolated worktree", "ask the | ||
| operator", "emit a heartbeat"). This file is the **one place** that maps those | ||
| capabilities to concrete tools per harness. | ||
|
|
||
| Skills should express needs generically and consult this map; **degrade | ||
| gracefully when a capability is absent** — every row has a generic fallback that | ||
| works on a bare harness with nothing but a shell. | ||
|
|
||
| | Capability | Claude Code | Antigravity | Generic fallback | | ||
| |---|---|---|---| | ||
| | **Spawn a sub-agent** | `Agent` (`subagent_type`) | `invoke_subagent` / `define_subagent` | run the work inline yourself in one shell | | ||
| | **Cheap vs strong model tier** | Haiku / Sonnet / Opus | `/models` (Flash / Pro) | one model for everything; just spend it sparingly | | ||
| | **Background / detached run** | `run_in_background` | `manage_task` / `manage_subagents` | `nohup … &` and poll a file/marker | | ||
| | **Scheduled wakeup / timer** | `ScheduleWakeup` | `schedule` | `sleep` between checks, or re-poll each turn | | ||
| | **Durable state** | task list | Artifacts / `write_to_file` | a plain notes file on disk | | ||
| | **Isolated worktree** | `EnterWorktree` | `run_command` + `git worktree` | `git worktree add` + a branch | | ||
| | **Ask the operator** | `AskUserQuestion` | `ask_question` | ask in chat | | ||
| | **Heartbeat / keepalive** | progress line, no early "done" | progress line, no early "done" | print a `still working: …` line each tick | | ||
|
|
||
| The **runner host** holds the durable run state (`RESUME_STAMP` under | ||
| `~/matrix-runs/<stamp>/`), so even a bare harness — one shell, no sub-agents, no | ||
| scheduler — can drive and re-attach to a run by polling files. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,144 @@ | ||
| # Agents | ||
|
|
||
| An **agent harness** is the thing under test. It drives one AI agent against one | ||
| task prompt and hands back a typed result the rest of the benchmark can score. | ||
| Everything in this layer lives under `devops_bench/agents/`. | ||
|
|
||
| The base class is `AgentHarness` (`devops_bench/agents/base.py`). It owns two | ||
| concerns so subclasses never have to: the base `run()` stamps wall-clock | ||
| **latency** onto every result, and it wraps the agent in a **safety net** — any | ||
| crash inside the agent is caught and turned into an errored result, so one faulty | ||
| agent never aborts the whole benchmark. Subclasses implement a single method, | ||
| `_execute()`, which does the provider-specific work and returns an `AgentResult`. | ||
|
|
||
| ```text | ||
| agent.run(prompt) -> AgentResult # base: latency + safety net | ||
| └─ agent._execute(prompt) # subclass: build invocation, parse, return | ||
| ``` | ||
|
|
||
| ## Supported harnesses | ||
|
|
||
| Four harnesses ship today. Each self-registers under a canonical key. | ||
|
|
||
| | Key | Wraps | How it runs | Capabilities | | ||
| | --- | --- | --- | --- | | ||
| | `gemini` | The Google **Gemini CLI** binary | Headless subprocess; trajectory parsed from `--output-format stream-json` on stdout | MCP, skills, rules, allowed-tools | | ||
| | `openclaw` | The **Openclaw Agent CLI** | `openclaw agent --local` with per-run isolated state/config; trajectory via `openclaw sessions export-trajectory` | MCP, skills, rules | | ||
| | `antigravity` | The **Antigravity CLI** (`agy` binary) | Headless subprocess that keeps the real `HOME` so cached OAuth/ADC credentials work (see the trust-boundary note below); trajectory parsed from the transcript JSONL it writes, token usage read from the conversation DB | MCP, skills, rules | | ||
| | `api` | **In-process** model call | Calls `get_model(provider, model)` and runs a model-agnostic MCP tool-use loop (`max_turns`, default 50) | MCP (spawns a stdio server), skills (served as tools), rules (system instruction) | | ||
|
|
||
| > `oc` is just a shorthand alias for the `openclaw` CLI; this doc uses `openclaw` throughout. | ||
|
|
||
| > [!NOTE] | ||
| > The `gemini` key names the CLI **harness** — the program that drives the agent. | ||
| > It is not the gemini **model**. You can run the gemini *model* through the `api` | ||
| > harness, or run a non-gemini model through the `gemini` CLI, because the harness | ||
| > and the model are chosen independently (see [Harness vs model](#harness-vs-model)). | ||
| > The alias `gemini-cli` also resolves to `gemini`, and is the default agent type. | ||
|
|
||
| ## Harness vs model | ||
|
|
||
| A harness does **not** hardcode a model. It reads `AGENT_PROVIDER` and | ||
| `AGENT_MODEL` from its config and maps them onto whatever it drives. | ||
|
|
||
| Every harness resolves `AGENT_PROVIDER` through one shared contract | ||
| (`devops_bench/core/model_providers.py`), so the same `AGENT_*` config behaves | ||
| identically across them. The `api` harness uses it to pick the adapter family and | ||
| backend for `get_model(provider, model)` and runs the tool-use loop in-process. | ||
| The CLI harnesses (`gemini`, `openclaw`) use it to route `AGENT_API_KEY` onto the | ||
| binary's provider-specific env var(s) and pass the model through: the Gemini CLI | ||
| gets `GEMINI_MODEL`, and openclaw gets a `--model provider/id` flag. Either way, | ||
| the model is a runtime input, never baked into the harness. | ||
|
|
||
| `antigravity` is the exception: it does not go through the shared contract. It | ||
| writes `AGENT_API_KEY` straight onto `GEMINI_API_KEY` and `GOOGLE_API_KEY` and | ||
| maps the model onto `GEMINI_MODEL` (`agents/cli/antigravity/agent.py`), so it is | ||
| Gemini-only in practice — pointing `AGENT_PROVIDER` at another provider will not | ||
| route it. | ||
|
|
||
| > [!WARNING] | ||
| > **`antigravity` runs with the operator's real `HOME`.** That is deliberate, so | ||
| > cached OAuth/ADC credentials keep working without a re-login, but it means the | ||
| > agent under test inherits read access to everything in that home directory — | ||
| > `~/.config/gcloud`, `~/.ssh`, shell history, other tools' tokens. Every other | ||
| > harness gets an isolated per-run state directory. Run untrusted agents under a | ||
| > dedicated account or an isolated `HOME`, and treat any credential reachable | ||
| > from that home as exposed to the agent. | ||
|
|
||
| For everything about providers, model ids, and how `get_model` resolves them, see | ||
| [Model providers](./model_providers.md). | ||
|
|
||
| ## Configuring a harness for an eval | ||
|
|
||
| Configuration is env-driven. The benchmark reads neutral `AGENT_*` variables and | ||
| each harness maps them onto its target. | ||
|
|
||
| **Selecting the harness** | ||
|
|
||
| | Variable | Default | Notes | | ||
| | --- | --- | --- | | ||
| | `BENCH_AGENT_TYPE` | `gemini-cli` (resolves to `gemini`) | The canonical key or an alias. The `--agent-type` flag overrides it. | | ||
|
|
||
| **Agent config** | ||
|
|
||
| | Variable | Default | Notes | | ||
| | --- | --- | --- | | ||
| | `AGENT_MODEL` | unset | Model id; flows to the harness's target. | | ||
| | `AGENT_PROVIDER` | unset | Provider key (e.g. `gemini`, `anthropic`, `google-vertex`). | | ||
| | `AGENT_API_KEY` | unset | Routed onto the provider's key env var(s) via the shared contract; omitted for keyless backends (Vertex/Bedrock ADC). | | ||
| | `AGENT_TARGET` | unset | Path to the CLI binary (`gemini` / `oc`). Ignored by `api`. | | ||
| | `AGENT_TIMEOUT_SEC` | `600` | Wall-clock budget for each external call. | | ||
| | `AGENT_MAX_TURNS` | harness default (50 for `api`) | Caps the `api` tool-use loop. | | ||
|
|
||
| **Capabilities** | ||
|
|
||
| | Variable | Default | Notes | | ||
| | --- | --- | --- | | ||
| | `BENCH_USE_MCP` | `true` | Master gate. `false` drops the MCP binding entirely. | | ||
| | `AGENT_MCP_SERVER` | unset | Shell-quoted argv for the MCP server (e.g. `"uv run k8s-mcp"`). | | ||
| | `AGENT_ALLOWED_TOOLS` | unset | CSV of pre-approved tool names. | | ||
| | `AGENT_SKILLS_PATHS` | unset | CSV of directories to discover `SKILL.md` files under. | | ||
| | `AGENT_RULES_TEXT` | unset | Operator-brief text handed to the agent. | | ||
|
|
||
| ### Example: gemini CLI with MCP + skills | ||
|
|
||
| ```bash | ||
| export BENCH_AGENT_TYPE=gemini | ||
| export AGENT_PROVIDER=gemini | ||
| export AGENT_MODEL=gemini-2.5-pro | ||
| export AGENT_API_KEY="$GEMINI_API_KEY" | ||
| export AGENT_TARGET=gemini | ||
|
|
||
| export BENCH_USE_MCP=true | ||
| export AGENT_MCP_SERVER="uv run k8s-mcp" | ||
| export AGENT_ALLOWED_TOOLS="list_clusters,get_pods" | ||
| export AGENT_SKILLS_PATHS="/opt/skills/devops,/opt/skills/k8s" | ||
| ``` | ||
|
|
||
| ### Example: api harness on Claude with MCP off | ||
|
|
||
| ```bash | ||
| export BENCH_AGENT_TYPE=api | ||
| export AGENT_PROVIDER=anthropic | ||
| export AGENT_MODEL=claude-sonnet-4-5 | ||
| export AGENT_API_KEY="$ANTHROPIC_API_KEY" | ||
|
|
||
| export BENCH_USE_MCP=false # no MCP server is spawned; tools are dropped | ||
| ``` | ||
|
|
||
| ## Capabilities | ||
|
|
||
| MCP tools, skills, and rules are the three augmentation axes, and they are | ||
| independent — an agent may run with any combination, or none. Each is expressed | ||
| as a structural Protocol (`SupportsMcp`, `SupportsSkills`, `SupportsRules` in | ||
| `devops_bench/agents/capabilities/`): a harness satisfies a Protocol simply by | ||
| assigning the matching binding attribute. **MCP** wires the agent to a tool | ||
| server, **skills** drop `SKILL.md` files the agent can discover, and **rules** | ||
| supply an operator brief. Setting `BENCH_USE_MCP=false` drops the MCP binding | ||
| entirely, so the agent sees no tools and the scorer agrees that none ran — skills | ||
| and rules are unaffected. | ||
|
|
||
| ## Adding your own harness | ||
|
|
||
| Want to wrap a different agent? See | ||
| [Add an agent harness](../how-to/add-an-agent-harness.md). | ||
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,161 @@ | ||
| # Add an agent harness | ||
|
|
||
| This guide walks through wrapping a new agent so the benchmark can drive it. The | ||
| contract is small: subclass `AgentHarness`, implement `_execute`, register the | ||
| class with `@AGENTS.register`, and add your module to the built-in import list. | ||
| That's it — no `cli.py` or `run.py` edits. | ||
|
|
||
| For the concepts (harness vs model, capabilities, configuration), read | ||
| [Agents](../components/agents.md) first. | ||
|
|
||
| ## The contract | ||
|
|
||
| | You do | Where | | ||
| | --- | --- | | ||
| | Subclass `AgentHarness` | `devops_bench/agents/base.py` | | ||
| | Implement `_execute(self, prompt, workspace_path=None) -> AgentResult` | your new module | | ||
| | Register with `@AGENTS.register("<key>")` | your new module | | ||
| | Add the module to `_BUILTIN_AGENT_MODULES` | `devops_bench/evalharness/default.py` | | ||
|
|
||
| ## Steps | ||
|
|
||
| ### 1. Create the module | ||
|
|
||
| Mirror an existing harness. For a CLI-backed agent, follow `gemini_cli` / | ||
| `openclaw`: | ||
|
|
||
| ```text | ||
| devops_bench/agents/cli/<name>/agent.py | ||
| ``` | ||
|
|
||
| For an in-process agent, follow `api`: | ||
|
|
||
| ```text | ||
| devops_bench/agents/<name>/agent.py | ||
| ``` | ||
|
|
||
| ### 2. Subclass `AgentHarness` and assign capability bindings | ||
|
|
||
| Call the base `__init__` with your config, then assign `self.mcp_servers`, | ||
| `self.skills`, and `self.rules` from `self.config.capabilities`. Those three | ||
| assignments are what make your harness structurally satisfy the capability | ||
| Protocols (`SupportsMcp` / `SupportsSkills` / `SupportsRules`) — no mixin needed. | ||
|
|
||
| ### 3. Implement only `_execute` | ||
|
|
||
| `_execute(self, prompt: str, workspace_path: Path | None = None) -> AgentResult` | ||
| is the single extension point. `run()` calls it positionally, so the second | ||
| parameter is required even if your harness ignores it. | ||
| Inside it: | ||
|
|
||
| - Build the invocation for your agent (argv, an API call, whatever it takes). | ||
| - Parse the agent's output into canonical `ToolCall` entries | ||
| (`devops_bench/agents/result.py`) for the trajectory. | ||
| - On a *known* failure (subprocess error, parse miss, timeout), record a message | ||
| on `AgentResult.errors` rather than dropping it silently. For a hard failure | ||
| with no usable output, return `AgentResult.errored(msg)`. | ||
| - Return an `AgentResult`. Leave `latency` at zero — the base `run()` fills it in. | ||
|
|
||
| > [!NOTE] | ||
| > Only handle your *known* errors. The base class already catches unexpected | ||
| > exceptions and converts them to an errored result, so you don't need a | ||
| > catch-all. | ||
|
|
||
| ### 4. Register the class | ||
|
|
||
| Decorate it with its canonical key: | ||
|
|
||
| ```python | ||
| @AGENTS.register("<key>") | ||
| class MyAgent(AgentHarness): | ||
| ... | ||
| ``` | ||
|
|
||
| ### 5. Wire it for import side-effects | ||
|
|
||
| Registration only fires when the module is imported, so add its path to | ||
| `_BUILTIN_AGENT_MODULES` in `devops_bench/evalharness/default.py`: | ||
|
|
||
| ```python | ||
| _BUILTIN_AGENT_MODULES: tuple[str, ...] = ( | ||
| "devops_bench.agents.cli.gemini_cli", | ||
| "devops_bench.agents.cli.openclaw", | ||
| "devops_bench.agents.api.agent", | ||
| "devops_bench.agents.<name>.agent", # <- your module | ||
| ) | ||
| ``` | ||
|
|
||
| The import loop tolerates `ImportError` / `MissingDependencyError`, so a harness | ||
| that needs an optional SDK won't break the host that lacks it. If you want a | ||
| friendlier selector name, add an entry to `_AGENT_TYPE_ALIASES` in the same file — | ||
| for example, mapping `gemini-cli` to `gemini`. | ||
|
|
||
| ### 6. Reuse the shared CLI helpers | ||
|
|
||
| For a CLI agent, don't re-implement capability plumbing. Reuse the helpers in | ||
| `devops_bench/agents/shared/cli_capabilities.py`: | ||
|
|
||
| - `build_mcp_servers(...)` — turns granted MCP bindings into a `{name: {command, args}}` launch map. | ||
| - `materialize_skills(...)` — copies discovered `SKILL.md` files into a skills directory and returns its path. | ||
|
|
||
| > [!IMPORTANT] | ||
| > These helpers stage the files, but they don't tell your agent where to find | ||
| > them. Your `_execute` is responsible for pointing the underlying tool at the | ||
| > staged locations — whether that's a CLI flag, a config file, or an environment | ||
| > variable (e.g. the Gemini CLI agent writes the MCP launch map into its settings | ||
| > and the openclaw agent exports its skills dir). Wire the path/env through in | ||
| > your harness, or the staged MCP servers and skills won't be picked up. | ||
|
|
||
| ### 7. Select it | ||
|
|
||
| Pick your harness with `BENCH_AGENT_TYPE=<key>` (or `--agent-type <key>`). No | ||
| other code changes are required — the registry resolves it at run time. | ||
|
|
||
| ## Skeleton | ||
|
|
||
| ```python | ||
| from pathlib import Path | ||
|
|
||
| from devops_bench.agents.base import AGENTS, AgentHarness | ||
| from devops_bench.agents.config import AgentConfig | ||
| from devops_bench.agents.result import AgentResult, ToolCall | ||
|
|
||
|
|
||
| @AGENTS.register("myagent") | ||
| class MyAgent(AgentHarness): | ||
| """Harness driving <the agent you wrap>.""" | ||
|
|
||
| def __init__(self, config: AgentConfig | None = None) -> None: | ||
| AgentHarness.__init__(self, config) | ||
| caps = self.config.capabilities | ||
| self.mcp_servers = caps.mcp_servers | ||
| self.skills = caps.skills | ||
| self.rules = caps.rules | ||
|
|
||
| def _execute(self, prompt: str, workspace_path: Path | None = None) -> AgentResult: | ||
|
coderabbitai[bot] marked this conversation as resolved.
|
||
| # 1. Build and run the invocation for `prompt`. | ||
| # 2. Parse output into canonical ToolCall entries. | ||
| trajectory: list[dict] = [ | ||
| ToolCall(name="example_tool", args={}).to_dict(), | ||
| ] | ||
| # 3. On a known failure, return AgentResult.errored("..."). | ||
| # 4. Return the result (leave latency at zero; the base stamps it). | ||
| return AgentResult(output="...", trajectory=trajectory) | ||
| ``` | ||
|
|
||
| ## Test it | ||
|
|
||
| Run a no-infra task with your harness selected. The `noop` deployer skips cluster | ||
| provisioning so you can confirm the harness drives the agent and returns a | ||
| trajectory end-to-end without standing up infrastructure: | ||
|
|
||
| ```bash | ||
| export BENCH_AGENT_TYPE=myagent | ||
| export BENCH_NO_INFRA=true | ||
| export AGENT_PROVIDER=myprovider | ||
| export AGENT_MODEL=mymodel | ||
| # run a single generation-only task and inspect results.json | ||
| ``` | ||
|
|
||
| Check the run's `results.json`: a clean run shows your parsed `trajectory`, a | ||
| populated `output`, and an empty `errors` list. | ||
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
Uh oh!
There was an error while loading. Please reload this page.