Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/).

## [Unreleased]
### Added
- Added the task-source adapter foundation (`src/clawbench/adapters/`): a shared `ClawBenchTask` type, an adapter registry with declared scoring layers and field-mapping warnings, an identity adapter for the bundled corpora, and a `clawbench-sources` CLI to list and inspect them. No change to how bundled tasks run. See [`docs/task-sources.md`](docs/task-sources.md).
- Added `scripts/export_openeval.py`, an additive script exporting a batch's `rescore-summary.json` as an [EvalPort](https://github.com/adhabnr-ux/evalport) `ResultSet` Thanks to [@adhabnr-ux](https://github.com/adhabnr-ux).
- Added a `--browser-runtime kernel` mode to the Harbor adapter that runs each task against one Kernel cloud browser, exposing only a credential-free CDP bridge to the agent, and finalizes the replay and deletes the browser during verification.

Expand Down
1 change: 1 addition & 0 deletions docs/cli.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,7 @@ Every ClawBench command. From a PyPI install run them directly (`clawbench-run
| `clawbench-rescore` | Re-judge trajectories you already have, without re-running agents. |
| `clawbench-reproduce` | Download published traces for one leaderboard row and check you reproduce it. |
| `clawbench-harbor-adapt` | Convert V2 into a Harbor dataset — see [`harbor.md`](harbor.md). |
| `clawbench-sources` | List and inspect task-source adapters — see [`task-sources.md`](task-sources.md). |
| `clawbench-edgebench-adapt`, `clawbench-edgebench-judge` | EdgeBench/SForge export — see [`edgebench.md`](edgebench.md). |

`./run.sh` from a source checkout is a shortcut for the TUI.
Expand Down
99 changes: 99 additions & 0 deletions docs/task-sources.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,99 @@
# Task sources

Every browser-agent benchmark encodes its tasks slightly differently — WebArena-style JSON, Mind2Web step traces, WebVoyager judge prompts. Task-source adapters convert those definitions into ClawBench's own task type so an external corpus can run through ClawBench's submission interception, five-layer recording, and judge pipeline without hand-converting files or forking the upstream repo.

Adapters are **import-only**: nothing writes back to an upstream format.

This page describes the foundation that is in place today — the shared task type, the registry, and the `clawbench-sources` CLI. Individual benchmark adapters land incrementally; see [issue #72](https://github.com/TIGER-AI-Lab/ClawBench/issues/72) for the sequence.

## Listing what is registered

```bash
uv run clawbench-sources
uv run clawbench-sources --json # same rows, machine-readable
```

```
SOURCE STATE UPSTREAM PIN PATH
clawbench-native bundled - - <repo>/test-cases
```

`STATE` is `bundled` when the tasks ship with ClawBench, `cached` when an external checkout is present, and `missing` when it still needs fetching.

```bash
uv run clawbench-sources show clawbench-native # status + field-mapping table
uv run clawbench-sources cases clawbench-native # every task the source exposes
uv run clawbench-sources cases clawbench-native --path test-cases/v2
```

A source can also be addressed as `<name>:<path>` to pin it to an explicit clone:

```bash
uv run clawbench-sources cases claw-eval:/srv/checkouts/claw-eval
```

Without a path, a source resolves under `$CLAWBENCH_SOURCES_DIR`, else `$XDG_CACHE_HOME/clawbench/sources`, else `~/.cache/clawbench/sources`. Set `CLAWBENCH_OFFLINE=1` to forbid network fetches; a source with no local checkout then fails loudly instead of cloning.

## The shared task type

Adapters produce `ClawBenchTask` (`src/clawbench/adapters/schema.py`), a superset of `test-cases/task.schema.json` plus provenance:

| Field | Meaning |
|---|---|
| `task_id` | ClawBench's identifier for the task |
| `source` | registered adapter name |
| `source_id` | the upstream benchmark's own identifier |
| `instruction` | prompt sent to the agent |
| `time_limit` | **minutes**, matching `task.json` and the container watchdog |
| `eval_schema` | interceptor config, when the source has a submission contract |
| `scoring_layers` | which scoring mechanisms this source can honour |
| `extra_info` / `judge_context` / `metadata` | carried through as in `task.json` |
| `warnings` | field-mapping gaps found at load time |

Most upstream schemas express time limits in seconds; adapters convert.

## Scoring layers

An adapter declares which layers its tasks can be judged by:

| Layer | Applies when |
|---|---|
| `submission_intercept` | the task has a final write request to intercept |
| `end_state_dom_match` | ClawBench's default judge pipeline applies |
| `step_trace_replay` | the upstream rubric is per-step (Mind2Web-style) |
| `goal_predicate` | the upstream rubric is a boolean goal function (WorkArena/BrowserGym) |
| `llm_judge_only` | the upstream rubric is a free-form judge prompt (WebVoyager) |

A layer a source cannot support scores `null` in the recording — never `0` — so leaderboard aggregation never confuses "the agent failed" with "this task was never scored on that axis". `ClawBenchTask.to_task_json()` refuses to render a native `task.json` for a task with no interception contract, rather than inventing one.

## Field-mapping warnings

When an adapter cannot map a field 1:1 it attaches an `AdapterWarning` to the task instead of dropping it silently. Each warning names the source, the task, the field, the fallback used, and the pinned upstream revision the mapping was written against:

```
[mind2web/t1] time_limit: upstream has no per-task limit (using 300s) [upstream abc1234]
```

Adapters pin an upstream commit or tag so a rename upstream cannot quietly change what a run measures.

## Writing an adapter

Subclass `AdapterBase`, declare the metadata, and register it:

```python
from clawbench.adapters import AdapterBase, ScoringLayer, register

@register
class MyBenchmarkAdapter(AdapterBase):
name = "my-benchmark"
upstream = "https://github.com/example/my-benchmark"
pinned_sha = "abc1234"
scoring_layers = (ScoringLayer.LLM_JUDGE_ONLY,)

def load(self, path):
... # -> list[ClawBenchTask]
```

Document the field mapping as a table in the module docstring — `clawbench-sources show <name>` prints it. `native.py` is the reference implementation.

Related: [`docs/cli.md`](cli.md) · [`docs/harbor.md`](harbor.md) · [`CONTRIBUTING.md`](../CONTRIBUTING.md)
1 change: 1 addition & 0 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -27,6 +27,7 @@ clawbench = "clawbench.tui:main"
clawbench-run = "clawbench.runner.run:main"
clawbench-batch = "clawbench.runner.batch:main"
clawbench-rescore = "clawbench.eval.rescore:main"
clawbench-sources = "clawbench.adapters.cli:main"
clawbench-analyze = "clawbench.eval.analyze:main"
clawbench-reproduce = "clawbench.eval.reproduce:main"
clawbench-harbor-adapt = "clawbench.eval.harbor_adapter:main"
Expand Down
52 changes: 52 additions & 0 deletions src/clawbench/adapters/__init__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,52 @@
"""Task-source adapters: run other benchmarks' tasks under ClawBench.

Every adapter converts one external benchmark's task definitions into
:class:`~clawbench.adapters.schema.ClawBenchTask`, so a research team can reuse
ClawBench's submission interception, five-layer recording, and judge pipeline
without hand-converting task files or forking the upstream repo.

Adapters are import-only and pin the upstream revision they were written
against. Fields with no 1:1 mapping surface as
:class:`~clawbench.adapters.schema.AdapterWarning` at load time; scoring layers
a source cannot support score ``null`` rather than 0, so "not scored" is never
mistaken for "failed".

``clawbench-sources`` lists what is registered. See ``docs/task-sources.md``.
"""

from clawbench.adapters._base import (
AdapterBase,
AdapterError,
SourceStatus,
get_adapter,
offline,
parse_source_spec,
register,
registered_sources,
source_cache_dir,
)
from clawbench.adapters.schema import (
AdapterWarning,
ClawBenchTask,
ExtraInfo,
ScoringLayer,
)

# Importing an adapter module is what registers it.
from . import native # noqa: F401 isort:skip

__all__ = [
"AdapterBase",
"AdapterError",
"AdapterWarning",
"ClawBenchTask",
"ExtraInfo",
"ScoringLayer",
"SourceStatus",
"get_adapter",
"offline",
"parse_source_spec",
"register",
"registered_sources",
"source_cache_dir",
]
154 changes: 154 additions & 0 deletions src/clawbench/adapters/_base.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,154 @@
"""Adapter base class and the source registry.

An adapter converts one external benchmark's task definitions into
:class:`~clawbench.adapters.schema.ClawBenchTask` values. It is import-only:
nothing here writes back to an upstream format.

Each adapter subclasses :class:`AdapterBase`, declares which scoring layers it
can honour, pins the upstream revision it was written against, and documents
its field mapping in its module docstring. Registration is by decorator:

@register
class MyAdapter(AdapterBase):
name = "my-benchmark"
...
"""

from __future__ import annotations

import os
from abc import ABC, abstractmethod
from dataclasses import dataclass
from pathlib import Path

from clawbench.adapters.schema import ClawBenchTask, ScoringLayer


class AdapterError(RuntimeError):
"""A source could not be loaded at all."""


@dataclass(frozen=True)
class SourceStatus:
"""What ``clawbench-sources list`` prints for one registered adapter."""

name: str
upstream: str | None
pinned_sha: str | None
scoring_layers: tuple[ScoringLayer, ...]
cache_dir: Path
cached: bool
bundled: bool


class AdapterBase(ABC):
"""Base class for every task-source adapter."""

#: Registry key, also the value accepted by ``--source``.
name: str = ""
#: Upstream repository this adapter reads, or ``None`` when the tasks ship
#: with ClawBench itself.
upstream: str | None = None
#: Upstream commit/tag the field mapping was written against. Pinning keeps
#: an upstream rename from silently changing what a run measures.
pinned_sha: str | None = None
#: Scoring layers this source's tasks can actually be judged by. Layers not
#: listed here score ``null``, never 0.
scoring_layers: tuple[ScoringLayer, ...] = ()

@property
def bundled(self) -> bool:
"""True when the source needs no external checkout."""
return self.upstream is None

@abstractmethod
def load(self, path: Path) -> list[ClawBenchTask]:
"""Convert every task under ``path`` into ClawBench tasks.

Implementations raise :class:`AdapterError` when ``path`` is not a
checkout of this source, and attach an
:class:`~clawbench.adapters.schema.AdapterWarning` to a task for each
field they could not map, rather than dropping the task silently.
"""

def default_path(self) -> Path:
"""Where this source is expected to live when ``--source`` gets no path."""
return source_cache_dir() / self.name

def status(self, path: Path | None = None) -> SourceStatus:
resolved = path or self.default_path()
return SourceStatus(
name=self.name,
upstream=self.upstream,
pinned_sha=self.pinned_sha,
scoring_layers=self.scoring_layers,
cache_dir=resolved,
cached=resolved.is_dir(),
bundled=self.bundled,
)


_REGISTRY: dict[str, AdapterBase] = {}


def register(adapter_cls: type[AdapterBase]) -> type[AdapterBase]:
"""Register an adapter class under its ``name``."""
name = adapter_cls.name
if not name:
raise ValueError(f"{adapter_cls.__name__} must define a non-empty name")
if name in _REGISTRY:
raise ValueError(f"duplicate adapter name: {name}")
_REGISTRY[name] = adapter_cls()
return adapter_cls


def registered_sources() -> tuple[str, ...]:
"""Every registered source name, in stable alphabetical order."""
return tuple(sorted(_REGISTRY))


def get_adapter(name: str) -> AdapterBase:
try:
return _REGISTRY[name]
except KeyError:
known = ", ".join(registered_sources()) or "(none)"
raise AdapterError(
f"unknown task source {name!r}; registered sources: {known}"
) from None


def source_cache_dir() -> Path:
"""Root for lazily fetched source checkouts.

Honours ``CLAWBENCH_SOURCES_DIR``, then ``XDG_CACHE_HOME``, then
``~/.cache``, so a shared machine can point several workspaces at one
checkout without re-cloning.
"""
if raw := os.environ.get("CLAWBENCH_SOURCES_DIR"):
return Path(raw).expanduser()
if raw := os.environ.get("XDG_CACHE_HOME"):
return Path(raw).expanduser() / "clawbench" / "sources"
return Path.home() / ".cache" / "clawbench" / "sources"


def parse_source_spec(spec: str) -> tuple[str, Path | None]:
"""Split ``--source`` into a registered name and an optional explicit path.

``"claw-eval"`` resolves to the adapter's default checkout location;
``"claw-eval:/path/to/repo"`` pins it to an explicit clone. Windows drive
letters are not mistaken for the separator.
"""
name, sep, raw_path = spec.partition(":")
if not sep or len(name) <= 1:
return spec, None
return name, Path(raw_path).expanduser()


def offline() -> bool:
"""True when ``CLAWBENCH_OFFLINE`` forbids network fetches."""
return os.environ.get("CLAWBENCH_OFFLINE", "").strip().lower() not in (
"",
"0",
"false",
"no",
)
Loading
Loading