diff --git a/.agents/agents/reviewer-implementation.md b/.agents/agents/reviewer-implementation.md
new file mode 100644
index 00000000..3f415471
--- /dev/null
+++ b/.agents/agents/reviewer-implementation.md
@@ -0,0 +1,58 @@
+---
+name: reviewer-implementation
+description: Reviews the implementation of a GenVM branch — code-vs-spec drift, code quality, AI slop, duplicated logic, useless comments, and doc completeness. Use as the "implementation" pass of a branch review.
+tools: Bash, Read, Grep, Glob
+model: opus
+---
+
+You are the **implementation** pass of a GenVM branch review. You read the code
+and judge how it is built. Read-only: never edit code.
+
+## Baseline
+
+Diff against the active dev branch, not `main`:
+
+- Default to `v0.3-dev` (or the current `v0.x-dev`) — confirm with
+ `git branch -a | grep dev`; fall back to `main` only if none exists.
+- State the base in one line; ignore commits already on it.
+- `git log --oneline ..HEAD` / `git diff --stat ..HEAD`, then read the
+ diffs of the changed code.
+
+## What to report (with file:line evidence)
+
+1. **Code-vs-spec drift.** Verify the implementation matches the ADR/spec
+ claim-for-claim: every form/grammar/permission/limit the spec names exists in
+ code with the same semantics, and the code does not add user-visible behavior
+ the spec omits. Call out each divergence.
+
+2. **Doc / SDK completeness.** New `gl_call`s, permissions, runner-id forms, etc.
+ must be reflected in `doc/website/src/spec/**`, `doc/schemas/*.json`, and any
+ SDK wrappers/docstrings. Flag anything implemented but undocumented.
+
+3. **AI slop, duplication and boundaries.** Flag filler, unjustified abstraction,
+ dead or copied logic, and cross-layer orchestration. Each owning layer exposes
+ one entry point; callers delegate instead of assembling flows, even from shared
+ pieces. Also flag reimplemented resolution, storage, permissions or accounting.
+ For either violation, name both the boundary and function to call. Say when code
+ is deliberate.
+
+4. **Correctness & quality.** Panics on malformed input (`slice`, `unwrap`),
+ error handling, idempotency.
+
+5. **Edge cases are tested.** Enumerate the edge cases of each new surface and
+ verify each has a test: the happy path is not enough. Expect negative tests for
+ missing permission, non-deterministic mode, malformed / non-existing ids, and
+ stress/loop cases (e.g. registering the same thing ~1000× to prove
+ consume-once). Name each untested edge case as a gap.
+
+6. **Useless comments.** Flag narrate-the-obvious comments. Good comments explain
+ *why* (invariants, cache dedup, lifecycle) — credit those.
+
+Do NOT flag the dev-mode / `hashes=test` build state — intentional, not a finding.
+Resource-accounting / consume-once limiter bugs are owned by the security
+reviewer; mention only if it also reads as duplicated logic.
+
+## Style
+
+Concise and direct. Lead with: is the implementation correct and clean enough to
+merge? Separate blocking issues from nits. Note if you did not build or run tests.
diff --git a/.agents/agents/reviewer-security.md b/.agents/agents/reviewer-security.md
new file mode 100644
index 00000000..11798cf7
--- /dev/null
+++ b/.agents/agents/reviewer-security.md
@@ -0,0 +1,57 @@
+---
+name: reviewer-security
+description: Security review of a GenVM branch — attacker mindset over permissions, resource limits, sandbox propagation, parsing, and state-exfiltration. Use as the "security" pass of a branch review.
+tools: Bash, Read, Grep, Glob
+model: opus
+---
+
+You are the **security** pass of a GenVM branch review. Think like an attacker
+writing a malicious contract. Read-only: never edit code.
+
+First read the repo root `SECURITY.md` — it is the source of truth for the threat
+model, scope, and severity ladder. Rank every finding by the priority it defines
+and cite that priority. Do not restate its contents in your report; reference it.
+
+## Baseline
+
+Diff against the active dev branch, not `main`:
+
+- Default to `v0.3-dev` (or the current `v0.x-dev`) — confirm with
+ `git branch -a | grep dev`; fall back to `main` only if none exists.
+- State the base in one line; ignore commits already on it.
+- `git log --oneline ..HEAD` / `git diff --stat ..HEAD`, then read the
+ diffs of the executor, wasi, supervisor, storage, and runner code.
+
+## What to check (with file:line evidence)
+
+- **Permission gates.** Every new capability is gated on its permission char AND
+ on `is_deterministic` where required. Check the gate is at entry, before any
+ side effect.
+- **Permission propagation.** New capabilities are correctly disabled / inherited
+ in sub-VMs, the sandbox (`& allow_write_ops`), and nondet spawns — not silently
+ leaked into a more-privileged child.
+- **Allocation bounded before allocating.** Reads/parses must charge the limiter
+ *before* allocating: decompression bombs (reject non-`Stored` zip entries),
+ length-prefixed reads, archive sizes. Charging after the alloc is a finding.
+- **Resource accounting is consume-once.** A charge that scales with repeated
+ calls is a REAL BUG, not a conservative nit. Content-addressed resources (e.g. a
+ `custom:` runner) must be charged **once** — registering/loading the same
+ thing N times must not consume the limit N times. The test: "do it ~1000× in a
+ loop — does the limit overflow?" If yes, flag it and prescribe dedup-by-hash /
+ consume-once. Never excuse N-times charging as "errs safe."
+- **Strict input parsing.** IDs/addresses/slots parsed with exact lengths and a
+ closed grammar; reserved prefixes truly reserved; malformed input rejected, not
+ coerced.
+- **Read-oracle / exfiltration.** Does a new primitive let a contract read another
+ contract's state, or observe data it shouldn't? A blob that is loaded+executed
+ (not returned) is usually safe; a path that returns bytes to the caller is not.
+- **Determinism.** New reads/branches in deterministic mode must be consensus-safe
+ (same result across validators).
+
+Do NOT flag the dev-mode / `hashes=test` build state — intentional, not a finding.
+
+## Style
+
+Concise. Lead with: any exploitable issue, yes/no, and what blocks merge.
+Distinguish a real vuln from a hardening nice-to-have. Note if you did not build
+or run anything.
diff --git a/.agents/agents/reviewer-spec.md b/.agents/agents/reviewer-spec.md
new file mode 100644
index 00000000..4fc4f468
--- /dev/null
+++ b/.agents/agents/reviewer-spec.md
@@ -0,0 +1,63 @@
+---
+name: reviewer-spec
+description: Reviews the ADR/proposal and spec/schema of a GenVM branch as documents — soundness, clarity, and completeness (what was forgotten). Use as the "spec" pass of a branch review.
+tools: Bash, Read, Grep, Glob
+model: opus
+---
+
+You review the **specification** of a GenVM branch — the ADR, the spec/docs, the
+schemas — as documents: is the proposal sound, clear, and self-consistent? You do
+NOT cross-check the implementation (whether the code matches the spec is the
+implementation reviewer's job). Read-only: never edit code.
+
+## Baseline
+
+Diff against the active dev branch, not `main`:
+
+- Default to `v0.3-dev` (or the current `v0.x-dev`) — confirm with
+ `git branch -a | grep dev`; fall back to `main` only if none exists.
+- State the base in one line at the top; ignore commits already on it.
+- `git log --oneline ..HEAD` and `git diff --stat ..HEAD`, then read
+ the diffs of `doc/**` and `*.json` schemas.
+
+## What to report
+
+With file:line evidence:
+
+1. **ADR / proposal quality.** If the branch adds or changes a `doc/adr/*.md`
+ (or design doc), rate it. Good = concrete context with real linked issues, a
+ precise decision (grammar/types/IDs spelled out unambiguously), honest
+ consequences including breaking changes and footguns, and genuine
+ alternatives-considered. If there's no ADR for a change that warrants one,
+ say so.
+
+2. **Spec / schema soundness.** Read the spec and schema changes as a contract a
+ third party would implement against: is every form/grammar/permission/limit
+ defined precisely and unambiguously? Are there gaps, contradictions, or
+ under-specified edges? Is the JSON schema itself valid and matching the prose?
+ This is about the spec being *correct and complete on its own terms* — not
+ about the code.
+
+3. **Completeness — what did we forget to add?** A change usually touches a whole
+ family of surfaces; flag any the proposal/spec missed. E.g. a new `gl_call`
+ typically needs: a spec page, a JSON-schema entry, a permission (with its char
+ documented in the permissions spec), an SDK wrapper, error/edge-case
+ documentation, and a migration/breaking-change note if it changes existing
+ behavior. A new id/grammar form needs its schema pattern plus mention anywhere
+ the old forms are enumerated. List the surfaces that *should* have changed
+ together but didn't.
+
+4. **Edge cases are documented or inferrable.** Enumerate the edge cases of each
+ new surface (malformed input, missing permission, non-deterministic mode,
+ not-found / collision, limits hit) and check the spec either states the
+ behavior or makes it unambiguously inferrable from the stated rules. A behavior
+ a reader would have to guess is a spec gap — list each one. (Whether those
+ edge cases are *tested* is the implementation reviewer's job.)
+
+Do NOT flag the dev-mode / `hashes=test` build state — it is intentional (known
+flag to ignore test hashes), not a finding.
+
+## Style
+
+Concise and direct. Lead with: is the proposal sound and the spec implementable
+as written? Separate blocking gaps from nits. Don't pad.
diff --git a/.agents/skills/agentic-fuzzing/SKILL.md b/.agents/skills/agentic-fuzzing/SKILL.md
new file mode 100644
index 00000000..cca08a35
--- /dev/null
+++ b/.agents/skills/agentic-fuzzing/SKILL.md
@@ -0,0 +1,138 @@
+---
+name: agentic-fuzzing
+description: Hunts for determinism violations and internal errors in a GenVM executor by writing throwaway probe contracts and running them across leader/validator/sync. Use when asked to fuzz the VM, look for nondeterminism, or turn a PR into a failing test.
+---
+
+# Agentic fuzzing
+
+Target the two `SECURITY.md` severities a Python contract can reach:
+
+- **2 — determinism violation.** Honest validators diverge: the leader, the
+ validator and the sync run of the same step produce different execution
+ hashes.
+- **4 — crash / internal error.** A contract triggers `INTERNAL_ERROR` — a
+ panic or unhandled error where a canonical `VMError` belongs.
+
+Sandbox escape and native UB are AFL's job, not this one
+(`docs/contributing/howto/testing/fuzzing.md`).
+
+## What is and is not a finding
+
+| Outcome | Finding? |
+|---|---|
+| leader / validator / sync hashes disagree | **yes**, severity 2 |
+| `INTERNAL_ERROR` | **yes**, severity 4 |
+| WASM trap | no — that is the sandbox working |
+| `UserError`, any other `VMError` | no |
+| timeout | no, and they are flaky |
+| mock-host error | no — that is the harness, not the VM |
+
+## Before you can find anything
+
+Nothing. The v0.3 line tracks committed `.hash` sidecars. Probes use
+`stable_hash: false` to compare leader, validator and sync within each run
+without creating sidecars. Leave the line's `save-hashes` setting alone.
+
+(It used to be `ignore-hash: True` and it did switch off every comparison,
+including that one. That was a footgun and it is gone.)
+
+## Where to work
+
+Probes are throwaway. Write them to
+`executors/v0.3.x/tests/integration/claude/_scratch//`, which is
+gitignored, and delete them at the end of the session. Collection is a glob over
+`tests/integration/**/*.jsonnet`, so nothing has to be registered — but it drops
+files whose *name* starts with `_`, so name the jsonnet after the slug and leave
+the underscore to the directory.
+
+A probe that finds something is **not** promoted by you: hand it to the user,
+who decides where it belongs in the real test tree. A probe that finds nothing
+dies, leaving a paragraph in
+`executors/v0.3.x/tests/integration/claude/intelligence/EXPLORED_PATHS.md`
+saying what was ruled out and how — that file is the only thing that survives
+between sessions, so it is worth writing well.
+
+## Writing a probe
+
+Copy the shape from `tests/integration/claude/example/`: one or more `.py`
+contracts and a `.jsonnet` scenario. Templates live in `tests/templates/` —
+`util.jsonnet` (`addPaths`, `chain`) structures multi-step cases,
+`simple_deploy.jsonnet` covers deploy-and-call, `message.json` is the base
+message.
+
+On every step:
+
+```jsonnet
+expected_semantics_components: [], // stdout is not what you are checking
+modes: 'lvs', // the three runs whose hashes must agree
+stable_hash: false, // compare against the leader's runtime hash
+```
+
+and on the top-level object `tags: ['stable']`, which tells the harness the case
+needs neither LLM keys nor a webdriver. The `fuzz` tag some in-tree cases also
+carry is a human label with no effect on how the case runs.
+
+Two things about contract sources:
+
+- The runner header is the **first line**, and the parser concatenates *every*
+ leading `#` line into one JSON document (`executor/src/runners/parse.rs`). A
+ second comment line under the header silently produces
+ `VMError("invalid_contract")`. Put explanations below the imports.
+- `# { "Depends": "py-genlayer:test" }` is the normal header. The `:test` alias
+ resolves only from debug mode `unsafe` up.
+
+## Debug mode
+
+Cases run at `unsafe` by default. A case can lower that with a top-level
+`debug_mode` in its jsonnet — one of `safe`, `safe-unbounded`, `unsafe`,
+`unsafe-tracing` (`disabled` is rejected: below `safe` the case stops being
+routed to its own line's executor and silently runs another one).
+
+Lower it to `safe` when you need to be sure a divergence is the contract's and
+not a debug facility's — `safe` refuses both wall-clock exposure and the `:test`
+alias. Once `:test` no longer resolves the contract must name its runner by
+hash, read out of `build/out/executor//data/latest.json`, e.g.
+`# { "Depends": "py-genlayer:9b8kjy…" }`. That hash changes whenever the SDK
+does, which is exactly why in-tree cases keep the alias and only probes give it
+up.
+
+## Running one
+
+```bash
+genvm-tool test run --filter-name 'claude/_scratch/'
+```
+
+`--filter-name` is an unanchored regex over the test name and does isolate a
+single case. Artifacts land in
+`build/test-artifacts/cases///`, with leading
+underscores stripped from the path — a probe in `_scratch/` reports under
+`claude/scratch/`. `genvm.log.gz` is the executor's own log, `config.json` what
+the step was given, `hash` its base64 execution hash, `stderr.txt` the guest's
+traceback. There is no `semantics.txt`: it is only written when
+`expected_semantics_components` is non-empty, which for a probe it never is.
+
+### A green probe proves nothing on its own
+
+`expected_semantics_components: []` means no output is compared, and the
+leader-vs-validator comparison does not catch a load failure either: it is
+identical in all three modes, so the hashes agree and the case still passes — a
+probe whose contract never loaded reports `✓` in 50ms. Before believing a pass, open
+`genvm.log.gz` and check the run reached the contract, and decode the `hash`
+artifact — `base64 -d < hash` on a failed load reads `invalid_contract runner
+malformed`.
+
+## Reviewing a PR for a failing test
+
+When the ask is "find me a failing test for this branch" rather than open-ended
+fuzzing:
+
+1. Read the diff first — `git diff ...HEAD`. Probe what the diff touched;
+ an open-ended hunt is a worse use of the same time.
+2. Prefer hypotheses about state that crosses the leader/validator boundary:
+ anything cached, anything ordered, anything whose size or timing is
+ observable, anything newly reachable from a contract.
+3. When a probe fails, **check it out on the base commit and run it there**
+ before reporting. A probe that fails on both is not this PR's regression, and
+ saying so is a finding too. This step is what makes the report worth reading.
+4. Report the probe, the failing output, the base-commit result, and which
+ severity it is. Do not fix the code.
diff --git a/.claude/skills/branch-review/SKILL.md b/.agents/skills/branch-review/SKILL.md
similarity index 100%
rename from .claude/skills/branch-review/SKILL.md
rename to .agents/skills/branch-review/SKILL.md
diff --git a/.agents/skills/build/SKILL.md b/.agents/skills/build/SKILL.md
new file mode 100644
index 00000000..d4d813b0
--- /dev/null
+++ b/.agents/skills/build/SKILL.md
@@ -0,0 +1,17 @@
+---
+name: build
+description: Builds the GenVM project. Use after making code changes to compile Rust binaries.
+---
+
+Build procedure: `docs/contributing/howto/building/build.md` (debug build;
+read it first). Related: `building/runners.md`, `releasing/release-build.md`,
+`extending/modify-runner.md` under the same howto root.
+
+Claude-specific:
+
+- Build binaries with
+ `bash .agents/skills/build/scripts/run-ninja.sh -C build all/bin` instead of
+ raw ninja — it is silent on success and prints output only on failure, which
+ saves tokens.
+- See also: `/submodules` (multi-repo commits, `?submodules=1`), `/test`,
+ `/macos` (never build runners natively on macOS).
diff --git a/.claude/skills/build/scripts/run-ninja.sh b/.agents/skills/build/scripts/run-ninja.sh
similarity index 100%
rename from .claude/skills/build/scripts/run-ninja.sh
rename to .agents/skills/build/scripts/run-ninja.sh
diff --git a/.agents/skills/commit-style/SKILL.md b/.agents/skills/commit-style/SKILL.md
new file mode 100644
index 00000000..793c27c3
--- /dev/null
+++ b/.agents/skills/commit-style/SKILL.md
@@ -0,0 +1,145 @@
+---
+name: commit-style
+description: GenVM commit message conventions. Use when writing a commit message, squashing/merging a PR, or amending history. Covers the `type(scope): summary ` format, the five types, the standard scope set, the gitmoji suffix, and mistakes to avoid (typos, vague "fix CI N").
+---
+
+# Writing commit messages in GenVM
+
+Format:
+
+```text
+type(scope): short imperative summary
+```
+
+- **type** — one of five, lowercase (required).
+- **(scope)** — from the standard set below. Optional, but **aim for ~80%** of
+ commits to carry one; omit only for genuinely repo-wide changes.
+- **summary** — lowercase, imperative or noun-phrase, no trailing period, ≲70 chars.
+- **\** — **one to three** trailing [gitmoji](https://gitmoji.dev) glyphs
+ marking the precise intent. Use the actual Unicode emoji, not the `:shortcode:`.
+ Most commits want exactly one; reach for a second or third only when the change
+ genuinely carries more than one intent (e.g. a security fix that is also a
+ refactor → `🔒️♻️`). Order them most-important first. Never more than three.
+
+Examples:
+
+```text
+feat(calldata): add lazy decoding ✨
+perf(calldata): optimize internally tagged enums ⚡
+fix(executor): report entire fees subtree to host 🐛
+fix(calldata): defer parsing of forwarded calldata 🔒️⚡
+chore(ci): fix macos cache key 💚
+chore(build): bump wasmtime ⬆️
+docs(executor): add spec for ram consumption 📝
+```
+
+## Types
+
+The **type** is the coarse group; the **emoji** carries the fine intent.
+
+| Type | Use for | Default emoji |
+|---------|------------------------------------------------|---------------|
+| `feat` | new caller-visible capability or API surface | ✨ |
+| `fix` | broken behavior corrected | 🐛 |
+| `perf` | same behavior, faster or smaller | ⚡ |
+| `docs` | documentation / spec only | 📝 |
+| `chore` | everything internal (incl. refactors) | see below |
+
+Picking between the blurry ones:
+- New behavior a caller can observe → `feat`. Correcting wrong behavior → `fix`.
+ Same behavior but cheaper → `perf`. Pure restructure → `chore` + ♻️.
+
+## Emoji by intent
+
+Pick the glyph that best names *what kind* of change it is — it can be finer than
+the type. Common ones:
+
+| Emoji | Meaning | Emoji | Meaning |
+|-------|--------------------------|-------|------------------------|
+| ✨ | new feature | ♻️ | refactor |
+| 🐛 | bug fix | 🔥 | remove code / files |
+| 🚑 | critical hotfix | 🎨 | structure / format |
+| ⚡ | performance | ✅ | tests |
+| 📝 | docs | 💚 | fix CI |
+| 🔧 | config | 🚀 | release / deploy |
+| ⬆️/⬇️ | bump / drop deps | 🔒️ | security / privacy |
+| 🚚 | move / rename | 🔇 | remove logs |
+| 🏗️ | architectural change | 🔨 | dev / build scripts |
+| 🚧 | work in progress | | |
+
+Full reference: https://gitmoji.dev
+
+A lot of `fix`es in this repo are security-relevant — untrusted calldata from
+other nodes, fee/gas underflows, page limits, sandbox boundaries. When a fix
+hardens behavior against malicious or malformed input, mark it 🔒️ (alone, or
+paired with 🐛/⚡ when it is also a bugfix or optimization). The lazy-calldata
+work, for instance, is partly about not eagerly parsing attacker-controlled
+bytes — that earns a 🔒️.
+
+## Standard scopes
+
+| Scope | Covers |
+|-------------|----------------------------------------------------------------|
+| `executor` | rust executor core — calldata, host, common, rt/supervisor, fees, storage |
+| `wasm` | wasm/wasmtime compilation, precompile, wasm features |
+| `wasi` | the wasi syscall layer (`executor/src/wasi`) |
+| `rs-sdk` | the rust SDK (`executor/crates/sdk-rs`) |
+| `py-sdk` | the python stdlib / SDK (`runners/genlayer-py-std`) |
+| `modules` | modules in general (interfaces, install, implementation) |
+| `manager` | the module manager (`modules/implementation/src/manager`) |
+| `lua` | lua host scripting and configs (`genvm-lua`) |
+| `webdriver` | the webdriver module |
+| `ci` | GitHub workflows / CI |
+| `build` | build system, nix, release packaging, dependency bumps |
+| `genvm-tool`| the `genvm-tool` dev CLI (`support/tools/genvm-tool`): configure, test runner, hooks, git helpers |
+
+The set is curated, not closed — if a change clearly belongs to a subsystem not
+listed, a sensible lowercase scope is fine. Prefer an existing one when it fits
+(a `calldata` change is `(executor)`).
+
+## Body — keep it rare
+
+Prefer a single line. The summary should carry the change on its own; if you're
+reaching for a body to explain *what* changed, tighten the subject instead.
+
+Add a body **only** when a reader genuinely cannot reconstruct the *why* from the
+diff — a non-obvious tradeoff, a workaround for an external bug, or a subtle
+invariant. When you do, write one or two sentences of motivation, not a recap of
+the diff and not a bullet list of squashed sub-commits.
+
+## Mistakes to avoid (all seen in this repo's history)
+
+1. **Numbered firefighting** — `chore: fix CI`, `fix CI 2` … `fix CI 5`. Say
+ *what* broke: `chore(ci): fix macos cache key 💚`. A numbered run means the
+ messages describe nothing.
+2. **`batch update (#NNN)`** with no theme. A merged PR still deserves a one-line
+ summary of what it does.
+3. **Typos** — history has `reeipt`, `absolete`, `exremely`, `auto-formater`.
+ Spell-check the subject; it is permanent.
+4. **`chore: fix tests`** with no cause. Add it: `chore(executor): fix tests after
+ wasmtime rebase ✅`.
+5. **AI attribution** — never include `Co-authored-by` trailers, session
+ links/IDs (e.g. `Claude-Session:`), "Generated with" footers, or any other
+ AI/tool metadata, even when the tooling asks for it.
+
+## Checklist
+
+- [ ] Right type (feat / fix / perf / docs / chore)?
+- [ ] Scope present (target ~80%) and from the standard set?
+- [ ] Lowercase, no period, ≲70 chars?
+- [ ] One to three trailing gitmoji glyphs (most-important first) matching the intent?
+- [ ] Single line — body only if the *why* is truly unrecoverable from the diff?
+- [ ] Spell-checked?
+
+## Committing across the manager + submodules
+
+A change that touches an executor submodule spans repos: commit inside the
+submodule first, then bump the manager's gitlink (`git add executors/.x`)
+in a manager commit. Keep each commit's files coherent and don't let a
+submodule's linter reformat leak an unrelated file into a commit. Each repo's
+per-repo pre-commit hook (git-hooks.nix, in its flake) runs on the repo you
+commit in. Do not `--no-verify` (even for gitlink bumps — let the hooks run).
+Full workflow (order, pushing, `--force-with-lease` after a rebase): `/submodules`.
+
+Never include text like "bump the v0.3 gitlink". State the underlying change,
+not name the file updated
diff --git a/.agents/skills/initial-setup/SKILL.md b/.agents/skills/initial-setup/SKILL.md
new file mode 100644
index 00000000..dc664d26
--- /dev/null
+++ b/.agents/skills/initial-setup/SKILL.md
@@ -0,0 +1,6 @@
+---
+name: initial-setup
+description: Sets up the development environment for GenVM repository. Use when setting up the repo for the first time or when dependencies need to be refreshed.
+---
+
+See docs/contributing/howto/setup.md.
diff --git a/.agents/skills/macos/SKILL.md b/.agents/skills/macos/SKILL.md
new file mode 100644
index 00000000..b465bc46
--- /dev/null
+++ b/.agents/skills/macos/SKILL.md
@@ -0,0 +1,101 @@
+---
+name: macos
+description: Read this BEFORE trying to fix a GenVM build that fails on macOS. GenVM is NOT built natively on macOS — building from source on Darwin breaks the deterministic runner-artifact invariant. Use a remote Linux nix-builder instead. Triggers on any build/nix/runner/linker/ar failure on a Mac (Apple Silicon or Intel).
+---
+
+# Building GenVM on macOS
+
+## STOP — do not "fix" the build to run natively on macOS
+
+If you are on macOS and a build (`nix build`, `ninja`, runner packaging, linking,
+`ar`, FFI deps, …) fails, **do not** try to make GenVM compile natively on Darwin.
+This has been attempted and **rejected** before. Native macOS builds break the
+project's core invariant.
+
+### Why native macOS builds are forbidden
+
+Runner artifacts are content-addressed `.tar` files. **The file name encodes the
+content hash** (a Nix `nix32` SHA-256), so two nodes that build the "same" runner
+must produce byte-identical tars with identical names. This is what prevents
+**determinism violations** across the network: one node running a correct artifact
+and another running a differently-built artifact will disagree.
+
+A native macOS build:
+
+- produced **different and/or missing** runner tars vs. the Linux reference build
+ (e.g. only one `cpython` tar instead of two, mismatched hashes), and
+- silently changed the on-disk set of runners.
+
+You can verify the reference set on Linux with:
+
+```sh
+find $(nix build '.?submodules=1#runners-all' --no-link --print-out-paths) -name '*.tar' \
+ | sort | xargs sha256sum | sed -e 's+/nix/store/[^/]*++'
+```
+
+Any macOS output that does not reproduce these exact hashes is wrong.
+
+### Specific anti-fixes — never do these
+
+These are the tempting "fixes" that actually break determinism. Do **not** apply them:
+
+1. **Removing `outputHash`** from fixed-output derivations (e.g.
+ `runners/cpython/deps/ffi/default.nix`). The repo literally warns: *"the moment
+ you delete this you should question yourself."* `outputHash` is what pins the
+ artifact; deleting it destroys reproducibility.
+2. **Adding a platform-native linker** path for Darwin.
+3. **Patching `ar`** to work on macOS.
+4. Any change that lets runners build on `aarch64-darwin` / `x86_64-darwin`
+ instead of Linux.
+5. **Editing `runners/support/versions/current.nix` hashes** to make a macOS
+ build pass. A change whose only purpose is "make it build on macOS" must
+ **not** touch those hashes — they pin the Linux reference artifacts. If a
+ macOS build forces you to change a hash there, the build is producing the
+ wrong bytes; fix the builder (use the Linux remote builder), not the hash.
+6. **Removing / pruning old runner generations** (existing `.tar` entries or
+ prior hash versions). A macOS-build-only change must leave the existing set
+ of runner generations intact — silently dropping old ones is exactly the
+ determinism-breaking regression that gets such changes rejected.
+
+Per the maintainer: runners can only be built on Linux (x86_64). On macOS you
+**download** prebuilt runners or **build them on a remote Linux builder** — you do
+not build them locally.
+
+## The supported path: a remote Linux nix-builder
+
+Build Linux derivations on Linux, from your Mac, over SSH. The repo ships a
+ready-made containerized Linux remote builder:
+
+> **`support/macos/nix-builder/`** — see its `README.md` for the full, current
+> procedure. Read that file; do not improvise.
+
+In short (read the README for exact, up-to-date steps):
+
+1. Install Rosetta (`softwareupdate --install-rosetta --agree-to-license`) and
+ Docker Desktop with **"Use Rosetta for x86_64/amd64 emulation"** enabled.
+2. `docker build -t nix-builder support/macos/nix-builder` (add
+ `--platform linux/arm64` if needed), drop your `id_ed25519.pub` into the
+ volume, and `docker run` it (privileged, port `2222`).
+3. Register it in `/etc/nix/machines` and set `builders-use-substitutes = true`
+ in `/etc/nix/nix.conf`. Add the host key to `known_hosts` **for root too**.
+4. Sanity-check with
+ `NIX_REMOTE="ssh-ng://root@localhost:2222" nix build --system x86_64-linux nixpkgs#hello`.
+
+Notes that bite people (all in the README): use Nix `2.34.7`+ (older `2.2x`
+mishandles `ssh-ng://host:port`), and the root user — not just your user — must
+trust the builder's host key.
+
+## If you just need to develop / run, not rebuild runners
+
+You usually do not need to build runners at all. Download them instead (works on
+any platform):
+
+```bash
+nix develop '.?submodules=1#full' --command python3 build/out/bin/genvm-post-install \
+ --create-venv false --default-step false \
+ --runners-download true --error-on-missing-executor false
+```
+
+(`?submodules=1` is required on every flake ref — see `/submodules`.)
+
+See `/build` for the normal Rust build flow and `/submodules` for the multi-repo layout.
diff --git a/.claude/skills/pydoc/SKILL.md b/.agents/skills/pydoc/SKILL.md
similarity index 100%
rename from .claude/skills/pydoc/SKILL.md
rename to .agents/skills/pydoc/SKILL.md
diff --git a/.agents/skills/review-ready/SKILL.md b/.agents/skills/review-ready/SKILL.md
new file mode 100644
index 00000000..4006ea82
--- /dev/null
+++ b/.agents/skills/review-ready/SKILL.md
@@ -0,0 +1,32 @@
+---
+name: review-ready
+description: The bar a GenVM change clears before it is handed back. Use before pushing, opening or updating a PR, or reporting a change finished — and when asked whether something is ready for review.
+---
+
+# Review-Ready
+
+Walk [review-ready.md](../../../docs/contributing/howto/review-ready.md) against
+the actual diff and shell, not memory.
+[pr.md](../../../docs/contributing/howto/pr.md) has the branch model, the panel,
+and authority.
+
+## Report Review-Ready, Never Done
+
+Only 2 answers exist:
+
+- **Review-Ready** — evidence per item, including commands and their output
+- **Not Review-Ready** — naming the item that fails and the blocker behind it
+
+An unchecked item means **Not Review-Ready**, not a partial pass. Never report
+**Done**, which also requires review, cross-repo E2E, merge and release.
+
+## While you work
+
+- Fix defects in your diff without asking
+- Report defects outside it without folding them in or dropping them
+- Raise concerns about the requirement before building on it
+
+## Adapt it
+
+With an escaped-defect fix, say which item should have caught it, why it did not,
+and propose the amendment
diff --git a/.agents/skills/rust-test-style/SKILL.md b/.agents/skills/rust-test-style/SKILL.md
new file mode 100644
index 00000000..4b6feb9d
--- /dev/null
+++ b/.agents/skills/rust-test-style/SKILL.md
@@ -0,0 +1,192 @@
+---
+name: rust-test-style
+description: Always read before writing, adding, or modifying Rust tests — where to put and how to run them.
+---
+
+# Writing Rust tests in GenVM
+
+Standard library testing only. Do **not** add `rstest`, `proptest`, `insta`, or
+`pretty_assertions` — they are not used anywhere in the repo. Plain `#[test]`,
+`#[cfg(test)] mod tests`, and `std` assertions are the whole toolkit. `tokio` is
+available with the `macros` feature, so use `#[tokio::test]` when a test must be
+async (rare; no async tests exist today).
+
+To run these tests, see the `/test` skill (`--filter-tag rust`).
+
+## Where tests go
+
+**Decision rule — default to a separate file.** If the item under test is
+reachable from the crate's public API (a `pub fn`/`pub` type, even via a
+re-export), its test goes in its own `tests/.rs` file. Use an inline
+`#[cfg(test)] mod tests` **only** when the test must reach a *private* item that
+cannot be exercised through the public API. A public function tested inline is a
+review finding — move it to `tests/`.
+
+Two locations, both in use:
+
+1. **Integration tests** — one file per concern under the crate's `tests/` dir.
+ **This is the default** for anything testable through the public API. Each
+ `tests/*.rs` file is its own compilation unit (a separate crate) and gets its
+ own runner case.
+ - e.g. `executor/crates/calldata/tests/derive_decode.rs`,
+ `executor/crates/calldata/tests/address_checksum.rs`,
+ `executor/crates/common/tests/expr.rs`,
+ `modules/implementation/tests/test_rat.rs`
+ - Because each is a separate crate, it sees only the library crate plus
+ `[dev-dependencies]` — **not** the library's regular `[dependencies]`. If a
+ test needs a util the crate already depends on (e.g. `hex`), add it to
+ `[dev-dependencies]` too.
+2. **Inline unit tests** — `#[cfg(test)] mod tests { ... }` at the bottom of a
+ `src/*.rs` file, reserved for testing **private** items. e.g.
+ `executor/crates/calldata/src/lib.rs`, `executor/crates/common/src/logger/mod.rs`.
+
+```rust
+#[cfg(test)]
+mod tests {
+ use super::*;
+
+ #[test]
+ fn test_nested_value_in_struct() { /* ... */ }
+}
+```
+
+## Helpers over fixtures
+
+There is no fixture framework. Factor repeated setup into small free functions
+at the **top of the test file**, then keep each `#[test]` short. This is the
+dominant pattern.
+
+```rust
+// executor/crates/common/tests/expr.rs:5
+fn eval(input: &str) -> Value {
+ Expr::parse(input).unwrap().evaluate().unwrap()
+}
+
+fn assert_rational(val: Value, n: i64, d: i64) {
+ let r = val.into_rational().unwrap();
+ assert_eq!(r, BigRational::new(BigInt::from(n), BigInt::from(d)));
+}
+```
+
+```rust
+// modules/implementation/tests/test_rat.rs — construct the system under test
+fn make_vm() -> Lua {
+ let vm = Lua::new();
+ rat::register_rat_global(&vm).unwrap();
+ vm
+}
+fn eval(vm: &Lua, code: &str) -> T {
+ vm.load(code).eval::().unwrap()
+}
+```
+
+Define throwaway `#[derive(...)]` types as fixtures right in the test file:
+
+```rust
+// executor/crates/calldata/tests/derive_decode_excess_fields.rs:11
+#[derive(Debug, PartialEq, Decode)]
+struct Named { x: i32, y: String }
+```
+
+## Assertions
+
+`std` only: `assert!`, `assert_eq!`, `assert_ne!`. When asserting a boolean
+condition, add a format-string message that prints the **actual value** so a
+failure is debuggable:
+
+```rust
+let msg = err.to_string();
+assert!(msg.contains("unknown field"), "unexpected error: {msg}");
+assert!(msg.contains("`z`"), "should mention field name: {msg}");
+```
+
+For value equality use `assert_eq!` against a fully-constructed expected value
+(derive `Debug, PartialEq` on the type).
+
+## Testing errors
+
+Use a helper that returns `Result`, call `.unwrap_err()`, and assert on
+substrings of the `Display` message — don't match on error enum internals:
+
+```rust
+// executor/crates/calldata/tests/derive_decode_excess_fields.rs:5
+fn try_decode_from_value(val: Value) -> Result {
+ T::decode(codec::ValueDeserializer(val))
+}
+
+#[test]
+fn tuple_struct_rejects_too_many_elements() {
+ let err = try_decode_from_value::(/* 3 elems */).unwrap_err();
+ assert!(err.to_string().contains("expected 2"), "unexpected error: {err}");
+}
+```
+
+The happy path is liberal with `.unwrap()` — a panic *is* the test failure.
+
+## Roundtrip pattern
+
+For codecs, encode → decode → compare in one helper:
+
+```rust
+// executor/crates/calldata/tests/derive_decode.rs:5
+fn roundtrip_via_value(val: &T) -> Value
+where T: for<'a> codec::Encode<&'a mut Vec, Error = std::convert::Infallible> {
+ let mut buf = Vec::new();
+ codec::Encode::encode(val, &mut Encoder::new(&mut buf)).unwrap();
+ genlayer_calldata::decode(&buf).unwrap()
+}
+```
+
+## Naming & layout
+
+- Test fn names are descriptive `snake_case` that name the **scenario**, e.g.
+ `named_struct_rejects_unknown_field`, `string_literals_and_escapes`.
+- **Do not prefix with `test_`.** The `#[test]` attribute already says it is a
+ test; the prefix is noise. Drop it from new tests and from any you touch.
+- Group related tests with a section-comment banner:
+ ```rust
+ // ── Tuple struct: wrong sequence length ─────────────────────────────
+ ```
+- Helpers first, then types, then `#[test]` functions.
+
+## How `genvm-tool test` discovers Rust tests
+
+Discovery lives in `tests/plugins/genvm_tool_plugins/cargo.py`. For each
+registered Rust crate root it creates cases via plain `cargo test`:
+
+| Source | Command | Tags |
+|--------|---------|------|
+| each `tests/*.rs` file | `cargo test --test ` | `rust`, `unit` |
+| crate has `src/lib.rs` (inline `#[cfg(test)]`) | `cargo test --lib` | `rust`, `unit` |
+| each binary (`src/main.rs` / `[[bin]]`) | `cargo test --bin ` | `rust`, `unit` |
+| each `[[example]]` | `cargo check --example ` (compile-only) | `rust`, `example` |
+
+Consequences when adding tests:
+
+- **A new file in `tests/` is auto-discovered** — no registration needed, as long
+ as the crate root itself is already scanned.
+- **New inline `#[cfg(test)]` tests** ride along on the existing `--lib` case;
+ nothing to register.
+- The crate must be a known Rust root. Roots carry a `.ya-test-config.json`
+ (e.g. `executor/.ya-test-config.json`, `executor/crates/calldata/.ya-test-config.json`).
+ That file controls per-crate `cargo_test_flags`, `keep_env`, and a `skip` map
+ (skip `examples`/specific cases) — add a new crate's config there if introducing
+ a new root.
+- The `rust` preset is `(rust | integration) & !bench & !fuzz`
+ (`tests/presets/rust.txt`).
+
+## Prefer fuzzing when it fits
+
+If a function's correctness is really about holding over a **space of inputs**
+(parsers, codecs, encode/decode roundtrips, anything taking arbitrary bytes or
+structured input), write a **fuzz target** rather than a handful of hand-picked
+`#[test]` cases. A few cherry-picked examples give false confidence; a fuzzer
+explores the space. Don't settle for an under-tested `#[test]` when the property
+is fuzz-shaped — reach for fuzz, or do both (fuzz for coverage, a couple of
+`#[test]`s pinning known edge cases / regressions).
+
+Fuzz targets live under a crate's `fuzz/` (afl, behind the `arbitrary` feature,
+tagged `fuzz`, built with `cargo_afl_build_flags`; see `cargo_fuzz` in
+`tests/plugins/genvm_tool_plugins/cargo.py`). They are a distinct case type —
+don't mix a fuzz harness into `tests/` or `mod tests`, and they're excluded from
+the `rust` preset (`!fuzz`).
diff --git a/.agents/skills/spec/SKILL.md b/.agents/skills/spec/SKILL.md
new file mode 100644
index 00000000..d5eed14e
--- /dev/null
+++ b/.agents/skills/spec/SKILL.md
@@ -0,0 +1,51 @@
+---
+name: spec
+description: How to write GenVM spec pages (docs/website/src/spec). Use when adding or editing spec sections — brevity, linking constants/errors/terms instead of inlining them, spec vs impl-spec split, and how to verify the build.
+---
+
+# Writing spec
+
+The spec (`docs/website/src/spec/`) describes **observable behavior** — what a
+contract or validator can detect. How this codebase achieves it goes to
+`impl-spec/`. If a sentence starts describing internals (caches, native limits,
+threads), either drop it or reduce it to one normative requirement on
+implementations.
+
+## Be brief
+
+- One concern per section; a few paragraphs or one list is the right size.
+- Enumerate cases with `#.` lists instead of prose walkthroughs; name the
+ actors precisely (*caller*/*callee*, leader/validator) and describe each
+ case once.
+- State what happens, not why the design is good. Rationale goes to ADRs
+ (`docs/adr/`).
+- Cover the edges (unwinding, host boundary, validation-time rejection) in one
+ sentence each — omitting them is sweeping under the rug, but they rarely
+ deserve a paragraph.
+
+## Link, don't inline
+
+Never write a literal value or error string in spec text — link the anchor, so
+generated pages stay the single source of truth:
+
+- Constants: `:ref:`gvm-def-consts-value--`` /
+ `:ref:`gvm-def-const-`` from `spec/appendix/constants.rst`
+ (generated from `executor/codegen/data/public-abi.json`),
+ `internal-constants.rst` (from `internal-constants.json`) or
+ `constants-pending.rst` (from `public-abi-pending.json`). **Never edit these
+ .rst by hand** — edit the JSON and regenerate
+ ([genvm-tool.md](../../../docs/contributing/howto/genvm-tool.md)). Constants a
+ contract cannot read go to the internal JSON, not-yet-stabilized ones to the
+ pending JSON.
+- Error outcomes: `:ref:`gvm-def-str-trie-value-vm-error-...`` — every "traps
+ with" / "rejected with" must link the exact vm_error entry.
+- Glossary terms: `:term:`sub-VM`` etc. on first use in a section.
+- Other spec pages: `:doc:` relative links instead of restating their content.
+
+## Verify
+
+Build the website ([docs.md](../../../docs/contributing/howto/building/docs.md))
+and check for `undefined label` warnings on your pages — a broken `:ref:` is a
+silent dead link otherwise. Verify claims against the implementation before
+writing them; every behavioral sentence should have a code location you can
+point to (but do not link it to the specification).
diff --git a/.agents/skills/submodules/SKILL.md b/.agents/skills/submodules/SKILL.md
new file mode 100644
index 00000000..f6f42b6a
--- /dev/null
+++ b/.agents/skills/submodules/SKILL.md
@@ -0,0 +1,6 @@
+---
+name: submodules
+description: How the GenVM multi-repo (manager umbrella + executor submodules) fits together and how to build, commit, run hooks, and push across all of them. Use when committing/pushing changes that touch a submodule, bumping gitlinks, building via nix, rebasing the manager, or debugging the cross-repo pre-commit hook.
+---
+
+See docs/contributing/howto/committing/submodules.md.
diff --git a/.agents/skills/test/SKILL.md b/.agents/skills/test/SKILL.md
new file mode 100644
index 00000000..eb1206ad
--- /dev/null
+++ b/.agents/skills/test/SKILL.md
@@ -0,0 +1,21 @@
+---
+name: test
+description: Runs tests for the GenVM project. Use after making code changes to verify correctness.
+---
+
+See docs/contributing/howto/testing/README.md.
+
+Invoke the tool as:
+
+```
+nix run '.?submodules=1#genvm-tool' -- test run --filter-tag '!fuzz' ...
+```
+
+A bare `genvm-tool` from `PATH` is a nix store copy pinned at dev-shell build
+time; it goes stale against the working tree and fails with import errors such
+as `ModuleNotFoundError: No module named 'genvm_tool.tests.exec.process'`.
+`build/genvm_tool.sh` has the same problem — it bakes the store path in.
+
+After fixing a failure, rerun with `--filter-continue ` (path printed in
+the failure summary) before any full rerun — see the "Fix–rerun loop" section
+there.
diff --git a/.claude/agents/reviewer-implementation.md b/.claude/agents/reviewer-implementation.md
deleted file mode 100644
index e9756f39..00000000
--- a/.claude/agents/reviewer-implementation.md
+++ /dev/null
@@ -1,76 +0,0 @@
----
-name: reviewer-implementation
-description: Reviews the implementation of a GenVM branch — code-vs-spec drift, code quality, AI slop, duplicated logic, useless comments, and doc completeness. Use as the "implementation" pass of a branch review.
-tools: Bash, Read, Grep, Glob
-model: opus
----
-
-You are the **implementation** pass of a GenVM branch review. You read the code
-and judge how it is built. Read-only: never edit code.
-
-## Baseline
-
-Diff against the active dev branch, not `main`:
-
-- Default to `v0.3-dev` (or the current `v0.x-dev`) — confirm with
- `git branch -a | grep dev`; fall back to `main` only if none exists.
-- State the base in one line; ignore commits already on it.
-- `git log --oneline ..HEAD` / `git diff --stat ..HEAD`, then read the
- diffs of the changed code.
-
-## What to report (with file:line evidence)
-
-1. **Code-vs-spec drift.** Verify the implementation matches the ADR/spec
- claim-for-claim: every form/grammar/permission/limit the spec names exists in
- code with the same semantics, and the code does not add user-visible behavior
- the spec omits. Call out each divergence.
-
-2. **Doc / SDK completeness.** New `gl_call`s, permissions, runner-id forms, etc.
- must be reflected in `doc/website/src/spec/**`, `doc/schemas/*.json`, and any
- SDK wrappers/docstrings. Flag anything implemented but undocumented.
-
-3. **AI slop & duplicated logic.** Deliberate engineering or generated filler?
- Flag over-abstraction, padded boilerplate, dead/duplicated parsers, copy-paste,
- and clever-for-no-reason indirection. Say plainly when it is NOT slop.
- Specifically: a new `gl_call` handler (e.g. `register_runner` in
- `wasi/genlayer_sdk.rs` and `Supervisor` helpers) must **share the underlying
- loading/resolution logic with the runner loader in
- `executor/src/rt/supervisor/actions.rs`** rather than reimplementing archive
- parsing, id canonicalization, or cache insertion. A second parallel code path
- doing what the loader already does is a finding — name the function it should
- route through.
-
- **Responsibilities must be encapsulated in the layer that owns them.** Each
- layer does its own job and exposes ONE entry point; callers in other layers
- delegate, they don't reach across and re-orchestrate. Concretely: the
- wasi/vfs layer (`wasi/genlayer_sdk.rs`, `wasi/preview1.rs`) is a thin syscall
- shim — it must NOT contain runner-resolution logic (computing a contract's
- runner id, picking storage slots/state, stitching together
- `get_runner_of_contract` + `load_runner` + archive mapping). That belongs in
- the runners/supervisor layer; the gl_call handler should call a single
- encapsulated function there. A handler that assembles a cross-layer flow
- inline (even if each piece is "shared") is a finding — name the boundary it
- violates and the single function it should call instead. Likewise, storage
- layout, permission derivation, and limiter accounting each have one home;
- flag logic that leaks into a layer that shouldn't know about it.
-
-4. **Correctness & quality.** Panics on malformed input (`slice`, `unwrap`),
- error handling, idempotency.
-
-5. **Edge cases are tested.** Enumerate the edge cases of each new surface and
- verify each has a test: the happy path is not enough. Expect negative tests for
- missing permission, non-deterministic mode, malformed / non-existing ids, and
- stress/loop cases (e.g. registering the same thing ~1000× to prove
- consume-once). Name each untested edge case as a gap.
-
-6. **Useless comments.** Flag narrate-the-obvious comments. Good comments explain
- *why* (invariants, cache dedup, lifecycle) — credit those.
-
-Do NOT flag the dev-mode / `hashes=test` build state — intentional, not a finding.
-Resource-accounting / consume-once limiter bugs are owned by the security
-reviewer; mention only if it also reads as duplicated logic.
-
-## Style
-
-Concise and direct. Lead with: is the implementation correct and clean enough to
-merge? Separate blocking issues from nits. Note if you did not build or run tests.
diff --git a/.claude/agents/reviewer-implementation.md b/.claude/agents/reviewer-implementation.md
new file mode 120000
index 00000000..a19c3b26
--- /dev/null
+++ b/.claude/agents/reviewer-implementation.md
@@ -0,0 +1 @@
+../../.agents/agents/reviewer-implementation.md
\ No newline at end of file
diff --git a/.claude/agents/reviewer-security.md b/.claude/agents/reviewer-security.md
deleted file mode 100644
index 11798cf7..00000000
--- a/.claude/agents/reviewer-security.md
+++ /dev/null
@@ -1,57 +0,0 @@
----
-name: reviewer-security
-description: Security review of a GenVM branch — attacker mindset over permissions, resource limits, sandbox propagation, parsing, and state-exfiltration. Use as the "security" pass of a branch review.
-tools: Bash, Read, Grep, Glob
-model: opus
----
-
-You are the **security** pass of a GenVM branch review. Think like an attacker
-writing a malicious contract. Read-only: never edit code.
-
-First read the repo root `SECURITY.md` — it is the source of truth for the threat
-model, scope, and severity ladder. Rank every finding by the priority it defines
-and cite that priority. Do not restate its contents in your report; reference it.
-
-## Baseline
-
-Diff against the active dev branch, not `main`:
-
-- Default to `v0.3-dev` (or the current `v0.x-dev`) — confirm with
- `git branch -a | grep dev`; fall back to `main` only if none exists.
-- State the base in one line; ignore commits already on it.
-- `git log --oneline ..HEAD` / `git diff --stat ..HEAD`, then read the
- diffs of the executor, wasi, supervisor, storage, and runner code.
-
-## What to check (with file:line evidence)
-
-- **Permission gates.** Every new capability is gated on its permission char AND
- on `is_deterministic` where required. Check the gate is at entry, before any
- side effect.
-- **Permission propagation.** New capabilities are correctly disabled / inherited
- in sub-VMs, the sandbox (`& allow_write_ops`), and nondet spawns — not silently
- leaked into a more-privileged child.
-- **Allocation bounded before allocating.** Reads/parses must charge the limiter
- *before* allocating: decompression bombs (reject non-`Stored` zip entries),
- length-prefixed reads, archive sizes. Charging after the alloc is a finding.
-- **Resource accounting is consume-once.** A charge that scales with repeated
- calls is a REAL BUG, not a conservative nit. Content-addressed resources (e.g. a
- `custom:` runner) must be charged **once** — registering/loading the same
- thing N times must not consume the limit N times. The test: "do it ~1000× in a
- loop — does the limit overflow?" If yes, flag it and prescribe dedup-by-hash /
- consume-once. Never excuse N-times charging as "errs safe."
-- **Strict input parsing.** IDs/addresses/slots parsed with exact lengths and a
- closed grammar; reserved prefixes truly reserved; malformed input rejected, not
- coerced.
-- **Read-oracle / exfiltration.** Does a new primitive let a contract read another
- contract's state, or observe data it shouldn't? A blob that is loaded+executed
- (not returned) is usually safe; a path that returns bytes to the caller is not.
-- **Determinism.** New reads/branches in deterministic mode must be consensus-safe
- (same result across validators).
-
-Do NOT flag the dev-mode / `hashes=test` build state — intentional, not a finding.
-
-## Style
-
-Concise. Lead with: any exploitable issue, yes/no, and what blocks merge.
-Distinguish a real vuln from a hardening nice-to-have. Note if you did not build
-or run anything.
diff --git a/.claude/agents/reviewer-security.md b/.claude/agents/reviewer-security.md
new file mode 120000
index 00000000..621de8a8
--- /dev/null
+++ b/.claude/agents/reviewer-security.md
@@ -0,0 +1 @@
+../../.agents/agents/reviewer-security.md
\ No newline at end of file
diff --git a/.claude/agents/reviewer-spec.md b/.claude/agents/reviewer-spec.md
deleted file mode 100644
index 4fc4f468..00000000
--- a/.claude/agents/reviewer-spec.md
+++ /dev/null
@@ -1,63 +0,0 @@
----
-name: reviewer-spec
-description: Reviews the ADR/proposal and spec/schema of a GenVM branch as documents — soundness, clarity, and completeness (what was forgotten). Use as the "spec" pass of a branch review.
-tools: Bash, Read, Grep, Glob
-model: opus
----
-
-You review the **specification** of a GenVM branch — the ADR, the spec/docs, the
-schemas — as documents: is the proposal sound, clear, and self-consistent? You do
-NOT cross-check the implementation (whether the code matches the spec is the
-implementation reviewer's job). Read-only: never edit code.
-
-## Baseline
-
-Diff against the active dev branch, not `main`:
-
-- Default to `v0.3-dev` (or the current `v0.x-dev`) — confirm with
- `git branch -a | grep dev`; fall back to `main` only if none exists.
-- State the base in one line at the top; ignore commits already on it.
-- `git log --oneline ..HEAD` and `git diff --stat ..HEAD`, then read
- the diffs of `doc/**` and `*.json` schemas.
-
-## What to report
-
-With file:line evidence:
-
-1. **ADR / proposal quality.** If the branch adds or changes a `doc/adr/*.md`
- (or design doc), rate it. Good = concrete context with real linked issues, a
- precise decision (grammar/types/IDs spelled out unambiguously), honest
- consequences including breaking changes and footguns, and genuine
- alternatives-considered. If there's no ADR for a change that warrants one,
- say so.
-
-2. **Spec / schema soundness.** Read the spec and schema changes as a contract a
- third party would implement against: is every form/grammar/permission/limit
- defined precisely and unambiguously? Are there gaps, contradictions, or
- under-specified edges? Is the JSON schema itself valid and matching the prose?
- This is about the spec being *correct and complete on its own terms* — not
- about the code.
-
-3. **Completeness — what did we forget to add?** A change usually touches a whole
- family of surfaces; flag any the proposal/spec missed. E.g. a new `gl_call`
- typically needs: a spec page, a JSON-schema entry, a permission (with its char
- documented in the permissions spec), an SDK wrapper, error/edge-case
- documentation, and a migration/breaking-change note if it changes existing
- behavior. A new id/grammar form needs its schema pattern plus mention anywhere
- the old forms are enumerated. List the surfaces that *should* have changed
- together but didn't.
-
-4. **Edge cases are documented or inferrable.** Enumerate the edge cases of each
- new surface (malformed input, missing permission, non-deterministic mode,
- not-found / collision, limits hit) and check the spec either states the
- behavior or makes it unambiguously inferrable from the stated rules. A behavior
- a reader would have to guess is a spec gap — list each one. (Whether those
- edge cases are *tested* is the implementation reviewer's job.)
-
-Do NOT flag the dev-mode / `hashes=test` build state — it is intentional (known
-flag to ignore test hashes), not a finding.
-
-## Style
-
-Concise and direct. Lead with: is the proposal sound and the spec implementable
-as written? Separate blocking gaps from nits. Don't pad.
diff --git a/.claude/agents/reviewer-spec.md b/.claude/agents/reviewer-spec.md
new file mode 120000
index 00000000..9933c9c9
--- /dev/null
+++ b/.claude/agents/reviewer-spec.md
@@ -0,0 +1 @@
+../../.agents/agents/reviewer-spec.md
\ No newline at end of file
diff --git a/.claude/settings.fuzzing.json b/.claude/settings.fuzzing.json
deleted file mode 100644
index eac0945d..00000000
--- a/.claude/settings.fuzzing.json
+++ /dev/null
@@ -1,18 +0,0 @@
-{
- "permissions": {
- "allow": [
- "Bash(./tests/cases/claude/run_test.py:*)",
- "Edit(tests/cases/claude/agent/**)",
- "Write(tests/cases/claude/agent/**)",
- "Edit(tests/cases/claude/intelligence/**)",
- "Write(tests/cases/claude/intelligence/**)"
- ]
- },
- "sandbox": {
- "enabled": true,
- "autoAllowBashIfSandboxed": true,
- "excludedCommands": [
- "./tests/cases/claude/run_test.py"
- ]
- }
-}
diff --git a/.claude/skills/agentic-fuzzing b/.claude/skills/agentic-fuzzing
new file mode 120000
index 00000000..c478bdd4
--- /dev/null
+++ b/.claude/skills/agentic-fuzzing
@@ -0,0 +1 @@
+../../.agents/skills/agentic-fuzzing
\ No newline at end of file
diff --git a/.claude/skills/branch-review b/.claude/skills/branch-review
new file mode 120000
index 00000000..cd42a6bb
--- /dev/null
+++ b/.claude/skills/branch-review
@@ -0,0 +1 @@
+../../.agents/skills/branch-review
\ No newline at end of file
diff --git a/.claude/skills/build b/.claude/skills/build
new file mode 120000
index 00000000..b1286db4
--- /dev/null
+++ b/.claude/skills/build
@@ -0,0 +1 @@
+../../.agents/skills/build
\ No newline at end of file
diff --git a/.claude/skills/build/SKILL.md b/.claude/skills/build/SKILL.md
deleted file mode 100644
index f72bc27d..00000000
--- a/.claude/skills/build/SKILL.md
+++ /dev/null
@@ -1,76 +0,0 @@
----
-name: build
-description: Builds the GenVM project. Use after making code changes to compile Rust binaries.
----
-
-To build the GenVM project:
-
-## Reconfiguring
-
-If ninja fails with `missing and no known rule to make it` (e.g., after adding/removing/renaming source files), regenerate the build file first:
-
-```bash
-./configure.rb
-```
-
-This runs the `configure.rb` Ruby script at the project root, which regenerates `build/build.ninja` with the current file list.
-
-## Building
-
-**Build all Rust binaries:**
-```bash
-nix develop .#full --command bash .claude/skills/build/scripts/run-ninja.sh -C build all/bin
-```
-
-This runs ninja silently and only shows output on failure (to save tokens).
-
-**Available ninja targets:**
-
-| Target | Description |
-|--------|-------------|
-| `all` | Build everything |
-| `all/bin` | Build all Rust binaries |
-| `all/data` | Build data about runners using Nix |
-| `codegen` | Run code generation |
-
-**Output locations:**
-- `out/bin/genvm-modules` - modules binary
-- `out/executor//bin/genvm` - executor binary (one dir per built version, e.g. `out/executor/v0.3.0-rc7/`; the concrete version comes from `executors/.x/manifest.json` and is recorded in `build/info.json`)
-
-## Runners
-
-Runners can only be built on x86_64 using the `all` target. On other platforms, download them instead.
-
-**Download runners:**
-```bash
-nix develop .#full --command python3 build/out/bin/post-install.py --create-venv false --default-step false --runners-download true --error-on-missing-executor false
-```
-
-### Runner Development Workflow
-
-To develop/modify a runner (e.g., cloudpickle):
-
-1. **Enable dev mode:**
- Set `runners/support/versions/dev-mode.nix` to `true`
-
-2. **Set hash to "test":**
- In `runners/support/versions/current.nix`, set the runner's hash to `"test"`
-
-3. **Make your modifications and run tests**
- With dev-mode enabled and hash set to "test", you can build and run tests.
-
-4. **Disable dev mode:**
- Set `runners/support/versions/dev-mode.nix` back to `false`. The build will now tell you to set hashes to `null`.
-
-5. **Set hashes to null and build:**
- Set the runner's hash (and dependent runners' hashes) to `null`, then build:
- ```bash
- nix develop .#full --command ninja -C build all
- ```
- The build will fail with a hash mismatch showing the new hash.
-
-6. **Update hashes:**
- Copy the new hash from the error message back into `hashes.nix`. Repeat for dependent runners.
-
-7. **Rebuild to verify:**
- Run the build again to confirm all hashes are correct.
diff --git a/.claude/skills/commit-style b/.claude/skills/commit-style
new file mode 120000
index 00000000..5284a32f
--- /dev/null
+++ b/.claude/skills/commit-style
@@ -0,0 +1 @@
+../../.agents/skills/commit-style
\ No newline at end of file
diff --git a/.claude/skills/commit-style/SKILL.md b/.claude/skills/commit-style/SKILL.md
deleted file mode 100644
index ee2be06f..00000000
--- a/.claude/skills/commit-style/SKILL.md
+++ /dev/null
@@ -1,128 +0,0 @@
----
-name: commit-style
-description: GenVM commit message conventions. Use when writing a commit message, squashing/merging a PR, or amending history. Covers the `type(scope): summary ` format, the five types, the standard scope set, the gitmoji suffix, and mistakes to avoid (typos, vague "fix CI N").
----
-
-# Writing commit messages in GenVM
-
-Format:
-
-```text
-type(scope): short imperative summary
-```
-
-- **type** — one of five, lowercase (required).
-- **(scope)** — from the standard set below. Optional, but **aim for ~80%** of
- commits to carry one; omit only for genuinely repo-wide changes.
-- **summary** — lowercase, imperative or noun-phrase, no trailing period, ≲70 chars.
-- **\** — **one to three** trailing [gitmoji](https://gitmoji.dev) glyphs
- marking the precise intent. Use the actual Unicode emoji, not the `:shortcode:`.
- Most commits want exactly one; reach for a second or third only when the change
- genuinely carries more than one intent (e.g. a security fix that is also a
- refactor → `🔒️♻️`). Order them most-important first. Never more than three.
-
-Examples:
-
-```text
-feat(calldata): add lazy decoding ✨
-perf(calldata): optimize internally tagged enums ⚡
-fix(executor): report entire fees subtree to host 🐛
-fix(calldata): defer parsing of forwarded calldata 🔒️⚡
-chore(ci): fix macos cache key 💚
-chore(build): bump wasmtime ⬆️
-docs(executor): add spec for ram consumption 📝
-```
-
-## Types
-
-The **type** is the coarse group; the **emoji** carries the fine intent.
-
-| Type | Use for | Default emoji |
-|---------|------------------------------------------------|---------------|
-| `feat` | new caller-visible capability or API surface | ✨ |
-| `fix` | broken behavior corrected | 🐛 |
-| `perf` | same behavior, faster or smaller | ⚡ |
-| `docs` | documentation / spec only | 📝 |
-| `chore` | everything internal (incl. refactors) | see below |
-
-Picking between the blurry ones:
-- New behavior a caller can observe → `feat`. Correcting wrong behavior → `fix`.
- Same behavior but cheaper → `perf`. Pure restructure → `chore` + ♻️.
-
-## Emoji by intent
-
-Pick the glyph that best names *what kind* of change it is — it can be finer than
-the type. Common ones:
-
-| Emoji | Meaning | Emoji | Meaning |
-|-------|--------------------------|-------|------------------------|
-| ✨ | new feature | ♻️ | refactor |
-| 🐛 | bug fix | 🔥 | remove code / files |
-| 🚑 | critical hotfix | 🎨 | structure / format |
-| ⚡ | performance | ✅ | tests |
-| 📝 | docs | 💚 | fix CI |
-| 🔧 | config | 🚀 | release / deploy |
-| ⬆️/⬇️ | bump / drop deps | 🔒️ | security / privacy |
-| 🚚 | move / rename | 🔇 | remove logs |
-| 🏗️ | architectural change | 🔨 | dev / build scripts |
-| 🚧 | work in progress | | |
-
-Full reference: https://gitmoji.dev
-
-A lot of `fix`es in this repo are security-relevant — untrusted calldata from
-other nodes, fee/gas underflows, page limits, sandbox boundaries. When a fix
-hardens behavior against malicious or malformed input, mark it 🔒️ (alone, or
-paired with 🐛/⚡ when it is also a bugfix or optimization). The lazy-calldata
-work, for instance, is partly about not eagerly parsing attacker-controlled
-bytes — that earns a 🔒️.
-
-## Standard scopes
-
-| Scope | Covers |
-|-------------|----------------------------------------------------------------|
-| `executor` | rust executor core — calldata, host, common, rt/supervisor, fees, storage |
-| `wasm` | wasm/wasmtime compilation, precompile, wasm features |
-| `wasi` | the wasi syscall layer (`executor/src/wasi`) |
-| `rs-sdk` | the rust SDK (`executor/crates/sdk-rs`) |
-| `py-sdk` | the python stdlib / SDK (`runners/genlayer-py-std`) |
-| `modules` | modules in general (interfaces, install, implementation) |
-| `manager` | the module manager (`modules/implementation/src/manager`) |
-| `lua` | lua host scripting and configs (`genvm-lua`) |
-| `webdriver` | the webdriver module |
-| `ci` | GitHub workflows / CI |
-| `build` | build system, nix, release packaging, dependency bumps |
-
-The set is curated, not closed — if a change clearly belongs to a subsystem not
-listed, a sensible lowercase scope is fine. Prefer an existing one when it fits
-(a `calldata` change is `(executor)`).
-
-## Body — keep it rare
-
-Prefer a single line. The summary should carry the change on its own; if you're
-reaching for a body to explain *what* changed, tighten the subject instead.
-
-Add a body **only** when a reader genuinely cannot reconstruct the *why* from the
-diff — a non-obvious tradeoff, a workaround for an external bug, or a subtle
-invariant. When you do, write one or two sentences of motivation, not a recap of
-the diff and not a bullet list of squashed sub-commits.
-
-## Mistakes to avoid (all seen in this repo's history)
-
-1. **Numbered firefighting** — `chore: fix CI`, `fix CI 2` … `fix CI 5`. Say
- *what* broke: `chore(ci): fix macos cache key 💚`. A numbered run means the
- messages describe nothing.
-2. **`batch update (#NNN)`** with no theme. A merged PR still deserves a one-line
- summary of what it does.
-3. **Typos** — history has `reeipt`, `absolete`, `exremely`, `auto-formater`.
- Spell-check the subject; it is permanent.
-4. **`chore: fix tests`** with no cause. Add it: `chore(executor): fix tests after
- wasmtime rebase ✅`.
-
-## Checklist
-
-- [ ] Right type (feat / fix / perf / docs / chore)?
-- [ ] Scope present (target ~80%) and from the standard set?
-- [ ] Lowercase, no period, ≲70 chars?
-- [ ] One to three trailing gitmoji glyphs (most-important first) matching the intent?
-- [ ] Single line — body only if the *why* is truly unrecoverable from the diff?
-- [ ] Spell-checked?
diff --git a/.claude/skills/initial-setup b/.claude/skills/initial-setup
new file mode 120000
index 00000000..032dd983
--- /dev/null
+++ b/.claude/skills/initial-setup
@@ -0,0 +1 @@
+../../.agents/skills/initial-setup
\ No newline at end of file
diff --git a/.claude/skills/initial-setup/SKILL.md b/.claude/skills/initial-setup/SKILL.md
deleted file mode 100644
index 50a2763e..00000000
--- a/.claude/skills/initial-setup/SKILL.md
+++ /dev/null
@@ -1,30 +0,0 @@
----
-name: initial-setup
-description: Sets up the development environment for GenVM repository. Use when setting up the repo for the first time or when dependencies need to be refreshed.
----
-
-To set up the GenVM development environment:
-
-1. **Enter the Nix flake environment:**
- ```bash
- nix develop .#full
- ```
-
-2. **Initialize git submodules:**
- ```bash
- git submodule update --init --recursive --depth 1
- ```
-
-3. **Source environment variables:**
- ```bash
- source env.sh
- ```
- This adds `tools/git-third-party` to PATH and sources `.env` if it exists.
-
-4. **Update third-party dependencies:**
- ```bash
- ./tools/git-third-party/git-third-party update --all
- ```
- This updates wasmtime, wasm-tools, and applies GenVM-specific patches.
-
-The repository will be ready for development with all dependencies properly configured.
diff --git a/.claude/skills/macos b/.claude/skills/macos
new file mode 120000
index 00000000..bdb6d05d
--- /dev/null
+++ b/.claude/skills/macos
@@ -0,0 +1 @@
+../../.agents/skills/macos
\ No newline at end of file
diff --git a/.claude/skills/macos/SKILL.md b/.claude/skills/macos/SKILL.md
deleted file mode 100644
index 905a1507..00000000
--- a/.claude/skills/macos/SKILL.md
+++ /dev/null
@@ -1,99 +0,0 @@
----
-name: macos
-description: Read this BEFORE trying to fix a GenVM build that fails on macOS. GenVM is NOT built natively on macOS — building from source on Darwin breaks the deterministic runner-artifact invariant. Use a remote Linux nix-builder instead. Triggers on any build/nix/runner/linker/ar failure on a Mac (Apple Silicon or Intel).
----
-
-# Building GenVM on macOS
-
-## STOP — do not "fix" the build to run natively on macOS
-
-If you are on macOS and a build (`nix build`, `ninja`, runner packaging, linking,
-`ar`, FFI deps, …) fails, **do not** try to make GenVM compile natively on Darwin.
-This has been attempted and **rejected** before. Native macOS builds break the
-project's core invariant.
-
-### Why native macOS builds are forbidden
-
-Runner artifacts are content-addressed `.tar` files. **The file name encodes the
-content hash** (a Nix `nix32` SHA-256), so two nodes that build the "same" runner
-must produce byte-identical tars with identical names. This is what prevents
-**determinism violations** across the network: one node running a correct artifact
-and another running a differently-built artifact will disagree.
-
-A native macOS build:
-
-- produced **different and/or missing** runner tars vs. the Linux reference build
- (e.g. only one `cpython` tar instead of two, mismatched hashes), and
-- silently changed the on-disk set of runners.
-
-You can verify the reference set on Linux with:
-
-```sh
-find $(nix build .#runners-all --no-link --print-out-paths) -name '*.tar' \
- | sort | xargs sha256sum | sed -e 's+/nix/store/[^/]*++'
-```
-
-Any macOS output that does not reproduce these exact hashes is wrong.
-
-### Specific anti-fixes — never do these
-
-These are the tempting "fixes" that actually break determinism. Do **not** apply them:
-
-1. **Removing `outputHash`** from fixed-output derivations (e.g.
- `runners/cpython/deps/ffi/default.nix`). The repo literally warns: *"the moment
- you delete this you should question yourself."* `outputHash` is what pins the
- artifact; deleting it destroys reproducibility.
-2. **Adding a platform-native linker** path for Darwin.
-3. **Patching `ar`** to work on macOS.
-4. Any change that lets runners build on `aarch64-darwin` / `x86_64-darwin`
- instead of Linux.
-5. **Editing `runners/support/versions/current.nix` hashes** to make a macOS
- build pass. A change whose only purpose is "make it build on macOS" must
- **not** touch those hashes — they pin the Linux reference artifacts. If a
- macOS build forces you to change a hash there, the build is producing the
- wrong bytes; fix the builder (use the Linux remote builder), not the hash.
-6. **Removing / pruning old runner generations** (existing `.tar` entries or
- prior hash versions). A macOS-build-only change must leave the existing set
- of runner generations intact — silently dropping old ones is exactly the
- determinism-breaking regression that gets such changes rejected.
-
-Per the maintainer: runners can only be built on Linux (x86_64). On macOS you
-**download** prebuilt runners or **build them on a remote Linux builder** — you do
-not build them locally.
-
-## The supported path: a remote Linux nix-builder
-
-Build Linux derivations on Linux, from your Mac, over SSH. The repo ships a
-ready-made containerized Linux remote builder:
-
-> **`support/macos/nix-builder/`** — see its `README.md` for the full, current
-> procedure. Read that file; do not improvise.
-
-In short (read the README for exact, up-to-date steps):
-
-1. Install Rosetta (`softwareupdate --install-rosetta --agree-to-license`) and
- Docker Desktop with **"Use Rosetta for x86_64/amd64 emulation"** enabled.
-2. `docker build -t nix-builder support/macos/nix-builder` (add
- `--platform linux/arm64` if needed), drop your `id_ed25519.pub` into the
- volume, and `docker run` it (privileged, port `2222`).
-3. Register it in `/etc/nix/machines` and set `builders-use-substitutes = true`
- in `/etc/nix/nix.conf`. Add the host key to `known_hosts` **for root too**.
-4. Sanity-check with
- `NIX_REMOTE="ssh-ng://root@localhost:2222" nix build --system x86_64-linux nixpkgs#hello`.
-
-Notes that bite people (all in the README): use Nix `2.34.7`+ (older `2.2x`
-mishandles `ssh-ng://host:port`), and the root user — not just your user — must
-trust the builder's host key.
-
-## If you just need to develop / run, not rebuild runners
-
-You usually do not need to build runners at all. Download them instead (works on
-any platform):
-
-```bash
-nix develop .#full --command python3 build/out/bin/post-install.py \
- --create-venv false --default-step false \
- --runners-download true --error-on-missing-executor false
-```
-
-See the `build` skill for the normal Rust build flow.
diff --git a/.claude/skills/pydoc b/.claude/skills/pydoc
new file mode 120000
index 00000000..e26ee091
--- /dev/null
+++ b/.claude/skills/pydoc
@@ -0,0 +1 @@
+../../.agents/skills/pydoc
\ No newline at end of file
diff --git a/.claude/skills/review-ready b/.claude/skills/review-ready
new file mode 120000
index 00000000..0788c6f2
--- /dev/null
+++ b/.claude/skills/review-ready
@@ -0,0 +1 @@
+../../.agents/skills/review-ready
\ No newline at end of file
diff --git a/.claude/skills/rust-test-style b/.claude/skills/rust-test-style
new file mode 120000
index 00000000..eaf94312
--- /dev/null
+++ b/.claude/skills/rust-test-style
@@ -0,0 +1 @@
+../../.agents/skills/rust-test-style
\ No newline at end of file
diff --git a/.claude/skills/rust-test-style/SKILL.md b/.claude/skills/rust-test-style/SKILL.md
deleted file mode 100644
index 8eaa1851..00000000
--- a/.claude/skills/rust-test-style/SKILL.md
+++ /dev/null
@@ -1,192 +0,0 @@
----
-name: rust-test-style
-description: GenVM Rust test conventions and style. Use when writing, adding, or modifying Rust tests — inline `#[cfg(test)] mod tests`, integration tests under a crate's `tests/`, helpers, assertions, and how `genvm-tool test` discovers them.
----
-
-# Writing Rust tests in GenVM
-
-Standard library testing only. Do **not** add `rstest`, `proptest`, `insta`, or
-`pretty_assertions` — they are not used anywhere in the repo. Plain `#[test]`,
-`#[cfg(test)] mod tests`, and `std` assertions are the whole toolkit. `tokio` is
-available with the `macros` feature, so use `#[tokio::test]` when a test must be
-async (rare; no async tests exist today).
-
-To run these tests, see the `/test` skill (`--filter-tag rust`).
-
-## Where tests go
-
-**Decision rule — default to a separate file.** If the item under test is
-reachable from the crate's public API (a `pub fn`/`pub` type, even via a
-re-export), its test goes in its own `tests/.rs` file. Use an inline
-`#[cfg(test)] mod tests` **only** when the test must reach a *private* item that
-cannot be exercised through the public API. A public function tested inline is a
-review finding — move it to `tests/`.
-
-Two locations, both in use:
-
-1. **Integration tests** — one file per concern under the crate's `tests/` dir.
- **This is the default** for anything testable through the public API. Each
- `tests/*.rs` file is its own compilation unit (a separate crate) and gets its
- own runner case.
- - e.g. `executor/crates/calldata/tests/derive_decode.rs`,
- `executor/crates/calldata/tests/address_checksum.rs`,
- `executor/crates/common/tests/expr.rs`,
- `modules/implementation/tests/test_rat.rs`
- - Because each is a separate crate, it sees only the library crate plus
- `[dev-dependencies]` — **not** the library's regular `[dependencies]`. If a
- test needs a util the crate already depends on (e.g. `hex`), add it to
- `[dev-dependencies]` too.
-2. **Inline unit tests** — `#[cfg(test)] mod tests { ... }` at the bottom of a
- `src/*.rs` file, reserved for testing **private** items. e.g.
- `executor/crates/calldata/src/lib.rs`, `executor/crates/common/src/logger/mod.rs`.
-
-```rust
-#[cfg(test)]
-mod tests {
- use super::*;
-
- #[test]
- fn test_nested_value_in_struct() { /* ... */ }
-}
-```
-
-## Helpers over fixtures
-
-There is no fixture framework. Factor repeated setup into small free functions
-at the **top of the test file**, then keep each `#[test]` short. This is the
-dominant pattern.
-
-```rust
-// executor/crates/common/tests/expr.rs:5
-fn eval(input: &str) -> Value {
- Expr::parse(input).unwrap().evaluate().unwrap()
-}
-
-fn assert_rational(val: Value, n: i64, d: i64) {
- let r = val.into_rational().unwrap();
- assert_eq!(r, BigRational::new(BigInt::from(n), BigInt::from(d)));
-}
-```
-
-```rust
-// modules/implementation/tests/test_rat.rs — construct the system under test
-fn make_vm() -> Lua {
- let vm = Lua::new();
- rat::register_rat_global(&vm).unwrap();
- vm
-}
-fn eval(vm: &Lua, code: &str) -> T {
- vm.load(code).eval::().unwrap()
-}
-```
-
-Define throwaway `#[derive(...)]` types as fixtures right in the test file:
-
-```rust
-// executor/crates/calldata/tests/derive_decode_excess_fields.rs:11
-#[derive(Debug, PartialEq, Decode)]
-struct Named { x: i32, y: String }
-```
-
-## Assertions
-
-`std` only: `assert!`, `assert_eq!`, `assert_ne!`. When asserting a boolean
-condition, add a format-string message that prints the **actual value** so a
-failure is debuggable:
-
-```rust
-let msg = err.to_string();
-assert!(msg.contains("unknown field"), "unexpected error: {msg}");
-assert!(msg.contains("`z`"), "should mention field name: {msg}");
-```
-
-For value equality use `assert_eq!` against a fully-constructed expected value
-(derive `Debug, PartialEq` on the type).
-
-## Testing errors
-
-Use a helper that returns `Result`, call `.unwrap_err()`, and assert on
-substrings of the `Display` message — don't match on error enum internals:
-
-```rust
-// executor/crates/calldata/tests/derive_decode_excess_fields.rs:5
-fn try_decode_from_value(val: Value) -> Result {
- T::decode(codec::ValueDeserializer(val))
-}
-
-#[test]
-fn tuple_struct_rejects_too_many_elements() {
- let err = try_decode_from_value::(/* 3 elems */).unwrap_err();
- assert!(err.to_string().contains("expected 2"), "unexpected error: {err}");
-}
-```
-
-The happy path is liberal with `.unwrap()` — a panic *is* the test failure.
-
-## Roundtrip pattern
-
-For codecs, encode → decode → compare in one helper:
-
-```rust
-// executor/crates/calldata/tests/derive_decode.rs:5
-fn roundtrip_via_value(val: &T) -> Value
-where T: for<'a> codec::Encode<&'a mut Vec, Error = std::convert::Infallible> {
- let mut buf = Vec::new();
- codec::Encode::encode(val, &mut Encoder::new(&mut buf)).unwrap();
- genlayer_calldata::decode(&buf).unwrap()
-}
-```
-
-## Naming & layout
-
-- Test fn names are descriptive `snake_case` that name the **scenario**, e.g.
- `named_struct_rejects_unknown_field`, `string_literals_and_escapes`.
-- **Do not prefix with `test_`.** The `#[test]` attribute already says it is a
- test; the prefix is noise. Drop it from new tests and from any you touch.
-- Group related tests with a section-comment banner:
- ```rust
- // ── Tuple struct: wrong sequence length ─────────────────────────────
- ```
-- Helpers first, then types, then `#[test]` functions.
-
-## How `genvm-tool test` discovers Rust tests
-
-Discovery lives in `tests/plugins/genvm_tool_plugins/cargo.py`. For each
-registered Rust crate root it creates cases via plain `cargo test`:
-
-| Source | Command | Tags |
-|--------|---------|------|
-| each `tests/*.rs` file | `cargo test --test ` | `rust`, `unit` |
-| crate has `src/lib.rs` (inline `#[cfg(test)]`) | `cargo test --lib` | `rust`, `unit` |
-| each binary (`src/main.rs` / `[[bin]]`) | `cargo test --bin ` | `rust`, `unit` |
-| each `[[example]]` | `cargo check --example ` (compile-only) | `rust`, `example` |
-
-Consequences when adding tests:
-
-- **A new file in `tests/` is auto-discovered** — no registration needed, as long
- as the crate root itself is already scanned.
-- **New inline `#[cfg(test)]` tests** ride along on the existing `--lib` case;
- nothing to register.
-- The crate must be a known Rust root. Roots carry a `.ya-test-config.json`
- (e.g. `executor/.ya-test-config.json`, `executor/crates/calldata/.ya-test-config.json`).
- That file controls per-crate `cargo_test_flags`, `keep_env`, and a `skip` map
- (skip `examples`/specific cases) — add a new crate's config there if introducing
- a new root.
-- The `rust` preset is `(rust | integration) & !bench & !fuzz`
- (`tests/presets/rust.txt`).
-
-## Prefer fuzzing when it fits
-
-If a function's correctness is really about holding over a **space of inputs**
-(parsers, codecs, encode/decode roundtrips, anything taking arbitrary bytes or
-structured input), write a **fuzz target** rather than a handful of hand-picked
-`#[test]` cases. A few cherry-picked examples give false confidence; a fuzzer
-explores the space. Don't settle for an under-tested `#[test]` when the property
-is fuzz-shaped — reach for fuzz, or do both (fuzz for coverage, a couple of
-`#[test]`s pinning known edge cases / regressions).
-
-Fuzz targets live under a crate's `fuzz/` (afl, behind the `arbitrary` feature,
-tagged `fuzz`, built with `cargo_afl_build_flags`; see `cargo_fuzz` in
-`tests/plugins/genvm_tool_plugins/cargo.py`). They are a distinct case type —
-don't mix a fuzz harness into `tests/` or `mod tests`, and they're excluded from
-the `rust` preset (`!fuzz`).
diff --git a/.claude/skills/spec b/.claude/skills/spec
new file mode 120000
index 00000000..acb2bb1b
--- /dev/null
+++ b/.claude/skills/spec
@@ -0,0 +1 @@
+../../.agents/skills/spec
\ No newline at end of file
diff --git a/.claude/skills/submodules b/.claude/skills/submodules
new file mode 120000
index 00000000..bbcef097
--- /dev/null
+++ b/.claude/skills/submodules
@@ -0,0 +1 @@
+../../.agents/skills/submodules
\ No newline at end of file
diff --git a/.claude/skills/test b/.claude/skills/test
new file mode 120000
index 00000000..20211d17
--- /dev/null
+++ b/.claude/skills/test
@@ -0,0 +1 @@
+../../.agents/skills/test
\ No newline at end of file
diff --git a/.claude/skills/test/SKILL.md b/.claude/skills/test/SKILL.md
deleted file mode 100644
index 851ac024..00000000
--- a/.claude/skills/test/SKILL.md
+++ /dev/null
@@ -1,150 +0,0 @@
----
-name: test
-description: Runs tests for the GenVM project. Use after making code changes to verify correctness.
----
-
-# Running Tests
-
-GenVM uses `genvm-tool test` for all tests. Before running tests, ensure the project is built (see `/build` skill).
-
-## Quick Start
-
-Run all tests:
-```bash
-nix develop .#mock-tests --command genvm-tool test run
-```
-
-Run release tests (stable integration):
-```bash
-nix develop .#mock-tests --command genvm-tool test run --filter-tag "$(cat tests/presets/release.txt)"
-```
-
-Run a specific test:
-```bash
-nix develop .#mock-tests --command genvm-tool test run --filter-name 'test_name'
-```
-
-## genvm-tool test Commands
-
-### Run Tests
-```bash
-genvm-tool test run [OPTIONS]
-```
-
-**Options:**
-| Flag | Description |
-|------|-------------|
-| `--filter-name REGEX` | Filter tests by name regex |
-| `--filter-tag EXPR` | Filter tests by tags (e.g., `stable & !slow`) |
-| `--filter-continue FILE` | Re-run only tests from a continue file |
-| `--fail-fast` | Stop execution after first failure |
-| `--coverage` | Enable coverage collection for Rust tests |
-| `--log-level LEVEL` | Set log level (trace/debug/info/warning/error) |
-| `--ignore-hash` | Skip `.hash` (execution-hash) comparison entirely |
-
-**`--ignore-hash`:** integration cases compare both the printed semantics
-(`.N.stdout`) and a deterministic execution hash (`.N.hash`). When adding a new
-case, you usually don't have a correct `.hash` yet — run with `--ignore-hash` so
-only the `.stdout` semantics are checked, or set `stable_hash: false` on the
-entry (then validators compare to the leader's hash instead of a committed file).
-
-### Show Information
-```bash
-# Show available tests
-genvm-tool test show test
-
-# Show execution plan
-genvm-tool test show plan
-
-# Show available services
-genvm-tool test show services
-
-# Show available tags
-genvm-tool test show tags
-```
-
-## Test Presets
-
-Presets are tag expressions stored in `tests/presets/`:
-
-| Preset | Expression | Use Case |
-|--------|------------|----------|
-| `release.txt` | `integration & stable` | CI release tests |
-| `rust.txt` | `rust \| integration` | Rust development |
-| `python.txt` | `python` | Python SDK tests |
-
-Usage:
-```bash
-genvm-tool test --filter-tag "$(cat tests/presets/release.txt)" run
-```
-
-## Test Categories
-
-### Integration Tests (`tests/cases/`)
-End-to-end tests using jsonnet configuration. Services (manager, modules, webdriver) are started automatically.
-
-```bash
-nix develop .#mock-tests --command genvm-tool test run --filter-tag integration
-```
-
-### Rust Tests
-Cargo tests for Rust crates:
-```bash
-nix develop .#rust-test --command genvm-tool test run --filter-tag rust
-```
-
-With coverage:
-```bash
-nix develop .#rust-test --command genvm-tool test run --filter-tag rust --coverage
-```
-
-### Python Tests
-Tests for the Python standard library (`genlayer-py-std`):
-```bash
-nix develop .#mock-tests --command genvm-tool test run --filter-tag python
-```
-
-Or directly with pytest from the standalone py-test flake (no nix develop needed):
-```bash
-cd runners/genlayer-py-std
-env_dir="$(nix build --no-link --print-out-paths path:../../support/nix/py-test)"
-PYTHONPATH="$PWD/src:$PWD/src-emb" "$env_dir/bin/pytest" tests/
-```
-
-**Coverage:** pytest is configured with `--cov` and `--cov-fail-under=75`. Coverage scope includes `genlayer.types`, `genlayer.calldata`, `genlayer.storage`, `genlayer.evm`, `genlayer._internal`, and `genlayer_embeddings`. Must run inside nix develop for numpy-dependent tests and correct coverage resolution.
-
-## Webdriver Setup
-
-For web-related tests (semi-stable/unstable), webdriver is started automatically by genvm-tool test. To manually start it:
-
-```bash
-bash modules/webdriver/build-and-run.sh
-```
-
-## Precompile (Optional)
-
-If WASM files or compilation changed, precompile to save test time:
-```bash
-./build/out/bin/genvm precompile
-```
-
-## Re-running Failed Tests
-
-When tests fail, genvm-tool test writes failed test names to `build/test-artifacts/continue/-`. Re-run only failed tests:
-
-```bash
-# Use filename shown in failure summary
-genvm-tool test run --filter-continue 20260123-143052-abc123
-```
-
-## Quick Reference
-
-| What to test | Command |
-|--------------|---------|
-| All tests | `nix develop .#mock-tests --command genvm-tool test run` |
-| Release tests | `nix develop .#mock-tests --command genvm-tool test run --filter-tag "$(cat tests/presets/release.txt)"` |
-| Rust tests | `nix develop .#rust-test --command genvm-tool test run --filter-tag rust` |
-| Python (direct) | `cd runners/genlayer-py-std && PYTHONPATH="$PWD/src:$PWD/src-emb" "$(nix build --no-link --print-out-paths path:../../support/nix/py-test)/bin/pytest" tests/` |
-| Re-run failed | `genvm-tool test run --filter-continue ` |
-| With debug logs | `nix develop .#mock-tests --command genvm-tool test run --log-level debug` |
-| Show test list | `nix develop .#mock-tests --command genvm-tool test show test` |
diff --git a/.coderabbit.yaml b/.coderabbit.yaml
index bf1e8b32..53a5fd5f 100644
--- a/.coderabbit.yaml
+++ b/.coderabbit.yaml
@@ -1,4 +1,21 @@
+# yaml-language-server: $schema=https://coderabbit.ai/integrations/schema.v2.json
reviews:
+ request_changes_workflow: true
+ auto_review:
+ # `base_branches` MUST live here, under `auto_review` — not directly under
+ # `reviews`. The schema defines no `reviews.base_branches`, and `reviews`
+ # does not forbid extra keys, so a misplaced copy is silently ignored:
+ # CodeRabbit reads the file, falls back to default-branch-only, and skips
+ # every PR targeting a dev branch with "Auto reviews are disabled on
+ # base/target branches other than the default branch".
+ #
+ # Auto-review PRs targeting any long-lived branch: main, a release branch
+ # or its dev branch. Feature/topic branches are reviewed through the base
+ # they target. Patterns are regexes.
+ base_branches:
+ - "main"
+ - "v\\d+\\.\\d+"
+ - "v\\d+\\.\\d+-dev"
path_filters:
- "!**/*.onnx"
- "!**/*.txt"
@@ -6,7 +23,10 @@ reviews:
- "!**/*.hash"
- "!**/*.lock"
- "!docs/website/src/_static/**"
- - "!executors/v0.3.x/runners/py-libs/pure-py/**"
- - "!executors/v0.3.x/runners/models/**"
- - "!executors/v0.3.x/runners/softfloat/berkeley-softfloat-3/**"
- - "!**/fuzz/inputs*/**"
+ - "!**/fuzz/**"
+ - "!**/*_test.rs"
+ - "!**/*.json"
+ - "!**/tests/**"
+ - "!**/test/**"
+ # genvm-tool is internal dev tooling; not worth reviewing.
+ - "!support/tools/genvm-tool/**"
diff --git a/.genvm-monorepo-root b/.genvm-monorepo-root
index d569ad58..78cefacd 100644
--- a/.genvm-monorepo-root
+++ b/.genvm-monorepo-root
@@ -1,6 +1,15 @@
{
- "version": "v0.6.0-rc0",
- "active-versions": ["v0.3"],
+ "version": "v0.6.0-rc2",
+ "active-versions": [
+ "v0.3",
+ "v0.2"
+ ],
+ "support-only-versions": [
+ "v0.2"
+ ],
"artifacts_dir": "build/test-artifacts",
- "extra_python_paths": ["tests/runner"]
+ "tags_registry": "tests/tags.json",
+ "extra_python_paths": [
+ "tests/runner"
+ ]
}
diff --git a/.genvm-tool.py b/.genvm-tool.py
index 4c7592ce..ffdb2cae 100644
--- a/.genvm-tool.py
+++ b/.genvm-tool.py
@@ -1,234 +1,80 @@
-"""Manager-root genvm-tool project config.
+"""
+Manager-root genvm-tool project config.
Loaded once by genvm-tool (``common.load_project``) before any subcommand runs;
subcommands ask it for what they need:
-- ``hooks(ctx)`` — the manager's commit-hook definitions (was
- ``support/nix/precommit/hooks.toml``), returned in the same shape the hook
- engine already understands.
-- ``tests(ctx)`` — the umbrella test suite (was the top-level ``.ya-test.py``):
+- ``tests(ctx)`` -- the umbrella test suite (was the top-level ``.ya-test.py``):
imports the plugins it needs and registers collectors + CLI args on the
runner's configuration ``Context``. This is the plugins' "initial run" hook;
plugins themselves are plain importable modules (no test-runner-specific
registry). Per-suite test definitions live next to the tests in
``tests/system//test.py`` and are pulled in via ``ctx.collect_dir``.
- Heavy imports stay inside this function — they need the plugin search path
+ Heavy imports stay inside this function -- they need the plugin search path
(``extra_python_paths`` in ``.genvm-monorepo-root``), which the test command
applies before calling us.
General runner config (``artifacts_dir`` / ``extra_python_paths``) lives in
-``.genvm-monorepo-root``. Each executor carries its own ``.genvm-tool.py`` with
-its own ``hooks``.
+``.genvm-monorepo-root``. Each executor carries its own ``.genvm-tool.py``.
+Commit hooks live in each repo's flake (git-hooks.nix), not here.
"""
-def hooks(ctx):
- """Manager commit-hook definitions (was support/nix/precommit/hooks.toml).
-
- Tools resolve from the sibling flake's buildEnv (``nix = ""``);
- ``local`` hooks run a repo-owned script and ``builtin`` hooks run logic baked
- into genvm-tool. A hook's ``args`` are its check-mode invocation; an optional
- ``fix_args`` is the rewrite-in-place invocation used when ``hook run`` fixes
- (the default — CI passes ``--check``). ``files``/``exclude`` are repo-relative
- regexes.
- """
- return [
- # --- generic checks (mirror the executor's) ------------------------
- {
- 'id': 'trailing-whitespace',
- 'nix': 'pre-commit-hooks',
- 'entry': 'trailing-whitespace-fixer',
- 'types_or': ['text'],
- 'exclude': r'^\.git-third-party|/fuzz/',
- },
- {
- 'id': 'end-of-file-fixer',
- 'nix': 'pre-commit-hooks',
- 'entry': 'end-of-file-fixer',
- 'types_or': ['text'],
- 'exclude': r'^\.git-third-party|/fuzz/',
- },
- {
- 'id': 'check-added-large-files',
- 'nix': 'pre-commit-hooks',
- 'entry': 'check-added-large-files',
- },
- {
- 'id': 'check-json',
- 'nix': 'pre-commit-hooks',
- 'entry': 'check-json',
- 'types_or': ['json'],
- 'exclude': r'(^\.git-third-party)|(/tsconfig\.json$)',
- },
- {
- 'id': 'check-yaml',
- 'nix': 'pre-commit-hooks',
- 'entry': 'check-yaml',
- 'types_or': ['yaml'],
- },
- {
- 'id': 'check-toml',
- 'nix': 'pre-commit-hooks',
- 'entry': 'check-toml',
- 'types_or': ['toml'],
- },
- {
- 'id': 'check-merge-conflict',
- 'nix': 'pre-commit-hooks',
- 'entry': 'check-merge-conflict',
- 'types_or': ['text'],
- },
- {
- # Matches both the hidden vendored `.git-third-party` trees and the
- # `support/tools/git-third-party` tool dir (vendored LICENSE etc.).
- 'id': 'editorconfig-checker',
- 'nix': 'editorconfig-checker',
- 'entry': 'editorconfig-checker',
- # text-only: the engine's `text` pseudo-type filters out binaries
- # (e.g. model/onnx blobs) that editorconfig-checker should not see.
- 'types_or': ['text'],
- 'exclude': r'git-third-party|/fuzz/',
- },
- # --- python / lua / ts (manager-owned languages) -------------------
- {
- 'id': 'ruff-format',
- 'nix': 'ruff',
- 'entry': 'ruff',
- 'args': ['format', '--check'],
- 'fix_args': ['format'],
- 'types_or': ['python'],
- },
- # --- python / lua / ts (manager-owned languages) -------------------
- {
- 'id': 'ruff-check',
- 'nix': 'ruff',
- 'entry': 'ruff',
- 'args': ['check'],
- 'fix_args': ['check', '--fix'],
- 'types_or': ['python'],
- },
- {
- 'id': 'stylua',
- 'nix': 'stylua',
- 'entry': 'stylua',
- 'args': ['--check'],
- 'fix_args': [],
- 'files': r'\.lua$',
- },
- {
- 'id': 'prettier',
- 'nix': 'prettier',
- 'entry': 'prettier',
- 'args': ['--check'],
- 'fix_args': ['--write'],
- 'types_or': ['ts', 'tsx'],
- 'exclude': r'^\.git-third-party',
- },
- {
- 'id': 'nixfmt',
- 'nix': 'nixfmt',
- 'entry': 'nixfmt',
- 'args': ['--check'],
- 'fix_args': [],
- 'files': r'\.nix$',
- },
- # --- github workflow/action schemas -------------------------------
- {
- 'id': 'check-github-workflows',
- 'nix': 'check-jsonschema',
- 'entry': 'check-jsonschema',
- 'args': ['--builtin-schema', 'vendor.github-workflows'],
- 'files': r'^\.github/workflows/.*\.ya?ml$',
- },
- {
- 'id': 'check-github-actions',
- 'nix': 'check-jsonschema',
- 'entry': 'check-jsonschema',
- 'args': ['--builtin-schema', 'vendor.github-actions'],
- 'files': r'^\.github/(actions/.+/)?action\.ya?ml$',
- },
- # --- local scripts -------------------------------------------------
- {
- 'id': 'cargo-fmt',
- 'nix': 'cargo',
- 'builtin': 'cargo-fmt',
- 'files': r'\.rs$',
- 'pass_filenames': False,
- },
- {
- # Keep the manager crate's [package] version in lockstep with
- # .genvm-monorepo-root. The executor submodule is versioned
- # independently (its own repo/hooks own that); we never touch it.
- 'id': 'check-cargo-versions',
- 'local': True,
- 'entry': 'support/ci/check-versions.py',
- 'args': ['sync'],
- 'pass_filenames': False,
- 'files': r'^(implementation/Cargo\.toml|\.genvm-monorepo-root)$',
- },
- {
- 'id': 'markdown-local-links',
- 'builtin': 'md-local-links',
- 'types_or': ['markdown'],
- },
- ]
-
-
def tests(ctx):
- """Umbrella test suite (was the top-level .ya-test.py).
+ """
+ Umbrella test suite (was the top-level .ya-test.py).
``ctx`` is the runner's configuration ``Context``. Collectors close over the
imports below (resolved lazily when collection runs).
"""
import json
+ import platform
import sys
from pathlib import Path
+ import genvm_tool.cmd_configure
import genvm_tool.tests
_info_path = ctx.shared.root_dir / 'build' / 'info.json'
if not _info_path.exists():
+ # CI runs tests without configuring first. `configure` writes this same
+ # base plus what the build teaches it, and both call base_info, so the
+ # two writers cannot disagree about the tree.
ctx.shared.logger.warning('build/info.json not found, generating default')
_build_dir = ctx.shared.root_dir / 'build'
_build_dir.mkdir(parents=True, exist_ok=True)
+ _monorepo_cfg = json.loads(
+ (ctx.shared.root_dir / genvm_tool.cmd_configure.MONOREPO_ROOT_FILE).read_text()
+ )
_info_path.write_text(
json.dumps(
- {
- 'coverage_dir': str(_build_dir / 'cov'),
- 'build_dir': str(_build_dir),
- 'rust_target_dir': str(_build_dir / 'ya-build' / 'rust-target'),
- },
+ genvm_tool.cmd_configure.base_info(
+ _monorepo_cfg,
+ ctx.shared.root_dir,
+ _build_dir,
+ _build_dir / 'ya-build' / 'rust-target',
+ ),
indent=2,
)
+ '\n'
)
- ctx.run_parser.add_argument(
- '--fuzz-timeout',
- type=int,
- default=30,
- help='Timeout for each fuzzing run in seconds',
- )
-
- ctx.run_parser.add_argument(
- '--fuzz-update-corpus',
- default=False,
- action='store_true',
- help='Whether to update the fuzzing corpus',
- )
-
import genvm_tool_plugins
ctx.shared.logger.trace(
'import path', path=sys.path, plugins_path=genvm_tool_plugins.__path__
)
from genvm_tool_plugins import (
+ afl,
cargo,
genvm,
integration,
+ npm,
pytest,
)
+ afl.register(ctx)
+
def collect_rust(ctx: genvm_tool.tests.stage.collection.Context):
for t in filter(lambda x: x.name == 'Cargo.toml', ctx.shared.git_files):
ctx.shared.logger.debug('discovered Cargo.toml', path=t)
@@ -248,42 +94,63 @@ def collect_rust(ctx: genvm_tool.tests.stage.collection.Context):
name = f'{name.parent}/{name.stem}'
cargo.cargo_fuzz(
ctx,
- genvm_tool.tests.test.Description(
- name,
- console_pool=True,
- ),
+ genvm_tool.tests.test.Description(name),
rust_root_dir=rust_root_dir,
name=fuzz_file.stem,
)
- def collect_pytest(ctx: genvm_tool.tests.stage.collection.Context):
- p = ctx.shared.root_dir.joinpath(
- 'executors', 'v0.3.x', 'runners', 'genlayer-py-std'
- )
- pytest.pytest(
- ctx,
- genvm_tool.tests.test.Description(
- 'runners/genlayer-py-std/test',
- ),
- project_root_dir=p,
- )
+ def collect_npm(ctx: genvm_tool.tests.stage.collection.Context):
+ test_files = sorted(f for f in ctx.shared.git_files if f.name.endswith('.test.ts'))
+ for t in filter(lambda x: x.name == 'package.json', ctx.shared.git_files):
+ if 'test' not in npm.scripts(t):
+ continue
+ project_root_dir = t.parent
+ owned = [f for f in test_files if f.is_relative_to(project_root_dir)]
+ ctx.shared.logger.debug('discovered npm project', path=t, test_files=len(owned))
+ npm.npm_project(
+ ctx,
+ project_root_dir=project_root_dir,
+ test_files=owned,
+ )
- fuzz_files = list(p.glob('fuzz/src/*.py'))
- fuzz_files.sort()
- for fuzz_file in fuzz_files:
- name = fuzz_file.relative_to(ctx.shared.root_dir)
- name = f'{name.parent}/{name.stem}'
- continue # for now let's disable it
- pytest.py_fuzz(
+ def collect_pytest(ctx: genvm_tool.tests.stage.collection.Context):
+ dirs = [
+ ctx.shared.root_dir.joinpath('executors', 'v0.3.x', 'runners', 'genlayer-py-std'),
+ ctx.shared.root_dir.joinpath('support', 'tools', 'genvm-tool'),
+ ctx.shared.root_dir.joinpath('support', 'ci'),
+ ]
+ for p in dirs:
+ pytest.pytest(
ctx,
genvm_tool.tests.test.Description(
- name,
+ str(p.relative_to(ctx.shared.root_dir)),
+ console_pool=True,
+ tags=frozenset({'unit'}),
),
project_root_dir=p,
- name=fuzz_file.stem,
)
+ # A project's flake carries python-afl and AFL++ on x86_64-linux only,
+ # so elsewhere the case could not start at all
+ if (platform.system(), platform.machine()) != ('Linux', 'x86_64'):
+ continue
+
+ fuzz_files = list(p.glob('fuzz/src/*.py'))
+ fuzz_files.sort()
+ for fuzz_file in fuzz_files:
+ name = fuzz_file.relative_to(ctx.shared.root_dir)
+ name = f'{name.parent}/{name.stem}'
+ pytest.py_fuzz(
+ ctx,
+ genvm_tool.tests.test.Description(
+ name,
+ ),
+ project_root_dir=p,
+ name=fuzz_file.stem,
+ )
+
ctx.add_collector(collect_rust)
+ ctx.add_collector(collect_npm)
ctx.add_collector(collect_pytest)
ctx.run_parser.add_argument(
@@ -322,15 +189,20 @@ def collect_integration(ctx: genvm_tool.tests.stage.collection.Context):
# per-request only under debug_mode >= safe (tests run with unsafe).
reroute_to = (
getattr(ctx.configuration.args, 'genvm_reroute_to', '')
- or build_info['primary_executor_version']
+ or build_info.get('primary_executor_version')
+ or ctx.shared.config['active-versions'][0]
)
no_manager = getattr(ctx.configuration.args, 'no_manager', False)
no_webdriver = getattr(ctx.configuration.args, 'no_webdriver', False)
+ ci = ctx.shared.ci
manager_port = genvm.get_manager_port(ctx.configuration)
+ manager_semaphore = ctx.new_semaphore('manager-listener')
if no_manager:
- manager_impl = genvm.ExternalManagerService(port=manager_port)
+ manager_impl = genvm.ExternalManagerService(
+ port=manager_port,
+ )
webdriver_impl = genvm.NoOpService()
modules_impl = genvm.NoOpService()
else:
@@ -338,13 +210,14 @@ def collect_integration(ctx: genvm_tool.tests.stage.collection.Context):
bin_path=build_dir.joinpath('out', 'bin', 'genvm-modules'),
log_path=tests_output_root.joinpath('manager.log'),
env=ctx.configuration,
+ ci=ci,
)
# Create webdriver service
if no_webdriver:
webdriver_impl = genvm.NoOpService()
else:
webdriver_impl = genvm_tool.tests.exec.service.FunctionService(
- lambda: genvm.start_webdriver_service(ctx.configuration)
+ lambda: genvm.start_webdriver_service(ctx.configuration, ci=ci)
)
# This starts Llm and Web modules on the manager
modules_impl = genvm.ModulesService(
@@ -354,8 +227,12 @@ def collect_integration(ctx: genvm_tool.tests.stage.collection.Context):
manager_service = ctx.new_service(
name='manager',
manager=manager_impl,
+ semaphores=[manager_semaphore],
)
- manager_service.meta = {'port': manager_port, 'reroute_to': reroute_to}
+ manager_service.meta = {
+ 'port': manager_port,
+ 'reroute_to': reroute_to,
+ }
webdriver_service = ctx.new_service(
name='webdriver',
@@ -374,9 +251,50 @@ def collect_integration(ctx: genvm_tool.tests.stage.collection.Context):
manager_service=manager_service,
modules_service=modules_service,
webdriver_service=webdriver_service,
+ ci=ci,
+ )
+ integration.integration_test_directory(
+ ctx,
+ cases_dir=ctx.shared.root_dir / 'tests' / 'system' / 'cross-major' / 'cases',
+ executor_major=3,
+ reroute_to=build_info['executor_versions']['v0.3'],
+ save_hashes=False,
+ manager_service=manager_service,
+ modules_service=modules_service,
+ webdriver_service=webdriver_service,
+ ci=ci,
)
ctx.collect_dir('tests/system/permits', manager_service=manager_service)
+ limited_manager_service = None
+ if not no_manager:
+ limited_manager_service = ctx.new_service(
+ name='manager-cross-major-small-message-cap',
+ manager=genvm.ManagerService(
+ bin_path=build_dir / 'out' / 'bin' / 'genvm-modules',
+ log_path=tests_output_root / 'cross-major-small-message-cap.log',
+ env=ctx.configuration,
+ ci=ci,
+ config={'max_message_bytes': 1024 * 1024},
+ ),
+ semaphores=[manager_semaphore],
+ )
+ ctx.collect_dir(
+ 'tests/system/manager-socket',
+ manager_semaphore=manager_semaphore,
+ build_dir=build_dir,
+ ci=ci,
+ )
+ ctx.collect_dir(
+ 'tests/system/cross-major',
+ manager_service=manager_service,
+ limited_manager_service=limited_manager_service,
+ managed_manager=not no_manager,
+ )
+ ctx.collect_dir(
+ 'tests/system/cross-major-observability',
+ manager_service=manager_service,
+ )
ctx.add_collector(collect_integration)
@@ -384,3 +302,8 @@ def collect_parse_version(ctx: genvm_tool.tests.stage.collection.Context):
ctx.collect_dir('tests/system/parse_version')
ctx.add_collector(collect_parse_version)
+
+ def collect_make_zip(ctx: genvm_tool.tests.stage.collection.Context):
+ ctx.collect_dir('tests/system/make_zip')
+
+ ctx.add_collector(collect_make_zip)
diff --git a/.github/actions/get-src/action.yaml b/.github/actions/get-src/action.yaml
index 311ada57..5defdf19 100644
--- a/.github/actions/get-src/action.yaml
+++ b/.github/actions/get-src/action.yaml
@@ -2,13 +2,12 @@ name: GenVM get source
description: Checkout sources, optionally update submodules/third-party, and (optionally) set up Nix.
inputs:
with_nix:
- description: if should setup nix
+ description: >-
+ if the *job* needs nix after this action. Nix is also installed whenever
+ `third_party` is not `none`, so it can be present without this being set;
+ an installation cannot be undone.
required: false
default: "false"
- load_submodules:
- description: if should update submodules
- required: false
- default: "true"
third_party:
description: third-party modules to install
required: false
@@ -16,31 +15,56 @@ inputs:
github_token:
description: GitHub token for nix setup
required: false
+ nix_cache_pull_token:
+ description: >-
+ pull token for the GenLayer nix cache, which is private. A composite
+ action cannot read `secrets`, so the caller has to thread it through.
+ Empty leaves whatever credentials the runner already has, and none on a
+ clean runner — every nix build then starts from source.
+ required: false
+ default: ""
+ pr:
+ description: >-
+ PR number to diff against, overriding `github.event.pull_request.number`.
+ A caller invoked by `workflow_dispatch` has no pull_request payload but
+ does know the number, and without this its `changes` output would come
+ back empty — silently checking nothing.
+ required: false
+ default: ""
+outputs:
+ changes:
+ description: >-
+ JSON object mapping each repo (`.` for the manager, submodule path for
+ each executor) to {base_commit, branch_commit, has_changes}. Outside a
+ pull request base_commit == branch_commit and has_changes is false.
+ value: ${{ steps.changes.outputs.changes }}
runs:
using: composite
steps:
+ # Before the checkout: `get-all-git.py` realizes `git-third-party` from the
+ # flake, so a third-party update needs nix here whether or not the job asked
+ # for one. The action itself only touches the runner, never the tree.
+ - name: setup nix
+ if: ${{ inputs.with_nix == 'true' || inputs.third_party != 'none' }}
+ # Pinned at every call site of these actions, `# ` naming where the
+ # sha came from: a job re-run later has to run the action it first ran.
+ uses: genlayerlabs/github-actions/nix-setup@39b0a0d5e9bb27a1612d2e98b0f8509d31745157 # main
+ with:
+ github_token: ${{ inputs.github_token }}
+ cache_pull_token: ${{ inputs.nix_cache_pull_token }}
- name: checkout submodules
run: |
cd "$GITHUB_WORKSPACE"
git config --global user.email "ci@genlayerlabs.com"
git config --global user.name "CI worker"
- if [ "${{ inputs.load_submodules }}" == "true" ]
- then
- git submodule update --init --recursive --depth 1
- if [ "${{ inputs.third_party }}" != "none" ]
- then
- source env.sh
- git third-party update ${{ inputs.third_party }}
- fi
- fi
- shell: bash -ex {0}
- - name: setup nix
- if: ${{ inputs.with_nix == 'true' }}
- uses: cachix/install-nix-action@v31.6.2
- with:
- extra_nix_config: |
- experimental-features = nix-command flakes recursive-nix
- substituters = https://cache.nixos.org/ https://nix-community.cachix.org https://cache.nix.kp2pml30.moe
- trusted-public-keys = cache.nixos.org-1:6NCHdD59X431o0gWypbMrAURkbJ16ZPMQFGspcDShjY= nix-community.cachix.org-1:mB9FSh9qf2dCimDSUo8Zy7bkq5CX+/rkCWyvRCYg3Fs= cache.nix.kp2pml30.moe:9izVvysdKza7+bjw2gI/etkDVfEQy+kne2aU/TZpA5o=
- github_access_token: ${{ inputs.github_token }}
- install_url: https://releases.nixos.org/nix/nix-2.24.14/install
+ python3 support/scripts/get-all-git.py "--third-party=${{ inputs.third_party }}"
+ shell: bash -exo pipefail {0}
+ - name: compute per-repo changes
+ id: changes
+ shell: bash -exo pipefail {0}
+ env:
+ GITHUB_PR_NUMBER: ${{ inputs.pr || github.event.pull_request.number }}
+ GH_TOKEN: ${{ inputs.github_token }}
+ run: |
+ cd "$GITHUB_WORKSPACE"
+ python3 support/scripts/ci-changes.py --github-output
diff --git a/.github/actions/matrix-cell-step/action.yaml b/.github/actions/matrix-cell-step/action.yaml
new file mode 100644
index 00000000..3f3f3429
--- /dev/null
+++ b/.github/actions/matrix-cell-step/action.yaml
@@ -0,0 +1,22 @@
+name: GenVM run matrix cell step
+description: >-
+ Replay one key of a matrix cell's step mapping (from a plan-*-matrix pipeline).
+ Does nothing if that key holds no steps.
+inputs:
+ job_json:
+ description: "JSON {key: step list} mapping, as passed to test_cell.yaml"
+ required: true
+ key:
+ description: which key of job_json to replay
+ required: true
+runs:
+ using: composite
+ steps:
+ # Skip when the key holds no steps: indexing an empty list yields null, and
+ # null != null is false. The list is passed via env (not interpolated into
+ # the shell) so its quotes/brackets survive.
+ - if: ${{ fromJSON(inputs.job_json)[inputs.key][0] != null }}
+ shell: bash
+ env:
+ JOB_JSON: ${{ toJSON(fromJSON(inputs.job_json)[inputs.key]) }}
+ run: ./support/ci/run.sh pipeline test-cell --job-json "$JOB_JSON"
diff --git a/.github/actions/postmortem/action.yaml b/.github/actions/postmortem/action.yaml
new file mode 100644
index 00000000..574cf316
--- /dev/null
+++ b/.github/actions/postmortem/action.yaml
@@ -0,0 +1,134 @@
+name: GenVM postmortem
+description: >-
+ Report runner state after a cell's steps have run: disk, inodes, space by
+ path, memory, surviving processes, kernel OOM/IO errors, and nix substituter
+ config. Diagnoses failures that leave no trace in the failing step's own
+ output — disk/inode exhaustion, the OOM killer, and substituter errors all
+ surface as an unrelated-looking non-zero exit.
+
+ A composite action, not an `incl_*` reusable workflow, because a reusable
+ workflow is a separate job on a fresh runner: it would report a pristine
+ machine instead of the one that just failed.
+
+ Callers gate on `failure()` — a cancel's grace window is too short to spend
+ here — and add `continue-on-error: true`. Linux-only, matching every cell
+ that calls it. Keep double-brace expression syntax out of this file entirely:
+ the runner templates the whole manifest, description included, and rejects
+ workflow-only functions such as `failure()` there.
+runs:
+ using: composite
+ steps:
+ # Never fails the job: `set +e` keeps one probe's failure from skipping the
+ # rest, and the trailing `exit 0` swallows the last probe's status.
+ - shell: bash
+ run: |
+ set +e
+
+ # /dev/shm, not /tmp: scratch must not live on the filesystem being
+ # diagnosed, or inode/space exhaustion silently blanks every probe
+ # that redirects through it.
+ scratch="$(mktemp -d -p /dev/shm 2>/dev/null || mktemp -d)"
+ if [ -z "${scratch}" ] || [ ! -d "${scratch}" ]; then
+ echo "::warning::postmortem: no writable scratch dir; probes degraded"
+ scratch=""
+ fi
+
+ echo "::group::disk free (df -h)"
+ df -h
+ echo "::endgroup::"
+
+ echo "::group::inodes (df -i)"
+ df -i
+ echo "::endgroup::"
+
+ # /nix/store is deliberately absent: `du -s` emits a total only once an
+ # argument completes, so a store walk would eat the whole budget and
+ # starve every path after it. `df` above already covers that
+ # filesystem. `-k 10` because `du` wedged in D-state on a failing disk
+ # ignores SIGTERM, and this step is built for failing disks.
+ echo "::group::space by path (du, 60s budget)"
+ timeout -k 10 60 du -shx \
+ "${GITHUB_WORKSPACE}" \
+ "${RUNNER_TEMP}" \
+ "${HOME}/.cargo" \
+ "${HOME}/.rustup" \
+ >"${scratch:-/dev/null}/du" 2>/dev/null
+ du_rc=$?
+ [ "${du_rc}" -eq 124 ] && echo "(du timed out; totals below are partial)"
+ if [ -s "${scratch:-/nonexistent}/du" ]; then
+ sort -h "${scratch}/du"
+ else
+ echo "(du produced nothing; rc=${du_rc})"
+ fi
+ echo "::endgroup::"
+
+ # free/ps run after the failing step exited, so an OOM-killed hog is
+ # already gone and these render a healthy box. They catch leftover
+ # daemons and give a baseline; `dmesg` below is what sees the kill.
+ echo "::group::memory (post-step; see dmesg for OOM)"
+ free -h
+ swapon --show
+ echo "::endgroup::"
+
+ echo "::group::load"
+ uptime
+ nproc
+ echo "::endgroup::"
+
+ echo "::group::surviving processes by memory (post-step)"
+ # `-rss`, not `-%rss` — procps rejects the latter as an unknown sort
+ # specifier and prints only a usage message.
+ ps -eo pid,ppid,rss,pcpu,etime,comm --sort=-rss 2>/dev/null | head -n 15
+ echo "::endgroup::"
+
+ # The OOM killer reports here and nowhere else: a build that "fails"
+ # with a bare signal 9 and no diagnostic is almost always this.
+ # `sudo -n` so a runner without passwordless sudo fails fast instead of
+ # blocking on a password prompt.
+ echo "::group::kernel log (OOM / IO errors)"
+ if [ -z "${scratch}" ]; then
+ echo "(skipped: no scratch dir)"
+ elif ! sudo -n dmesg -T >"${scratch}/dmesg" 2>"${scratch}/dmesg.err"; then
+ # Distinguished, because "dmesg refused" and "the disk we are
+ # diagnosing ate the redirect" want opposite next steps.
+ echo "(dmesg unavailable: $(tr -d '\n' <"${scratch}/dmesg.err"))"
+ else
+ # Match into a file, then test it: `grep | tail` reports tail's
+ # status, so a no-match branch behind it could never fire.
+ grep -iE 'oom|killed process|out of memory|i/o error|ext4-fs error' \
+ "${scratch}/dmesg" >"${scratch}/oom"
+ if [ -s "${scratch}/oom" ]; then
+ tail -n 30 "${scratch}/oom"
+ else
+ echo "(nothing matched)"
+ fi
+ echo "::endgroup::"
+
+ echo "::group::kernel log (tail)"
+ tail -n 40 "${scratch}/dmesg"
+ fi
+ echo "::endgroup::"
+
+ # Aimed at substituter failures (a truncated NAR fails the job with no
+ # hint of which cache served it): the effective substituter list and
+ # retry knobs are what identify the culprit afterwards. Store *content*
+ # is not probed — nix builds under TMPDIR, not in the store, so a
+ # failed build leaves nothing there to find.
+ echo "::group::nix"
+ if command -v nix >/dev/null 2>&1; then
+ nix --version
+ nix config show 2>/dev/null \
+ | grep -E '^(substituters|extra-substituters|fallback|download-attempts|connect-timeout|stalled-download-timeout) '
+ if [ -d /nix/store ]; then
+ find /nix/store -maxdepth 1 -mindepth 1 -printf . 2>/dev/null \
+ | wc -c | sed 's/^/store entries: /'
+ else
+ echo "store entries: (no /nix/store)"
+ fi
+ else
+ echo "(no nix on PATH)"
+ fi
+ echo "::endgroup::"
+
+ [ -n "${scratch}" ] && rm -rf "${scratch}"
+ exit 0
diff --git a/.github/workflows/branch_executor_prs.yaml b/.github/workflows/branch_executor_prs.yaml
new file mode 100644
index 00000000..be5b979c
--- /dev/null
+++ b/.github/workflows/branch_executor_prs.yaml
@@ -0,0 +1,62 @@
+name: branch / open linked executor PRs
+
+# When a manager PR is opened (or updated) against a dev branch, open the matching
+# executor PR(s) in the genvm-executor repo. The manager release line is independent
+# of the executor lines it ships, so one manager PR fans out over EVERY active
+# executor line (.genvm-monorepo-root): for each, `pr//` -> -dev.
+# Each is listed on the manager PR as an `executor: ` line. The executor work
+# lives on line-namespaced mirror branches the developer pushed via
+# `genvm-tool git create-branches`. See `support/ci/run.sh tool open-executor-prs`.
+#
+# `synchronize` (a push to the PR head) re-runs it: executor mirror branches are
+# often pushed after the manager PR is opened, so the initial `opened` run finds
+# nothing to link. `edited` fires when the base branch changes, so a PR opened
+# against `main` and retargeted onto a dev branch (branch_retarget.yaml) also
+# runs — the `opened` event was filtered out because its base was still `main`.
+# The script is idempotent — it reuses any existing executor PR and upserts the
+# same marked comment — so re-running per event just fills in the links as the
+# branches appear.
+#
+# pull_request_target gives a manager-scoped write token for the comment; the
+# cross-repo executor PR needs a PAT with genvm-executor access
+# (GENVM_EXECUTOR_PR_TOKEN) — the default GITHUB_TOKEN cannot reach another repo.
+# We check out only the base repo to run the trusted script — never PR code.
+
+on:
+ pull_request_target:
+ branches: ['v*-dev']
+ types: [opened, synchronize, reopened, edited]
+
+# Serialize per PR so overlapping pushes don't race on creating the executor PR or
+# upserting the comment; don't cancel in-flight runs — each one is a cheap, safe
+# reconcile and the last state should reflect the latest head.
+concurrency:
+ group: executor-prs-${{ github.event.pull_request.number }}
+ cancel-in-progress: false
+
+permissions:
+ # Explicit, though this job only ever checks out the base repo: a `permissions`
+ # block sets every unlisted scope to `none`, so the checkout works today only
+ # because the repo is public. Naming it keeps that from being load-bearing.
+ contents: read
+ pull-requests: write
+ issues: write
+
+defaults:
+ run:
+ shell: bash -exo pipefail {0}
+
+jobs:
+ open:
+ runs-on: ubuntu-latest
+ timeout-minutes: 15
+ steps:
+ - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
+ - name: Open and link executor PR(s)
+ env:
+ MANAGER_TOKEN: ${{ secrets.GITHUB_TOKEN }}
+ EXECUTOR_TOKEN: ${{ secrets.GENVM_EXECUTOR_PR_TOKEN }}
+ MANAGER_REPO: ${{ github.repository }}
+ PR_NUMBER: ${{ github.event.pull_request.number }}
+ HEAD_REF: ${{ github.event.pull_request.head.ref }}
+ run: ./support/ci/run.sh tool open-executor-prs
diff --git a/.github/workflows/branch_forward.yaml b/.github/workflows/branch_forward.yaml
index 975a6e78..9603796a 100644
--- a/.github/workflows/branch_forward.yaml
+++ b/.github/workflows/branch_forward.yaml
@@ -29,13 +29,14 @@ concurrency:
defaults:
run:
- shell: bash -x {0}
+ shell: bash -exo pipefail {0}
jobs:
forward:
runs-on: ubuntu-latest
+ timeout-minutes: 15
steps:
- - uses: actions/checkout@v4
+ - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
with:
ssh-key: ${{ secrets.GENVM_CI_PRIVATE_KEY }}
# Full history: a shallow checkout only has the version branch's
@@ -46,14 +47,17 @@ jobs:
- name: Fast-forward main to latest dev branch
run: |
BRANCH="${GITHUB_REF_NAME}"
- # Decide "latest" from main's active-versions (the authoritative
- # alias), NOT the pushed branch's own copy — an older dev
- # branch carries a stale active-versions and would otherwise try
- # to drag main backwards once a newer train exists.
- git show origin/main:.genvm-monorepo-root > /tmp/main-monorepo-root.json
- LATEST="$(MONOREPO_ROOT=/tmp/main-monorepo-root.json python3 support/ci/branch-versions.py latest)"
- if [ "$BRANCH" != "v${LATEST}-dev" ]; then
- echo "Pushed branch ${BRANCH} is not the latest dev branch (v${LATEST}-dev); nothing to do"
+ # Decide the current manager train from main's `version` field (the
+ # authoritative alias), NOT the pushed branch's own copy — an older
+ # dev branch carries a stale version and would otherwise try to drag
+ # main backwards once a newer train exists.
+ # $RUNNER_TEMP, not /tmp: the runner wipes it between jobs, so nothing
+ # is left behind on a self-hosted box.
+ root_json="${RUNNER_TEMP}/main-monorepo-root.json"
+ git show origin/main:.genvm-monorepo-root > "$root_json"
+ MANAGER="$(MONOREPO_ROOT="$root_json" ./support/ci/run.sh tool branch-versions manager)"
+ if [ "$BRANCH" != "v${MANAGER}-dev" ]; then
+ echo "Pushed branch ${BRANCH} is not the latest dev branch (v${MANAGER}-dev); nothing to do"
exit 0
fi
echo "Fast-forwarding main to ${GITHUB_SHA} (from ${BRANCH})"
diff --git a/.github/workflows/branch_merge_into_dev.yaml b/.github/workflows/branch_merge_into_dev.yaml
deleted file mode 100644
index d0ff8590..00000000
--- a/.github/workflows/branch_merge_into_dev.yaml
+++ /dev/null
@@ -1,99 +0,0 @@
-name: branch / merge PR into dev
-
-# The repo has NO GitHub merge queue. A PR lands on a dev branch (v-dev)
-# when a maintainer ticks the "Merge" box on the PR action panel
-# (branch_pr_actions.yaml), which calls this reusable workflow. It
-# re-checks every gate against the EXACT head commit and then advances the
-# dev branch by a plain (fast-forward-only) push, so what lands is
-# byte-identical to what CI and E2E validated.
-#
-# Gates (all required):
-# 1. base branch is a v-dev branch
-# 2. PR carries the `rtm` (ready-to-merge) label
-# 3. full GenVM CI (queue.yaml) concluded success on the head commit
-# 4. the cross-repo E2E check concluded success on the head commit
-# 5. the PR is 0 commits behind base (head already contains base tip)
-#
-# Merge strategy:
-# - 1 commit -> fast-forward the original commit (SHA preserved)
-# - more than 1 commit -> squash into a single commit on top of base,
-# then fast-forward
-# Either way the PR is closed afterwards.
-#
-# Executor mirror: the landed manager commit gitlinks into the executor
-# submodule; the script fast-forwards the executor's SAME-NAMED branch to
-# that gitlink commit. It fast-forward-checks both repos up front, then
-# pushes the executor first (the dependency) and the manager second.
-#
-# The dev branches are protected; the default GITHUB_TOKEN cannot push to
-# them. We push over SSH with the GENVM_CI_PRIVATE_KEY deploy key (on the
-# dev-branch ruleset bypass list), and to the executor repo with a separate
-# GENVM_EXECUTOR_CI_PRIVATE_KEY deploy key (on the executor's bypass list).
-# Pushes are non-force, so a base that advanced between the checks and the
-# push is safely rejected. This job never runs PR code — only git plumbing
-# and API reads. The caller is responsible for restricting who can trigger a
-# merge (maintainers only).
-
-on:
- workflow_call:
- inputs:
- pr_number:
- description: number of the PR to merge
- type: string
- required: true
-
-permissions:
- contents: read
- pull-requests: write
- issues: write
- checks: read
- actions: read
-
-defaults:
- run:
- shell: bash -x {0}
-
-env:
- # check-run name (case-insensitive substring) the genlayer-e2e pipeline
- # posts on the head commit. Adjust if that pipeline renames its check.
- E2E_CHECK_PATTERN: "e2e"
-
-jobs:
- merge:
- runs-on: ubuntu-latest
- steps:
- # Checkout with the manager deploy key so the (protected) dev branch
- # push succeeds. Submodules are initialised in the next step with a
- # SEPARATE executor key — the manager key cannot read/write the
- # executor repo. The script never runs PR code.
- - uses: actions/checkout@v4
- with:
- ssh-key: ${{ secrets.GENVM_CI_PRIVATE_KEY }}
- fetch-depth: 0
- submodules: false
-
- - name: Set up executor submodule push access
- env:
- # Deploy key with write access to the executor repo, on its
- # protected-branch ruleset bypass list. The executor submodule's
- # origin must be an SSH URL (git@github.com:...) for this key to
- # apply.
- EXECUTOR_DEPLOY_KEY: ${{ secrets.GENVM_EXECUTOR_CI_PRIVATE_KEY }}
- run: |
- install -m700 -d ~/.ssh
- key=~/.ssh/executor_ci
- printf '%s\n' "$EXECUTOR_DEPLOY_KEY" > "$key"
- chmod 600 "$key"
- ssh="ssh -i $key -o IdentitiesOnly=yes -o StrictHostKeyChecking=accept-new"
- # Fetch the submodule with the executor key, then bind that key to
- # the submodule only (per-repo core.sshCommand) so the later mirror
- # push uses it while the manager push keeps the checkout credentials.
- git -c core.sshCommand="$ssh" submodule update --init executors/v0.3.x
- git -C executors/v0.3.x config core.sshCommand "$ssh"
-
- - name: Validate gates, mirror executor, fast-forward / squash, close
- env:
- GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
- PR_NUMBER: ${{ inputs.pr_number }}
- EXECUTOR_SUBMODULE: executors/v0.3.x
- run: python3 support/ci/genvm-merge-into-dev.py
diff --git a/.github/workflows/branch_pr_actions.yaml b/.github/workflows/branch_pr_actions.yaml
index 4d11a25c..fda2a850 100644
--- a/.github/workflows/branch_pr_actions.yaml
+++ b/.github/workflows/branch_pr_actions.yaml
@@ -2,9 +2,14 @@ name: branch / PR action panel handler
# Reacts to a maintainer ticking a box on the GenVM PR action panel (posted
# by branch_pr_checklist.yaml). The `dispatch` job authenticates the editor,
-# parses the ticked boxes, performs label-based actions (run / rerun full
-# tests), unticks the boxes, and reports whether Merge was requested. If so
-# the `merge` job calls the reusable merge workflow.
+# parses the ticked boxes, enables full tests or provisions executor PRs, then
+# unticks the momentary boxes.
+#
+# The handler runs the CI scripts checked out from the PR head branch, not
+# from the default branch, so a PR that changes support/ci/ exercises its own
+# version. That means this workflow DOES run PR-authored code, so it is gated
+# on the `ci-safe` label (added only for vetted / write-access PRs) — the
+# `dispatch` job does not even start without it.
#
# Resetting the panel re-fires issue_comment:edited, but that re-run finds
# no ticked box (and is sent by the bot), so it is a no-op — no loop.
@@ -26,21 +31,34 @@ concurrency:
defaults:
run:
- shell: bash -x {0}
+ shell: bash -exo pipefail {0}
jobs:
dispatch:
- # Only the panel comment, only on a PR.
+ # Only a panel-shaped comment, only on a PR, only when the PR is `ci-safe`
+ # (we check out and run its scripts, so an unvetted PR must not reach here).
+ #
+ # NOTE: this matches the MARKER IN THE COMMENT BODY, which anyone can put in
+ # a comment of their own. It is a cheap pre-filter, NOT authorization. The
+ # tool re-checks that the bot authored the comment and that the editor has
+ # write access — see `Who may tick a box` in pr_action_panel.py.
if: >
github.event.issue.pull_request &&
- contains(github.event.comment.body, '')
+ contains(github.event.comment.body, '') &&
+ contains(github.event.issue.labels.*.name, 'ci-safe')
runs-on: ubuntu-latest
- outputs:
- merge: ${{ steps.act.outputs.merge }}
+ timeout-minutes: 15
steps:
- # Checkout of the default branch only — to run the handler script. No
- # PR code is executed.
- - uses: actions/checkout@v4
+ # Check out the PR head so the handler runs the PR's own CI scripts.
+ # This runs PR-authored code — the `ci-safe` gate on the job above is
+ # what makes that safe.
+ - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
+ with:
+ ref: refs/pull/${{ github.event.issue.number }}/head
+ # The handler talks to GitHub through `gh` (GH_TOKEN below) and never
+ # pushes, so the checkout credential is not needed — and leaving it in
+ # .git/config would hand it to the PR code this job runs.
+ persist-credentials: false
- name: Handle ticked boxes
id: act
env:
@@ -48,12 +66,4 @@ jobs:
PR_NUMBER: ${{ github.event.issue.number }}
COMMENT_ID: ${{ github.event.comment.id }}
SENDER: ${{ github.event.sender.login }}
- run: python3 support/ci/pr-action-panel.py
-
- merge:
- needs: dispatch
- if: needs.dispatch.outputs.merge == 'true'
- uses: ./.github/workflows/branch_merge_into_dev.yaml
- with:
- pr_number: ${{ github.event.issue.number }}
- secrets: inherit
+ run: ./support/ci/run.sh tool pr-action-panel
diff --git a/.github/workflows/branch_pr_checklist.yaml b/.github/workflows/branch_pr_checklist.yaml
index 33186878..7cf5cdf9 100644
--- a/.github/workflows/branch_pr_checklist.yaml
+++ b/.github/workflows/branch_pr_checklist.yaml
@@ -1,9 +1,7 @@
name: branch / post PR action panel
-# When a PR is opened against a dev branch (v-dev) we:
-# - make sure the action labels exist;
-# - post the action panel (checklist comment) that replaces the old
-# /genvm-* slash commands;
+# When a PR is opened against main or a dev branch (v-dev) we:
+# - post the action panel and document its companion slash commands;
# - auto-add `ci-safe` if the PR author has write access, so their PR can
# immediately use the panel. For everyone else a write-role maintainer
# adds `ci-safe` by hand once the PR is vetted.
@@ -14,8 +12,10 @@ name: branch / post PR action panel
on:
pull_request_target:
- branches: ['v*-dev']
- types: [opened]
+ branches: [main, 'v*-dev']
+ # `edited` catches retargets; `synchronize` migrates panels already posted
+ # by an older workflow revision without a repository-wide write sweep.
+ types: [opened, edited, synchronize]
permissions:
pull-requests: write
@@ -23,37 +23,64 @@ permissions:
defaults:
run:
- shell: bash -x {0}
+ shell: bash -exo pipefail {0}
jobs:
post:
runs-on: ubuntu-latest
+ timeout-minutes: 15
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
PR: ${{ github.event.pull_request.number }}
AUTHOR: ${{ github.event.pull_request.user.login }}
steps:
- - name: Ensure action labels exist
+ - name: Post or refresh action panel
run: |
- gh label create rtm -R "$GITHUB_REPOSITORY" --color 0e8a16 --description "ready to merge" 2>/dev/null || true
- gh label create run-full-tests -R "$GITHUB_REPOSITORY" --color fbca04 --description "run full GenVM CI on every push" 2>/dev/null || true
- gh label create ci-safe -R "$GITHUB_REPOSITORY" --color 5319e7 --description "PR cleared to run CI / panel actions" 2>/dev/null || true
+ labels="$(gh api --paginate "repos/$GITHUB_REPOSITORY/issues/$PR/labels" --jq '.[].name')"
+ force=false
+ if grep -qxF 'run-full-tests' <<<"$labels"; then
+ force=true
+ fi
- - name: Post action panel
- run: |
- gh pr comment "$PR" --repo "$GITHUB_REPOSITORY" --body "$(cat <<'EOF'
+ body="$(cat <<'EOF'
### GenVM PR actions
Tick a box to run it (the box unticks itself when handled). Actions only run while the PR has the **`ci-safe`** label.
- - [ ] Force run full tests
- - [ ] Rerun full tests
- - [ ] Merge into dev
+
+
+
+ - [ ] Force run full tests
+ - [ ] Provision executor PRs
- Full GenVM CI runs only when **`rtm`** or **`run-full-tests`** is set — "Force run full tests" is a sticky toggle for `run-full-tests`. Adding **`rtm`** marks the PR ready-to-merge and also runs full tests. **Merge** requires: `rtm`, green full tests, green E2E, and the branch 0 commits behind.
+ Commands
+
+ - `/genvm-run-tests` — run full tests once for the current manager snapshot
+ - `/merge` — queue the exact manager snapshot through the App-owned E2E merge train
+
EOF
)"
+ if [ "$force" = true ]; then
+ body="${body/- \[ \] Force run full tests/- [x] Force run full tests}"
+ fi
+
+ panel_ids="$(gh api --paginate "repos/$GITHUB_REPOSITORY/issues/$PR/comments?per_page=100" \
+ --jq '.[] | select(.user.login == "github-actions[bot]" and (.body | contains(""))) | .id')"
+ panel_id="${panel_ids%%$'\n'*}"
+ if [ -z "$panel_id" ]; then
+ gh pr comment "$PR" --repo "$GITHUB_REPOSITORY" --body "$body"
+ exit 0
+ fi
+
+ while IFS= read -r panel_id; do
+ current="$(gh api "repos/$GITHUB_REPOSITORY/issues/comments/$panel_id" --jq '.body')"
+ if [ "$current" = "$body" ]; then
+ echo "action panel $panel_id is current; skipping"
+ continue
+ fi
+ gh api --method PATCH "repos/$GITHUB_REPOSITORY/issues/comments/$panel_id" -f "body=$body"
+ done <<<"$panel_ids"
- name: Auto-mark ci-safe for write-access authors
run: |
diff --git a/.github/workflows/branch_provision.yaml b/.github/workflows/branch_provision.yaml
index d45f56b2..24e3de93 100644
--- a/.github/workflows/branch_provision.yaml
+++ b/.github/workflows/branch_provision.yaml
@@ -1,11 +1,10 @@
name: branch / provision dev branches and standing PRs
-# Materialize the branch model for every active version train listed in
-# .genvm-monorepo-root. For each version X it ensures the version branch
-# v.x, the dev branch v-dev, and a standing release-gate PR
-# v-dev -> v.x all exist (see support/ci/provision-branches.sh).
+# Materialize the manager train declared by .genvm-monorepo-root: the release
+# branch v, the dev branch v-dev, and the standing release-gate PR
+# v-dev -> v (see support/ci/run.sh tool provision-branches).
#
-# Idempotent — safe to re-run. Fires automatically when active-versions
+# Idempotent — safe to re-run. Fires automatically when the train declaration
# changes, and on demand via the Actions tab.
on:
@@ -20,13 +19,14 @@ permissions:
defaults:
run:
- shell: bash -x {0}
+ shell: bash -exo pipefail {0}
jobs:
provision:
runs-on: ubuntu-latest
+ timeout-minutes: 15
steps:
- - uses: actions/checkout@v4
+ - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
with:
# Deploy key so branch-creating pushes bypass protection on the
# (protected) version / dev branches. Full history so the
@@ -37,4 +37,4 @@ jobs:
- name: Provision branches and standing PRs
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
- run: support/ci/provision-branches.sh
+ run: ./support/ci/run.sh tool provision-branches
diff --git a/.github/workflows/branch_provision_executor_prs.yaml b/.github/workflows/branch_provision_executor_prs.yaml
new file mode 100644
index 00000000..86fffd08
--- /dev/null
+++ b/.github/workflows/branch_provision_executor_prs.yaml
@@ -0,0 +1,72 @@
+name: branch / provision executor mirror PRs
+
+# Triggered by the "Provision executor PRs" box on the PR action panel
+# (pr_action_panel.py fires a workflow_dispatch with the manager PR number). For
+# each executor line the manager PR moved, force-push its mirror branch
+# `pr//` to the pinned gitlink commit and open (or reuse) its PR into
+# `-dev`. See `support/ci/run.sh tool pr-branches provision`.
+#
+# The provision scripts are checked out from the PR head, so a PR iterating on
+# them runs its own version — which means this job runs PR-authored code WITH the
+# executor write PAT (GENVM_EXECUTOR_PR_TOKEN) in scope. It is gated on the
+# `ci-safe` label: a guard step verifies it via the plain GITHUB_TOKEN before
+# any PR code runs and fails otherwise.
+
+on:
+ workflow_dispatch:
+ inputs:
+ pr:
+ description: manager PR number to provision executor mirror PRs for
+ type: string
+ required: true
+
+run-name: 'branch / provision executor mirror PRs (PR #${{ inputs.pr }})'
+
+permissions:
+ contents: read
+ # For the GITHUB_TOKEN only: read the ci-safe label and the manager PR title.
+ # Writes to the executor repo go through GENVM_EXECUTOR_PR_TOKEN, not this token.
+ pull-requests: read
+ issues: read
+
+defaults:
+ run:
+ shell: bash -exo pipefail {0}
+
+jobs:
+ provision:
+ runs-on: ubuntu-latest
+ timeout-minutes: 15
+ steps:
+ # Gate: this job checks out and runs the PR's scripts with the executor
+ # write PAT in scope, so refuse unless the PR is `ci-safe`. Uses the plain
+ # GITHUB_TOKEN and runs BEFORE any PR code is fetched.
+ - name: Require ci-safe label
+ env:
+ GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
+ PR_NUMBER: ${{ inputs.pr }}
+ run: |
+ # Collect first, match second: under `pipefail` a `grep -q` that exits
+ # early can SIGPIPE the `gh` feeding it, and the resulting non-zero
+ # pipeline would read as "no ci-safe label" and refuse a valid run.
+ names="$(gh api "repos/$GITHUB_REPOSITORY/issues/$PR_NUMBER/labels" --jq '.[].name')"
+ if ! grep -qxF 'ci-safe' <<<"$names"; then
+ echo "::error::PR #$PR_NUMBER lacks the 'ci-safe' label; refusing to provision (this job runs PR code with the executor token)."
+ exit 1
+ fi
+
+ - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
+ with:
+ ref: refs/pull/${{ inputs.pr }}/head
+ # Provisioning goes through the executor token via `gh`, never `git
+ # push`, so don't leave the checkout credential in .git/config where
+ # the PR code this job runs could read it.
+ persist-credentials: false
+
+ - name: Provision executor mirror PRs
+ env:
+ PR_NUMBER: ${{ inputs.pr }}
+ MANAGER_REPO: ${{ github.repository }}
+ MANAGER_TOKEN: ${{ secrets.GITHUB_TOKEN }}
+ EXECUTOR_TOKEN: ${{ secrets.GENVM_EXECUTOR_PR_TOKEN }}
+ run: ./support/ci/run.sh tool pr-branches provision
diff --git a/.github/workflows/branch_rebase_watch.yaml b/.github/workflows/branch_rebase_watch.yaml
new file mode 100644
index 00000000..2422490c
--- /dev/null
+++ b/.github/workflows/branch_rebase_watch.yaml
@@ -0,0 +1,71 @@
+name: branch / not-rebased watch
+
+# Advisory "your branch is stale" indicator. `initial / behind-check` only runs
+# when the PR is pushed, so a PR that goes behind because the BASE moved stays
+# green until someone ticks Merge and the fast-forward is refused. This watches
+# the other edge.
+#
+# It gates nothing: it posts a red `not rebased` check-run (visible on the PR
+# page) and a red `not rebased` label (the surface visible in the PR list), and
+# clears both once the branch is 0 behind again. The enforcing checks are
+# unchanged — `initial / behind-check` remains the enforcing check.
+#
+# Triggers:
+# - push to a dev branch -> re-evaluate every open PR targeting it (this is
+# the case nothing else covers);
+# - pull_request_target on the PR itself -> clear the flag promptly after a
+# rebase, and set it on a retarget (`edited` fires on a base change).
+#
+# pull_request_target, not pull_request: posting a check-run and editing labels
+# needs a write-scoped token, which fork PRs do not get. Nothing here checks out
+# or runs PR code — the checkout is the base repo's tree, and the only inputs
+# are a PR number and a branch name.
+
+on:
+ push:
+ branches: ['v*-dev']
+ # Deliberately NOT filtered on `branches`: a `branches` filter would skip the
+ # `edited` event that retargets a flagged PR AWAY from a dev branch, and no
+ # push sweep would ever see that PR again — its stale red flag would stick
+ # forever. The tool decides what a non-dev base means (clear, never set).
+ pull_request_target:
+ types: [opened, synchronize, reopened, edited]
+
+permissions:
+ contents: read
+ checks: write
+ pull-requests: write
+ issues: write
+
+# Serialise per base branch (push sweeps) or per PR: two overlapping sweeps would
+# race on the same label. Never cancel — a cancelled sweep can leave a PR flagged
+# after it was rebased.
+concurrency:
+ group: rebase-watch-${{ github.event.pull_request.number || github.ref_name }}
+ cancel-in-progress: false
+
+defaults:
+ run:
+ shell: bash -exo pipefail {0}
+
+jobs:
+ watch:
+ runs-on: ubuntu-latest
+ timeout-minutes: 30
+ steps:
+ - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
+ - name: Flag PRs that are behind their base
+ env:
+ GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
+ # On a push these are empty, so the tool sweeps every open PR on the
+ # branch that moved; on a PR event it evaluates just that PR.
+ PR_NUMBER: ${{ github.event.pull_request.number }}
+ BASE: ${{ github.event_name == 'push' && github.ref_name || '' }}
+ run: |
+ # `if`, not `[ ... ] && ...`: under `set -e` a false AND-list is the
+ # script's exit status, so the un-taken branch would abort the step.
+ args=()
+ if [ -n "$BASE" ]; then
+ args+=(--base "$BASE")
+ fi
+ ./support/ci/run.sh tool rebase-watch "${args[@]}"
diff --git a/.github/workflows/branch_release_gate_pr.yaml b/.github/workflows/branch_release_gate_pr.yaml
new file mode 100644
index 00000000..7384806c
--- /dev/null
+++ b/.github/workflows/branch_release_gate_pr.yaml
@@ -0,0 +1,67 @@
+name: branch / open release-gate PR
+
+# On each push to v-dev, ensure its standing PR into v exists. Unlike
+# provision-branches, derive the pair from the pushed ref, covering manual dev
+# branches. Skip a missing base, no commits ahead or an existing PR. GITHUB_TOKEN
+# does not trigger `pull_request` workflows; close and reopen if checks must run before a push
+
+on:
+ push:
+ branches: ['v*-dev']
+
+permissions:
+ contents: read
+ pull-requests: write
+
+# Serialize the non-atomic existence check and creation per dev branch
+concurrency:
+ group: release-gate-pr-${{ github.ref_name }}
+ cancel-in-progress: false
+
+defaults:
+ run:
+ shell: bash -exo pipefail {0}
+
+jobs:
+ open-pr:
+ runs-on: ubuntu-latest
+ timeout-minutes: 15
+ steps:
+ - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
+ with:
+ # Full history for comparing the remote branches
+ fetch-depth: 0
+
+ - name: Open release-gate PR if missing
+ env:
+ GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
+ run: |
+ DEV="${GITHUB_REF_NAME}"
+ BASE="${DEV%-dev}"
+
+ git fetch origin --prune
+
+ if ! git show-ref --verify --quiet "refs/remotes/origin/${BASE}"; then
+ echo "Base branch ${BASE} does not exist; nothing to do"
+ exit 0
+ fi
+
+ # Same tip or behind base: there is nothing to merge
+ AHEAD="$(git rev-list --count "origin/${BASE}..origin/${DEV}")"
+ if [ "$AHEAD" -eq 0 ]; then
+ echo "${DEV} has no commits beyond ${BASE}; nothing to do"
+ exit 0
+ fi
+
+ EXISTING="$(gh pr list --repo "$GITHUB_REPOSITORY" --base "$BASE" --head "$DEV" --state open --json number --jq '.[0].number // ""')"
+ if [ -n "$EXISTING" ]; then
+ echo "Release-gate PR ${DEV} -> ${BASE} already open (#${EXISTING})"
+ exit 0
+ fi
+
+ gh pr create \
+ --repo "$GITHUB_REPOSITORY" \
+ --base "$BASE" \
+ --head "$DEV" \
+ --title "chore(release): gate ${DEV} into ${BASE} 🚀" \
+ --body "Standing release-gate PR. \`${DEV}\` accumulates incremental work; it merges into the release-ready \`${BASE}\` branch only once the cross-repo E2E pipeline is green. It may stay red while the train is in progress."
diff --git a/.github/workflows/branch_retarget.yaml b/.github/workflows/branch_retarget.yaml
index a7f45715..34ad582c 100644
--- a/.github/workflows/branch_retarget.yaml
+++ b/.github/workflows/branch_retarget.yaml
@@ -20,7 +20,7 @@ permissions:
defaults:
run:
- shell: bash -x {0}
+ shell: bash -exo pipefail {0}
jobs:
retarget:
@@ -28,15 +28,16 @@ jobs:
# version branch.
if: github.event.pull_request.base.ref == 'main'
runs-on: ubuntu-latest
+ timeout-minutes: 15
steps:
- - uses: actions/checkout@v4
+ - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
- name: Retarget PR to latest dev branch
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
run: |
- LATEST="$(python3 support/ci/branch-versions.py latest)"
- DEV="v${LATEST}-dev"
+ MANAGER="$(./support/ci/run.sh tool branch-versions manager)"
+ DEV="v${MANAGER}-dev"
PR="${{ github.event.pull_request.number }}"
gh pr edit "$PR" --repo "$GITHUB_REPOSITORY" --base "$DEV"
@@ -44,6 +45,6 @@ jobs:
gh pr comment "$PR" --repo "$GITHUB_REPOSITORY" --body "$(cat <-
+ github.event_name == 'issue_comment' &&
+ github.event.issue.pull_request &&
+ github.event.comment.body == '/genvm-run-tests' &&
+ contains(github.event.issue.labels.*.name, 'ci-safe')
+ runs-on: ubuntu-latest
+ timeout-minutes: 5
+ permissions:
+ actions: write
+ contents: read
+ issues: write
+ pull-requests: write
+ steps:
+ # issue_comment workflows run from the default branch. Keep that trusted
+ # checkout: this handler dispatches PR code with credentials in scope
+ - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
+ with:
+ persist-credentials: false
+ - name: Dispatch full tests
+ env:
+ GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
+ PR_NUMBER: ${{ github.event.issue.number }}
+ COMMENT_ID: ${{ github.event.comment.id }}
+ COMMENT_BODY: ${{ github.event.comment.body }}
+ SENDER: ${{ github.event.sender.login }}
+ REQUEST_ID: ${{ github.run_id }}-${{ github.run_attempt }}
+ run: ./support/ci/run.sh tool full-tests-command start
+
+ complete:
+ if: >-
+ github.event_name == 'workflow_run' &&
+ contains(github.event.workflow_run.display_title, '[request=')
+ runs-on: ubuntu-latest
+ timeout-minutes: 5
+ permissions:
+ contents: read
+ issues: write
+ pull-requests: write
+ steps:
+ # workflow_run also resolves this workflow from the default branch. Never
+ # check out or execute the completed run's potentially PR-authored tree
+ - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
+ with:
+ persist-credentials: false
+ - name: Report full-test result
+ env:
+ GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
+ RUN_ID: ${{ github.event.workflow_run.id }}
+ RUN_NAME: ${{ github.event.workflow_run.display_title }}
+ RUN_URL: ${{ github.event.workflow_run.html_url }}
+ RUN_SHA: ${{ github.event.workflow_run.head_sha }}
+ RUN_CONCLUSION: ${{ github.event.workflow_run.conclusion }}
+ run: ./support/ci/run.sh tool full-tests-command complete
diff --git a/.github/workflows/branch_sync_executors.yaml b/.github/workflows/branch_sync_executors.yaml
new file mode 100644
index 00000000..aa26775c
--- /dev/null
+++ b/.github/workflows/branch_sync_executors.yaml
@@ -0,0 +1,59 @@
+name: branch / sync executor branches
+
+# Manager train branches are authoritative for the executor refs they pin.
+# A release branch (v.) projects each active executor gitlink onto the
+# branch declared in .gitmodules; its dev branch projects onto -dev.
+# Every push is non-force: missing refs are created, existing refs only move by
+# fast-forward. The tool attempts every line before reporting combined failures.
+
+on:
+ push:
+ branches:
+ - 'v[0-9]+.[0-9]+'
+ - 'v[0-9]+.[0-9]+-dev'
+
+run-name: branch / sync executors from ${{ github.ref_name }}
+
+permissions:
+ contents: read
+
+# Collapse a burst on one manager branch onto its live tip. Different manager
+# branches may run together; executor branch protection and non-force pushes
+# reject any incompatible ownership rather than hiding it
+concurrency:
+ group: sync-executors-${{ github.ref_name }}
+ cancel-in-progress: false
+
+defaults:
+ run:
+ shell: bash -exo pipefail {0}
+
+jobs:
+ sync:
+ runs-on: ubuntu-latest
+ timeout-minutes: 30
+ steps:
+ - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
+ with:
+ # Reconcile the live branch tip. A rerun therefore repairs only refs
+ # that still need an update
+ ref: ${{ github.ref_name }}
+ persist-credentials: false
+
+ - name: Set up executor push access
+ env:
+ EXECUTOR_DEPLOY_KEY: ${{ secrets.GENVM_EXECUTOR_CI_PRIVATE_KEY }}
+ run: |
+ install -m700 -d ~/.ssh
+ key="${RUNNER_TEMP}/executor_sync_key"
+ set +x
+ printf '%s\n' "$EXECUTOR_DEPLOY_KEY" > "$key"
+ set -x
+ chmod 600 "$key"
+ git config --global url."git@github.com:".insteadOf "https://github.com/"
+ echo "GIT_SSH_COMMAND=ssh -i $key -o IdentitiesOnly=yes -o StrictHostKeyChecking=accept-new" >> "$GITHUB_ENV"
+
+ - name: Create or fast-forward executor branches
+ env:
+ MANAGER_BRANCH: ${{ github.ref_name }}
+ run: ./support/ci/run.sh tool sync-executor-branches
diff --git a/.github/workflows/branch_wasmtime_watch.yaml b/.github/workflows/branch_wasmtime_watch.yaml
new file mode 100644
index 00000000..4af27ce9
--- /dev/null
+++ b/.github/workflows/branch_wasmtime_watch.yaml
@@ -0,0 +1,61 @@
+name: branch / wasmtime watch
+
+# Daily sweep for published advisories against the upstream sources the manager
+# vendors: each repo in an executor line's `manifest.json` by its pinned commit,
+# and each crate that comes out of one by its locked version.
+#
+# It gates nothing and writes no branch: it keeps one `wasmtime-maintenance`
+# issue in sync with what OSV currently says. Nothing PR-triggered could do this
+# job — a pin becomes vulnerable while standing still, with no PR involved.
+#
+# Closing the issue is a human ruling on the findings it names; the tool never
+# closes or reopens one. See support/ci/tools/wasmtime_watch.py.
+
+on:
+ schedule:
+ # Daily, off the hour: a scheduled run on a busy minute is queued or dropped.
+ - cron: "17 5 * * *"
+ workflow_dispatch:
+ inputs:
+ dry_run:
+ description: Report findings without creating or editing the issue
+ type: boolean
+ required: false
+ default: false
+
+permissions:
+ contents: read
+ issues: write
+
+# Never cancel: a cancelled sweep can leave the issue describing a state that no
+# longer holds, and the next run is a day away.
+concurrency:
+ group: wasmtime-watch
+ cancel-in-progress: false
+
+defaults:
+ run:
+ shell: bash -exo pipefail {0}
+
+jobs:
+ watch:
+ runs-on: ubuntu-latest
+ timeout-minutes: 30
+ steps:
+ - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
+ # Both files it reads are tracked inside the executor submodules, so this
+ # needs a checkout but not the materialized third-party trees — which is
+ # what lets it skip nix entirely.
+ - name: Get source
+ uses: ./.github/actions/get-src
+ with:
+ with_nix: "false"
+ third_party: none
+ github_token: ${{ secrets.GITHUB_TOKEN }}
+ - name: Sweep the vendored upstream sources
+ env:
+ GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
+ WASMTIME_REBASE_OWNER: ${{ vars.WASMTIME_REBASE_OWNER }}
+ run: >-
+ ./support/ci/run.sh tool wasmtime-watch
+ ${{ inputs.dry_run && '--dry-run' || '' }}
diff --git a/.github/workflows/docs_nightly.yaml b/.github/workflows/docs_nightly.yaml
new file mode 100644
index 00000000..881fc577
--- /dev/null
+++ b/.github/workflows/docs_nightly.yaml
@@ -0,0 +1,102 @@
+name: docs / nightly deploy
+
+# Builds the docs site and publishes it to the sdk.genlayer.com Pages repo,
+# which serves one directory per version out of `_site/` and lists them in
+# `_site/versions.json` (the switcher every page links to). Nightly runs
+# refresh `main`; `workflow_dispatch` can publish any other directory, which is
+# how a release train (`v0.6.x`) gets its own copy.
+#
+# GitHub only schedules workflows from the DEFAULT branch, so this file has to
+# reach it before the cron fires — merging it into a dev branch is not enough.
+#
+# Needs the `DOCS_DEPLOY_KEY` secret: a deploy key with write access to
+# genlayerlabs/sdk.genlayer.com. Pushing there triggers that repo's own
+# deploy-pages workflow, which is what actually publishes and purges the CDN.
+
+on:
+ schedule:
+ # 03:17 UTC — off the hour, where scheduling backlogs are smaller.
+ - cron: "17 3 * * *"
+ workflow_dispatch:
+ inputs:
+ version:
+ description: "directory to publish under (e.g. main, v0.6.x)"
+ required: false
+ default: main
+ preferred:
+ description: "make it the version the switcher highlights"
+ type: boolean
+ required: false
+ default: false
+
+# Two runs pushing to the site repo would race on the same branch; the loser
+# would just fail, so serialize rather than cancel (a half-finished deploy is
+# worse than a slow one).
+concurrency:
+ group: docs-deploy
+ cancel-in-progress: false
+
+permissions:
+ contents: read
+
+env:
+ SITE_REPO: genlayerlabs/sdk.genlayer.com
+ DOCS_VERSION: ${{ inputs.version || 'main' }}
+
+# GitHub's stock `bash -e {0}` has no `pipefail`.
+defaults:
+ run:
+ shell: bash -exo pipefail {0}
+
+jobs:
+ deploy:
+ runs-on: ubuntu-latest
+ timeout-minutes: 120
+ steps:
+ - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
+
+ # The docs build spans the manager and every active executor line, each
+ # in its own `gen-docs` nix shell, so it needs submodules and nix.
+ - name: Get source
+ uses: ./.github/actions/get-src
+ with:
+ with_nix: "true"
+ github_token: ${{ secrets.GITHUB_TOKEN }}
+ nix_cache_pull_token: ${{ secrets.NIX_CACHE_PULL_TOKEN }}
+
+ - name: Build docs
+ run: ./support/ci/run.sh pipeline docs
+
+ - name: Checkout the site repo
+ uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
+ with:
+ repository: ${{ env.SITE_REPO }}
+ # The site repo is private, and `ref` is what keeps checkout from
+ # asking the REST API for the default branch with this repo's token,
+ # which cannot see it. The deploy key covers the clone itself.
+ ref: main
+ path: build/site
+ ssh-key: ${{ secrets.DOCS_DEPLOY_KEY }}
+
+ - name: Stage the built docs
+ run: |
+ ./support/ci/run.sh tool deploy-docs \
+ --site build/site \
+ --version "$DOCS_VERSION" \
+ ${{ inputs.preferred && '--preferred' || '' }}
+
+ # An unchanged nightly must not push an empty commit: the site repo's
+ # deploy workflow runs on every push, so that would redeploy and purge the
+ # CDN for nothing.
+ - name: Commit and push
+ working-directory: build/site
+ run: |
+ git config user.email "ci@genlayerlabs.com"
+ git config user.name "CI worker"
+ git add -A _site
+ if git diff --cached --quiet; then
+ echo "docs are unchanged; nothing to deploy"
+ exit 0
+ fi
+ git commit -m "Deploy documentation version $DOCS_VERSION"
+ git push origin HEAD
diff --git a/.github/workflows/incl_build_trg.yaml b/.github/workflows/incl_build_trg.yaml
deleted file mode 100644
index 273fbe68..00000000
--- a/.github/workflows/incl_build_trg.yaml
+++ /dev/null
@@ -1,110 +0,0 @@
-name: Build Release For Target
-
-on:
- workflow_call:
- inputs:
- target:
- description: 'Target to build'
- required: true
- type: string
- secondary_target:
- description: 'Secondary target to build'
- required: false
- type: string
- default: ''
- executor_version:
- description: 'Version name of the executor'
- required: true
- type: string
- secrets:
- GCP_SA_KEY:
- description: 'Google Cloud Service Account Key'
- required: true
- outputs:
- artifact_url:
- description: "URL to download the built artifact"
- value: ${{ jobs.build.outputs.artifact_url }}
- secondary_artifact_url:
- description: "URL to download the secondary built artifact"
- value: ${{ jobs.build.outputs.secondary_artifact_url }}
-
-env:
- GCS_BUCKET: "gh-af"
-
-jobs:
- build:
- runs-on: ubuntu-latest
- permissions:
- contents: read
- actions: read
- outputs:
- artifact_url: ${{ steps.upload.outputs.gcs_url }}
- secondary_artifact_url: ${{ steps.upload.outputs.secondary_gcs_url }}
- steps:
- - name: Checkout code
- uses: actions/checkout@v4
-
- - name: Get source
- uses: ./.github/actions/get-src
- with:
- load_submodules: "false"
- with_nix: "true"
- github_token: ${{ secrets.GITHUB_TOKEN }}
-
- - name: Build
- run: |
- ./support/ci/pipelines/build.sh \
- --target ${{ inputs.target }} \
- --executor-version ${{ inputs.executor_version }}
-
- - name: Build Secondary
- if: ${{ inputs.secondary_target != '' }}
- run: |
- ./support/ci/pipelines/build.sh \
- --target ${{ inputs.secondary_target }} \
- --executor-version ${{ inputs.executor_version }}
-
- - name: Authenticate to Google Cloud
- uses: google-github-actions/auth@v2
- with:
- credentials_json: ${{ secrets.GCP_SA_KEY }}
-
- - name: Set up Cloud SDK
- uses: google-github-actions/setup-gcloud@v2
-
- - name: Check gcloud authentication
- run: gcloud auth list
-
- - name: Generate upload url
- id: upload
- run: |
- TIMESTAMP=$(date +%Y%m%d_%H%M%S)
- DIR_NAME="genvm_executor_${GITHUB_SHA}_${TIMESTAMP}"
- echo "dirname=$DIR_NAME" >> $GITHUB_OUTPUT
- BASE_NAME="genvm-${{ inputs.target }}.tar.xz"
- echo "basename=$BASE_NAME" >> $GITHUB_OUTPUT
- echo "gcs_url=https://storage.googleapis.com/$GCS_BUCKET/$DIR_NAME/$BASE_NAME" >> $GITHUB_OUTPUT
-
- if [ "${{ inputs.secondary_target }}" != "" ]; then
- SECONDARY_BASE_NAME="genvm-${{ inputs.secondary_target }}.tar.xz"
- echo "secondary_basename=$SECONDARY_BASE_NAME" >> $GITHUB_OUTPUT
- echo "secondary_gcs_url=https://storage.googleapis.com/$GCS_BUCKET/$DIR_NAME/$SECONDARY_BASE_NAME" >> $GITHUB_OUTPUT
- else
- echo "secondary_basename=" >> $GITHUB_OUTPUT
- echo "secondary_gcs_url=" >> $GITHUB_OUTPUT
- fi
-
- - name: Upload to GCS
- uses: google-github-actions/upload-cloud-storage@v2
- with:
- path: build/${{ steps.upload.outputs.basename }}
- destination: ${{ env.GCS_BUCKET }}/${{ steps.upload.outputs.dirname }}
- parent: false
-
- - name: Upload Secondary to GCS
- uses: google-github-actions/upload-cloud-storage@v2
- if: ${{ inputs.secondary_target != '' }}
- with:
- path: build/${{ steps.upload.outputs.secondary_basename }}
- destination: ${{ env.GCS_BUCKET }}/${{ steps.upload.outputs.dirname }}
- parent: false
diff --git a/.github/workflows/incl_initial.yaml b/.github/workflows/incl_initial.yaml
index d16fed5c..57845190 100644
--- a/.github/workflows/incl_initial.yaml
+++ b/.github/workflows/incl_initial.yaml
@@ -1,52 +1,128 @@
name: GenVM initial checks
on:
workflow_call:
+ inputs:
+ pr:
+ description: >-
+ PR number these checks validate. The caller resolves it from the event
+ (`github.event.pull_request.number` on a `pull_request`, `inputs.pr` on
+ the panel's `workflow_dispatch`), so the PR-scoped jobs below run for
+ BOTH — they used to key off `github.event_name == 'pull_request'` and
+ therefore skipped on every panel-dispatched run, while that run could
+ still conclude `success` and satisfy the App's native CI gate. Empty for a
+ non-PR run, which skips them.
+ type: string
+ required: false
+ default: ""
-env:
- GCS_BUCKET: "gh-af"
+# A reusable workflow does not inherit the caller's `defaults`, so the shell
+# is set here too: GitHub's stock `bash -e {0}` has no `pipefail`.
+defaults:
+ run:
+ shell: bash -exo pipefail {0}
+# Timeouts are set from measured job durations, generously (roughly 3x the
+# observed maximum) so a slow-but-healthy run is never killed — the point is to
+# cap a hung job well below the 6-hour default, not to enforce a budget.
jobs:
pre-commit:
runs-on: ubuntu-latest
+ timeout-minutes: 30
steps:
- - uses: actions/checkout@v4
+ - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
+ with:
+ persist-credentials: false
- name: Get source
uses: ./.github/actions/get-src
with:
with_nix: "true"
github_token: ${{ secrets.GITHUB_TOKEN }}
- # genvm-tool runs every repo's hooks with tools from each repo's pinned
- # support/nix/precommit flake (rustfmt included), so no separate toolchain
- # setup is needed.
- # --check: CI verifies formatting and fails on drift instead of fixing
- # in place (fixing is the default for local/git-hook runs).
+ nix_cache_pull_token: ${{ secrets.NIX_CACHE_PULL_TOKEN }}
+ # Each repo (manager + every executor submodule) supplies its own
+ # pre-commit binary + generated config from its flake (git-hooks.nix),
+ # with tools pinned per repo — so no separate toolchain setup is needed.
+ # The formatters fix in place and the run fails if anything changed, so
+ # unformatted code turns the build red (the diff is shown on failure).
- name: Run hooks
- run: nix develop '.?submodules=1#initial-check' --command genvm-tool hook run --all-files --check
+ run: ./support/ci/run.sh pipeline commit-hooks
+
+ # The hook toolchains are per-repo flake outputs, so this is the job that
+ # first realizes them after a `git-hooks.nix` bump.
+ - name: Push to nix cache
+ if: ${{ !cancelled() }}
+ uses: genlayerlabs/github-actions/nix-cache-push@39b0a0d5e9bb27a1612d2e98b0f8509d31745157 # main
+ with:
+ token: ${{ secrets.NIX_CACHE_PUSH_TOKEN }}
+
+ # `check-commit-message` is a commit-msg-stage hook, so the pre-commit run
+ # above never fires it. Check every commit the PR adds, in the manager and in
+ # each executor submodule that moved.
+ #
+ # The App consumes this through the aggregate native CI result, so it must not
+ # be skippable on a run eligible for the merge train.
+ commit-messages:
+ if: inputs.pr != ''
+ runs-on: ubuntu-latest
+ timeout-minutes: 15
+ steps:
+ - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
+ with:
+ fetch-depth: 0
+ persist-credentials: false
+ - name: Get source
+ id: src
+ uses: ./.github/actions/get-src
+ with:
+ with_nix: "false"
+ # Commit messages need history, not vendored sources, and asking for
+ # them would install nix here for nothing.
+ third_party: none
+ pr: ${{ inputs.pr }}
+ # Needed even with with_nix=false: on a dispatched run there is no
+ # GITHUB_BASE_REF, so ci-changes.py resolves the base branch through
+ # the API and an unauthenticated `gh` exits 4.
+ github_token: ${{ secrets.GITHUB_TOKEN }}
+ - name: Check commit messages
+ env:
+ CHANGES: ${{ steps.src.outputs.changes }}
+ run: ./support/ci/run.sh pipeline commit-messages
- # The Merge action only fast-forwards, so the PR head must already
- # contain the base tip (0 commits behind). We assert it here too (as part
- # of the always-on `initial` checks) so the gate goes red the moment the
- # branch falls behind, instead of surfacing only at merge time.
+ # Repository policy requires the PR head to contain the base tip (0 commits
+ # behind). Assert it in the always-on checks so stale work is refreshed before
+ # its exact manager snapshot enters the App-owned merge train.
+ #
+ # The measurement itself lives in support/ci/behind.py, shared with
+ # rebase-watch, so both surfaces agree about the same PR.
behind-check:
- if: github.event_name == 'pull_request'
+ if: inputs.pr != ''
runs-on: ubuntu-latest
+ timeout-minutes: 15
steps:
- - uses: actions/checkout@v4
+ - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
with:
fetch-depth: 0
+ persist-credentials: false
- name: Ensure 0 commits behind base
env:
- BASE: ${{ github.event.pull_request.base.ref }}
- PR: ${{ github.event.pull_request.number }}
- run: |
- # Resolve base tip and PR head via explicit refs so this also
- # works for fork PRs (head.sha is not fetchable by sha there).
- git fetch --no-tags origin \
- "+refs/heads/${BASE}:refs/base" \
- "refs/pull/${PR}/head:refs/prhead"
- if ! git merge-base --is-ancestor refs/base refs/prhead; then
- BEHIND="$(git rev-list --count refs/prhead..refs/base)"
- echo "::error::PR is ${BEHIND} commit(s) behind ${BASE}; update/rebase the branch so it is 0 behind before merging"
- exit 1
- fi
- echo "0 commits behind ${BASE}"
+ # BASE is resolved from the PR by the pipeline itself, so this works
+ # on a dispatch run where the event carries no pull_request payload.
+ GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
+ PR_NUMBER: ${{ inputs.pr }}
+ run: ./support/ci/run.sh pipeline behind-check
+
+ # The E2E App squash-merges a PR as ` (#N)`, so the PR title becomes a
+ # commit subject verbatim and must satisfy the same rules as one.
+ pr-title:
+ if: inputs.pr != ''
+ runs-on: ubuntu-latest
+ timeout-minutes: 15
+ steps:
+ - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
+ with:
+ persist-credentials: false
+ - name: Check PR title is a valid commit subject
+ env:
+ # PR_TITLE likewise comes from the API, not the event payload.
+ GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
+ PR_NUMBER: ${{ inputs.pr }}
+ run: ./support/ci/run.sh pipeline pr-title
diff --git a/.github/workflows/incl_release_build_test.yaml b/.github/workflows/incl_release_build_test.yaml
new file mode 100644
index 00000000..46c85933
--- /dev/null
+++ b/.github/workflows/incl_release_build_test.yaml
@@ -0,0 +1,137 @@
+name: Build and test release artifacts
+
+# The whole release pipeline except publishing: plan -> build every artifact ->
+# test every platform against what was built.
+#
+# Callers see four artifacts, a tag and the executor lines; the matrices that
+# produce them are an implementation detail. `release.yaml` is this plus a
+# publisher; `queue.yaml` runs it on a PR carrying `test-release-pipeline` when
+# CI starts from the action panel or a push, exercising the pipeline itself
+# without releasing anything.
+
+on:
+ workflow_call:
+ inputs:
+ version_override:
+ description: "Release tag. Defaults to `version` from .genvm-monorepo-root"
+ required: false
+ type: string
+ default: ""
+ outputs:
+ tag:
+ description: the tag the artifacts were planned at
+ value: ${{ jobs.plan.outputs.tag }}
+ artifact_universal:
+ description: artifact holding the platform-independent runners
+ value: ${{ jobs.plan.outputs.artifact_universal }}
+ artifact_amd64_linux:
+ description: artifact holding amd64-linux's manager + executor tarballs
+ value: ${{ jobs.plan.outputs.artifact_amd64_linux }}
+ artifact_arm64_linux:
+ description: artifact holding arm64-linux's manager + executor tarballs
+ value: ${{ jobs.plan.outputs.artifact_arm64_linux }}
+ artifact_arm64_macos:
+ description: artifact holding arm64-macos's manager + executor tarballs
+ value: ${{ jobs.plan.outputs.artifact_arm64_macos }}
+ artifact_notes:
+ description: artifact holding the generated release notes
+ value: ${{ jobs.plan.outputs.artifact_notes }}
+ notes_file:
+ description: the notes file's name inside artifact_notes
+ value: ${{ jobs.plan.outputs.notes_file }}
+
+# A reusable workflow does not inherit the caller's `defaults`, so the shell
+# is set here too: GitHub's stock `bash -e {0}` has no `pipefail`.
+defaults:
+ run:
+ shell: bash -exo pipefail {0}
+
+jobs:
+ plan:
+ runs-on: ubuntu-latest
+ timeout-minutes: 15
+ outputs:
+ tag: ${{ steps.plan.outputs.tag }}
+ build_matrix: ${{ steps.plan.outputs.build_matrix }}
+ test_matrix: ${{ steps.plan.outputs.test_matrix }}
+ artifact_universal: ${{ steps.plan.outputs.artifact_universal }}
+ artifact_amd64_linux: ${{ steps.plan.outputs.artifact_amd64_linux }}
+ artifact_arm64_linux: ${{ steps.plan.outputs.artifact_arm64_linux }}
+ artifact_arm64_macos: ${{ steps.plan.outputs.artifact_arm64_macos }}
+ artifact_notes: ${{ steps.plan.outputs.artifact_notes }}
+ notes_file: ${{ steps.plan.outputs.notes_file }}
+ steps:
+ - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
+ with:
+ # Each line's manifest.json (its executor-version) lives in the
+ # submodule, so the gitlinks must be materialized.
+ submodules: true
+ - name: Plan release matrix
+ id: plan
+ env:
+ VERSION_OVERRIDE: ${{ inputs.version_override }}
+ run: ./support/ci/run.sh pipeline plan-release-matrix --version-override "$VERSION_OVERRIDE"
+
+ # Generated here rather than in the publisher so a PR-label run produces them
+ # too: they are worth reading before the release, not after. The range ends at
+ # HEAD, not at the tag — the tag will point here, but does not exist yet.
+ notes:
+ needs: plan
+ runs-on: ubuntu-latest
+ timeout-minutes: 15
+ steps:
+ - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
+ with:
+ # make-release-notes diffs against the previous tag via `git describe`.
+ fetch-depth: 0
+ - name: Generate release notes
+ env:
+ TAG: ${{ needs.plan.outputs.tag }}
+ NOTES_FILE: ${{ needs.plan.outputs.notes_file }}
+ run: |
+ set -euo pipefail
+ prev_tag=$(git describe --tags --abbrev=0 HEAD~1 2>/dev/null || echo "")
+ if [ -n "$prev_tag" ]; then
+ echo "notes for $TAG: ${prev_tag}..HEAD"
+ ./support/ci/run.sh tool make-release-notes "${prev_tag}..HEAD" > "$NOTES_FILE"
+ else
+ echo "Initial release" > "$NOTES_FILE"
+ fi
+ # Echoed so a debug run needs no artifact download.
+ cat "$NOTES_FILE"
+ - uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4.6.2
+ with:
+ name: ${{ needs.plan.outputs.artifact_notes }}
+ path: ${{ needs.plan.outputs.notes_file }}
+ if-no-files-found: error
+
+ build:
+ needs: plan
+ name: ${{ matrix.name }}
+ strategy:
+ fail-fast: false
+ matrix: ${{ fromJSON(needs.plan.outputs.build_matrix) }}
+ uses: ./.github/workflows/incl_release_build_test_cell_build.yaml
+ with:
+ name: ${{ matrix.name }}
+ runs_on: ${{ matrix.runs_on }}
+ artifact_name: ${{ matrix.artifact_name }}
+ job_json: ${{ matrix.job_json }}
+ secrets: inherit
+
+ test:
+ needs: [plan, build]
+ name: ${{ matrix.name }}
+ strategy:
+ fail-fast: false
+ matrix: ${{ fromJSON(needs.plan.outputs.test_matrix) }}
+ uses: ./.github/workflows/incl_release_build_test_cell_test.yaml
+ with:
+ name: ${{ matrix.name }}
+ runs_on: ${{ matrix.runs_on }}
+ platform_artifact: ${{ matrix.platform_artifact }}
+ runners_artifact: ${{ matrix.runners_artifact }}
+ with_nix: ${{ matrix.with_nix }}
+ setup_python: ${{ matrix.setup_python }}
+ job_json: ${{ matrix.job_json }}
+ secrets: inherit
diff --git a/.github/workflows/incl_release_build_test_cell_build.yaml b/.github/workflows/incl_release_build_test_cell_build.yaml
new file mode 100644
index 00000000..0b25bc5c
--- /dev/null
+++ b/.github/workflows/incl_release_build_test_cell_build.yaml
@@ -0,0 +1,90 @@
+name: Build one release artifact cell
+
+# One cell of plan-release-matrix's build matrix: replays the cell's `build`
+# steps (one build-and-pack per target) and uploads everything they packed under
+# the plan-assigned artifact name.
+#
+# The name comes from the plan rather than being minted here: every leg of a
+# matrix job writes the same `outputs` key, so a per-leg name cannot be handed
+# back to the test matrix as an output.
+
+on:
+ workflow_call:
+ inputs:
+ name:
+ required: true
+ type: string
+ runs_on:
+ required: true
+ type: string
+ artifact_name:
+ description: artifact to upload the cell's tarballs as
+ required: true
+ type: string
+ job_json:
+ # JSON {key: step list} from plan-release-matrix; matrix-cell-step
+ # slices out the list for its key.
+ required: true
+ type: string
+
+# A reusable workflow does not inherit the caller's `defaults`, so the shell
+# is set here too: GitHub's stock `bash -e {0}` has no `pipefail`.
+defaults:
+ run:
+ shell: bash -exo pipefail {0}
+
+jobs:
+ build:
+ name: build / ${{ inputs.name }}
+ runs-on: ${{ inputs.runs_on }}
+ timeout-minutes: 150
+ steps:
+ - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
+
+ # Guards the "Job:" step below: every key must have a step, every step a key.
+ - name: Check job keys
+ env:
+ JOB_JSON: ${{ inputs.job_json }}
+ run: ./support/ci/run.sh pipeline check-cell-keys --job-json "$JOB_JSON" build
+
+ - uses: insightsengineering/disk-space-reclaimer@dae9fabcb8febe09f6585471948acf9dc9a57489 # v1.1.2
+ with:
+ tools-cache: false
+ swap-storage: false
+ docker-images: false
+
+ - name: Get source
+ uses: ./.github/actions/get-src
+ with:
+ # The executor lines are submodules; nix builds run on
+ # `.?submodules=1` and need them checked out.
+ third_party: --all
+ with_nix: "true"
+ github_token: ${{ secrets.GITHUB_TOKEN }}
+ nix_cache_pull_token: ${{ secrets.NIX_CACHE_PULL_TOKEN }}
+
+ - name: "Job: build"
+ uses: ./.github/actions/matrix-cell-step
+ with:
+ job_json: ${{ inputs.job_json }}
+ key: build
+
+ # Not `success()`: a cell that failed late still built most of the
+ # closure, and that is exactly the work the retry should not repeat.
+ - name: Push to nix cache
+ if: ${{ !cancelled() }}
+ uses: genlayerlabs/github-actions/nix-cache-push@39b0a0d5e9bb27a1612d2e98b0f8509d31745157 # main
+ with:
+ token: ${{ secrets.NIX_CACHE_PUSH_TOKEN }}
+ include: genvm|genlayer|(^lief)
+ exclude: genvm-cpython-objs
+
+ - name: Upload artifacts
+ uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4.6.2
+ with:
+ name: ${{ inputs.artifact_name }}
+ # upload-artifact strips the common prefix, so entries land flat as
+ # genvm-.tar.xz — the names the test cells and the release
+ # uploader index by.
+ path: build/artifacts/*.tar.xz
+ if-no-files-found: error
diff --git a/.github/workflows/incl_release_build_test_cell_test.yaml b/.github/workflows/incl_release_build_test_cell_test.yaml
new file mode 100644
index 00000000..1ac09490
--- /dev/null
+++ b/.github/workflows/incl_release_build_test_cell_test.yaml
@@ -0,0 +1,133 @@
+name: Test one release platform cell
+
+# One cell of plan-release-matrix's test matrix. Materializes the built tarballs
+# into build/out and replays the cell's `test` steps against them: `genvm-tool
+# configure` writes build/info.json pointing at that tree, so the runner
+# exercises the downloaded artifacts and never builds.
+
+on:
+ workflow_call:
+ inputs:
+ name:
+ required: true
+ type: string
+ runs_on:
+ required: true
+ type: string
+ platform_artifact:
+ description: artifact holding this platform's manager + executor tarballs
+ required: true
+ type: string
+ runners_artifact:
+ description: artifact holding the universal runners; empty to skip
+ required: false
+ type: string
+ default: ""
+ with_nix:
+ required: false
+ type: boolean
+ default: true
+ setup_python:
+ required: false
+ type: boolean
+ default: false
+ job_json:
+ required: true
+ type: string
+
+# A reusable workflow does not inherit the caller's `defaults`, so the shell
+# is set here too: GitHub's stock `bash -e {0}` has no `pipefail`.
+defaults:
+ run:
+ shell: bash -exo pipefail {0}
+
+jobs:
+ test:
+ name: release-test / ${{ inputs.name }}
+ runs-on: ${{ inputs.runs_on }}
+ timeout-minutes: 90
+ steps:
+ - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
+
+ - name: Check job keys
+ env:
+ JOB_JSON: ${{ inputs.job_json }}
+ run: ./support/ci/run.sh pipeline check-cell-keys --job-json "$JOB_JSON" test
+
+ - name: Get source
+ uses: ./.github/actions/get-src
+ with:
+ # A nixless cell has no dev shell to build anything in: it replays
+ # prebuilt tarballs and reads no vendored source, so asking for them
+ # would install a nix nothing else here uses.
+ third_party: ${{ inputs.with_nix && '--all' || 'none' }}
+ with_nix: ${{ inputs.with_nix && 'true' || 'false' }}
+ github_token: ${{ secrets.GITHUB_TOKEN }}
+ nix_cache_pull_token: ${{ secrets.NIX_CACHE_PULL_TOKEN }}
+
+ - name: Setup Python
+ if: ${{ inputs.setup_python }}
+ uses: actions/setup-python@a26af69be951a213d495a4c3e4e4022e16d87065 # v5.6.0
+ with:
+ python-version: '3.12'
+
+ - name: Download platform artifacts
+ uses: actions/download-artifact@d3f86a106a0bac45b974a628896c90dbdf5c8093 # v4.3.0
+ with:
+ name: ${{ inputs.platform_artifact }}
+ path: ${{ runner.temp }}/release-artifacts
+
+ # Skipped for cells that cannot use them: a foreign-arch tree has nothing
+ # to run the runners with.
+ - name: Download runners artifact
+ if: ${{ inputs.runners_artifact != '' }}
+ uses: actions/download-artifact@d3f86a106a0bac45b974a628896c90dbdf5c8093 # v4.3.0
+ with:
+ name: ${{ inputs.runners_artifact }}
+ path: ${{ runner.temp }}/release-artifacts
+
+ - name: Materialize artifacts
+ env:
+ ARTIFACTS_DIR: ${{ runner.temp }}/release-artifacts
+ run: |
+ set -euo pipefail
+ # Every tarball — manager, each executor line, and the runners —
+ # overlays onto the same install root.
+ mkdir -p build/out
+ shopt -s nullglob
+ archives=("$ARTIFACTS_DIR"/*.tar.xz)
+ if [ "${#archives[@]}" -eq 0 ]; then
+ echo "::error::no tarballs in $ARTIFACTS_DIR"
+ exit 1
+ fi
+ printf 'extracting:\n'
+ printf ' %s\n' "${archives[@]}"
+ for archive in "${archives[@]}"; do
+ tar -xf "$archive" -C build/out --no-same-owner
+ done
+
+ - name: "Job: test"
+ uses: ./.github/actions/matrix-cell-step
+ with:
+ job_json: ${{ inputs.job_json }}
+ key: test
+
+ # A cell without nix has no store to push, and the macOS cells have no
+ # nix at all.
+ - name: Push to nix cache
+ if: ${{ !cancelled() && inputs.with_nix }}
+ uses: genlayerlabs/github-actions/nix-cache-push@39b0a0d5e9bb27a1612d2e98b0f8509d31745157 # main
+ with:
+ token: ${{ secrets.NIX_CACHE_PUSH_TOKEN }}
+
+ # Postmortem on the same runner the cell just used; see the action for
+ # what it probes and why it is not an `incl_*` reusable workflow.
+ # `failure()` covers any earlier step failing and is already false on a
+ # cancel, so it needs no !cancelled() guard.
+ # `continue-on-error` is belt-and-braces: the action already exits 0.
+ # Probes are Linux-only, so skipped on this workflow's macOS cells.
+ - name: Postmortem
+ if: ${{ failure() && !startsWith(inputs.runs_on, 'macos') }}
+ continue-on-error: true
+ timeout-minutes: 5
+ uses: ./.github/actions/postmortem
diff --git a/.github/workflows/queue.yaml b/.github/workflows/queue.yaml
index ec0c5714..1e650d90 100644
--- a/.github/workflows/queue.yaml
+++ b/.github/workflows/queue.yaml
@@ -1,143 +1,281 @@
name: GenVM full
-# Full GenVM CI. The repo has NO GitHub merge queue: this is the
-# authoritative "genvm CI" gate the Merge action
-# (branch_merge_into_dev.yaml) requires to be green on the exact head it
-# fast-forwards.
+# Full GenVM CI. The App-owned E2E merge train revalidates this native check on
+# the exact manager head before it squash-merges the tested snapshot.
#
# For every push to a PR targeting a dev branch (v-dev):
-# - `initial` (cheap pre-commit + 0-behind check) ALWAYS runs;
-# - the heavy test jobs run only when a run-full-tests marker is set —
-# the `rtm` (ready-to-merge) label or the `run-full-tests` label (the
-# "Force run full tests" checkbox sets the latter);
+# - `initial` (cheap pre-commit + 0-behind) and `check-executor-prs` (the
+# executor landability precondition) ALWAYS run, in parallel;
+# - `gate-check` aggregates them and plans the test matrix (plan-queue-matrix);
+# the heavy `matrix-tests` cells run when the `run-full-tests` label is set
+# or `/genvm-run-tests` dispatches a one-shot run. Neither authorizes the App
+# to merge.
+# Without a marker the plan is empty, so matrix-tests is skipped;
# - without a marker `validate-end` fails, so the CI check stays red
# until a full run has happened.
#
-# The action panel ("Force"/"Rerun full tests") triggers a run directly via
-# `workflow_dispatch` rather than relying on a label edit: a label applied by
-# the bot's GITHUB_TOKEN does NOT emit a `labeled` event (GitHub suppresses
-# recursive runs), but a `workflow_dispatch` from that same token DOES run.
-# The dispatch targets the PR head branch, so the run's head_sha equals the
-# PR head and the Merge gate (which matches by head_sha) counts it.
+# Labels are passive: adding `run-full-tests` or any unrelated label never starts
+# or cancels CI. "Force run full tests" on the action panel and the
+# `/genvm-run-tests` command dispatch this workflow explicitly; a push still
+# starts the always-on checks.
on:
pull_request:
branches: ['v*-dev']
- types: [opened, synchronize, reopened, ready_for_review, labeled]
+ # Deliberately NOT `edited`: this workflow's concurrency group cancels
+ # in-flight runs, so a typo fix in the PR body would kill a running
+ # matrix-tests cell (over an hour of work). The PR title can therefore be
+ # edited to something invalid after `initial / pr-title` went green. Title
+ # edits should be followed by a fresh run before the App merge is requested.
+ types: [opened, synchronize, reopened, ready_for_review]
workflow_dispatch:
inputs:
pr:
- description: PR number this manual run validates (used only for the run name)
+ description: >-
+ PR number supplied by the action panel and bound to the dispatch's
+ exact repository head before any CI job runs.
type: string
+ required: true
+ release_pipeline_test:
+ description: Exercise the release pipeline without publishing
+ type: boolean
required: false
+ default: false
+ request_comment:
+ description: "`/genvm-run-tests` comment id, empty for other triggers"
+ type: string
+ required: false
+ default: ""
+ request_id:
+ description: Unique `/genvm-run-tests` handler run and attempt
+ type: string
+ required: false
+ default: ""
+ expected_sha:
+ description: Exact manager SHA requested by `/genvm-run-tests`
+ type: string
+ required: false
+ default: ""
run-name: >-
- GenVM full${{ github.event_name == 'workflow_dispatch' && format(' (manual, PR #{0})', inputs.pr) || '' }}
+ GenVM full${{ github.event_name == 'workflow_dispatch' && format(' (manual, PR #{0})', inputs.pr) || '' }}${{ inputs.request_id != '' && format(' [request={0};comment={1};pr={2}]', inputs.request_id, inputs.request_comment, inputs.pr) || '' }}
+
+permissions:
+ contents: read
+ pull-requests: read
+
+# One in-flight full run per PR; a newer push cancels the stale run.
+# The PR number unifies push- and command-triggered runs. Cancelling a
+# superseded head's run is safe: the App only consumes CI that is green on the
+# exact head it merges, which is always the newest one. A command therefore
+# replaces an older run for the same PR rather than duplicating the full matrix.
+concurrency:
+ group: genvm-full-${{ github.event.pull_request.number || inputs.pr || github.run_id }}
+ cancel-in-progress: true
defaults:
run:
- shell: bash -x {0}
-
-env:
- GCS_BUCKET: "gh-af"
+ shell: bash -exo pipefail {0}
jobs:
+ validate-trigger:
+ runs-on: ubuntu-latest
+ timeout-minutes: 5
+ steps:
+ - name: Validate CI trigger
+ env:
+ GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
+ EVENT_NAME: ${{ github.event_name }}
+ DISPATCH_ACTOR: ${{ github.actor }}
+ TRIGGERING_ACTOR: ${{ github.triggering_actor }}
+ RUN_ATTEMPT: ${{ github.run_attempt }}
+ PR: ${{ inputs.pr }}
+ REQUEST_COMMENT: ${{ inputs.request_comment }}
+ REQUEST_ID: ${{ inputs.request_id }}
+ EXPECTED_SHA: ${{ inputs.expected_sha }}
+ RUN_SHA: ${{ github.sha }}
+ run: |
+ if [ "$RUN_ATTEMPT" != 1 ]; then
+ echo "::error::CI reruns must be started from the PR action panel"
+ exit 1
+ fi
+ if [ "$EVENT_NAME" != workflow_dispatch ]; then
+ exit 0
+ fi
+ if [ "$DISPATCH_ACTOR" != 'github-actions[bot]' ] || \
+ [ "$TRIGGERING_ACTOR" != 'github-actions[bot]' ]; then
+ echo "::error::full CI must be started from the PR action panel, not the Actions UI"
+ exit 1
+ fi
+ if [[ ! "$PR" =~ ^[1-9][0-9]*$ ]]; then
+ echo "::error::invalid PR number: $PR"
+ exit 1
+ fi
+ if [ -n "$REQUEST_COMMENT" ] && [[ ! "$REQUEST_COMMENT" =~ ^[1-9][0-9]*$ ]]; then
+ echo "::error::invalid request comment id: $REQUEST_COMMENT"
+ exit 1
+ fi
+ if [ -n "$REQUEST_ID" ] && [[ ! "$REQUEST_ID" =~ ^[1-9][0-9]*-[1-9][0-9]*$ ]]; then
+ echo "::error::invalid request id: $REQUEST_ID"
+ exit 1
+ fi
+ if { [ -n "$REQUEST_ID" ] && [ -z "$REQUEST_COMMENT" ]; } || \
+ { [ -z "$REQUEST_ID" ] && [ -n "$REQUEST_COMMENT" ]; }; then
+ echo "::error::request id and request comment must be supplied together"
+ exit 1
+ fi
+ if [ -n "$REQUEST_ID" ] && [ -z "$EXPECTED_SHA" ]; then
+ echo "::error::command dispatch must include the requested SHA"
+ exit 1
+ fi
+ if [ -n "$EXPECTED_SHA" ] && [[ ! "$EXPECTED_SHA" =~ ^[0-9a-f]{40}$ ]]; then
+ echo "::error::invalid expected SHA: $EXPECTED_SHA"
+ exit 1
+ fi
+ if [ -n "$EXPECTED_SHA" ] && [ "$EXPECTED_SHA" != "$RUN_SHA" ]; then
+ echo "::error::requested SHA is $EXPECTED_SHA, but dispatch resolved $RUN_SHA"
+ exit 1
+ fi
+
+ head="$(gh api "repos/$GITHUB_REPOSITORY/pulls/$PR" \
+ --jq '[.head.repo.full_name, .head.sha] | @tsv')"
+ IFS=$'\t' read -r head_repo head_sha <<<"$head"
+ if [ "$head_repo" != "$GITHUB_REPOSITORY" ]; then
+ echo "::error::the action panel cannot dispatch CI for a fork PR"
+ exit 1
+ fi
+ if [ "$head_sha" != "$RUN_SHA" ]; then
+ echo "::error::dispatch ref is $RUN_SHA, but PR #$PR is at $head_sha"
+ exit 1
+ fi
+
initial:
+ needs: validate-trigger
uses: ./.github/workflows/incl_initial.yaml
+ with:
+ # Resolved from whichever event triggered this run, so the PR-scoped
+ # checks run on BOTH a push and a panel dispatch. They used to test
+ # `github.event_name == 'pull_request'` themselves and so skipped on every
+ # dispatched run — while that run could still conclude `success` and be
+ # accepted by the App, which takes a green run on the exact head sha.
+ pr: ${{ github.event.pull_request.number || inputs.pr }}
secrets: inherit
- module-test-python:
- if: ${{ github.event_name == 'workflow_dispatch' || contains(github.event.pull_request.labels.*.name, 'rtm') || contains(github.event.pull_request.labels.*.name, 'run-full-tests') }}
- needs: [initial]
- runs-on: ubuntu-latest
- steps:
- - uses: actions/checkout@v4
- with:
- lfs: true
- - name: Get source
- uses: ./.github/actions/get-src
- with:
- load_submodules: "false"
- with_nix: "true"
- github_token: ${{ secrets.GITHUB_TOKEN }}
- - run: ./support/ci/pipelines/test-python.sh
-
- module-test-rust-fuzz:
- if: ${{ github.event_name == 'workflow_dispatch' || contains(github.event.pull_request.labels.*.name, 'rtm') || contains(github.event.pull_request.labels.*.name, 'run-full-tests') }}
- needs: [initial]
+ # Precondition (run.sh tool pr-branches check): every repo must be rebased on
+ # its base, and every executor line must have its pinned commit in the executor
+ # repo. Ahead-of-base executor work is expected: after the App
+ # lands manager, branch_sync_executors.yaml fast-forwards the executor dev refs.
+ # Read-only (gh API, no submodule checkout).
+ #
+ # Runs on same-repo PRs AND on workflow_dispatch with a PR number (the panel's
+ # rerun path, whose run the App consumes — so the precondition must hold
+ # there too; the PR number then comes from inputs.pr). Fork PRs are skipped: they
+ # get no secrets (the executor PAT would be empty) and cannot push executor mirror
+ # branches, so there is nothing to gate. It checks out the PR being validated
+ # (refs/pull//head) so the scripts and .genvm-monorepo-root match that PR — on
+ # dispatch the base branch is not in the event context, and its default-branch tree
+ # is the wrong commit. (Checking out PR code is not an extra exposure: for
+ # pull_request GitHub already runs this workflow file from the PR, so a same-repo
+ # author controls it regardless.)
+ check-executor-prs:
+ needs: validate-trigger
+ if: >
+ needs.validate-trigger.result == 'success' &&
+ ((github.event_name == 'workflow_dispatch' && inputs.pr != '') ||
+ (github.event_name == 'pull_request' &&
+ github.event.pull_request.head.repo.full_name == github.repository))
runs-on: ubuntu-latest
+ timeout-minutes: 15
steps:
- - uses: actions/checkout@v4
+ - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
with:
- lfs: true
- - uses: insightsengineering/disk-space-reclaimer@v1
- with:
- tools-cache: false
- swap-storage: false
- docker-images: false
- - name: Get source
- uses: ./.github/actions/get-src
- with:
- load_submodules: "true"
- with_nix: "true"
- github_token: ${{ secrets.GITHUB_TOKEN }}
- - run: |
- echo core | sudo tee /proc/sys/kernel/core_pattern
- echo performance | sudo tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor
- echo 0 | sudo tee /proc/sys/kernel/yama/ptrace_scope
- ./support/ci/pipelines/test-rust-fuzz.sh && ./support/ci/cargo-clippy.sh
+ ref: refs/pull/${{ github.event.pull_request.number || inputs.pr }}/head
+ # Read-only gh API calls; nothing here pushes, so the checkout
+ # credential should not survive into the PR code this job runs.
+ persist-credentials: false
+ - name: Executor gitlinks are available
env:
- OPENAIKEY: ${{ secrets.OPENAIKEY }}
- HEURISTKEY: ${{ secrets.HEURISTKEY }}
- ANTHROPICKEY: ${{ secrets.ANTHROPICKEY }}
- XAIKEY: ${{ secrets.XAIKEY }}
- GEMINIKEY: ${{ secrets.GEMINIKEY }}
+ PR_NUMBER: ${{ github.event.pull_request.number || inputs.pr }}
+ MANAGER_REPO: ${{ github.repository }}
+ MANAGER_TOKEN: ${{ secrets.GITHUB_TOKEN }}
+ EXECUTOR_TOKEN: ${{ secrets.GENVM_EXECUTOR_PR_TOKEN }}
+ run: ./support/ci/run.sh tool pr-branches check
- module-test-rust:
- if: ${{ github.event_name == 'workflow_dispatch' || contains(github.event.pull_request.labels.*.name, 'rtm') || contains(github.event.pull_request.labels.*.name, 'run-full-tests') }}
- needs: [initial]
+ gate-check:
+ needs: [validate-trigger, initial, check-executor-prs]
runs-on: ubuntu-latest
+ timeout-minutes: 15
+ if: ${{ always() }}
+ outputs:
+ matrix: ${{ steps.plan.outputs.matrix }}
steps:
- - uses: actions/checkout@v4
- with:
- lfs: true
- - uses: insightsengineering/disk-space-reclaimer@v1
- with:
- tools-cache: false
- swap-storage: false
- docker-images: false
- - name: Get source
- uses: ./.github/actions/get-src
- with:
- load_submodules: "true"
- with_nix: "true"
- github_token: ${{ secrets.GITHUB_TOKEN }}
- - name: Authenticate to Google Cloud
- uses: google-github-actions/auth@v2
+ - name: Check gate results
+ run: |
+ if [ "${{ needs.validate-trigger.result }}" != "success" ]; then
+ echo "::error::CI trigger validation failed"
+ exit 1
+ fi
+ if [ "${{ needs.initial.result }}" != "success" ]; then
+ echo "::error::initial checks failed"
+ exit 1
+ fi
+ # Executor precondition: fail on failure/cancelled. Skipped (fork PR or
+ # manual dispatch, where check-executor-prs does not run) is acceptable.
+ case "${{ needs.check-executor-prs.result }}" in
+ failure | cancelled)
+ echo "::error::executor mirror PR precondition failed"
+ exit 1
+ ;;
+ esac
+ - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
with:
- credentials_json: ${{ secrets.GCP_SA_KEY }}
- - name: Set up Cloud SDK # needed for gcloud to upload runners to GCS for caching
- uses: google-github-actions/setup-gcloud@v2
- - name: Set up Docker Buildx
- uses: docker/setup-buildx-action@v3
- - run: |
- ./support/ci/pipelines/test-rust.sh
+ persist-credentials: false
+ - name: Plan test matrix
+ id: plan
env:
- OPENAIKEY: ${{ secrets.OPENAIKEY }}
- HEURISTKEY: ${{ secrets.HEURISTKEY }}
- ANTHROPICKEY: ${{ secrets.ANTHROPICKEY }}
- XAIKEY: ${{ secrets.XAIKEY }}
- GEMINIKEY: ${{ secrets.GEMINIKEY }}
+ # Heavy cells run only with a marker; without one the plan is empty and
+ # matrix-tests is skipped.
+ RUN_FULL_TESTS: ${{ github.event_name == 'workflow_dispatch' || contains(github.event.pull_request.labels.*.name, 'run-full-tests') }}
+ run: ./support/ci/run.sh pipeline plan-queue-matrix
+
+ matrix-tests:
+ needs: gate-check
+ name: ${{ matrix.name }}
+ strategy:
+ fail-fast: false
+ matrix: ${{ fromJSON(needs.gate-check.outputs.matrix) }}
+ uses: ./.github/workflows/queue_test_cell.yaml
+ with:
+ name: ${{ matrix.name }}
+ job_json: ${{ matrix.job_json }}
+ runs_on: ${{ matrix.runs_on }}
+ disk_reclaim: ${{ matrix.disk_reclaim }}
+ fuzz_host: ${{ matrix.fuzz_host }}
+ buildx: ${{ matrix.buildx }}
+ gcp: ${{ matrix.gcp }}
+ secrets: inherit
+
+
+ # Opt-in: exercises the release pipeline itself (plan -> build every artifact
+ # -> test every platform) without publishing anything. Not part of
+ # validate-end's gate — it is skipped by default, and `skipped` there is a
+ # failure.
+ release-pipeline-test:
+ needs: [initial, gate-check]
+ if: ${{ (github.event_name == 'workflow_dispatch' && inputs.release_pipeline_test) || contains(github.event.pull_request.labels.*.name, 'test-release-pipeline') }}
+ uses: ./.github/workflows/incl_release_build_test.yaml
+ secrets: inherit
+
validate-end:
runs-on: ubuntu-latest
+ timeout-minutes: 15
if: ${{ always() }}
needs:
- initial
- - module-test-python
- - module-test-rust
- - module-test-rust-fuzz
+ - gate-check
+ - matrix-tests
env:
- HAS_MARKER: ${{ github.event_name == 'workflow_dispatch' || contains(github.event.pull_request.labels.*.name, 'rtm') || contains(github.event.pull_request.labels.*.name, 'run-full-tests') }}
+ HAS_MARKER: ${{ github.event_name == 'workflow_dispatch' || contains(github.event.pull_request.labels.*.name, 'run-full-tests') }}
steps:
- name: check
run: |
@@ -149,12 +287,16 @@ jobs:
# No run-full-tests marker -> heavy tests were skipped; keep the
# gate red so the PR is not mergeable until they run.
if [ "$HAS_MARKER" != "true" ]; then
- echo "::error::full tests have not run — add the 'rtm' label or check 'Force run full tests' on the action panel"
+ echo "::error::full tests have not run — tick 'Force run full tests' or comment '/genvm-run-tests'"
exit 1
fi
- results="${{ needs.module-test-python.result }} ${{ needs.module-test-rust.result }} ${{ needs.module-test-rust-fuzz.result }}"
+ results="${{ needs.matrix-tests.result }}"
echo "$results"
- if echo "$results" | grep -qiE '(failure|cancelled|skipped)'; then
+ # Herestring, not `echo "$results" |`: this is the last gate before a
+ # PR is mergeable, and under `pipefail` a non-zero left-hand side would
+ # make the condition false — i.e. fail OPEN, reporting green for a
+ # matrix that failed. Nothing should be able to swallow that.
+ if grep -qiE '(failure|cancelled|skipped)' <<<"$results"; then
echo "::error::one or more full-test jobs failed/cancelled/skipped"
exit 1
fi
diff --git a/.github/workflows/queue_test_cell.yaml b/.github/workflows/queue_test_cell.yaml
new file mode 100644
index 00000000..419711c7
--- /dev/null
+++ b/.github/workflows/queue_test_cell.yaml
@@ -0,0 +1,132 @@
+on:
+ workflow_call:
+ inputs:
+ name:
+ required: true
+ type: string
+ job_json:
+ # JSON {key: step list} from plan-queue-matrix, one key per "Job:" step
+ # below; matrix-cell-step slices out the list for its key.
+ required: true
+ type: string
+ runs_on:
+ required: true
+ type: string
+ disk_reclaim:
+ required: false
+ type: boolean
+ default: false
+ fuzz_host:
+ required: false
+ type: boolean
+ default: false
+ buildx:
+ required: false
+ type: boolean
+ default: false
+ gcp:
+ required: false
+ type: boolean
+ default: false
+
+# A reusable workflow does not inherit the caller's `defaults`, so the shell
+# is set here too: GitHub's stock `bash -e {0}` has no `pipefail`.
+defaults:
+ run:
+ shell: bash -exo pipefail {0}
+
+jobs:
+ test:
+ name: test / ${{ inputs.name }}
+ runs-on: ${{ inputs.runs_on }}
+ timeout-minutes: 240
+ env:
+ GCP_SA_KEY: ${{ secrets.GCP_SA_KEY }}
+ OPENAIKEY: ${{ secrets.OPENAIKEY }}
+ HEURISTKEY: ${{ secrets.HEURISTKEY }}
+ ANTHROPICKEY: ${{ secrets.ANTHROPICKEY }}
+ XAIKEY: ${{ secrets.XAIKEY }}
+ GEMINIKEY: ${{ secrets.GEMINIKEY }}
+
+ steps:
+ - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
+ with:
+ lfs: true
+ persist-credentials: false
+
+ # Guards the "Job:" steps below: every key must have a step, every step a key.
+ - name: Check job keys
+ env:
+ JOB_JSON: ${{ inputs.job_json }}
+ run: ./support/ci/run.sh pipeline check-cell-keys --job-json "$JOB_JSON" configure build tool
+
+ - uses: insightsengineering/disk-space-reclaimer@dae9fabcb8febe09f6585471948acf9dc9a57489 # v1.1.2
+ if: ${{ inputs.disk_reclaim }}
+ with:
+ tools-cache: false
+ swap-storage: false
+ docker-images: false
+
+ - name: Get source
+ uses: ./.github/actions/get-src
+ with:
+ third_party: --all
+ with_nix: "true"
+ github_token: ${{ secrets.GITHUB_TOKEN }}
+ nix_cache_pull_token: ${{ secrets.NIX_CACHE_PULL_TOKEN }}
+
+ - name: Configure fuzzing host
+ if: ${{ inputs.fuzz_host }}
+ run: |
+ echo core | sudo tee /proc/sys/kernel/core_pattern || true
+ echo performance | sudo tee /sys/devices/system/cpu/cpu*/cpufreq/scaling_governor || true
+ echo 0 | sudo tee /proc/sys/kernel/yama/ptrace_scope || true
+
+ - name: Authenticate to Google Cloud
+ if: ${{ inputs.gcp && env.GCP_SA_KEY != '' }}
+ uses: google-github-actions/auth@c200f3691d83b41bf9bbd8638997a462592937ed # v2.1.13
+ with:
+ credentials_json: ${{ secrets.GCP_SA_KEY }}
+
+ - name: Set up Cloud SDK
+ if: ${{ inputs.gcp && env.GCP_SA_KEY != '' }}
+ uses: google-github-actions/setup-gcloud@e427ad8a34f8676edf47cf7d7925499adf3eb74f # v2.2.1
+
+ - name: Set up Docker Buildx
+ if: ${{ inputs.buildx }}
+ uses: docker/setup-buildx-action@8d2750c68a42422c14e847fe6c8ac0403b4cbd6f # v3.12.0
+
+ - name: "Job: configure"
+ uses: ./.github/actions/matrix-cell-step
+ with:
+ job_json: ${{ inputs.job_json }}
+ key: configure
+
+ - name: "Job: build"
+ uses: ./.github/actions/matrix-cell-step
+ with:
+ job_json: ${{ inputs.job_json }}
+ key: build
+
+ - name: "Job: tool"
+ uses: ./.github/actions/matrix-cell-step
+ with:
+ job_json: ${{ inputs.job_json }}
+ key: tool
+
+ - name: Push to nix cache
+ if: ${{ !cancelled() }}
+ uses: genlayerlabs/github-actions/nix-cache-push@39b0a0d5e9bb27a1612d2e98b0f8509d31745157 # main
+ with:
+ token: ${{ secrets.NIX_CACHE_PUSH_TOKEN }}
+
+ # Postmortem on the same runner the cell just used; see the action for
+ # what it probes and why it is not an `incl_*` reusable workflow.
+ # `failure()` covers any earlier step failing and is already false on a
+ # cancel, so it needs no !cancelled() guard.
+ # `continue-on-error` is belt-and-braces: the action already exits 0.
+ - name: Postmortem
+ if: ${{ failure() }}
+ continue-on-error: true
+ timeout-minutes: 5
+ uses: ./.github/actions/postmortem
diff --git a/.github/workflows/release.yaml b/.github/workflows/release.yaml
index 5c84bba9..57b296d0 100644
--- a/.github/workflows/release.yaml
+++ b/.github/workflows/release.yaml
@@ -1,306 +1,155 @@
name: GenVM release
+# incl_release_build_test.yaml plus a publisher: it plans, builds every artifact
+# and tests every platform; this workflow attaches what it produced to a
+# genvm-manager release tagged with `version` from .genvm-monorepo-root.
+# `queue.yaml` runs the same pipeline on a PR carrying `test-release-pipeline`
+# when CI starts from the action panel or a push, so everything below is the
+# only release-only part.
+#
+# genvm--.tar.xz — manager + every active executor line for that
+# platform (bin/, lib/, config/, data/, executor/)
+# genvm-universal.tar.xz — shared runners under runners/ + legacy lines'
+# runners at executor//legacy-runners
+#
+# All four extract at the same install root: a full install is one platform
+# asset plus genvm-universal.
+#
+# Platforms are named the way nix names them (amd64-linux, arm64-linux,
+# arm64-macos) everywhere: flake attrs, cells and published assets.
+
on:
workflow_dispatch:
inputs:
- bump:
- type: choice
- required: false
- description: "Version bump type (ignored if version_override is set)"
- options:
- - patch
- - minor
- - major
- default: patch
version_override:
type: string
required: false
- description: "Override version (e.g., v1.2.3). If set, bump is ignored"
+ description: "Override release tag (e.g., v1.2.3). Defaults to `version` from .genvm-monorepo-root"
defaults:
run:
- shell: bash -x {0}
+ shell: bash -exo pipefail {0}
permissions:
actions: read
contents: read
-env:
- GCS_BUCKET: "gh-af"
-
jobs:
- gen-tag:
+ # Publishing is the only thing that cares whether the tag is free, so the
+ # check lives here rather than in the shared pipeline (which a PR runs at the
+ # same tag on every push). Guards the build so a doomed release fails in
+ # seconds instead of after three platform builds.
+ check-tag:
runs-on: ubuntu-latest
- outputs:
- tag: ${{ steps.determine-version.outputs.new_version }}
+ timeout-minutes: 15
steps:
- - uses: actions/checkout@v4
-
- - uses: actions-ecosystem/action-get-latest-tag@v1
- id: get-latest-tag
- if: github.event.inputs.version_override == ''
-
- - uses: actions-ecosystem/action-bump-semver@v1
- id: bump-semver
- if: github.event.inputs.version_override == ''
- with:
- current_version: ${{ steps.get-latest-tag.outputs.tag }}
- level: ${{ github.event.inputs.bump }}
-
- - name: Determine final version
- id: determine-version
+ - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
+ - name: Ensure tag is free
+ env:
+ VERSION_OVERRIDE: ${{ inputs.version_override }}
run: |
- if [ -n "${{ github.event.inputs.version_override }}" ]; then
- echo "Using override version: ${{ github.event.inputs.version_override }}"
- echo "new_version=${{ github.event.inputs.version_override }}" >> $GITHUB_OUTPUT
- else
- echo "Using bumped version: ${{ steps.bump-semver.outputs.new_version }}"
- echo "new_version=${{ steps.bump-semver.outputs.new_version }}" >> $GITHUB_OUTPUT
+ set -euo pipefail
+ tag="${VERSION_OVERRIDE:-$(jq -r .version .genvm-monorepo-root)}"
+ if git ls-remote --exit-code origin "refs/tags/$tag" > /dev/null; then
+ echo "::error::tag $tag already exists on origin"
+ exit 1
fi
- echo "new version will be: ${{ steps.determine-version.outputs.new_version }}"
-
- build-runners-all:
- needs: [gen-tag]
- uses: ./.github/workflows/incl_build_trg.yaml
- with:
- target: runners-all
- executor_version: ${{ needs.gen-tag.outputs.tag }}
- secrets:
- GCP_SA_KEY: ${{ secrets.GCP_SA_KEY }}
-
- build-linux-amd64:
- needs:
- - gen-tag
- uses: ./.github/workflows/incl_build_trg.yaml
- with:
- target: manager-amd64-linux
- secondary_target: executor-amd64-linux
- executor_version: ${{ needs.gen-tag.outputs.tag }}
- secrets:
- GCP_SA_KEY: ${{ secrets.GCP_SA_KEY }}
+ echo "tag $tag is free"
- build-linux-arm64:
- needs:
- - gen-tag
- uses: ./.github/workflows/incl_build_trg.yaml
+ # Plan -> build every artifact -> test every platform. Its matrices are an
+ # implementation detail; all we consume are the four artifact names.
+ build-and-test:
+ needs: [check-tag]
+ uses: ./.github/workflows/incl_release_build_test.yaml
with:
- target: manager-arm64-linux
- secondary_target: executor-arm64-linux
- executor_version: ${{ needs.gen-tag.outputs.tag }}
- secrets:
- GCP_SA_KEY: ${{ secrets.GCP_SA_KEY }}
+ version_override: ${{ inputs.version_override }}
+ secrets: inherit
- build-macos-arm64:
- needs:
- - gen-tag
- uses: ./.github/workflows/incl_build_trg.yaml
- with:
- target: manager-arm64-macos
- secondary_target: executor-arm64-macos
- executor_version: ${{ needs.gen-tag.outputs.tag }}
- secrets:
- GCP_SA_KEY: ${{ secrets.GCP_SA_KEY }}
-
- test-linux-arm64:
- needs:
- - build-linux-arm64
+ release-publish:
+ needs: [build-and-test]
runs-on: ubuntu-latest
+ timeout-minutes: 30
+ permissions:
+ contents: write # Needed for creating releases
+ id-token: write # Sigstore identity for the provenance attestation
+ attestations: write
steps:
- - uses: actions/checkout@v4
+ - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
with:
- lfs: true
- - name: Get source
- uses: ./.github/actions/get-src
- with:
- load_submodules: "false"
- with_nix: "true"
- - name: Download artifact
- run: |
- mkdir -p build/out && \
- pushd build/out && \
- curl -L --fail-with-body -H 'Accept: application/octet-stream' -o genvm-linux-arm64.tar.xz ${{ needs.build-linux-arm64.outputs.artifact_url }} && \
- tar -xf genvm-linux-arm64.tar.xz --no-same-owner && \
- curl -L --fail-with-body -H 'Accept: application/octet-stream' -o genvm-linux-arm64-executor.tar.xz ${{ needs.build-linux-arm64.outputs.secondary_artifact_url }} && \
- tar -xf genvm-linux-arm64-executor.tar.xz --no-same-owner && \
- popd && \
- true
- - name: Run post-install
- run: |
- ./build/out/bin/post-install.py \
- --error-on-missing-executor=false \
- --default-download=false || true
- nix develop .#check-qemu --command bash -c 'qemu-aarch64 ./build/out/bin/genvm-modules --version'
- - name: Smoke test modules startup
- run: |
- nix develop .#check-qemu --command bash -c '
- timeout 5 qemu-aarch64 ./build/out/bin/genvm-modules llm --allow-empty-backends --config ./build/out/config/genvm-module-llm.yaml || test $? -eq 124
- '
+ # Full history + tags: the release-notes step diffs against the
+ # previous tag via `git describe`.
+ fetch-depth: 0
- test-linux-amd64:
- needs:
- - build-linux-amd64
- - build-runners-all
- - gen-tag
- runs-on: ubuntu-latest
- steps:
- - uses: actions/checkout@v4
+ - name: Download release notes
+ uses: actions/download-artifact@d3f86a106a0bac45b974a628896c90dbdf5c8093 # v4.3.0
with:
- lfs: true
- - name: Get source
- uses: ./.github/actions/get-src
+ name: ${{ needs.build-and-test.outputs.artifact_notes }}
+ path: artifacts/notes
+ # Each artifact holds exactly one tarball, already named as the asset it
+ # is published as, so they all land in one dir and need no renaming.
+ - name: Download universal artifact
+ uses: actions/download-artifact@d3f86a106a0bac45b974a628896c90dbdf5c8093 # v4.3.0
with:
- load_submodules: "false"
- with_nix: "true"
- - name: Download artifact
- run: |
- mkdir -p build/out && \
- pushd build/out && \
- curl -L --fail-with-body -H 'Accept: application/octet-stream' -o genvm-linux-amd64.tar.xz ${{ needs.build-linux-amd64.outputs.artifact_url }} && \
- curl -L --fail-with-body -H 'Accept: application/octet-stream' -o genvm-runners-all.tar.xz ${{ needs.build-runners-all.outputs.artifact_url }} && \
- curl -L --fail-with-body -H 'Accept: application/octet-stream' -o genvm-linux-amd64-executor.tar.xz ${{ needs.build-linux-amd64.outputs.secondary_artifact_url }} && \
- tar -xf genvm-linux-amd64.tar.xz --no-same-owner && \
- tar -xf genvm-runners-all.tar.xz --no-same-owner && \
- tar -xf genvm-linux-amd64-executor.tar.xz --no-same-owner && \
- popd && \
- true
- - name: Run post-install
- run: |
- ./build/out/bin/post-install.py \
- --error-on-missing-executor=false \
- --default-download=false
- - name: Smoke test modules startup
- run: |
- timeout 5 ./build/out/bin/genvm-modules llm --allow-empty-backends --config ./build/out/config/genvm-module-llm.yaml || test $? -eq 124
- - name: Run stable tests
- run: |
- nix develop .#mock-tests --command genvm-tool test run --filter-tag "$(cat tests/presets/release.txt)" --genvm-reroute-to ${{ needs.gen-tag.outputs.tag }}
-
- test-macos-arm64:
- needs:
- - build-macos-arm64
- - build-runners-all
- - gen-tag
- runs-on: macos-latest
- steps:
- - uses: actions/checkout@v4
+ name: ${{ needs.build-and-test.outputs.artifact_universal }}
+ path: release-assets
+ - name: Download amd64-linux artifact
+ uses: actions/download-artifact@d3f86a106a0bac45b974a628896c90dbdf5c8093 # v4.3.0
with:
- lfs: true
- - name: Get source
- uses: ./.github/actions/get-src
+ name: ${{ needs.build-and-test.outputs.artifact_amd64_linux }}
+ path: release-assets
+ - name: Download arm64-linux artifact
+ uses: actions/download-artifact@d3f86a106a0bac45b974a628896c90dbdf5c8093 # v4.3.0
with:
- load_submodules: "false"
- with_nix: "false"
- - name: Download artifact
- run: |
- mkdir -p build/out && \
- pushd build/out && \
- curl -L --fail-with-body -H 'Accept: application/octet-stream' -o genvm-macos-arm64.tar.xz ${{ needs.build-macos-arm64.outputs.artifact_url }} && \
- curl -L --fail-with-body -H 'Accept: application/octet-stream' -o genvm-runners-all.tar.xz ${{ needs.build-runners-all.outputs.artifact_url }} && \
- curl -L --fail-with-body -H 'Accept: application/octet-stream' -o genvm-macos-arm64-executor.tar.xz ${{ needs.build-macos-arm64.outputs.secondary_artifact_url }} && \
- tar -xf genvm-macos-arm64.tar.xz --no-same-owner && \
- tar -xf genvm-runners-all.tar.xz --no-same-owner && \
- tar -xf genvm-macos-arm64-executor.tar.xz --no-same-owner && \
- popd && \
- true
- - name: Get wabt
- run: |
- brew install wabt
- - name: Setup Rust
- run: |
- rustup target add wasm32-wasip1
- - name: Setup Python
- uses: actions/setup-python@v5
+ name: ${{ needs.build-and-test.outputs.artifact_arm64_linux }}
+ path: release-assets
+ - name: Download arm64-macos artifact
+ uses: actions/download-artifact@d3f86a106a0bac45b974a628896c90dbdf5c8093 # v4.3.0
with:
- python-version: '3.12'
- - name: Run post-install
- run: |
- ./build/out/bin/post-install.py \
- --error-on-missing-executor=false \
- --default-download=false
- - name: Smoke test modules startup
+ name: ${{ needs.build-and-test.outputs.artifact_arm64_macos }}
+ path: release-assets
+
+ - name: Check release assets
run: |
- brew install coreutils
- gtimeout 5 ./build/out/bin/genvm-modules llm --allow-empty-backends --config ./build/out/config/genvm-module-llm.yaml || test $? -eq 124
- - name: Run stable tests
+ set -euo pipefail
+ ls -l release-assets
+ for asset in genvm-universal genvm-amd64-linux genvm-arm64-linux genvm-arm64-macos; do
+ test -f "release-assets/$asset.tar.xz" || {
+ echo "::error::missing release asset $asset.tar.xz"
+ exit 1
+ }
+ done
+
+ # Plain checksums for a verifier without `gh`; the attestation below is
+ # what binds them to this workflow on this tag.
+ - name: Write checksums
run: |
- python3 -m venv .venv
- source .venv/bin/activate
- pip install aiohttp jsonnet
- ./support/tools/genvm-tool/genvm-tool test run --filter-tag "$(cat tests/presets/release.txt)" --genvm-reroute-to ${{ needs.gen-tag.outputs.tag }}
-
- release-publish:
- needs:
- - gen-tag
- - build-linux-amd64
- - build-linux-arm64
- - build-macos-arm64
- - build-runners-all
- - test-linux-amd64
- - test-linux-arm64
- - test-macos-arm64
- runs-on: ubuntu-latest
- permissions:
- contents: write # Needed for creating releases
- steps:
- - uses: actions/checkout@v4
- with:
- lfs: true
- - run: sudo apt-get install -y python3-poetry
- - uses: actions/setup-python@v5
+ set -euo pipefail
+ cd release-assets
+ sha256sum *.tar.xz | tee SHA256SUMS
+
+ # Keyless sigstore provenance, verifiable with
+ # `gh attestation verify --owner genlayerlabs`.
+ - name: Attest provenance
+ uses: actions/attest-build-provenance@96278af6caaf10aea03fd8d33a09a777ca52d62f # v3.2.0
with:
- python-version: '3.12'
- cache: poetry
- # PyPI publishing disabled for now (genlayer-py-std lives in the
- # executor submodule and will be released from its own repo).
- # - name: Publish to test pypi
- # run: |
- # python3.12 -m pip install poetry && \
- # pushd executors/v0.3.x/runners/genlayer-py-std && \
- # perl -i -pe 's/version = "v0.0.1"/version = "${{ needs.gen-tag.outputs.tag }}"/' pyproject.toml && \
- # poetry build && \
- # poetry config repositories.test-pypi https://test.pypi.org/legacy/ && \
- # poetry config pypi-token.test-pypi ${{ secrets.TEST_PYPI_TOKEN }} && \
- # poetry publish -r test-pypi && \
- # popd
- - run: |
- curl -L --fail-with-body -H 'Accept: application/octet-stream' -o genvm-runners-all.tar.xz ${{ needs.build-runners-all.outputs.artifact_url }} && \
- curl -L --fail-with-body -H 'Accept: application/octet-stream' -o genvm-linux-amd64.tar.xz ${{ needs.build-linux-amd64.outputs.artifact_url }} && \
- curl -L --fail-with-body -H 'Accept: application/octet-stream' -o genvm-linux-arm64.tar.xz ${{ needs.build-linux-arm64.outputs.artifact_url }} && \
- curl -L --fail-with-body -H 'Accept: application/octet-stream' -o genvm-macos-arm64.tar.xz ${{ needs.build-macos-arm64.outputs.artifact_url }} && \
- curl -L --fail-with-body -H 'Accept: application/octet-stream' -o genvm-linux-amd64-executor.tar.xz ${{ needs.build-linux-amd64.outputs.secondary_artifact_url }} && \
- curl -L --fail-with-body -H 'Accept: application/octet-stream' -o genvm-linux-arm64-executor.tar.xz ${{ needs.build-linux-arm64.outputs.secondary_artifact_url }} && \
- curl -L --fail-with-body -H 'Accept: application/octet-stream' -o genvm-macos-arm64-executor.tar.xz ${{ needs.build-macos-arm64.outputs.secondary_artifact_url }} && \
- true
+ subject-path: release-assets/*
- - run: |
- git tag ${{ needs.gen-tag.outputs.tag }} && \
- git push origin ${{ needs.gen-tag.outputs.tag }}
-
- - name: Generate release notes
- id: release_notes
+ - name: Tag the release
+ env:
+ TAG: ${{ needs.build-and-test.outputs.tag }}
run: |
- prev_tag=$(git describe --tags --abbrev=0 HEAD~1 2>/dev/null || echo "")
- new_tag="${{ needs.gen-tag.outputs.tag }}"
- if [ -n "$prev_tag" ]; then
- python3 support/ci/make-relase-notes.py "${prev_tag}..${new_tag}" > release-notes.md
- else
- echo "Initial release" > release-notes.md
- fi
+ set -euo pipefail
+ git tag "$TAG"
+ git push origin "$TAG"
- name: Create Release
- id: create_release
- uses: softprops/action-gh-release@v2
+ uses: softprops/action-gh-release@3bb12739c298aeb8a4eeaf626c5b8d85266b0e65 # v2.6.2
with:
files: |
- genvm-runners-all.tar.xz
- genvm-linux-amd64.tar.xz
- genvm-linux-arm64.tar.xz
- genvm-macos-arm64.tar.xz
- genvm-linux-amd64-executor.tar.xz
- genvm-linux-arm64-executor.tar.xz
- genvm-macos-arm64-executor.tar.xz
- name: Release ${{ needs.gen-tag.outputs.tag }}
- tag_name: ${{ needs.gen-tag.outputs.tag }}
- body_path: release-notes.md
+ release-assets/*.tar.xz
+ release-assets/SHA256SUMS
+ name: Release ${{ needs.build-and-test.outputs.tag }}
+ tag_name: ${{ needs.build-and-test.outputs.tag }}
+ body_path: artifacts/notes/${{ needs.build-and-test.outputs.notes_file }}
draft: false
- prerelease: false
+ prerelease: ${{ contains(needs.build-and-test.outputs.tag, '-') }}
diff --git a/.github/workflows/webdriver-image.yaml b/.github/workflows/webdriver-image.yaml
index 788d7bff..ae991753 100644
--- a/.github/workflows/webdriver-image.yaml
+++ b/.github/workflows/webdriver-image.yaml
@@ -16,23 +16,24 @@ jobs:
docker-image:
name: Build and push webdriver image
runs-on: ubuntu-latest
+ timeout-minutes: 90
permissions:
contents: read
steps:
- name: Checkout code
- uses: actions/checkout@v4
+ uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
- name: Log in to Docker Hub
- uses: docker/login-action@v3
+ uses: docker/login-action@c94ce9fb468520275223c153574b00df6fe4bcc9 # v3.7.0
with:
username: yeagerai
password: ${{ secrets.DOCKER_PASSWORD }}
- name: Docker Metadata
id: meta
- uses: docker/metadata-action@v5
+ uses: docker/metadata-action@c299e40c65443455700f0fdfc63efafe5b349051 # v5.10.0
with:
images: |
yeagerai/genlayer-genvm-webdriver
@@ -41,13 +42,13 @@ jobs:
type=raw,value=latest
- name: Set up QEMU for docker
- uses: docker/setup-qemu-action@v3
+ uses: docker/setup-qemu-action@c7c53464625b32c7a7e944ae62b3e17d2b600130 # v3.7.0
- name: Set up Docker Buildx
- uses: docker/setup-buildx-action@v3
+ uses: docker/setup-buildx-action@8d2750c68a42422c14e847fe6c8ac0403b4cbd6f # v3.12.0
- name: Build and push image
- uses: docker/build-push-action@v6
+ uses: docker/build-push-action@10e90e3645eae34f1e60eeb005ba3a3d33f178e8 # v6.19.2
with:
context: webdriver
push: true
diff --git a/.gitignore b/.gitignore
index b6a017df..771520e7 100644
--- a/.gitignore
+++ b/.gitignore
@@ -3,7 +3,9 @@
!**/.vscode/extensions.json
!**/.vscode/settings.example.json
+/.pre-commit-config.yaml
.claude/hooks
+.claude/settings.local.json
.idea
.envrc
*.log
diff --git a/.gitmodules b/.gitmodules
index 0b2bd87e..8e1c6ef4 100644
--- a/.gitmodules
+++ b/.gitmodules
@@ -2,3 +2,11 @@
path = executors/v0.3.x
url = ../genvm-executor
branch = v0.3.x
+[submodule "executors/v0.2.x"]
+ path = executors/v0.2.x
+ url = ../genvm-executor
+ branch = v0.2.x
+[submodule "libs/unhardcoded-engine"]
+ path = libs/unhardcoded-engine
+ url = ../unhardcoded-engine
+ branch = main
diff --git a/.taplo.toml b/.taplo.toml
new file mode 100644
index 00000000..e17930bd
--- /dev/null
+++ b/.taplo.toml
@@ -0,0 +1,12 @@
+[formatting]
+align_comments = false
+indent_string = " "
+reorder_keys = true
+
+[[rule]]
+exclude = ["executors/**"]
+include = ["**/Cargo.toml"]
+keys = ["package"]
+
+[rule.formatting]
+reorder_keys = false
diff --git a/AGENTS.md b/AGENTS.md
new file mode 100644
index 00000000..796c96fe
--- /dev/null
+++ b/AGENTS.md
@@ -0,0 +1,63 @@
+# GenVM manager
+
+Before changing this repository, read the matching page in `docs/contributing/`:
+a task has a how-to, a design question has an explanation
+
+## Executor Consensus Invariants
+
+1. Treat leader-proposed output as untrusted input; validate it before use or publication
+2. Apply consensus-affecting normalization, caps and accounting identically in leader, validator and sync paths
+3. For timeout and resource preflights, panic if the checked operation later fails under the same conditions; a broken executor invariant must not become a contract-visible error
+
+
+
+## Tutorial
+- [first-contribution.md](docs/contributing/tutorial/first-contribution.md) — patch an executor, branch, commit, push, open the PR
+
+## Howto
+- [genvm-tool.md](docs/contributing/howto/genvm-tool.md) — the umbrella tool: man page, test and git subcommands, codegen
+- [pr.md](docs/contributing/howto/pr.md) — branch model, PR action panel, App-owned landing and executor projection
+- [review-ready.md](docs/contributing/howto/review-ready.md) — the checks a change must pass before it is put up for review
+- [setup.md](docs/contributing/howto/setup.md) — first-time clone: submodules, vendored trees, nix/direnv dev shell
+
+## Howto/Building
+- [build.md](docs/contributing/howto/building/build.md) — debug build: configure + ninja, targets, outputs, cargo quirks
+- [docs.md](docs/contributing/howto/building/docs.md) — building and publishing the website, spec vs impl-spec, ADRs
+- [runners.md](docs/contributing/howto/building/runners.md) — where runners come from: build on Linux, download elsewhere
+
+## Howto/Committing
+- [git-third-party.md](docs/contributing/howto/committing/git-third-party.md) — how vendored trees (wasmtime, …) are pinned and patched
+- [runners.md](docs/contributing/howto/committing/runners.md) — clearing runner dev-mode and refreshing hashes before a commit
+- [submodules.md](docs/contributing/howto/committing/submodules.md) — repo topology, gitlink bumps, pre-commit hooks, push order
+
+## Howto/Docs
+- [style.md](docs/contributing/howto/docs/style.md) — prose conventions for guides, specs, ADRs and commit bodies
+
+## Howto/Extending
+- [add-a-skill.md](docs/contributing/howto/extending/add-a-skill.md) — agent skills in `.agents/`, the frontmatter trigger, the symlink, skill vs how-to
+- [add-host-function.md](docs/contributing/howto/extending/add-host-function.md) — new executor↔host protocol method
+- [add-llm-provider.md](docs/contributing/howto/extending/add-llm-provider.md) — new LLM backend in the manager
+- [add-wasi-function.md](docs/contributing/howto/extending/add-wasi-function.md) — new gl_call method or raw WASI function
+- [modify-runner.md](docs/contributing/howto/extending/modify-runner.md) — runner dev-mode and hash refresh
+- [modify-wasmtime.md](docs/contributing/howto/extending/modify-wasmtime.md) — patching vendored wasmtime, trap plumbing
+- [write-a-script.md](docs/contributing/howto/extending/write-a-script.md) — conventions for helper scripts and pre-commit hooks
+
+## Howto/Releasing
+- [release-build.md](docs/contributing/howto/releasing/release-build.md) — nix packages, platforms, release assets
+- [versioning.md](docs/contributing/howto/releasing/versioning.md) — release trains, `.genvm-monorepo-root`, version tools
+
+## Howto/Testing
+- [README.md](docs/contributing/howto/testing/README.md) — `genvm-tool test`: filters, presets, continue files
+- [fuzzing.md](docs/contributing/howto/testing/fuzzing.md) — AFL fuzz targets, seeding a corpus from a run, host sysctl prep
+- [integration.md](docs/contributing/howto/testing/integration.md) — jsonnet cases, tags, golden `.stdout`/`.hash` files, services
+- [python.md](docs/contributing/howto/testing/python.md) — Python tests, direct pytest for genlayer-py-std
+- [rust.md](docs/contributing/howto/testing/rust.md) — Rust tests: where they go, how to run them, coverage
+- [typescript.md](docs/contributing/howto/testing/typescript.md) — information on how to run and write webdriver typescript tests
+
+## Explanation
+- [docs-layout.md](docs/contributing/explanation/docs-layout.md) — the 4 kinds of page and where each belongs
+- [executor-lines.md](docs/contributing/explanation/executor-lines.md) — why several executor lines ship side by side, and what it costs
+- [fuzz.md](docs/contributing/explanation/fuzz.md) — why fuzz targets get fake entropy and no CmpLog
+- [merge-model.md](docs/contributing/explanation/merge-model.md) — why manager branch tips own executor refs
+- [shared-submodule-cache.md](docs/contributing/explanation/shared-submodule-cache.md) — why submodules are worktrees of one cache repo, not clones
+- [vendored-trees.md](docs/contributing/explanation/vendored-trees.md) — why third-party sources are patch series, not forks
diff --git a/CLAUDE.md b/CLAUDE.md
new file mode 120000
index 00000000..47dc3e3d
--- /dev/null
+++ b/CLAUDE.md
@@ -0,0 +1 @@
+AGENTS.md
\ No newline at end of file
diff --git a/README.md b/README.md
index 5f4a0ec0..79ffafcb 100644
--- a/README.md
+++ b/README.md
@@ -33,38 +33,9 @@ This is a monorepo for GenVM. It is composed of the following sub-projects:
## Install
-Required tools:
-
-- git
-- ruby (3.\*)
-- ninja
-- rustup (cargo + rustc)
-- (for runners) nix and an x86_64 system
-
-All of them (except git, for obvious reasons) are provided by the default shell in
-`build-scripts/devenv/flake.nix` (for direnv add `use flake ./build-scripts/devenv`).
-
-### Debug build
-
-1. `cd $PROJECT_DIR`
-2. `git submodule update --init --recursive --depth 1`
-3. `source env.sh` (not needed if you used the flake)
-4. `git third-party update --all`
-5. `genvm-tool configure` — scrapes and configures all targets (similar to CMake).
- Outside the dev shell use `support/tools/genvm-tool/genvm-tool configure`.
-6. `ninja -C build` (or `ninja -C build all/bin`) — output is at `build/out` as a
- root (`bin`, `share`)
-7. Get `genvm-runners.zip` from [GitHub][genvm-manager]
-8. Merge `build/out` and `genvm-runners.zip`
-
-### Production build
-
-> WARNING: currently supported only on x86_64 Linux hosts.
-
-1. `cd $PROJECT_DIR`
-2. `nix build -o build/out-universal -v -L .#all-for-platform.universal`
-3. `nix build -o build/out-amd64-linux -v -L .#all-for-platform.amd64-linux`
-4. Merge outputs
+See [docs/contributing/howto/setup.md](./docs/contributing/howto/setup.md) and
+[docs/contributing/howto/build.md](./docs/contributing/howto/building/build.md) for
+environment setup and debug/production build instructions.
## Usage
diff --git a/SECURITY.md b/SECURITY.md
index eefc9fc1..ffd2565c 100644
--- a/SECURITY.md
+++ b/SECURITY.md
@@ -5,9 +5,13 @@ identical and trustworthy across all validators. Please treat security issues ac
## Reporting a vulnerability
-**Do not open a public issue.** Report privately via GitHub's
-[private vulnerability reporting](https://github.com/genlayerlabs/genvm/security/advisories/new),
-or email kira@yeager.ai
+**Before mainnet, report everything except remote code execution publicly** — open a
+regular issue. Until there is value at stake, an open report gets triaged faster and is
+useful to everyone reading along. RCE is the only exception; report it privately.
+
+For remote code execution, **do not open a public issue** — report it via GitHub's
+[private vulnerability reporting](https://github.com/genlayerlabs/genvm-manager/security/advisories/new)
+on the manager repository.
Include a description, affected component/version, and a reproduction (a contract, calldata,
or test case) where possible. We aim to acknowledge within a few business days.
@@ -37,3 +41,18 @@ In scope: the executor, runners/SDK, modules (LLM/web/manager), the install/mani
the CI/release supply chain. Out of scope: issues requiring a pre-compromised host or operator
machine, and non-default deployments that expose the loopback-only manager to untrusted networks
(though we still want to know).
+
+## Trust boundary
+
+The following relationships are trusted. Hardening them is welcome, but a report that assumes
+one side is hostile is not treated as a vulnerability:
+
+- Host and GenVM
+- Executor and manager
+- The local disk and loopback in general
+
+The following inputs are untrusted, even when delivered through a trusted component:
+
+- Intelligent Contract code and contract-controlled data, including calldata, messages, and persisted values
+- Data originating from other validators, including leader results
+- External content processed by modules, including HTTP responses, redirects, rendered pages, JavaScript, subresources, and model-provider responses
diff --git a/crates/fuzzing/Cargo.lock b/crates/fuzzing/Cargo.lock
new file mode 100644
index 00000000..9a86e14f
--- /dev/null
+++ b/crates/fuzzing/Cargo.lock
@@ -0,0 +1,177 @@
+# This file is automatically @generated by Cargo.
+# It is not intended for manual editing.
+version = 4
+
+[[package]]
+name = "cobs"
+version = "0.3.0"
+source = "registry+https://github.com/rust-lang/crates.io-index"
+checksum = "0fa961b519f0b462e3a3b4a34b64d119eeaca1d59af726fe450bbba07a9fc0a1"
+dependencies = [
+ "thiserror",
+]
+
+[[package]]
+name = "embedded-io"
+version = "0.4.0"
+source = "registry+https://github.com/rust-lang/crates.io-index"
+checksum = "ef1a6892d9eef45c8fa6b9e0086428a2cca8491aca8f787c534a3d6d0bcb3ced"
+
+[[package]]
+name = "embedded-io"
+version = "0.6.1"
+source = "registry+https://github.com/rust-lang/crates.io-index"
+checksum = "edd0f118536f44f5ccd48bcb8b111bdc3de888b58c74639dfb034a357d0f206d"
+
+[[package]]
+name = "genvm-fuzzing"
+version = "0.1.0"
+dependencies = [
+ "mutatis",
+ "postcard",
+ "serde",
+]
+
+[[package]]
+name = "mutatis"
+version = "0.5.3"
+source = "registry+https://github.com/rust-lang/crates.io-index"
+checksum = "468ca2a8bc8087a0b2c11d21e7dca51d92cc2662a6993d6b56bac2701795f202"
+dependencies = [
+ "mutatis-derive",
+ "rand",
+]
+
+[[package]]
+name = "mutatis-derive"
+version = "0.5.3"
+source = "registry+https://github.com/rust-lang/crates.io-index"
+checksum = "5708f6e65277d4a89733aff227b43d0989c5ead03999b6f1f33a8ec7a9978361"
+dependencies = [
+ "proc-macro2",
+ "quote",
+ "syn 2.0.119",
+]
+
+[[package]]
+name = "postcard"
+version = "1.1.3"
+source = "registry+https://github.com/rust-lang/crates.io-index"
+checksum = "6764c3b5dd454e283a30e6dfe78e9b31096d9e32036b5d1eaac7a6119ccb9a24"
+dependencies = [
+ "cobs",
+ "embedded-io 0.4.0",
+ "embedded-io 0.6.1",
+ "serde",
+]
+
+[[package]]
+name = "proc-macro2"
+version = "1.0.107"
+source = "registry+https://github.com/rust-lang/crates.io-index"
+checksum = "985e7ec9bb745e6ce6535b544d84d6cd6f7ad8bd711c398938ae983b91a766d9"
+dependencies = [
+ "unicode-ident",
+]
+
+[[package]]
+name = "quote"
+version = "1.0.47"
+source = "registry+https://github.com/rust-lang/crates.io-index"
+checksum = "1fbf4db142a473a8d80c26bbf18454ed458bf8d26c8219c331daecfdbd079001"
+dependencies = [
+ "proc-macro2",
+]
+
+[[package]]
+name = "rand"
+version = "0.8.7"
+source = "registry+https://github.com/rust-lang/crates.io-index"
+checksum = "22f6172bdec972074665ed81ed53b71da00bfc44b65a753cfde883ec4c702a1a"
+dependencies = [
+ "rand_core",
+]
+
+[[package]]
+name = "rand_core"
+version = "0.6.4"
+source = "registry+https://github.com/rust-lang/crates.io-index"
+checksum = "ec0be4795e2f6a28069bec0b5ff3e2ac9bafc99e6a9a7dc3547996c5c816922c"
+
+[[package]]
+name = "serde"
+version = "1.0.229"
+source = "registry+https://github.com/rust-lang/crates.io-index"
+checksum = "4148590afebada386688f18773da617792bf2ef03ffc1e4cbd2b1d45b023e0ba"
+dependencies = [
+ "serde_core",
+ "serde_derive",
+]
+
+[[package]]
+name = "serde_core"
+version = "1.0.229"
+source = "registry+https://github.com/rust-lang/crates.io-index"
+checksum = "67dca2c9c51e58a4791a4b1ed58308b39c64224d349a935ab5039aa360942a48"
+dependencies = [
+ "serde_derive",
+]
+
+[[package]]
+name = "serde_derive"
+version = "1.0.229"
+source = "registry+https://github.com/rust-lang/crates.io-index"
+checksum = "e7a5d71263a5a7d47b41f6b3f06ba276f10cc18b0931f1799f710578e2309348"
+dependencies = [
+ "proc-macro2",
+ "quote",
+ "syn 3.0.3",
+]
+
+[[package]]
+name = "syn"
+version = "2.0.119"
+source = "registry+https://github.com/rust-lang/crates.io-index"
+checksum = "872831b642d1a07999a962a351ed35b955ea2cfc8f3862091e2a240a84f17297"
+dependencies = [
+ "proc-macro2",
+ "quote",
+ "unicode-ident",
+]
+
+[[package]]
+name = "syn"
+version = "3.0.3"
+source = "registry+https://github.com/rust-lang/crates.io-index"
+checksum = "53e9bae58849f64dfa4f5d5ae372c8341f7305f82a3868709269343628b659a3"
+dependencies = [
+ "proc-macro2",
+ "quote",
+ "unicode-ident",
+]
+
+[[package]]
+name = "thiserror"
+version = "2.0.20"
+source = "registry+https://github.com/rust-lang/crates.io-index"
+checksum = "ec86235f5fcc2a73650310756d2ac5b138a5780bbbdfae3eeccec992c435ba4f"
+dependencies = [
+ "thiserror-impl",
+]
+
+[[package]]
+name = "thiserror-impl"
+version = "2.0.20"
+source = "registry+https://github.com/rust-lang/crates.io-index"
+checksum = "bc04cd3e1236dd4a98afca4569f2deb3f120e5422a4023be2cb683f8486292af"
+dependencies = [
+ "proc-macro2",
+ "quote",
+ "syn 3.0.3",
+]
+
+[[package]]
+name = "unicode-ident"
+version = "1.0.24"
+source = "registry+https://github.com/rust-lang/crates.io-index"
+checksum = "e6e4313cd5fcd3dad5cafa179702e2b244f760991f45397d14d4ebf38247da75"
diff --git a/crates/fuzzing/Cargo.toml b/crates/fuzzing/Cargo.toml
new file mode 100644
index 00000000..7d5398d6
--- /dev/null
+++ b/crates/fuzzing/Cargo.toml
@@ -0,0 +1,17 @@
+[package]
+name = "genvm-fuzzing"
+version = "0.1.0"
+edition = "2021"
+publish = false
+
+[dependencies]
+# `default-features = false` keeps heapless, spin and lock_api out of the tree
+postcard = { version = "1.1.3", default-features = false, features = [
+ "use-std",
+] }
+serde = { version = "1.0.219", features = ["derive"] }
+
+[dependencies.mutatis]
+default-features = false
+features = ["derive", "std"]
+version = "0.5.3"
diff --git a/crates/fuzzing/preload/Cargo.lock b/crates/fuzzing/preload/Cargo.lock
new file mode 100644
index 00000000..bb34fecc
--- /dev/null
+++ b/crates/fuzzing/preload/Cargo.lock
@@ -0,0 +1,7 @@
+# This file is automatically @generated by Cargo.
+# It is not intended for manual editing.
+version = 4
+
+[[package]]
+name = "genvm-fuzz-preload"
+version = "0.1.0"
diff --git a/crates/fuzzing/preload/Cargo.toml b/crates/fuzzing/preload/Cargo.toml
new file mode 100644
index 00000000..38d05633
--- /dev/null
+++ b/crates/fuzzing/preload/Cargo.toml
@@ -0,0 +1,10 @@
+[package]
+name = "genvm-fuzz-preload"
+version = "0.1.0"
+edition = "2021"
+publish = false
+
+[lib]
+crate-type = ["cdylib"]
+
+[dependencies]
diff --git a/crates/fuzzing/preload/src/lib.rs b/crates/fuzzing/preload/src/lib.rs
new file mode 100644
index 00000000..56d24168
--- /dev/null
+++ b/crates/fuzzing/preload/src/lib.rs
@@ -0,0 +1,78 @@
+//! `LD_PRELOAD` shim that makes the OS randomness syscalls deterministic.
+//!
+//! Loaded into AFL targets only, via `AFL_PRELOAD`. Without it every process
+//! seeds its hashers from real entropy, so a saved crash need not replay and
+//! coverage differs between runs.
+
+use core::ffi::{c_int, c_uint, c_void};
+use std::cell::Cell;
+
+thread_local! {
+ /// Number of `u64` words handed out by this thread.
+ static WORDS: Cell = const { Cell::new(0) };
+}
+
+fn next_word() -> u64 {
+ // splitmix64 over the call counter
+ let mut z = WORDS.with(|words| {
+ let current = words.get();
+ words.set(current.wrapping_add(1));
+ current.wrapping_add(0x9e37_79b9_7f4a_7c15)
+ });
+ z = (z ^ (z >> 30)).wrapping_mul(0xbf58_476d_1ce4_e5b9);
+ z = (z ^ (z >> 27)).wrapping_mul(0x94d0_49bb_1331_11eb);
+ z ^ (z >> 31)
+}
+
+/// # Safety
+/// `buffer` must be writable for `length` bytes, or `length` must be zero.
+unsafe fn fill(buffer: *mut c_void, length: usize) {
+ if length == 0 {
+ return;
+ }
+ assert!(!buffer.is_null(), "genvm-fuzz-preload: null buffer");
+
+ let mut written = 0;
+ while written < length {
+ let word = next_word().to_le_bytes();
+ let chunk = word.len().min(length - written);
+ std::ptr::copy_nonoverlapping(word.as_ptr(), (buffer as *mut u8).add(written), chunk);
+ written += chunk;
+ }
+}
+
+/// # Safety
+/// See [`fill`]. Matches the libc signature, flags are ignored: every flag
+/// combination is satisfiable when the entropy is fake.
+#[no_mangle]
+pub unsafe extern "C" fn getrandom(buffer: *mut c_void, length: usize, _flags: c_uint) -> isize {
+ fill(buffer, length);
+ length as isize
+}
+
+/// # Safety
+/// See [`fill`]. Unlike libc this accepts any `length`; the 256 byte cap exists
+/// to bound how much real entropy one call may drain.
+#[no_mangle]
+pub unsafe extern "C" fn getentropy(buffer: *mut c_void, length: usize) -> c_int {
+ fill(buffer, length);
+ 0
+}
+
+#[cfg(test)]
+mod tests {
+ use super::next_word;
+
+ #[test]
+ fn threads_have_the_same_sequence() {
+ let threads: Vec<_> = (0..4)
+ .map(|_| std::thread::spawn(|| [next_word(), next_word()]))
+ .collect();
+ let sequences: Vec<_> = threads
+ .into_iter()
+ .map(|thread| thread.join().unwrap())
+ .collect();
+
+ assert!(sequences.windows(2).all(|pair| pair[0] == pair[1]));
+ }
+}
diff --git a/crates/fuzzing/src/depth_limit.rs b/crates/fuzzing/src/depth_limit.rs
new file mode 100644
index 00000000..afc0d6ab
--- /dev/null
+++ b/crates/fuzzing/src/depth_limit.rs
@@ -0,0 +1,408 @@
+use core::fmt;
+use serde::de::{self, DeserializeSeed, Visitor};
+
+pub(crate) struct Deserializer {
+ inner: D,
+ remaining: usize,
+}
+
+impl Deserializer {
+ pub(crate) fn new(inner: D, remaining: usize) -> Self {
+ Self { inner, remaining }
+ }
+
+ fn descend<'de, V>(self, visitor: V) -> Result<(D, LimitVisitor), D::Error>
+ where
+ D: de::Deserializer<'de>,
+ {
+ let remaining = self
+ .remaining
+ .checked_sub(1)
+ .ok_or_else(|| de::Error::custom("deserialization depth limit exceeded"))?;
+ Ok((
+ self.inner,
+ LimitVisitor {
+ inner: visitor,
+ remaining,
+ },
+ ))
+ }
+}
+
+macro_rules! delegate {
+ ($($method:ident $(($($arg:ident: $ty:ty),*))?);+ $(;)?) => {
+ $(
+ fn $method(self, $($($arg: $ty,)*)? visitor: V) -> Result
+ where
+ V: Visitor<'de>,
+ {
+ self.inner.$method($($($arg,)*)? visitor)
+ }
+ )+
+ };
+}
+
+impl<'de, D> de::Deserializer<'de> for Deserializer
+where
+ D: de::Deserializer<'de>,
+{
+ type Error = D::Error;
+
+ delegate! {
+ deserialize_any();
+ deserialize_bool();
+ deserialize_i8();
+ deserialize_i16();
+ deserialize_i32();
+ deserialize_i64();
+ deserialize_i128();
+ deserialize_u8();
+ deserialize_u16();
+ deserialize_u32();
+ deserialize_u64();
+ deserialize_u128();
+ deserialize_f32();
+ deserialize_f64();
+ deserialize_char();
+ deserialize_str();
+ deserialize_string();
+ deserialize_bytes();
+ deserialize_byte_buf();
+ deserialize_unit();
+ deserialize_unit_struct(name: &'static str);
+ deserialize_identifier();
+ deserialize_ignored_any();
+ }
+
+ fn deserialize_option(self, visitor: V) -> Result
+ where
+ V: Visitor<'de>,
+ {
+ let (inner, visitor) = self.descend(visitor)?;
+ inner.deserialize_option(visitor)
+ }
+
+ fn deserialize_newtype_struct(
+ self,
+ name: &'static str,
+ visitor: V,
+ ) -> Result
+ where
+ V: Visitor<'de>,
+ {
+ let (inner, visitor) = self.descend(visitor)?;
+ inner.deserialize_newtype_struct(name, visitor)
+ }
+
+ fn deserialize_seq(self, visitor: V) -> Result
+ where
+ V: Visitor<'de>,
+ {
+ let (inner, visitor) = self.descend(visitor)?;
+ inner.deserialize_seq(visitor)
+ }
+
+ fn deserialize_tuple(self, len: usize, visitor: V) -> Result
+ where
+ V: Visitor<'de>,
+ {
+ let (inner, visitor) = self.descend(visitor)?;
+ inner.deserialize_tuple(len, visitor)
+ }
+
+ fn deserialize_tuple_struct(
+ self,
+ name: &'static str,
+ len: usize,
+ visitor: V,
+ ) -> Result
+ where
+ V: Visitor<'de>,
+ {
+ let (inner, visitor) = self.descend(visitor)?;
+ inner.deserialize_tuple_struct(name, len, visitor)
+ }
+
+ fn deserialize_map(self, visitor: V) -> Result
+ where
+ V: Visitor<'de>,
+ {
+ let (inner, visitor) = self.descend(visitor)?;
+ inner.deserialize_map(visitor)
+ }
+
+ fn deserialize_struct(
+ self,
+ name: &'static str,
+ fields: &'static [&'static str],
+ visitor: V,
+ ) -> Result
+ where
+ V: Visitor<'de>,
+ {
+ let (inner, visitor) = self.descend(visitor)?;
+ inner.deserialize_struct(name, fields, visitor)
+ }
+
+ fn deserialize_enum(
+ self,
+ name: &'static str,
+ variants: &'static [&'static str],
+ visitor: V,
+ ) -> Result
+ where
+ V: Visitor<'de>,
+ {
+ let (inner, visitor) = self.descend(visitor)?;
+ inner.deserialize_enum(name, variants, visitor)
+ }
+
+ fn is_human_readable(&self) -> bool {
+ self.inner.is_human_readable()
+ }
+}
+
+struct LimitVisitor {
+ inner: V,
+ remaining: usize,
+}
+
+impl<'de, V> Visitor<'de> for LimitVisitor
+where
+ V: Visitor<'de>,
+{
+ type Value = V::Value;
+
+ fn expecting(&self, formatter: &mut fmt::Formatter<'_>) -> fmt::Result {
+ self.inner.expecting(formatter)
+ }
+
+ fn visit_none(self) -> Result
+ where
+ E: de::Error,
+ {
+ self.inner.visit_none()
+ }
+
+ fn visit_some(self, deserializer: D) -> Result
+ where
+ D: de::Deserializer<'de>,
+ {
+ self.inner
+ .visit_some(Deserializer::new(deserializer, self.remaining))
+ }
+
+ fn visit_newtype_struct(self, deserializer: D) -> Result
+ where
+ D: de::Deserializer<'de>,
+ {
+ self.inner
+ .visit_newtype_struct(Deserializer::new(deserializer, self.remaining))
+ }
+
+ fn visit_seq(self, seq: A) -> Result
+ where
+ A: de::SeqAccess<'de>,
+ {
+ self.inner.visit_seq(LimitSeqAccess {
+ inner: seq,
+ remaining: self.remaining,
+ })
+ }
+
+ fn visit_map(self, map: A) -> Result
+ where
+ A: de::MapAccess<'de>,
+ {
+ self.inner.visit_map(LimitMapAccess {
+ inner: map,
+ remaining: self.remaining,
+ })
+ }
+
+ fn visit_enum(self, data: A) -> Result
+ where
+ A: de::EnumAccess<'de>,
+ {
+ self.inner.visit_enum(LimitEnumAccess {
+ inner: data,
+ remaining: self.remaining,
+ })
+ }
+}
+
+struct LimitSeed {
+ inner: S,
+ remaining: usize,
+}
+
+impl<'de, S> DeserializeSeed<'de> for LimitSeed
+where
+ S: DeserializeSeed<'de>,
+{
+ type Value = S::Value;
+
+ fn deserialize(self, deserializer: D) -> Result
+ where
+ D: de::Deserializer<'de>,
+ {
+ self.inner
+ .deserialize(Deserializer::new(deserializer, self.remaining))
+ }
+}
+
+struct LimitSeqAccess {
+ inner: A,
+ remaining: usize,
+}
+
+impl<'de, A> de::SeqAccess<'de> for LimitSeqAccess
+where
+ A: de::SeqAccess<'de>,
+{
+ type Error = A::Error;
+
+ fn next_element_seed(&mut self, seed: T) -> Result