Repository navigation
[BREAKING] [FIX]: Evaluate probe verdicts over final traces - #149
Draft
Spencer Schoenberg (spencrr) wants to merge 8 commits into
Draft
Spencer Schoenberg (spencrr) wants to merge 8 commits into
Spencer Schoenberg (spencrr) wants to merge 8 commits into
Conversation
|
Azure Pipelines: There may be pipelines that require an authorized user to comment /azp run to run. |
3 tasks done
Spencer Schoenberg (spencrr)
force-pushed
the
dev/spencrr/trace-probe-cadence
branch
from
August 8, 2026 02:31
cb71f74 to
bc06529
Compare
Spencer Schoenberg (spencrr)
force-pushed
the
dev/spencrr/trace-probe-cadence
branch
2 times, most recently
from
August 27, 2026 17:30
2e9304b to
2cdbab0
Compare
Spencer Schoenberg (spencrr)
force-pushed
the
dev/spencrr/trace-probe-cadence
branch
2 times, most recently
from
September 8, 2026 22:46
9f80bdf to
a491cdb
Compare
Spencer Schoenberg (spencrr)
force-pushed
the
dev/spencrr/trace-probe-cadence
branch
4 times, most recently
from
October 1, 2026 00:59
c69240a to
9b4eea1
Compare
3 tasks done
Require the same observability and manifest before reusing an online judgment. Copy optional evidence through the shared tolerant renderer so malformed supporting text cannot discard an established verdict. Keep terminal evidence and operand lists independent from online records.
Rename evaluate_terminal_async to evaluate_final_trace_async so the public runner helper matches Result.final_trace_evaluation. Document the trace execution helpers where they are introduced.
Rename TraceRun.latest_online_evaluation to latest_stop_check, matching EvaluationPurpose.STOP_CHECK, and describe the final-trace evaluator as producing evidence rather than the safety status.
Retire the list reducer after probe execution adopts terminal evaluation. Keep explicit response scopes, online evidence separation, zero-turn errors and xdist v3 semantics consistent across tests, exports and extension guidance.
Call evaluate_final_trace_async and describe probe verdict evidence as final-trace evaluation. Record a compatible trace-contract decision against the merged v2 base; Result fields and schemas are unchanged.
Explain the per-turn to final-trace verdict change, its single-prompt blast radius, trial sampling impact, and resolver replacement. Remove private xdist envelope details already covered by schema-drift rejection and the mixed-version limitation.
Spencer Schoenberg (spencrr)
force-pushed
the
dev/spencrr/trace-probe-cadence
branch
from
October 8, 2026 02:14
9b4eea1 to
82a86a9
Compare
Spencer Schoenberg (spencrr)
added a commit
that referenced
this pull request
Oct 9, 2026
<!-- This repository loosely follows the "Conventional Commits" specification for commit messages. See https://www.conventionalcommits.org/ for more information. --> <!-- Squash-merge commit messages must match: ^\[(FEAT|FIX|REFACTOR|STYLE|TEST|DOCS|CI|MAINT|META|REVERT)\]( \[BREAKING\])?:\s.+\(#\d+\) --> <!-- GitHub appends (#N) automatically during squash-merge; just use the [TAG]: description format for your PR title. --> <!-- If your PR is not yet ready for review, please mark it as [DRAFT]. --> ## Description <!-- What does this PR do? Provide a brief summary of the changes. --> <!-- Mention any relevant issues or pull requests with #<issue_number> --> <!-- Tag any relevant reviewers or teams using @<username> --> Adds the linear trace runner shared by attack and probe strategies. `run_trace_async` drives a conversation with an optional online `stop_when` check, keeps separate raw and annotated turn histories, and records `TraceEndReason`. `evaluate_final_trace_async` evaluates the final trace once. It reuses the latest online evaluation only when the evaluator, raw turns, manifest, and observability level are all identical. The runner does not own session lifetime, polarity, cleanup, or exception conversion; those remain strategy and `BaseExecution` responsibilities. Driver history receives a shallow copy, evaluator contexts contain annotation-free turns, and reused evaluator evidence is copied defensively. `EvaluationRecord`, `TraceRun`, `run_trace_async`, and `evaluate_final_trace_async` are exported from `rampart.core` for custom strategies and documented in the API reference. Built-in probes adopt them in #149 and XPIA in #150. ## Breaking changes <!-- If none, write "None". If breaking, describe the impact and migration path. --> None. This PR adds APIs and does not change existing strategies. ## Checklist - [x] `pre-commit run --all-files` passes - [x] Tests added or updated for changes <!-- Please describe what tests were added or updated --> — termination reasons, raw and annotated histories, stop behavior, exact-context reuse, changed observability and manifest, optional-evidence copying, post-run mutation, exceptions, and turn budgets - [x] Documentation updated
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Behavioral probes now evaluate the verdict once over the final trace instead of after every response. This fixes sequence and prior-condition probes, which cannot be judged correctly from partial prefixes. Probes no longer stop on a detected verdict by default; pass
stop_whenwhen an online condition should end the trace early.Final verdict evidence is stored on
Result.final_trace_evaluation, online stop evidence stays on turns withEvaluationPurpose.STOP_CHECK, andResult.trace_end_reasonrecords why the trace ended. Final evaluation runs while the session is still active.Because probe status semantics change, the private xdist envelope moves to
rampart.xdist.v3. Controllers reject v2 worker payloads instead of mixing per-turn and final-trace statuses. The trace-contract declaration records a compatible change againstmain: fingerprinted files change only by removing the list-basedresolve_as_probehelper and updating docstrings.Depends on #148.
Breaking changes
max_turnsunlessstop_whenis set.stop_when, adaptive drivers receive no online evaluator feedback and may run tomax_turns.resolve_as_probe(eval_results=...)is removed; useresolve_probe_verdict(evaluation=...).The upgrade note in
docs/probes/behavioral.mdcovers these changes.Checklist
pre-commit run --all-filespasses