Skip to content

[BREAKING] [FIX]: Evaluate probe verdicts over final traces - #149

Draft
Spencer Schoenberg (spencrr) wants to merge 8 commits into
microsoft:mainfrom
spencrr:dev/spencrr/trace-probe-cadence
Draft

Spencer Schoenberg (spencrr) wants to merge 8 commits into
microsoft:mainfrom
spencrr:dev/spencrr/trace-probe-cadence

Conversation

@spencrr

@spencrr Spencer Schoenberg (spencrr) commented Aug 4, 2026 •

Copy link
Copy Markdown
Contributor

Description

Behavioral probes now evaluate the verdict once over the final trace instead of after every response. This fixes sequence and prior-condition probes, which cannot be judged correctly from partial prefixes. Probes no longer stop on a detected verdict by default; pass stop_when when an online condition should end the trace early.

Final verdict evidence is stored on Result.final_trace_evaluation, online stop evidence stays on turns with EvaluationPurpose.STOP_CHECK, and Result.trace_end_reason records why the trace ended. Final evaluation runs while the session is still active.

Because probe status semantics change, the private xdist envelope moves to rampart.xdist.v3. Controllers reject v2 worker payloads instead of mixing per-turn and final-trace statuses. The trace-contract declaration records a compatible change against main: fingerprinted files change only by removing the list-based resolve_as_probe helper and updating docstrings.

Depends on #148.

Breaking changes

  • Probe verdicts are computed once over the final trace. Single-prompt probes with deterministic evaluators keep their verdicts. Multi-turn probes can resolve differently because the evaluator's scope now applies to the whole trace, which runs up to max_turns unless stop_when is set.
  • For multi-turn probes, stochastic evaluators such as LLM judges are sampled once per run instead of once per turn, so trial pass rates can shift.
  • Without stop_when, adaptive drivers receive no online evaluator feedback and may run to max_turns.
  • resolve_as_probe(eval_results=...) is removed; use resolve_probe_verdict(evaluation=...).
  • Controllers and workers must run the same RAMPART version; v2 xdist payloads are rejected.

The upgrade note in docs/probes/behavioral.md covers these changes.

Checklist

  • pre-commit run --all-files passes
  • Tests added or updated for changes — one-call verdict cadence, sequence and prior-condition probes, temporal scopes, explicit stopping and reuse, driver feedback, budgets, session ordering, summaries, error results, and the v3 xdist envelope
  • Documentation updated

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
There may be pipelines that require an authorized user to comment /azp run to run.

@spencrr
Spencer Schoenberg (spencrr) force-pushed the dev/spencrr/trace-probe-cadence branch from cb71f74 to bc06529 Compare August 8, 2026 02:31
@spencrr
Spencer Schoenberg (spencrr) force-pushed the dev/spencrr/trace-probe-cadence branch 2 times, most recently from 2e9304b to 2cdbab0 Compare August 27, 2026 17:30
@spencrr
Spencer Schoenberg (spencrr) force-pushed the dev/spencrr/trace-probe-cadence branch 2 times, most recently from 9f80bdf to a491cdb Compare September 8, 2026 22:46
@spencrr
Spencer Schoenberg (spencrr) force-pushed the dev/spencrr/trace-probe-cadence branch 4 times, most recently from c69240a to 9b4eea1 Compare October 1, 2026 00:59
@spencrr Spencer Schoenberg (spencrr) changed the title [FIX]: Evaluate probe verdicts over final traces [FIX] [BREAKING]: Evaluate probe verdicts over final traces Oct 1, 2026
Require the same observability and manifest before reusing an online judgment. Copy optional evidence through the shared tolerant renderer so malformed supporting text cannot discard an established verdict. Keep terminal evidence and operand lists independent from online records.
Rename evaluate_terminal_async to evaluate_final_trace_async so the public runner helper matches Result.final_trace_evaluation. Document the trace execution helpers where they are introduced.
Rename TraceRun.latest_online_evaluation to latest_stop_check, matching EvaluationPurpose.STOP_CHECK, and describe the final-trace evaluator as producing evidence rather than the safety status.
Retire the list reducer after probe execution adopts terminal evaluation. Keep explicit response scopes, online evidence separation, zero-turn errors and xdist v3 semantics consistent across tests, exports and extension guidance.
Call evaluate_final_trace_async and describe probe verdict evidence as final-trace evaluation. Record a compatible trace-contract decision against the merged v2 base; Result fields and schemas are unchanged.
Explain the per-turn to final-trace verdict change, its single-prompt blast radius, trial sampling impact, and resolver replacement. Remove private xdist envelope details already covered by schema-drift rejection and the mixed-version limitation.
@spencrr
Spencer Schoenberg (spencrr) force-pushed the dev/spencrr/trace-probe-cadence branch from 9b4eea1 to 82a86a9 Compare October 8, 2026 02:14
@spencrr Spencer Schoenberg (spencrr) changed the title [FIX] [BREAKING]: Evaluate probe verdicts over final traces [BREAKING] [FIX]: Evaluate probe verdicts over final traces Oct 8, 2026
Spencer Schoenberg (spencrr) added a commit that referenced this pull request Oct 9, 2026
<!-- This repository loosely follows the "Conventional Commits"
specification for commit messages. See
https://www.conventionalcommits.org/ for more information. -->
<!-- Squash-merge commit messages must match:
^\[(FEAT|FIX|REFACTOR|STYLE|TEST|DOCS|CI|MAINT|META|REVERT)\](
\[BREAKING\])?:\s.+\(#\d+\) -->
<!-- GitHub appends (#N) automatically during squash-merge; just use the
[TAG]: description format for your PR title. -->
<!-- If your PR is not yet ready for review, please mark it as [DRAFT].
-->

## Description

<!-- What does this PR do? Provide a brief summary of the changes. -->
<!-- Mention any relevant issues or pull requests with #<issue_number>
-->
<!-- Tag any relevant reviewers or teams using @<username> -->

Adds the linear trace runner shared by attack and probe strategies.
`run_trace_async` drives a conversation with an optional online
`stop_when` check, keeps separate raw and annotated turn histories, and
records `TraceEndReason`. `evaluate_final_trace_async` evaluates the
final trace once. It reuses the latest online evaluation only when the
evaluator, raw turns, manifest, and observability level are all
identical.

The runner does not own session lifetime, polarity, cleanup, or
exception conversion; those remain strategy and `BaseExecution`
responsibilities. Driver history receives a shallow copy, evaluator
contexts contain annotation-free turns, and reused evaluator evidence is
copied defensively.

`EvaluationRecord`, `TraceRun`, `run_trace_async`, and
`evaluate_final_trace_async` are exported from `rampart.core` for custom
strategies and documented in the API reference. Built-in probes adopt
them in #149 and XPIA in #150.

## Breaking changes
<!-- If none, write "None". If breaking, describe the impact and
migration path. -->

None. This PR adds APIs and does not change existing strategies.

## Checklist

- [x] `pre-commit run --all-files` passes
- [x] Tests added or updated for changes <!-- Please describe what tests
were added or updated --> — termination reasons, raw and annotated
histories, stop behavior, exact-context reuse, changed observability and
manifest, optional-evidence copying, post-run mutation, exceptions, and
turn budgets
- [x] Documentation updated

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant