Skip to content

fix(dispatch): carry inter-agent chain depth through loops, schedule triggers, session turns and events (#2973) - #3120

Merged
vybe merged 3 commits into
devfrom
feature/2973-chain-depth-laundering
Sep 30, 2026
Merged

vybe merged 3 commits into
devfrom
feature/2973-chain-depth-laundering

Conversation

@obasilakis

Copy link
Copy Markdown
Contributor

Summary

  • Follow-up to bug: agent-to-agent chat chains have no depth guard — a runaway A→B→A bounce is bounded only by capacity parks and hop timeouts #2806. An agent near inter_agent_max_chain_depth could reset its depth to 0 by starting a loop, manually triggering a schedule, driving a chat-session turn, or emitting an event its subscription answers. All four now inherit the caller's depth and are refused at the max with the same named 403 inter_agent_depth_exceeded.
  • Event depth travels as a signed chain_depth claim in the EVT-001 loopback JWT, so an emit on another agent's behalf and system agent.task.* completions (finished row's depth + 1) carry it too. _chain_caller falls back to vouched_source_agent.
  • New per source→subscriber hourly event-dispatch budget (CSO 2026-09-29 Finding 3, from the issue thread): ops setting event_dispatch_max_fires_per_hour, default 120, one high-priority notification per window, fail-open on Redis errors.

Changes

  • routers/{loops,schedules,sessions,event_subscriptions}.py — call enforce_inter_agent_depth; refusal propagates to the new app-level handler (error_handlers.inter_agent_depth_exceeded, registered in main.py).
  • services/loop_service.py, db/loops.py + agent_loops.chain_depth (SQLite migration loop_chain_depth, Alembic 0084, schema.py, tables.py).
  • src/scheduler/* — ExecutionOrigin.chain_depth (validated 1–1000), stamped on insert, kept on retry.
  • services/event_dispatch_service.py — loopback claim, terminal-event depth, dispatch budget; dependencies.py / models.User.loopback_chain_depth; services/dispatch_admission_service.py — claim read before the root early-return.
  • services/task_execution_service.execute_task(chain_depth=).
  • services/monitoring_alerts.py — budget alert; config.py / settings_service.py — new ops setting.
  • MCP: client.ts::depthRefusalFromError; run_agent_loop, trigger_agent_schedule, emit_event return the refusal as a non-retryable result.
  • Docs: requirements/core-agent.md §9.1.1, requirements/mcp.md, architecture/{execution,database,api-endpoints}.md, feature flows, IA-04 gaps, learnings fragment, CSO diff report.

Test Plan

Deferred: webhook tokens, agent-created cron schedules and self-reminders still start a depth-0 root — #3116 (needs a design call on whether a stored trigger carries its creator's depth forever).

Fixes #2973

🤖 Generated with Claude Code

…triggers, session turns and events (#2973)

Follow-up to #2806. Loop starts, manual schedule triggers, chat-session turns
and agent event emits by an agent principal now inherit the caller's chain
depth and are refused past inter_agent_max_chain_depth with the named 403
inter_agent_depth_exceeded, instead of starting a fresh depth-0 root.

- Loops: depth stored on agent_loops.chain_depth (SQLite migration +
  Alembic 0084) and stamped on every iteration row.
- Schedule trigger: depth forwarded to the scheduler; retries keep it.
- Session turns: depth passed through run_resumable_turn to execute_task.
- Events: depth signed into the EVT-001 loopback JWT (chain_depth claim),
  minted whether or not the source is vouched; task-completion events carry
  the finished row's depth + 1 (the max when the row is unreadable);
  _chain_caller falls back to vouched_source_agent.
- Per source->subscriber hourly event-dispatch budget in
  trigger_subscription (ops setting event_dispatch_max_fires_per_hour,
  default 120, one alert per window, fail-open on Redis errors).
- App-level handler maps InterAgentDepthExceeded to the #2806 403 body.
- MCP run_agent_loop, trigger_agent_schedule and emit_event render the
  refusal as a non-retryable result.

Mutation-checked: reverting each fix turns its test red in
tests/unit/test_2973_depth_new_roots.py, test_2973_event_dispatch_budget.py
and src/mcp-server/src/tools/depth-refusal.test.ts.

Webhook tokens, agent-created cron schedules and self-reminders remain
roots; tracked in #3116.

Fixes #2973

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@obasilakis
obasilakis force-pushed the feature/2973-chain-depth-laundering branch from 3d4278c to 6b02181 Compare September 30, 2026 13:25
…ons, budget alert write (#2973)

Review follow-ups on #3120: the cold retry in run_resumable_turn keeps
chain_depth (mutation-checked), a missing session answers 404 before the
depth 403, and the budget alert writes a high-priority notification on the
subscriber. Doc counts updated for the four covered paths.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@obasilakis obasilakis added the ui PR touches the frontend UI — triggers Playwright e2e tests label Sep 30, 2026
 single-flight scan

The dispatch-budget alert flag is a once-guard (no release, no token, TTL
only), the same non-lock nx shape as the heartbeat and voip entries.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@vybe

vybe commented Sep 30, 2026

Copy link
Copy Markdown
Contributor

merge-train: this PR is on today's train with #3122, #3119 and #3085. The one Tier 1 blocker was the test_1920 allowlist, and your 5c0bc209d already fixes it.

Two gaps to consider. Neither blocks this train. Each was found by a mutation that stayed green:

  • main.py handler registration isn't pinned. Deleting app.add_exception_handler(InterAgentDepthExceeded, …) (around main.py:1217) leaves all 48 new tests green, because they build their own app. In production, every depth refusal would then be a 500 instead of the named 403. One assertion covers it: InterAgentDepthExceeded in main.app.exception_handlers.
  • The scheduler's manual-trigger hand-off isn't pinned. Changing chain_depth=origin.chain_depth in src/scheduler/main.py to None leaves 2973, 1968 and 1970 green. Extending the handler-driven test in test_1968 to send a chain_depth and assert it on the stored row would cover it.

Also worth a decision, from /cso --diff at low severity: anyone with shared access to agent X can use up X→S's hourly fire budget through POST /api/agents/X/emit-event. That silently skips X's real events for up to an hour. You could key that route's budget on the emitting principal, or record it in #3116.

This PR goes in ahead of #3085, and #3085 will re-parent onto your 0084_agent_loops_chain_depth.

@vybe vybe left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

merge-train: batch validated on train/20260930-1429 (#3126, all gates green)

@vybe
vybe merged commit f93a124 into dev Sep 30, 2026
28 checks passed
vybe pushed a commit that referenced this pull request Sep 30, 2026
…5_ent720_email_identity (#3085)

#3120 landed 0084_agent_loops_chain_depth off the same parent as this PR's
0084_ent720_email_identity. Disjoint tables, so the revision is re-parented:
0085_ent720_email_identity <- 0084_agent_loops_chain_depth. SQLite list keeps
both entries, dev's first — the same resolution train #3126 was gated on.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ui PR touches the frontend UI — triggers Playwright e2e tests

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants