Skip to content

feat(dashboard): readiness stamp on the agents list and the fleet grid (abilityai/trinity-enterprise#527) - #3038

Merged
vybe merged 10 commits into
devfrom
feature/ent527-readiness-on-lists
Sep 29, 2026
Merged

vybe merged 10 commits into
devfrom
feature/ent527-readiness-on-lists

Conversation

@dolho

@dolho dolho commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Rider on abilityai/trinity-enterprise#527, from the operator ruling of 2026-09-24 (ent#560's closure): readiness is a role-companion property, and the agents list and fleet grid should show the readiness stamp, not only the role card.

  • Backend: GET /api/agents attaches readiness: {status, changed_at, source} per agent.
    • It comes from one batched read, db.get_role_readiness_for_agents, following the display-label pattern, so there is no N+1 on the fleet's hottest endpoint.
    • An agent with no stamp carries null and shows nothing — never a guessed calibrating. Whether an agent is a companion at all is in its template.yaml, which a list must not read.
    • The list says what and when, never who. changed_by is an email or the ent#689 rollout sentinel. The role card is owner-scoped and keeps the person.
  • Frontend: one predicate, utils/readinessBadge.js, is used by both surfaces. It carries the role card's words and variants: ready = success, calibrating = warning, each with a dot. The tooltip names a rollout stamp ("carried over when the readiness gate shipped") and what calibrating holds back ("its scheduled brief is paused…").
    • Agents list: both layouts. The badge sits after the pressure badge, and the documented fixed order is updated.
    • Fleet grid tile: after the runtime badge.

Changes

  • src/backend/db/role_readiness.py: get_role_readiness_for_agents
  • src/backend/database.py: the facade method
  • src/backend/routers/agents.py: attaches readiness in list_agents_endpoint
  • src/frontend/src/utils/readinessBadge.js: the new shared predicate
  • src/frontend/src/components/AgentListPanel.vue, AgentTile.vue: render it
  • docs/memory/requirements/core-agent.md §5.36: the rider

Test plan

  • tests/unit/test_ent527_readiness_on_lists.py (7), on a real DB:
    • only stamped agents are returned
    • a rollout stamp is named as such
    • who flipped it is never included
    • an unknown state is skipped
    • an empty list makes no query
    • the endpoint attaches the result of the one read, and None when unstamped
  • readinessBadge.spec.js (6), readinessOnGrid.spec.js (3, mounts the real tile; 2 are red without the tile change)
  • Full vitest, including the ratchets (3,689); npm run build; 282 related backend tests
  • Live check on a local instance: API and UI, list and grid, light and dark (see comment)

Related to abilityai/trinity-enterprise#527. That issue closes at release.

🤖 Generated with Claude Code

…d (trinity-enterprise#527 rider)

Operator ruling 2026-09-24 (ent#560's closure): readiness is a
role-companion property, and the owner's stamp should be visible on
the agents list and the fleet grid, not only on the role card.

- GET /api/agents attaches readiness {status, changed_at, source} from
  one batched read (get_role_readiness_for_agents, the display-label
  pattern, no N+1). An unstamped agent carries null and shows nothing:
  whether it is a companion is in its template.yaml, which a list must
  not read, so there is no guessed calibrating. The list says what and
  when, never who; the owner-scoped role card keeps the person.
- One predicate for both surfaces (utils/readinessBadge.js), with the
  role card's words and variants (ready = success, calibrating =
  warning, dot). The tooltip names a rollout stamp and what calibrating
  holds back.
- The list places it after the pressure badge (documented order
  updated) on both layouts; the grid tile places it after the runtime
  badge.

Tests: 7 backend (real DB + the endpoint), 6 helper, 3 mounted tile
(2 red without the tile change). Full vitest and ratchets green.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@dolho dolho added the ui PR touches the frontend UI — triggers Playwright e2e tests label Sep 28, 2026
@dolho

dolho commented Sep 28, 2026

Copy link
Copy Markdown
Contributor Author

Live check (local instance, 37 agents, SQLite): this branch's diff was applied to a running stack (backend with --reload, frontend on the Vite dev server), then reverted.

API. GET /api/agents returns the readiness key on all 37 agents, with a value only on the three stamped ones:

  • growth-companion: ready, source owner (a real stamp that was already there)
  • pm-demo: calibrating, source owner
  • dev-bot: ready, source rollout (the rollout:ent#689 sentinel)

changed_by never appears.

UI. 2 themes × 2 views × 4 agents, found by data-agent in a single logged-in session. All 16 results match expectations:

Agent List Grid Tooltip
pm-demo calibrating (warning) calibrating Calibrating since 2026-09-28 (UTC) — its scheduled brief is paused until its owner marks it ready
growth-companion ready (success) ready Ready — marked ready by its owner since 2026-09-25 (UTC)
dev-bot ready ready Ready — carried over when the readiness gate shipped since …
sidekick (no stamp) nothing nothing —

On the grid tile the badge sits after the runtime icon. In the list it goes on the secondary line under the name. Both themes read correctly.

The two test stamps (pm-demo, dev-bot) have been removed from the instance.

@dolho

dolho commented Sep 28, 2026

Copy link
Copy Markdown
Contributor Author

/review Report

Branch: feature/ent527-readiness-on-lists → dev (merge-base diff)
Files changed: 11 (+344 / −2)
Scope: CLEAN. The intent is the ent#527 rider (ruling 2026-09-24): show the owner's readiness stamp on the agents list and the fleet grid. The diff delivers exactly that — one batched backend read, the field on GET /api/agents, one shared predicate, and the badge on the list and on the tile. Nothing unrelated.
Plan completion: rider "stamp visible on the agents list and the fleet grid" → DONE:

  • routers/agents.py list_agents_endpoint attaches readiness
  • AgentListPanel.vue: lg and md secondary lines
  • AgentTile.vue: nameline

Execution coverage (Step 2.5)

Changed symbol / test Executed by Live consumer Verdict
RoleReadinessOperations.get_role_readiness_for_agents test_ent527_readiness_on_lists.py::TestBatchedRead (real DB via db_harness) routers/agents.py::list_agents_endpoint ✅ executed
readiness on GET /api/agents TestListEndpoint (calls the real endpoint function) stores/network.js visibleAgents → list and grid ✅ executed
utils/readinessBadge.js::readinessBadge readinessBadge.spec.js (6) AgentTile.vue, AgentListPanel.vue ✅ executed
AgentTile.vue badge readinessOnGrid.spec.js (mounts the real tile; 2/3 red with the tile change reverted) dashboard grid ✅ executed
AgentListPanel.vue badge (both rows) — (live check only) dashboard list ⚠️ no automated test, see I2

No source-text assertions in the changed tests.

Critical findings

None.

Informational findings

[I1] Staleness: a readiness flip does not reach an open dashboard until a reload (confidence 8/10)
File: src/frontend/src/stores/network.js:1373–1395
Evidence: the 30 s poll replaces agents.value only when the set of names changed:

const hasChanges =
  newAgents.length !== agents.value.length ||
  !newAgents.every(a => currentAgentNames.has(a.name))
if (hasChanges) { agents.value = newAgents … }

An owner who marks a companion ready in the Workspace keeps seeing calibrating on an already-open dashboard until a full reload or a filter change. The flip also emits no WS event (client_portal/role_card.flip_readiness). The same pre-existing gap applies to tags; this PR adds a field that is expected to change.
Suggestion: on each poll, patch readiness in place for agents already present (cheap, and no node rebuild), or add readiness to the change signature. A one-line vitest over the poll body would pin it.

[I2] Test gap: the list rows are not mounted (confidence 7/10)
File: src/frontend/src/components/AgentListPanel.vue (row-secondary-lg, row-secondary-md)
The helper is covered, and the grid tile is mounted, but the two list placements are proven only by the live check. A regression there (a mistyped agent.readiness, the badge dropped from one breakpoint) would ship green.
Suggestion: extend readinessOnGrid.spec.js, or add a sibling that mounts AgentListPanel with one stamped and one unstamped agent and asserts one badge per stamped row on both secondary lines.

[I3] Doc staleness: the list badge contract (confidence 9/10)
File: docs/memory/feature-flows/dashboard-list-view.md:100 and D14 (:212–228)
Evidence: "fixed order slug · pressure · runtime · tags · +N … A future badge goes here in that order or it does not go in the row." The component comment now says slug · pressure · readiness · runtime · tags · +N, but the flow doc (the contract D14 points to) was not updated.
Suggestion: add readiness to the order in both places. It complies with D14 on every other point: lg/md only, never base, not in the name cell, flex-shrink-0. The grid-view flow (dashboard-grid-view.md) can gain one line for the nameline badge.

Clean categories

  • SQL / data safety: one SELECT … WHERE agent_name IN (:names) via SQLAlchemy Core, parameterised. The input is de-duplicated, and an empty list short-circuits with no query (tested).
  • Auth: no new endpoint. The field rides the existing get_current_user + get_accessible_agents scope, so a viewer sees stamps only for agents they can already list. MCP list_agents passes the same objects to the same accessible set.
  • Disclosure: changed_by is dropped at the DB layer (tested: test_never_carries_who), so neither the owner's email nor the rollout sentinel crosses the list. The rollout case is reported as source: rollout.
  • Concurrency: read-only; no writes, no shared state.
  • Enum completeness: READINESS_STATES = ("calibrating", "ready") is used by the DB filter; an unknown stored state is dropped rather than rendered (tested). The helper accepts exactly those two values.
  • Performance: one query per list call (the display-label pattern), no N+1, no container reads. Readiness for unstamped companions is deliberately not derived, because that would need a template read per agent.
  • Frontend: BaseBadge primitive with semantic variants; the raw-color ratchet is unchanged; no v-html; the badge renders only when a stamp exists, so no new empty or loading state.
  • Enterprise-docs guard: requirement §5.36 addition checked; it names no paid module.

Summary

  • Critical: 0
  • Informational: 3. I1 (poll staleness) and I3 (flow-doc order) are worth fixing before merge; I2 is a test-coverage follow-up.
  • Scope: clean

🤖 Generated with Claude Code

@AndriiPasternak31 AndriiPasternak31 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The backend side is right: one batched, parameterised read, no new endpoint, changed_by dropped before it leaves the DB layer, and the field rides the existing get_accessible_agents scope. My main concern is the calibrating tooltip. It claims a pause that the role card deliberately refuses to claim. Two of your own /review findings are also still open on this head, since the PR is a single commit. None of this blocks, but I'd fix 1–3 before merge.

  1. [should-fix] src/frontend/src/utils/readinessBadge.js:31: the calibrating tooltip always says "its scheduled brief is paused until its owner marks it ready". The role card only says this when _brief_held() is true (src/backend/client_portal/role_card.py:334-352), meaning the agent has an enabled seat-delivery schedule and autonomy is on. The comment there spells out why: with autonomy off, "paused until you mark it ready" would promise a flip that starts nothing. Scenario: an owner stamps a companion calibrating that has no Workspace-delivery schedule, or whose autonomy is off. The role card shows no pause line, but the list and grid tooltip says its brief is paused. The owner marks it ready and nothing happens. This is the list-vs-card disagreement that the file header says the shared predicate exists to prevent. Fix: the list rows already carry autonomy_enabled, so suppress the clause when that is false. For the schedule half, either add a batched brief_held bool next to readiness (one query over agent_schedules for the same agent_names), or drop the clause and let the role card own that sentence.

  2. [should-fix] docs/memory/feature-flows/dashboard-list-view.md:100 and :225-228 (your /review I3, still open). D14 is the contract for the row's meta strip: "a future badge goes here in that order or it does not go in the row". The component comment now reads slug · pressure · readiness · runtime · tags · +N, but the flow doc still says slug · pressure · runtime · tags · +N. The next person adding a badge will follow the doc and lose the order. Update both places, and add one line to dashboard-grid-view.md for the nameline badge (after runtime, before SYSTEM/SKILL RUNNER).

  3. [should-fix] src/frontend/src/stores/network.js:1383-1397 (your /review I1, still open). The 30s poll only replaces agents.value when the set of names changes, and a readiness flip emits no WS event. A dashboard left open in another tab keeps showing calibrating after the owner marks the agent ready in the Workspace, until a reload. Remounting the Dashboard refetches, so this only hits an already-open tab. That is still the case where an operator watches the fleet. tags has the same gap already, but this PR adds a field that is expected to change. Fix: on each poll, patch readiness in place for rows already present, with no node rebuild. Pin it with a small vitest over the poll body.

  4. [nit] src/frontend/tests/unit/readinessOnGrid.spec.js:4-6 (your I2). The header says the list rows "are covered by the shared helper's spec", but the helper spec covers the words, not the placement. Dropping the badge from row-secondary-md or mistyping agent.readiness in one row would still ship green. A mount of AgentListPanel with one stamped and one unstamped agent, asserting one readiness-badge per stamped row, would close this. Otherwise, reword the comment so it doesn't overclaim.

  5. [nit] src/frontend/src/utils/readinessBadge.js:27-29: the rollout tooltip reads "Ready — carried over when the readiness gate shipped since 2026-09-24 (UTC)". Here "since" attaches to "shipped". The role card puts the date first ("since … · carried over when the readiness gate shipped"). Use "Ready since 2026-09-24 (UTC) — carried over when the readiness gate shipped", and update the assertion in readinessBadge.spec.js.

What I checked

  • Read the diff, the ent#527 role card (role_card.py, PortalAgentRole.vue), the rollout seed migration, agent_cleanup.py (the stamp cascades on delete and follows a rename, so no stale badge on a recreated name), and stores/network.js fetch/poll. The field travels on the agent object, so all modes see it through visibleAgents.
  • Design contract: BaseBadge with semantic variants only. No raw palette classes or loading gates added. The badge sits inside the D14 lg/md strip only, flex-shrink-0. On the tile it is shorter than the name line, so tile rhythm doesn't move.
  • Auth/disclosure: no new route, and the stamp is visible only on agents the caller can already list. Non-owner viewers already see the status on the role card; the list omits who flipped it.
  • Frontend on a clean npm ci: npm run test:unit gives 177 files / 3689 tests passed, including the colour and loading-gate ratchets. npm run build passes.
  • Backend: test_ent527_readiness_on_lists.py, test_ent527_role_card.py, test_ent689_readiness_gate.py and test_2889_readiness_classifier.py give 142 passed. The display-label and cleanup tests give 35 passed.
  • gh pr checks 3038: all required checks green (pytest, e2e, schema-parity, pg-migrations, CodeQL).

dolho and others added 3 commits September 29, 2026 10:05
…ne is held

PR #3038 review item 1: the list/grid calibrating tooltip always said "its
scheduled brief is paused until its owner marks it ready", while the role
card says so only when _brief_held() is true (an enabled seat-delivery
schedule AND autonomy on). With autonomy off or no Workspace-delivery
schedule the owner would mark it ready and nothing would start.

- services/role_readiness_gate: is_seat_delivery_schedule + brief_is_held,
  the one predicate; role_card._brief_held now uses it.
- GET /api/agents attaches brief_held next to readiness: one batched,
  projected read over agent_schedules (get_workspace_delivery_schedules_for_agents),
  only for calibrating stamps with autonomy on; fail-soft to no claim.
- readinessBadge(readiness, briefHeld): the pause clause only when held,
  otherwise "not yet marked ready by its owner". Tile and list pass
  agent.brief_held.
- Nit: ready tooltips put the date first, as the role card does
  ("Ready since … (UTC) — carried over when the readiness gate shipped").
- Nit: AgentListPanel is now mounted in the spec — one badge per stamped
  row on both the lg and md secondary lines, none on the unstamped row.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…board

PR #3038 review item 3: the 30s poll replaced agents only when the set of
names changed, and a readiness flip emits no WS event, so an open tab kept
showing calibrating until a reload. On an unchanged name set the poll now
patches readiness and brief_held on the rows already present (only when
the value changed) — no row replacement, no node rebuild. Pinned by a
vitest over the real poll body (fake timers, same row objects, same nodes).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rid nameline

PR #3038 review item 2: dashboard-list-view.md (diagram and the D14
secondary-line contract) now reads slug · pressure · readiness · runtime ·
tags · +N; dashboard-grid-view.md gains the nameline badge order (runtime ·
readiness · SYSTEM / SKILL RUNNER) and the brief_held tooltip rule.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@dolho

dolho commented Sep 29, 2026

Copy link
Copy Markdown
Contributor Author

Thanks @AndriiPasternak31. All five items are addressed in three commits on top of 04872ee.

# Review item Fix Commit
1 Calibrating tooltip claims a paused brief unconditionally I went with your preferred option. There is now one predicate in services/role_readiness_gate.py (is_seat_delivery_schedule, brief_is_held), and role_card._brief_held uses it too. GET /api/agents attaches brief_held next to readiness. It comes from one batched, projected read over agent_schedules (get_workspace_delivery_schedules_for_agents), and only for calibrating stamps with autonomy on. If the read fails, the row makes no claim. readinessBadge(readiness, briefHeld) says "paused" only when the brief is held. Otherwise the tooltip reads "Calibrating since … — not yet marked ready by its owner". 0d73c64
2 Flow docs meta-strip order dashboard-list-view.md now reads slug · pressure · readiness · runtime · tags · +N in both the diagram and the D14 contract. dashboard-grid-view.md gains the nameline order: runtime · readiness · SYSTEM / SKILL RUNNER. 592899d
3 30s poll never refreshes readiness When the set of agent names is unchanged, the poll now patches readiness and brief_held in place on the rows already there. It writes only values that changed, and it neither replaces rows nor rebuilds nodes. A new readinessPoll.spec.js pins this over the real poll body (fake timers; the rows and nodes are the same objects before and after). 687fc22
4 nit: list rows not mounted readinessOnGrid.spec.js now mounts AgentListPanel with a held calibrating row, a ready row and an unstamped row. It asserts one badge per stamped row on both row-secondary-lg and row-secondary-md, none on the unstamped row, plus the variants and the held tooltip. 0d73c64
5 nit: "since" attaches to "shipped" The ready tooltips now put the date first, as the role card does: "Ready since 2026-09-24 (UTC) — carried over when the readiness gate shipped". The owner case matches: "Ready since … — marked ready by its owner". 0d73c64

Verification

  • Tests were written first. The new backend tests failed before the change (18 red), and so did the new and updated vitest cases.
  • Backend: test_ent527_readiness_on_lists.py, test_ent527_role_card.py, test_ent689_readiness_gate.py and test_2889_readiness_classifier.py, plus the display-label, cleanup, models-centralized and auth-wiring suites: 302 passed. The existing _brief_held role-card tests pass unchanged on the shared predicate.
  • Frontend: npm run test:unit gives 178 files / 3695 tests passed, with the colour, loading-gate and source-text ratchets green. npm run build passes.

🤖 Generated with Claude Code

dolho and others added 3 commits September 29, 2026 15:37
Resolve tests/registry.json: dev's file plus this branch's ent#527 entry.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Resolve tests/registry.json: dev's file plus this branch's entries.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Resolve tests/registry.json: dev's file plus this branch's ent#527 entry.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@obasilakis obasilakis left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving. Reviewed and validated: all five review items addressed; new backend/frontend tests pass locally and each fix has a test that goes red with the fix reverted; full CI green on b9461fb including regression diff.

Non-blocking: get_role_readiness_for_agents in list_agents_endpoint is unguarded, so a failed read 500s the whole list — same as the display-label read beside it.

vybe and others added 2 commits September 29, 2026 18:53
…echanical, per the merge-train note on the PR

get_role_readiness_for_agents was the one unguarded decoration read on
GET /api/agents: a DB fault 500'd the whole list. Degrade to no stamp
(briefs_held_for_list already degrades to no pause), plus a test that
drives the endpoint with the read raising.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@vybe

vybe commented Sep 29, 2026

Copy link
Copy Markdown
Contributor

merge-train: two commits pushed to this branch (mechanical, before assembling the train):

  1. c143582aa — merged origin/dev (clean).
  2. 0c3f81abc — routers/agents.py list_agents_endpoint: db.get_role_readiness_for_agents is now wrapped; a DB fault logs a warning and lists every agent with readiness: null instead of 500-ing all of GET /api/agents. This matches briefs_held_for_list, which already degrades to "no pause". Added TestListEndpoint::test_a_failed_readiness_read_lists_without_stamps, which drives the endpoint with the read raising (it fails without the guard, passes with it; the suite is 26/26).

Other validation notes, not blocking: the PR body's Changes list predates the follow-up commits (brief_held, get_workspace_delivery_schedules_for_agents, patchReadinessInPlace); and brief_held is now visible to shared (non-owner) viewers of the list — a single boolean, noting it in case that wasn't intended.

…ss-on-lists

# Conflicts:
#	tests/registry.json

@vybe vybe left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

merge-train: batch validated on train/20260929-1807 (#3093)

@vybe
vybe merged commit 4aef586 into dev Sep 29, 2026
25 of 26 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ui PR touches the frontend UI — triggers Playwright e2e tests

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants