Skip to content

DO NOT MERGE — merge train B: 3468,3467,3437,3478,3477,3339,3486,3412 - #3504

Closed
vybe wants to merge 57 commits into
devfrom
train/20261009-2324-b
Closed

vybe wants to merge 57 commits into
devfrom
train/20261009-2324-b

Conversation

@vybe

@vybe vybe commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

Integration surface for #3468, #3467, #3437, #3478, #3477, #3339, #3486, #3412. Never merged; members merge individually once green.

webmixgamer and others added 30 commits October 8, 2026 14:38
…val (ent#754 C1)

Requirements first (Rule #1): requirements/skills.md §22 becomes the merged
Skills tab (own + shared, Run, Requires approval), superseding PLAYBOOK-001;
the AC5 canon-role deviation is named and linked to trinity-enterprise#848.

GET /api/skills now reports per skill:
- source: "platform" when the platform's .trinity-skill.json marker is in the
  directory (the same isfile test the backend uses), else "agent";
- dir: the directory name, which the gate fingerprint resolves before the
  frontmatter name;
- approval: the closed set {recommended}, read with the backend contract's
  precedence (trinity: block first, then the flat key).

All three are informational only. A behavioural parity table runs the same
documents through both parsers.

Refs Abilityai/trinity-enterprise#754

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…(ent#754 C2)

The /playbooks proxy existed three times (agent page, public link, MCP
connector). It now lives once, in services/agent_skills_listing.py
(Invariant #1), and every live success keeps the listing in Redis under
agent:skills_list:{name} with no TTL (the user's ruling on Q1).

- GET /api/agents/{name}/playbooks?last_known=true is opt-in. When the agent
  is stopped or unreachable it serves the copy, labelled
  last_known: {captured_at, reason}. With no copy, or without the param, the
  original errors stand, so the other consumers are unaffected.
- Only a real listing is kept, and a live empty list replaces the copy. A
  failed read never overwrites it. A listing over 256 KB is not kept and
  drops the older copy.
- Redis is fail-open on every path.
- The keyspace is registered in CLEARED_KEYSPACES and cleared on teardown
  and on the create path, never on a stop or start.
- The public link keeps to its pre-existing per-skill fields, so the new
  source/dir/approval never reach an anonymous visitor.

Refs Abilityai/trinity-enterprise#754

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…forces it (ent#754 C3)

GET /api/agents/{name}/skill-gates gains two additive reads (no new route):

- approvers: [{kind, reachable, viewer_fills}] for every kind the install
  resolves. viewer_fills is decided by the same two functions enforce uses
  for self-approval (requester_from_principal plus the casefolded approver
  list, through a new non-raising approver_people wrapper). So the card's
  "you approve this" agrees with what a run does, and a parity test runs
  enforce for owner, viewer, admin and agent key. Booleans only, no person
  data.
- ?probe=true adds hook, the agent's /health -> skill_gate_hook (ent#752).
  It is one direct read with a 3 s timeout and no breaker bookkeeping, and is
  honoured only for a person who may manage the agent's skills (the #3052
  precedent). A 200 without the field is "predates", meaning an image older
  than the hook. No answer is "unknown", never "not ok".

Refs Abilityai/trinity-enterprise#754

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…54 C4)

The executions list and detail gain gate_self_approved and
gate_self_approved_by_viewer. They are derived from ent#752's existing
self_approved records in skill_gate_requests (UNIQUE
dispatched_execution_id), so there is no new column and no migration (UC1,
approved at the plan gate).

- List: one batch read per page (get_self_approved_runs, scoped to the agent
  and the ids). The PERF-001 column list the portal shares is untouched.
- Viewer compare: the record's requester key vs the caller's casefolded
  email. It is done server-side and only for a person principal, so the
  email never leaves the server and a machine key is never "you".
- The read never fails the executions read. A record whose write failed
  (fail-open by design) reads false, which matches what the in-agent hook did.

Driven end to end: the record is written through the real
record_self_approval and read back through the real routes.

Refs Abilityai/trinity-enterprise#754

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…uires approval (ent#754 C5-C6)

The Playbooks tab (the agent's own skills, with Run) and the library-only
Skills tab (assignment) become one Skills tab of fixed-size cards, built to
the approved design.

- Own skills: the agent's .claude/skills/, live. When the agent is stopped,
  its last-known list is shown with Run disabled and a stopped banner, never
  an empty tab.
- Shared skills: library assignments. The picker moves into an "Assign
  skills" dialog, sets show as status chips with a "Manage sets" dialog
  (AgentSkillSets unchanged), and the #2914 conflict, delivery warnings and
  superseded text open from the card's note line.
- Every card has one top-left badge area. The author's automation value is
  relabelled by meaning (runs unattended / asks mid-run / start by hand, Q4)
  on a new BaseBadge size="sm".
- Gate line, shown to everyone: names the approver kind, never a person,
  and says "you approve this" to the person who fills it.
- Owner or admin (not on a ghost or the system agent): Requires approval
  toggle plus approver picker; a kind nobody fills is shown but can't be
  selected. A gate with no matching skill is kept on its own card, with Clear.
- Run sends POST /task {"/<name>", async_mode} (the shared client times out
  at 30 s). A 202 shows the approval notice and nothing opens, a refusal goes
  to the error toast, and any other failure stays on the card.
- An in-agent enforcement warning appears for the owner when the hook reports
  anything but ok (predates, missing, altered, unsupported runtime).
- A self-approved run is marked on its Tasks row and its execution page.
- ?tab=playbooks resolves to skills (TAB_ALIASES, now in utils/agentTabs.js).
  Every rule is in utils/skillCards.js; data is in stores/skills.js and the
  new agent-scoped stores/skillGates.js.

PlaybooksPanel.vue and SkillsPanel.vue are deleted. Their specs are ported,
not dropped, and the mutation checks still go red. The two ratchet entries
are removed by hand; the raw-colour baseline is not regenerated.

Refs Abilityai/trinity-enterprise#754

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The chat and Workspace / menus head their lists "Skills" (Q5: both). The
composer placeholders, the Workspace overflow and empty lines, the
exposed-skills panel, the connector and sharing copy, and the dashboard's
update button follow. API and field names (/playbooks, exposed_playbooks,
run_playbook, use-playbook) are unchanged, as the AC says.

Refs Abilityai/trinity-enterprise#754

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…#754 C9)

This is the 10-03 eyeball case itself. The person who fills a gated skill's
approver role asked for it in the Workspace chat, the gate self-approved, and
the reply landed with no sign a gate applied.

Since #3166 every Workspace message names the turn that wrote it, so the
history read now marks each agent reply whose run has an ent#752
self_approved record. This is the same derivation as the Tasks marker: no
new column, one read per thread, a casefolded viewer compare, and no email
added. The reply poll's narrow read carries it too, so a reply that has just
landed shows "Ran without approval: you are the approver" under the bubble
without a reload (assistantRow and replyFromHistory carry the two fields).

The hotspot gets a two-line call in get_history; the logic lives in
skill_gate_map_service.annotate_self_approved_turns. This is its own commit,
so it can be cut. portalReplyRateable.spec.js pins the mapper shapes and the
reattach call, and is extended rather than loosened.

Refs Abilityai/trinity-enterprise#754

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ead store field (ent#754)

From the feature-flow sync's code read:
- When the sets read failed, the set-chips line vanished, so a viewer (who
  can't open the sets dialog) never learned it. The line now says "Couldn't
  read this agent's sets" with a Retry, and adds "Some sets need
  credentials" when any set lacks them (the plan's error registry).
- skillGates.warnings was written and asserted but never rendered. The map
  is re-read after every write and its reachable:false is already what the
  card's "nobody fills it yet" says, so the state is removed.

Refs Abilityai/trinity-enterprise#754

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…er docs (ent#754 C8)

- Feature flows:
  - new skills-tab.md (the whole slice);
  - playbooks-tab.md becomes a pointer;
  - skill-assignment, skill-injection, library-page, skill-gate,
    playbook-autocomplete, workspace-composer-typeahead, mcp-connector,
    public-agent-links, tasks-tab, execution-detail-page and run-agent-loop
    are brought up to date;
  - the index gets one row plus one Recent Updates row.
- Architecture:
  - backend (agent_skills_listing, the keyspace, the map service's reads);
  - api-endpoints (/playbooks?last_known, the executions flags,
    skill-gates approvers / probe);
  - frontend (the Skills tab paragraph);
  - agent-runtime (the SkillInfo facts);
  - workspace (the self-approved reply).
- Requirements: §22.3 now names SkillsTab.
- Screenshot manifest: sources point at SkillsTab. The image itself still
  shows the old tab and needs a localhost capture.
- User docs: 14 pages move from the Playbooks tab to the Skills tab, plus a
  new "Requiring approval for a skill" section. Anchors and FAQ headings
  are kept.

The disclosure-guard pattern finds 0 hits. All 3,084 unit tests that read
the docs pass.

Refs Abilityai/trinity-enterprise#754

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ting (ent#754)

Two of the listing's lifecycle properties were pinned by reading source
(AST call lists), not by running the code. Both now execute:

- the create path, through the #1484 characterization harness that runs
  the real create_agent_internal: the name's listing is dropped beside the
  breakers and before the container exists (order clear -> forget -> run);
- the breaker sweep a start runs: the real clear_agent_breakers leaves the
  copy in place.

Mutation: removing forget(config.name) from crud.py turns
test_ent754_create_drops_a_recycled_names_skills_listing red; adding a
forget to clear_agent_breakers turns test_a_start_does_not_clear_the_copy
red. Both restored byte-identical.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…the next (ent#754)

AgentDetail is KeepAlive'd and the merged Skills tab is open to everyone,
so the same tab instance carries from one agent to the next. Two things
leaked across that switch:

- The tab's own verb state: the last outcome line ("Unassigned x."), card
  errors, a pending-sync emphasis, open dialogs, the filter, and a run still
  starting. A run that answered after the switch opened the NEW agent's
  Tasks tab on the previous agent's execution id, or toasted there. The
  state now resets on a switch, and Run / Unassign / Sync drop an answer for
  an agent the page has left.
- The skills store's reads and writes (load, sync, save, set assign and
  unassign) applied whatever answered last. A slow load for the previous
  agent overwrote the next agent's assignments, so an Unassign there could
  PUT the previous agent's list onto it. Each answer is now checked against
  the agent it was asked for (a load also against a newer load), a write's
  follow-up read targets that agent, and the per-agent busy flags reset
  with the agent. saveAssignments answers null when superseded, so the
  assign dialog doesn't report a save on the wrong agent.

Plan §3 promised these guards ("stale-response guards keyed on agent");
only the gate store had them.

Mutation: dropping the reset turns "the previous agent's outcome line and
card errors do not follow" red; dropping the Run guard turns "a run that
answers after the switch…" red; dropping the load guard turns the store's
"a load" red. Restored byte-identical.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…swered (ent#754)

The Shared section counted itself loaded on the platform-wide library
status, which survives an agent switch, and every card was drawn before the
gate map answered. So, until (or instead of) a real answer:

- a switched-to agent read "No shared skills yet" with an Assign button
  whose draft started empty, and saving it replaced the real list;
- a library skill the agent holds was shown under Own as "from library: the
  next sync removes it";
- a failed assignments or library read was invisible ("The library is
  configured but has no skills yet");
- every approval toggle read "off", and turning one on wrote the default
  approver over an existing gate. A failed gate-map read hid every gate line
  from everyone, so the AC "a gated row shows its gate to everyone" failed
  silently.

The store now records sharedLoaded (status, assignments, the library list,
then the sets read) per agent. The sections draw once the reads they are
built from have answered or failed: Own needs the list, the assignments and
the gate map; Shared needs the assignments and the map. A skill therefore
never moves between sections, and the sets and enforcement lines land in the
same paint as the cards. A failed read says so: Shared gets LoadFailed with
Retry; a failed gate map gets a line for everyone, with Retry, and the
toggles are held. Assign, Manage sets and Sync wait for the assignments.

Also:
- a running agent that isn't answering yet offers "Check again", since
  nothing else re-asks;
- a failed refresh of the own list keeps the list and names it (the
  stale-banner copy, with Retry);
- a 200 that is not a listing is a failure, never an empty list (plan §3,
  row 5).

Mutation:
- the old sharedView gate turns the failed-library case red;
- the unguarded leftover rule turns the pure case and the failed-assignments
  mount red;
- dropping the malformed-list check turns its case red;
- an Own view that doesn't wait turns the two waiting cases red;
- unheld shared toggles turn the pure case red.
All restored byte-identical.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rrent default (ent#754)

- A skill whose frontmatter name differs from its directory can be gated
  under the directory; the card found that gate, but every write used the
  name. Turning approval off sent DELETE /skill-gates/<name>, so the toggle
  sprang back on, and changing the approver created a second gate. An
  existing gate is now changed and cleared under its own key. Only a new
  gate takes the name: that is what a request types, so the request-time
  check sees it, and the in-agent check matches either key.
- The kind that turning approval on sends was fixed when the card was
  drawn. After the map was re-read (every write re-reads it), a card could
  still send a kind nobody fills: the bodiless-reset case S7 was meant to
  prevent. The default is now computed from the current map; a pick holds
  the select only until its write settles.
- Every approval switch was announced as "Requires approval". Each is now
  named for its skill. BaseToggle lets an explicit aria-label name the
  switch beside a visible label; its only aria-label caller has no label,
  so it is unchanged.
- The enforcement line names every hook state in words. The raw code
  ("(not_root_owned)") is on hover only.
- The "via <sets>" badge, the one that can grow long, comes last with its
  full text on hover, so it can't push "name conflict" or "failed" out of
  the fixed badge area.
- Five gate-store fields that nothing outside its own spec read are gone.

Mutation: the name-keyed own gate turns the pure key case and both
dir-keyed mount cases red; a default frozen at draw time turns "the default
follows the map when it is re-read" red. Restored byte-identical.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… the safe action (ent#754)

Design principle 19: destructive actions restate the consequence and focus
the safe action first.

- A card's Unassign removed the library skill in one click, and with it the
  skill's explicit approval gate (the server drops it on unassign). It now
  asks first ("Unassign /x?") and says the approval requirement goes with
  it when the skill is gated. The dialog opens on Cancel. The old tab's
  untick-then-Save flow had no confirm either, but it took a second
  deliberate step.
- The details dialog put initial focus on "Unassign library skill". The
  Manage sets dialog, which now holds AgentSkillSets, put it on the first
  "Unassign set". Both buttons are now marked data-destructive, so focus
  lands on a safe control.

Mutation: wiring Unassign straight to the write turns the confirm case
(and the switch case that goes through it) red; dropping either
data-destructive turns its dialog's focus case red. Restored byte-identical.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…lter (ent#754)

Two lines carried their behaviour with no test that could fail without
them. Each surviving mutation was found in review.

- `?probe=true` is honoured only for a person. can_manage_agent_skills
  alone admits a system-scoped key, and an agent key that holds
  skills.manage on an agent its owner owns. Forcing is_person_principal to
  True left all 26 gate-map tests green. Both keys are now driven through
  the real route: hook stays null and no /health call is made.
- A run someone else approved is dispatched under the same UNIQUE
  dispatched_execution_id as a self-approved one, and only the
  self_approved state means "ran without approval". Dropping that filter
  left all 9 execution and Workspace tests green. A dispatched record,
  written through the real create, claim and transition calls, now reads
  (False, False) on the list and the detail.

Mutation: each turns its new test red (3 cases), restored byte-identical.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…gs and the library contract (ent#754)

Listing proxy (services/agent_skills_listing.py and its three readers):
- A container removed between the lookup and the reload raised Docker's
  NotFound inside each route's catch-all. Its text (the Docker API URL and
  the container id) reached the caller, including the anonymous public
  link. It is now a plain not_found: 404 "Agent not found" on the agent
  page, 503 "Agent is not running" on the public link.
- Every httpx transport failure is now "unreachable". A connection dropped
  mid-answer (ReadError, RemoteProtocolError) was a 500 that last_known
  could not cover.
- The connector keeps its own "Agent error: …" wording. A body that is not
  a listing is a 500 again, never "no skills".
- Two `except HTTPException: raise` branches nothing could reach are gone.

Self-approved flags:
- The batch read is chunked (500 ids per statement). The page is the
  caller's unbounded `limit`, and SQLite and PostgreSQL both cap bound
  parameters; one oversized IN would raise and fail open to "no run is
  marked".
- Workspace: a machine principal (a system key on the platform session,
  is_person false) sees the fact but is never "you", as on the executions
  path. The history route passes principal.is_person.
- Workspace: the synchronous fallback (POST …/chat) now answers both flags.
  `_persist_reply` computes them, PortalChatResponse declares them, and the
  client reads either spelling through portalUtils.turnGateFlags. That
  closes the gap C9 had named, as plan C9 said: "the completion payload
  carries the flag".

Hook probe: an unexpected failure (not httpx, not JSON) is still `unknown`,
but it is now logged at debug with exc_info, so a fault that recurs on every
call leaves a trace.

Library: GET /skills/library builds SkillInfo field by field and never
named `approval`, so it was always null over REST. A Shared card's "author
recommends approval" (AC8) had no source for a stopped agent or an older
image.

Mutation: each fix reverted turns its own new test red (11 cases across
the cache, gate-map, executions, Workspace and #2580 suites), restored
byte-identical. 55 neighbouring backend files: 1,238 passed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…orkspace architecture (ent#754)

- requirements §22.1: the gate-write key and the live default approver;
  nothing drawn from an unanswered read, with each failure named; Unassign
  asks first; an agent switch carries nothing over.
- skills-tab.md: a "Loading, failures and agent switches" subsection, the
  new and extended test rows, the Workspace fallback gap removed from the
  known limits, a revision row.
- architecture/workspace.md: the synchronous fallback answers the
  self-approved pair, and "you" is for a person only.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- A tab that outlives an agent switch carries its component state, every
  async answer and its "loaded" flag across unless each is keyed to the
  agent. This extends the 10-07 write-then-reload fragment to reads,
  view state and the loaded flag.
- A field declared on a response model is still dropped by a route that
  builds the model field by field. That is the third shape of the #2580 /
  ent#2320 allowlist class.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Operator ruling (review round): beside a run's source_user_email,
"ran without approval" tells the reader that this person fills the gate's
approver kind. On enterprise that names a role holder. The gate map already
withholds `set_by` from machine keys for this reason (#715). An agent,
MCP, connector or system key now reads both flags false:

- self_approved_flags returns no flags for a non-person principal
  (executions list and detail, so MCP list_executions / get_execution);
- annotate_self_approved_turns takes viewer_is_person, and the Workspace
  history read and the synchronous reply pass the caller's.

People who can see the agent (owner, admin, shared users) still see the
marker, and "you" only on their own runs.

Mutation: dropping either branch turns its machine-key test red (3 cases:
executions, Workspace history route, #2580 sync reply). Restored
byte-identical.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…w round (ent#754)

- requirements §22.1, skills-tab.md and architecture/workspace.md state the
  people-only rule for the self-approved marker.
- 75 `path:line` citations in docs/memory are remapped from the C8-era
  files (f51472b) to the current ones:
  - 44 explicit citations, each resolved to its full path first, so a
    shared basename (models.py, service.py) is never guessed;
  - 26 relative `:N` citations, resolved to the file named before them in
    the same section;
  - five citations into the listing service, where that heuristic picked
    the wrong file, corrected by hand.
  Citations into other files, and those already out of range, are
  unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…g internal (ent#754)

cso-diff 2026-10-08, finding 1 (LOW, independently verified). The listing
service reloads the container, and the public link has done so only since
this branch. Only Docker's NotFound was mapped, so any other Docker failure
(a daemon 500, a dropped socket) reached the routes' catch-all. Its text
went out in the 500: the Docker transport and API version, the full
container id, and the daemon's own message. That reached the
unauthenticated public link too. A caller cannot cause the fault, but the
leak was new.

- fetch_live maps any other reload failure to `unreachable` (503 "Could
  not read the agent's state", logged with the cause). A kept copy is
  served for it, as for an agent that is not answering.
- The public link's catch-all answers a fixed sentence and logs the cause.
  An anonymous caller never gets an exception's own text.

Mutation: narrowing the new except turns the two daemon-fault cases red;
echoing the exception again turns the public fixed-sentence case red.
Restored byte-identical. Public-link suites: 135 passed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Daily gate, diff scope 51b01c7..7ab9439:
- one LOW finding: Docker daemon text echoed by the listing's reload,
  including on the anonymous public link. Independently verified, fixed
  in-branch (852fcc4), mutation-proven;
- 11 security guard suites green (277 tests);
- no secrets, invisible Unicode or v-html in the added lines;
- coverage gaps named: no Docker pass, no PostgreSQL run, and the live
  end-to-end run pending /verify-local.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…(ent#754)

Run on the merged Skills tab sends `async_mode: true`, so the Tasks tab
opens on the run while it is still going. TasksPanel read its rows once
on mount and its 5 s poll refreshed only the queue chip, so the row read
"running" until Refresh (found on the localhost eyeball).

- The poll re-reads the list while any loaded row is queued, running or
  pending_retry, and stops when none is. It never stacks a read on one
  still out, and it holds off while a task typed into the panel is
  awaited (its local row stands for it; a read meanwhile listed it twice).
- Only the latest read is applied, so a slow read for the agent the
  KeepAlive'd page left is dropped.
- The highlighted row is opened and scrolled to once, on the first read
  that answers. Before, every load re-opened it.
- A row whose status changes drops the details read while it ran; the
  open one re-reads them in place, so its result shows where it is.

Mutation-proven (tasksPanelRunRefresh.spec.js, 14 tests): dropping the
re-read reddens 8, and each guard reverted reddens its own test
(highlight once, typed-task hold-off, latest-read guard, no stacking,
open-row re-read, cache drop, the in-flight status set).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… card row (ent#754)

The Skills card's approver picker sits in a 40px row beside a 20px
toggle and sm buttons. The `field` recipe is a form field, 38.3px in
Chromium, edge to edge of that row; `ghost` is the composer's 44px box
(it overflowed the row) and drops its border, so a disabled picker read
as faint text. The approved card design draws a compact bordered select.

`size="sm"` sizes the `field` recipe to BaseButton sm's box: 12.5 ink,
padding 4x10, ghost's 12px chevron at right 8px (28.8px measured).
Both sizes come from one builder in fieldClasses.js, so they differ in
the box only, and FIELD_CLASS (BaseInput, BaseTextarea) is byte-identical.
`ghost` takes no size. Cataloged in design-system.md §5 and the contract.

Mutation-proven (baseSelectSize.spec.js): ignoring the size, or keeping
the md chevron, reddens the sm test.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
From the localhost eyeball:
- The Own meta line read "From .claude/skills, .claude/skills". The
  agent server scans /home/developer/.claude/skills and ~/.claude/skills,
  one folder in the container (HOME=/home/developer, USER developer).
  The store now names each folder once (every image), and the agent
  server scans and names it once (new images).
- The Unassign confirm read as already done ("is removed"). It now says
  the library skill will be removed, along with its approval requirement
  when the card is gated.
- The approver picker uses BaseSelect size="sm" (previous commit), so it
  sits inside the 40px approval row; disabled, it is still a bordered box.
- "Own skills 0" sat above a kept-gate card: the count now counts every
  card the section lists.
- The not-synced line again says statuses appear after a sync (the
  Shared cards' delivery badges come from the last sync).

Flow docs follow, and their stale citations are corrected (onEditRun,
onSetGate/onClearGate, showOwnerRow/showManage, and skill-assignment's
Unassign, Manage sets and Sync now rows).

Mutation-proven: each change reverted reddens its test
(skillsTab.mount.spec.js: folder once, both Unassign wordings, the
picker's size and its ghost alternative, Own count, sync hint;
test_ent754_agent_server_skillinfo.py: the folder scanned once).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The Tasks tab's own flow states the eyeball-round poll: re-read while a
row is queued / running / pending_retry, only the latest read applied,
the highlighted row opened once, a settled row's details re-read. The
loadExecutions excerpt and its line citation follow the code.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ent#754)

`test_the_attribution_reaches_both_rows` pins `portal_chat`'s call to
`_persist_reply(..., voice_call_id, execution_id)` with its closing
paren. The review round passes `viewer_is_person=` after the execution
id, so the literal no longer matched and the test failed in
/verify-local's stage 1 (it was outside the neighbour runs). The pin now
ends at the execution id, as the same test's `_persist_user_turn` pin
already does for #3265's `attachments=`.

Still bites: dropping the call id or the execution id from the reply
write reddens it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Brings the branch up from 51b01c7 to f6bedcd (6 commits: ent#838
a2a trusted networks, #3357 Work card, #3306 slot admission, #3308
subprocess ownership, ent#841 Workspace chat tabs, #3311 sanitizer).
No textual conflicts.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ent#841 and ent#838 moved lines in PortalConversation.vue, portalUtils.js,
client_portal/router.py and service.py, database.py, db_models.py and
db/schema.py. The 15 citations this branch wrote into those files point
at the same code again (each checked: the cited line's text is the
same before and after the merge).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Brings the branch up to ed59049 (12 commits, incl. #3406 public-link
sessions, #3422 `user-invocable:` in the library, #3418 safe_yaml).

One conflict, in docs/user-docs/automation/skills-and-playbooks.md: this
PR's prose kept, with the `user-invocable:` key spelling (the hyphen is
the only spelling the agent server reads). The Skills flow's Run table
now names the listing field and the frontmatter key apart.

PublicChat.vue merged cleanly with #3406 (this PR's "/ for skills"
placeholder kept).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
webmixgamer and others added 18 commits October 9, 2026 11:37
…tchers

From the #3412 review round (trinity-enterprise#754): an interval opened
on mount kept running while the KeepAlive'd page was away, and three
watchers on one agent switch each re-read, one of them for the agent
being left.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The 28 click-through phases dated from January and none matched the
product: they drove a Terminal tab that is hidden, an /agents page that
is now a redirect, eight GitHub-template agents of which three exist,
and several pages that are gone. Two of them deleted the fixture agents.

- Rewrite the 23 surviving phases against current source. Every phase is
  now independent, uses only the three local fixture agents, restores
  what it changes, carries no credentials, and moves anything that needs
  a second user, a mailbox, an external repo or a fresh database under a
  Manual-only heading.
- Retire phases 7, 8, 10, 14 and 16 as short stubs that say why and what
  covers them now.
- Add phases 29-37 for surfaces that had no scenario: Workspace (chats;
  inbox, rooms, projects), Operations, Library, the newer agent tabs,
  Canvas and shared canvas links, Settings keys and retention,
  onboarding and app chrome, mobile admin.
- Rewrite README.md and INDEX.md; drop references to a
  run_test_phases.py runner that never existed.

Source-verified only: no phase has been run in a browser since the
rewrite, and the index marks each one accordingly.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Whole-system sweep report and state: 1 critical, 27 medium, 62 low;
24 public bugs filed (#3443-#3466), 2 existing issues commented.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ifted (#3442, #3272, #2420)

A full suite run against a freshly rebuilt stack failed 7 of 10 tiers; one
product bug came out of it (#3441) and the rest was test rot:

- scheduler fixture: add `chain_depth` and `claim_token` to the hand-mirrored
  `schedule_executions` DDL (70 failures -> 243 passing)
- test_deploy_writable_templates: restore the parent-package attributes the
  re-import rebinds, not only the sys.modules entries, so later files stop
  patching one module object while the code reads another (#3272)
- test_subscriptions: delete only the suite's own `test-` subscriptions and
  skip the no-subscription case when the operator has real ones, instead of
  wiping every subscription on the instance; test_subscription_auto_switch
  restores the setting it found (#2420, first two criteria)
- test_circuit_breaker: assert the platform-alert seam the dormant alert
  moved to in #3246
- test_1661_sanitizer_linear: pin the agent sanitizer's known-value cache so
  an ambient TRINITY_TEST_PASSWORD cannot redact the key under test
- test_ip_rate_limit_fix: give the fastapi stub the `status` module
  services.rate_limiter now imports
- test_operator_queue: answer with a decision the item actually offers
- test_parallel_task: `queued` is a legitimate first status

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…n, so history survives a reload (#3443)

get_or_create_session stores an identifier verbatim unless it is an email,
but get_session_by_identifier lowercased its argument unconditionally. An
anonymous identifier is a mixed-case secrets.token_urlsafe(16) token, so the
lookup behind GET /api/public/history/{token} and DELETE
/api/public/session/{token} never matched the row the chat turn had just
written: history came back empty after every reload, and "New conversation"
cleared nothing. Where two tokens differed only in case, the mixed-case
visitor was handed the lowercase visitor's session.

The lookup now matches the exact string first and reaches a row through its
lowercased form only when that row is identifier_type = 'email', which keeps
email lookups case-insensitive. One SQLAlchemy Core implementation serves
both SQLite and PostgreSQL; both route callers already pass the anonymous
token unmodified, so no caller changes.

Tests: tests/unit/test_3443_public_chat_history_lookup.py. With the fix
reverted, three go red:
test_mixed_case_anonymous_token_finds_its_own_session,
test_anonymous_tokens_differing_only_in_case_stay_separate and
test_anonymous_lookup_is_not_case_insensitive.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…never the operator's error (#3461)

A public link's failed turn rendered the execution row's own `error` in
the chat banner: the agent server's diagnosis ("Execution failed with no
output (exit code 1): ...") followed by remediation addressed to the
operator. The page is read by anonymous visitors.

Two layers, because the page was only the last hop:

- `GET /api/public/executions/{token}/{id}/status` is unauthenticated and
  returned `execution.error` verbatim for a `failed` row. It now answers
  the fixed line the synchronous public path already answers; the detail
  stays on the row for the operator. A cancel reason (#679) and a skill
  gate's notice (trinity#3274) are written for the visitor and still pass
  through.
- `PublicChat.vue` no longer renders the status route's `error` for a
  failed turn at all. It shows one plain line and puts the visitor's
  message back in the input, so sending again is the retry.

Tests, each red with its fix reverted and green with it restored:
- src/frontend/tests/unit/publicChatFailureMessage.spec.js (mounted; 4/4
  red under mutation, incl. "never renders the execution row's own error
  text")
- tests/unit/test_3461_public_status_failure_text.py (2 red under
  mutation; the cancel and gate pass-through cases stay green)
- tests/unit/test_679_public_poll_cancel.py: the `failed` case pinned the
  raw string and now pins the fixed line.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…it fails; the key is a rung, not a reassignment (#3470)

SUB-003 remediated once per turn (one switch or one API-key fallback, one
re-issue) and three surfaces then handed the retry back to the person. Now
each refused attempt asks `advance_subscription_walk` for the next credential
— another subscription, then the platform key — and the SAME execution_id is
re-issued with it as a request-scoped `auth_override` the agent applies to
that one spawn (`X-Trinity-Auth-Override` echoes it; an old image falls back
to the #792 commit-and-retry, one restart max, no key rung). Nothing moves
until a trial serves; the serving subscription is committed once with a CAS
on the assignment the turn started from. Bounded by the remaining turn budget
(#2789, the only ceiling) over a 30 s floor and by the walk's own rung cap.

Part 2: `fallback_to_api_key` no longer clears the assignment, flips
`use_platform_api_key` or restarts the container. The key is tried by override
only, skipped while recently refused (auth/billing 2h, a 429 5 min), never
offered to an agent with `use_platform_api_key` explicitly off; a key refusal
is one more exhausted candidate and the assignment is preserved throughout.

Agent server: `AuthOverride` (SecretStr) on /api/task + /api/chat,
`build_execution_env(…, drop=)`, header on success and HTTPException, value
staged for redaction. Backend: the walk core in subscription_auto_switch, both
dispatch loops, the portal ladder, `db.set_execution_subscription` (SUB-004
re-attribution, no schema change), `is_container_fault` so a container-level
503 never walks, two EXEMPT dedupe keyspaces. Scheduler classifier mirror
re-synced. Docs + two learnings entries.

Tests: tests/unit/test_3470_subscription_walk.py (51); test_792/test_2789
re-seamed on the walk; test_2638's one-remediation test reframed.

Fixes #3470

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
….manage + create_project / import_project (trinity-enterprise#588)

Re-scope of ent#588 (ruling 2026-09-29, ent#661): on Trinity, /project-init
creates the project through the platform. ent#661 planned the agent right as
"a new capability beside skills.manage; changing members or visibility stays
person-only" — this is that capability, on the #3236 seam:

- db/capability_grants: `projects.manage` joins the closed set — start a
  Workspace project ON BEHALF OF the agent's owner (owner = creator and
  first member, agent = steward, members-only) and import one of the
  agent's OWN folder projects. People and visibility are never grantable.
- dependencies: its named refusal (`project_management_not_permitted`)
  with the ask instruction.
- Settings → Permissions to change itself: a fifth toggle, "Start projects".
- MCP: `create_project` and `import_project` (license-blind, like the rest
  of projects.ts), each saying the grant is needed and what the owner keeps.

The enforcement and the create/import themselves are in the enterprise
projects module (abilityai/trinity-enterprise, same issue), which checks
projects.manage before any read. Until this lands on dev the enterprise
routes refuse every agent — the capability is not in the closed set — so
the merge order is this PR, then the enterprise PR.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ts.manage (ent#588) — mechanical, per the merge-train note on the PR
# Conflicts:
#	tests/registry.json
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ui PR touches the frontend UI — triggers Playwright e2e tests

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants