Skip to content

System Health follows the runtime switcher for all 32 runtimes: no OpenClaw cards elsewhere, no invented sub-agent success rate - #5982

Merged
vivekchand merged 5 commits into
mainfrom
fix/system-health-runtime-scoped
Sep 15, 2026
Merged

vivekchand merged 5 commits into
mainfrom
fix/system-health-runtime-scoped

Conversation

@vivekchand

@vivekchand vivekchand commented Sep 14, 2026 •

Copy link
Copy Markdown
Owner

Problem

With the runtime switcher on Claude Code, Overview showed OpenClaw-only health: "OpenClaw Gateway :18789", a "Healthy" heartbeat, Cron Jobs, "Connect a channel", gateway config diagnostics and an empty "Is your agent alive?" card. Sub-Agents read "0 runs, 100% success": /api/system-health counted session ids containing "subagent" across every runtime, stamped each one a success, and defaulted an idle node to 100.

The same was true for every non-OpenClaw runtime, and the capability map that decides what a runtime has had drifted: the hosted dashboard, which has only the static map, hid sub-agents for 13 runtimes that emit them, showed them for Exo, which does not, and had no entry at all for Muse Code, OpenWorker, qm and Replit (so the sidebar showed them every tab, OpenClaw's included).

Fix

  • Runtime-scoped System Health (loadSystemHealth, _shRuntimeScope): sections follow the runtime's capabilities. Gateway service and vitals, heartbeat (both Overview heartbeat cards too), version regression, diagnostics, inference and security read from openclaw.json, and delegation chains need GATEWAY_RPC. Crons need CRONS, channels and channel ingest need CHANNELS, sub-agents need SUBAGENTS or real runs. Disk, sandbox, daemon and handler latency are machine-wide and always shown. A line under the title says what the panel covers; the node-wide reliability trend hides under one runtime. Gating runs before the fetch, so a slow request cannot paint OpenClaw's error state under another runtime. A runtime switch, and the local /api/agents capability override landing, both re-render the panel.
  • Capability map synced for all 32 runtimes with the adapters' _base_capabilities() in clawmetry-pro 0.7.28 (the version the local daemon runs, read from the installed wheel): SUBAGENTS added for Codex, Copilot, Cursor, DeepAgents, Devin, Goose, Grok, n8n, NanoClaw, OpenCode, Pi, PicoClaw, Qwen Code; removed for Exo; entries added for Muse Code, OpenWorker, qm, Replit. COST flags unchanged.
  • Honest sub-agent numbers (subagent_health_block): reads query_subagent_stats_by_runtime per runtime (?runtime=); its prefix list covers all 30 non-OpenClaw runtimes. successPct is null until a run finished ("N/A, No finished runs"), and "unavailable" when the store cannot be read. A legacy payload without completed/failed counts never renders its percentage.

Result per runtime

Runtime System Health shows
OpenClaw, NemoClaw machine-wide + OpenClaw gateway/heartbeat/diagnostics, Crons, Channels, Sub-Agents
Antigravity, Claude Code, Cline, Codex, Copilot, Cursor, DeepAgents, DeepSeek Harness, Devin, Gemini CLI, Goose, Grok, Hermes, Kimi, Muse Code, n8n, NanoClaw, OpenCode, OpenHands, OpenWorker, Pi, PicoClaw, qm, Qwen Code machine-wide + Sub-Agents
Aider, Exo, Grok Bot, Lovable, OpenExecutive, Replit machine-wide only (Sub-Agents appears if runs exist)

Verified

  • tests/test_system_health_runtime_scope.py (24 tests): backend block, template wrappers, gating, heartbeat cards, a node run of _shRuntimeScope against the real map, and a guard that every shipped runtime has an entry and only OpenClaw/NemoClaw claim gateway, cron or channel capabilities.
  • The matrix above computed by node over the shipped app.js for all 32 FREE_RUNTIMES | PAID_RUNTIMES.
  • Served page on a local build with every System Health endpoint answering full OpenClaw data: under Claude Code only machine-wide cards and Sub-Agents render, and only /api/system-health?runtime=claude_code + /api/handler-latency are fetched; under OpenClaw the full set renders; both heartbeat cards hide under Claude Code. No console errors.
  • Real sub-agent rows on this machine exist for claude_code, codex, copilot, opencode, pi, qwen_code; each of those now shows the card.
  • check_ac_coverage --check, lint_daemon_allowlist, check_py39_annotations, node --check pass. test_runtime_tab_capability_parity fails identically on untouched origin/main (OpenClaw now declares INPUTS/REASONING), unrelated.

Cloud

The hosted dashboard serves this app.js, so the gating and the synced map apply there. Cloud's own /api/system-health still sends successPct: 100 without completed/failed counts; the card ignores that value, and fixing the cloud payload is a follow-up in clawmetry-cloud.

Requirements: https://factory.8090.ai/project/b415065f-ab2f-4f53-8864-0c009fd098cb/requirements/d518c6c3-eb50-4b0f-9ed0-59440380b7bf (AC-OBS-002.3), https://factory.8090.ai/project/b415065f-ab2f-4f53-8864-0c009fd098cb/requirements/1a471e48-24c5-4a19-8e40-bf35c24f7084 (AC-GOV-001.3)

🤖 Generated with Claude Code

https://claude.ai/code/session_01QpcVFD77MdQbNCsvrBSLDW

@8090-software-factory

Copy link
Copy Markdown

✅ Drift Bot (ClawMetry): no drift detected

Drift Bot analyzed the changed files against this project's blueprints and requirements and found no drift.

@github-actions

github-actions Bot commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor

Visual diff

Comparing 5a5df63d35d4 (head) against the PR base branch.

26 of 70 comparison(s) flagged (>1% pixel diff).

View Before After Diff
desktop overview before after diff · 0.01%
desktop flow before after diff · 0.02%
desktop brain ⚠️ before after diff · 100.00%
desktop usage ⚠️ before after diff · 100.00%
desktop crons before after diff · 0.00%
desktop memory before after diff · 0.03%
desktop security before after diff · 0.18%
desktop subagents ⚠️ before after diff · 100.00%
desktop transcripts before after diff · 0.01%
desktop logs ⚠️ before after diff · 5.41%
desktop skills before after diff · 0.38%
desktop models before after diff · 0.00%
desktop approvals ⚠️ before after diff · 3.21%
desktop alerts before after diff · 0.24%
desktop notifications before after diff · 0.00%
desktop limits ⚠️ before after diff · 1.45%
desktop history before after diff · 0.27%
desktop channels ⚠️ before after diff · 1.43%
desktop harness before after diff · 0.00%
desktop inventory before after diff · 0.00%
desktop nemoclaw ⚠️ before after diff · 100.00%
desktop guard before after diff · 0.26%
desktop signals before after diff · 0.00%
desktop policy before after diff · 0.01%
desktop selfevolve before after diff · 0.01%
desktop swimlane before after diff · 0.67%
desktop tool-catalog before after diff · 0.01%
desktop tracing ⚠️ before after diff · 2.44%
desktop turn-anatomy before after diff · 0.36%
desktop version-impact ⚠️ before after diff · 1.91%
desktop context-economics before after diff · 0.37%
desktop agents ⚠️ before after diff · 1.79%
desktop evals before after diff · 0.23%
desktop bench before after diff · 0.00%
desktop trail before after diff · 0.37%
mobile overview ⚠️ before after diff · 1.20%
mobile flow ⚠️ before after diff · 5.53%
mobile brain before after diff · 0.01%
mobile usage before after diff · 0.13%
mobile crons ⚠️ before after diff · 1.69%
mobile memory ⚠️ before after diff · 4.24%
mobile security before after diff · 0.70%
mobile subagents before after diff · 0.00%
mobile transcripts before after diff · 0.01%
mobile logs ⚠️ before after diff · 1.66%
mobile skills before after diff · 0.01%
mobile models before after diff · 0.01%
mobile approvals before after diff · 0.00%
mobile alerts ⚠️ before after diff · 100.00%
mobile notifications before after diff · 0.02%
mobile limits ⚠️ before after diff · 1.69%
mobile history before after diff · 0.99%
mobile channels ⚠️ before after diff · 100.00%
mobile harness ⚠️ before after diff · 4.03%
mobile inventory before after diff · 0.01%
mobile nemoclaw ⚠️ before after diff · 100.00%
mobile guard before after diff · 0.00%
mobile signals ⚠️ before after diff · 4.25%
mobile policy ⚠️ before after diff · 1.61%
mobile selfevolve before after diff · 0.01%
mobile swimlane before after diff · 0.01%
mobile tool-catalog before after diff · 0.00%
mobile tracing ⚠️ before after diff · 1.09%
mobile turn-anatomy ⚠️ before after diff · 1.02%
mobile version-impact before after diff · 0.01%
mobile context-economics before after diff · 0.01%
mobile agents ⚠️ before after diff · 1.69%
mobile evals before after diff · 0.01%
mobile bench before after diff · 0.01%
mobile trail before after diff · 0.01%

Folder: 5a5df63d35d4. Full PNGs also attached as a workflow artefact.

Generated by visual-diff bot. Pixel diffs >1% flagged; eyeball the table before merging. This check is non-blocking — fail = bot bug, not a code problem.

@8090-software-factory

Copy link
Copy Markdown

✅ Drift Bot (ClawMetry): no drift detected

Drift Bot analyzed the changed files against this project's blueprints and requirements and found no drift.

github-actions Bot pushed a commit that referenced this pull request Sep 14, 2026

Copy link
Copy Markdown
Owner Author

E2E Gate timed out — not a code failure; needs a re-push to trigger a fresh run.

The E2E Gate (34850096106) polled commit f44af37 for 3600 s and timed out at 15:13 UTC because five checks finished slightly after the window:

Check Completed at
Store invariants 15:15 (2 min late)
E2E Browser Tests 15:18 (5 min late)
MOAT Keystone 15:22 (9 min late)
pip install (ubuntu py3.11) ~15:20
Entitlement API tests 15:25 (12 min late)

All five passed cleanly — the gate just hit the 1-hour ceiling before they reported in. All other checks including Syntax & Lint and Drift Bot were green. scripts/gen_module_map.py --check, check_ac_coverage.py --check, sync_runtime_count.py --check, and lint_daemon_allowlist.py all pass locally on this branch.

A trivial push to the branch will start a fresh CI run; on a quieter runner slot the gate should complete well inside the 60-minute window.


Generated by Claude Code

Copy link
Copy Markdown
Owner Author

Auto-rebase attempted by the PR-mergeability janitor; conflicts span >5 files (add/add conflicts in .github/workflows/ci.yml, CHANGELOG.md, CLAUDE.md, clawmetry/local_store.py, clawmetry/sync.py, clawmetry/static/js/app.js, and 30+ more files). Cannot resolve mechanically — manual rebase needed.


Generated by Claude Code

@8090-software-factory

Copy link
Copy Markdown

✅ Drift Bot (ClawMetry): no drift detected

Drift Bot analyzed the changed files against this project's blueprints and requirements and found no drift.

Copy link
Copy Markdown
Owner Author

Resolved merge conflict in CHANGELOG.md (kept both the PR's System Health runtime-scope entry and main's enterprise-readiness release entry) and pushed a merge commit to bring the branch up to current main.


Generated by Claude Code

github-actions Bot pushed a commit that referenced this pull request Sep 14, 2026

Copy link
Copy Markdown
Owner Author

✨ auto-fixed: merged latest main into branch (was behind)


Generated by Claude Code

@8090-software-factory

Copy link
Copy Markdown

✅ Drift Bot (ClawMetry): no drift detected

Drift Bot analyzed the changed files against this project's blueprints and requirements and found no drift.

Copy link
Copy Markdown
Owner Author

Blocked on required review — skipping (auto-mergeability sweep). @vivekchand please approve when ready.


Generated by Claude Code

github-actions Bot pushed a commit that referenced this pull request Sep 14, 2026
@vivekchand
vivekchand force-pushed the fix/system-health-runtime-scoped branch from 784f5ae to aed367c Compare September 14, 2026 22:37
@vivekchand vivekchand changed the title System Health follows the runtime switcher: no OpenClaw gateway under Claude Code, no invented sub-agent success rate System Health follows the runtime switcher for all 32 runtimes: no OpenClaw cards elsewhere, no invented sub-agent success rate Sep 14, 2026
@8090-software-factory

Copy link
Copy Markdown

✅ Drift Bot (ClawMetry): no drift detected

Drift Bot analyzed the changed files against this project's blueprints and requirements and found no drift.

vivekchand and others added 3 commits September 15, 2026 00:11
…cess rate

With the switcher on Claude Code the Overview System Health panel listed
"OpenClaw Gateway :18789", a Healthy heartbeat, Cron Jobs and "Connect a
channel", none of which Claude Code has. Sub-Agents read "0 runs, 100%":
the endpoint counted session ids containing "subagent" across every runtime
and stamped each one a success.

- loadSystemHealth scopes sections to the runtime's declared _CM_RT_CAPS:
  OpenClaw-family checks (gateway service and vitals, heartbeat, version
  regression, diagnostics, openclaw.json inference/security, delegation
  chains) need GATEWAY_RPC; crons need CRONS; channels and ingest need
  CHANNELS; sub-agents need SUBAGENTS. Disk, sandbox, daemon and latency
  stay. Gating runs before the fetch, so a slow request never paints
  OpenClaw's error state under another runtime. A scope line says what
  the panel covers; the node-wide reliability trend hides under one runtime.
- Switching runtime reloads the panel instead of waiting 30s.
- /api/system-health reads query_subagent_stats_by_runtime with ?runtime=
  and returns runs/completed/failed/running; successPct is null until a
  run finished, and "unavailable" when the store cannot be read.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QpcVFD77MdQbNCsvrBSLDW
"Is your agent alive?" and the Overview header heartbeat card both read
OpenClaw's 30-minute HEARTBEAT_OK session, so under Claude Code they sat
at "waiting..." with an empty check-in. loadHeartbeat, the template
renderer and loadSystemHealth (on a runtime switch) now hide them unless
the runtime declares GATEWAY_RPC.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QpcVFD77MdQbNCsvrBSLDW
… runtimes

System Health scopes from _CM_RT_CAPS. A local install overrides that map
from /api/agents, but the hosted dashboard has only the static copy, and it
had drifted from what the adapters declare in clawmetry-pro 0.7.28:

- SUBAGENTS added for codex, copilot, cursor, deepagents, devin, goose,
  grok, n8n, nanoclaw, opencode, pi, picoclaw and qwen_code; removed for
  exo, which declares none.
- muse_code, openworker, qm and replit had no entry, so the sidebar showed
  them every tab, OpenClaw's included.

The Sub-Agents card also shows whenever a runtime has runs, so a stale map
never hides real children, and _cmLoadDeclaredCaps re-renders System Health
when the local override lands. New guard: every FREE|PAID runtime needs an
entry, and only OpenClaw/NemoClaw may claim GATEWAY_RPC, CRONS or CHANNELS.
COST flags are unchanged (computed per install for some adapters).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QpcVFD77MdQbNCsvrBSLDW
@vivekchand
vivekchand force-pushed the fix/system-health-runtime-scoped branch from aed367c to 5a5df63 Compare September 15, 2026 00:11

Copy link
Copy Markdown
Owner Author

Auto-rebase pushed; CI now running. If still not green in 10min, may need manual attention.


Generated by Claude Code

@8090-software-factory

Copy link
Copy Markdown

✅ Drift Bot (ClawMetry): no drift detected

Drift Bot analyzed the changed files against this project's blueprints and requirements and found no drift.

github-actions Bot pushed a commit that referenced this pull request Sep 15, 2026
…cope param and _cmMarkLoaded

Takes both sides of the app.js conflict:
- ?runtime= query param for scoped system health (this PR)
- _cmMarkLoaded('systemHealth') call (#5935, landed on main)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011Ai4CGH9XcWK1Jc3wy1a62
@8090-software-factory

Copy link
Copy Markdown

✅ Drift Bot (ClawMetry): no drift detected

Drift Bot analyzed the changed files against this project's blueprints and requirements and found no drift.

…his PR)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011Ai4CGH9XcWK1Jc3wy1a62
@8090-software-factory

Copy link
Copy Markdown

✅ Drift Bot (ClawMetry): no drift detected

Drift Bot analyzed the changed files against this project's blueprints and requirements and found no drift.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants