Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 5 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,6 +24,11 @@
- **What:** the Sessions item moved from above the Monitoring label to directly under Agents (Home, Agents, Sessions, Activity, Cost, Models, Context usage). It is still the page the dashboard opens on and keeps the default highlight; the tab id and deep links are unchanged.
- **Verified:** `tests/test_beginner_nav_phase_a.py` and `tests/test_trail_tab_template.py` pin the new order.

### Fixed: Home showed other runtimes' cards under the runtime switcher (2026-09-15)
- **Why:** on app.clawmetry.com with Codex selected, Home still listed OpenClaw's "Gateway :18789" in System Health (the hosted snapshot names it plain "Gateway", which the OpenClaw filter missed), Run Health drew a `claude_code` row, and 30-day activity, "How independent is your agent?", Anomaly Detection, Reliability and the session quality tile counted every runtime on the machine. The snapshot's Run Health slice had no Codex row at all: the newest 60 sessions on that machine were all Claude Code.
- **What:** each of those cards now asks for the selected runtime (`?runtime=` on `/api/activity-heatmap`, `/api/health-timeline`, `/api/autonomy`, `/api/evals/summary`, each echoing the runtime it counted) and hides rather than show a whole-machine answer under one runtime's name. Anomaly Detection keeps only that runtime's sessions; Reliability, which has no per-runtime form, is hidden under a runtime. The daemon adds each runtime's own recent sessions to the Run Health slice and ships `autonomyByRuntime` (both reused for a few minutes, so a sync cycle does not re-read them). A runtime switch reloads Home immediately instead of waiting for the next refresh.
- **Verified:** `tests/test_overview_runtime_scope.py` (16 tests, all red on the previous code). The hosted chip row, gateway pill and autonomy/live-activity overrides are fixed in clawmetry-cloud.

### Release: enterprise readiness, batch 3 (2026-09-15)
- **Carries:** #5950 (fleet install for shared hosts and virtual desktops, refs #5942), #5965 (LiteLLM gateway spend by team, person and key, refs #5940), #5957 (the dashboard's first load no longer times out its own requests, refs #5935), #5996 (the hosted Cost Optimizer shows evidence-backed experiments, refs #5934; hosted rendering lands with clawmetry-cloud#2450 after this pin) #5967 (SECURITY.md: the DPA is not available and the sub-processor list is published; documentation only) and #6007 (Compliance tab shell, refs clawmetry-pro#250; the evaluation ships in clawmetry-pro 0.7.29). Any other change merged before this release carries its own entry below. Their entries follow.

Expand Down
29 changes: 23 additions & 6 deletions clawmetry/local_store.py
Original file line number Diff line number Diff line change
Expand Up @@ -16502,6 +16502,7 @@ def query_eval_summary(
self,
*,
window_hours: int = 24,
runtime: str | None = None,
) -> dict[str, Any]:
"""Aggregate scores over the recent window. Drives
``/api/evals/summary``.
Expand All @@ -16510,24 +16511,40 @@ def query_eval_summary(
``total`` is sessions touched in the window (scored OR not);
``scored`` is the subset with a numeric eval_score. The ratio
``scored/total`` surfaces coverage on the overview tile.

``runtime`` keeps only that runtime's sessions, bucketed by id prefix
like :func:`_runtime_of_session_id` (NemoClaw sessions carry OpenClaw
ids). ``None`` / ``"all"`` counts every runtime.
"""
try:
from datetime import datetime, timedelta, timezone
cutoff = (datetime.now(timezone.utc) - timedelta(hours=int(window_hours))).isoformat()
except Exception:
cutoff = ""
rt_sql = ""
rt_params: list[Any] = []
rt = str(runtime or "").strip().lower()
if rt and rt != "all":
placeholders = ", ".join(["?"] * len(_NON_OPENCLAW_RUNTIME_PREFIXES))
rt_sql = (
f" AND (CASE WHEN split_part(session_id, ':', 1) IN ({placeholders})"
f" THEN split_part(session_id, ':', 1) ELSE 'openclaw' END) = ?"
)
rt_params = list(_NON_OPENCLAW_RUNTIME_PREFIXES) + [
"openclaw" if rt == "nemoclaw" else rt
]
# Two queries — one for totals (scored + un-scored), one for the
# quantile/avg over the scored subset. Keeps the SQL readable
# without a CTE that would have to handle NULLs in two places.
try:
total_row = self._fetch(
"""
f"""
SELECT COUNT(*) AS total,
COUNT(eval_score) AS scored
FROM sessions
WHERE (? = '' OR COALESCE(last_active_at, started_at, '') >= ?)
WHERE (? = '' OR COALESCE(last_active_at, started_at, '') >= ?){rt_sql}
""",
[cutoff, cutoff],
[cutoff, cutoff] + rt_params,
)
except Exception as e:
log.warning("local store: eval summary totals failed: %s", e)
Expand All @@ -16541,15 +16558,15 @@ def query_eval_summary(
if scored > 0:
try:
stats = self._fetch(
"""
f"""
SELECT AVG(eval_score) AS avg_score,
quantile_cont(eval_score, 0.5) AS p50,
quantile_cont(eval_score, 0.1) AS p10
FROM sessions
WHERE eval_score IS NOT NULL
AND (? = '' OR COALESCE(last_active_at, started_at, '') >= ?)
AND (? = '' OR COALESCE(last_active_at, started_at, '') >= ?){rt_sql}
""",
[cutoff, cutoff],
[cutoff, cutoff] + rt_params,
)
if stats:
avg = float(stats[0][0] or 0.0)
Expand Down
67 changes: 58 additions & 9 deletions clawmetry/static/js/app.js
Original file line number Diff line number Diff line change
Expand Up @@ -1425,6 +1425,18 @@ async function loadAnomalyPanel() {
if (!panel) return;
var anomalies = data.anomalies || [];
var baselines = data.baselines || {};
// Under a selected runtime keep only that runtime's sessions. Node-wide
// aggregate rows (session_key "__error_rate__") and the node-wide
// baselines are not about this runtime, so they are left out.
var _anRt = (typeof _cmRuntimeFilter === 'function') ? _cmRuntimeFilter() : 'all';
var _anScoped = !!_anRt && _anRt !== 'all';
if (_anScoped) {
anomalies = anomalies.filter(function(a){
var sk = String(a && a.session_key || '');
return !!sk && sk.indexOf('__') !== 0 && _cmRuntimeOf({session_id: sk}) === _anRt;
});
baselines = {};
}
var active = anomalies.filter(function(a){ return !a.acknowledged; });

// Badge
Expand All @@ -1451,7 +1463,7 @@ async function loadAnomalyPanel() {
if (baselines.baseline_cost_7d > 0) blHtml += '<span style="background:var(--bg-hover);padding:3px 8px;border-radius:6px;color:var(--text-secondary);">Avg cost: $' + Number(baselines.baseline_cost_7d).toFixed(4) + '/session</span>';
if (baselines.baseline_tokens_7d > 0) blHtml += '<span style="background:var(--bg-hover);padding:3px 8px;border-radius:6px;color:var(--text-secondary);">Avg tokens: ' + Math.round(baselines.baseline_tokens_7d).toLocaleString() + '/session</span>';
if (baselines.baseline_sessions_per_day_7d > 0) blHtml += '<span style="background:var(--bg-hover);padding:3px 8px;border-radius:6px;color:var(--text-secondary);">Sessions/day: ' + Number(baselines.baseline_sessions_per_day_7d).toFixed(1) + '</span>';
blEl.innerHTML = blHtml || '<span style="color:var(--text-muted);">Collecting baseline data...</span>';
blEl.innerHTML = blHtml || (_anScoped ? '' : '<span style="color:var(--text-muted);">Collecting baseline data...</span>');
}

// Anomaly list
Expand Down Expand Up @@ -2738,6 +2750,14 @@ async function loadReliabilityCard() {
var detEl = document.getElementById('reliability-detail-lt');
var iconEl = document.getElementById('reliability-icon-lt');
if (!dirEl) return;
// The trend is built from this machine's daemon heartbeats plus every
// runtime's error events; it has no per-runtime form, so it is not shown
// under a selected runtime.
var _relRt = (typeof _cmRuntimeFilter === 'function') ? _cmRuntimeFilter() : 'all';
var _relScoped = !!_relRt && _relRt !== 'all';
var _relCard = document.getElementById('reliability-card-lt');
if (_relCard) _relCard.style.display = _relScoped ? 'none' : '';
if (_relScoped) return;
try {
var d = await fetchJsonWithTimeout('/api/reliability', 5000);
d = d || {};
Expand Down Expand Up @@ -2793,9 +2813,12 @@ async function loadAutonomy() {
}

try {
// Check-in gaps are per runtime: Codex's cadence is not Claude Code's.
var _auRt = (typeof _cmRuntimeFilter === 'function') ? _cmRuntimeFilter() : 'all';
var _auUrl = '/api/autonomy' + ((_auRt && _auRt !== 'all') ? '?runtime=' + encodeURIComponent(_auRt) : '');
var d = await (typeof fetchJsonWithTimeout === 'function'
? fetchJsonWithTimeout('/api/autonomy', 5000)
: fetch('/api/autonomy').then(function(r){return r.json();}));
? fetchJsonWithTimeout(_auUrl, 5000)
: fetch(_auUrl).then(function(r){return r.json();}));

if (d.score == null) {
labelEl.textContent = t("app.just_getting_started", null, "Just getting started");
Expand Down Expand Up @@ -3963,13 +3986,23 @@ async function loadHealthTimeline() {
var card = document.getElementById('health-timeline-card');
var body = document.getElementById('health-timeline-body');
if (!card || !body) return;
var rt = (typeof _cmRuntimeFilter === 'function') ? _cmRuntimeFilter() : 'all';
var scoped = !!rt && rt !== 'all';
var data;
try {
var resp = await fetch('/api/health-timeline');
var resp = await fetch('/api/health-timeline' + (scoped ? '?runtime=' + encodeURIComponent(rt) : ''));
if (!resp.ok) { card.style.display = 'none'; return; }
data = await resp.json();
} catch (e) { card.style.display = 'none'; return; }
var runtimes = (data && data.runtimes) || [];
// Under a selected runtime only its own row renders: an older server and
// the hosted snapshot both answer with every runtime they know. NemoClaw
// runs the OpenClaw adapter, so its sessions bucket as openclaw.
if (scoped) {
runtimes = runtimes.filter(function (r) {
return r && (r.runtime === rt || (rt === 'nemoclaw' && r.runtime === 'openclaw'));
});
}
if (!runtimes.length || !runtimes.some(function(r){ return (r.dots||[]).length; })) {
card.style.display = 'none';
return;
Expand Down Expand Up @@ -6805,8 +6838,11 @@ async function loadEvalSummary() {
function setTitleCheck(show) { if (checkEl) checkEl.style.display = show ? '' : 'none'; }
if (!avgEl) return;
try {
var data = await fetch('/api/evals/summary?window=24h').then(function(r){return r.json();}).catch(function(){return null;});
if (!data || typeof data.scored !== 'number') {
var _evRt = (typeof _cmRuntimeFilter === 'function') ? _cmRuntimeFilter() : 'all';
var _evQ = (_evRt && _evRt !== 'all') ? '&runtime=' + encodeURIComponent(_evRt) : '';
var data = await fetch('/api/evals/summary?window=24h' + _evQ).then(function(r){return r.json();}).catch(function(){return null;});
// A server that ignores ?runtime answers for the whole node.
if (!data || typeof data.scored !== 'number' || (_evQ && data.runtime !== _evRt)) {
setTitleCheck(false);
avgEl.textContent = '--';
if (covEl) covEl.textContent = '';
Expand Down Expand Up @@ -12883,6 +12919,9 @@ function _cmApplyRuntimeSelection(val) {
// Swap the Flow + Overview diagram to the selected runtime's topology.
try { if (typeof _applyRuntimeFlowDiagram === 'function') _applyRuntimeFlowDiagram(val); } catch (e) {}
// Reload the current tab so any runtime-aware view re-filters in place.
// loadAll coalesces calls 2 s apart; a switch must not be swallowed by that,
// or the Overview keeps the previous runtime's cards until the next refresh.
try { _loadAllLastFinishedMs = 0; } catch (e) {}
if (typeof switchTab === 'function' && _cmCurrentTab) switchTab(_cmCurrentTab);
// System Health refreshes on a 30s timer and is not part of loadAll, so
// re-scope it now or the previous runtime's checks linger.
Expand Down Expand Up @@ -17173,7 +17212,13 @@ async function loadSystemHealth() {
}
}
var services = Array.isArray(d.services) ? d.services : [];
if (!isOc) services = services.filter(function (s) { return !/openclaw/i.test(String(s && s.name || '')); });
// OpenClaw's gateway arrives as "OpenClaw Gateway" locally and as a bare
// "Gateway" from the hosted snapshot; both, and anything on its port,
// belong to OpenClaw alone.
if (!isOc) services = services.filter(function (s) {
var name = String(s && s.name || '').trim();
return !(/openclaw/i.test(name) || /^gateway$/i.test(name) || Number(s && s.port) === 18789);
});
var channels = (scope.has('CHANNELS') && Array.isArray(d.channels)) ? d.channels : [];
var disks = Array.isArray(d.disks) ? d.disks : [];
var crons = (d.crons && typeof d.crons === 'object') ? d.crons : {enabled: 0, ok24h: 0, failed: []};
Expand Down Expand Up @@ -17884,9 +17929,13 @@ async function loadActivityHeatmap() {
var grid = document.getElementById('activity-heatmap-grid');
if (!card || !grid) return;
var data;
try { data = await fetchJsonWithTimeout('/api/activity-heatmap', 5000); } catch(e) { return; }
var rt = (typeof _cmRuntimeFilter === 'function') ? _cmRuntimeFilter() : 'all';
var q = (rt && rt !== 'all') ? ('?runtime=' + encodeURIComponent(rt)) : '';
try { data = await fetchJsonWithTimeout('/api/activity-heatmap' + q, 5000); } catch(e) { card.style.display = 'none'; return; }
var days = (data && data.days) || [];
if (!days.length) return;
// A server that ignores ?runtime answers for the whole node; hide the card
// rather than draw every runtime's days under this one's name.
if (!days.length || (q && data.runtime !== rt)) { card.style.display = 'none'; return; }
var maxSessions = Math.max.apply(null, days.map(function(d){ return d.sessions || 0; }));
var shades = ['#12122a','#1a3a2a','#2a6a3a','#4a9a2a','#6adb3a'];
var html = '';
Expand Down
Loading
Loading