Repository navigation
[RELEASE] Enterprise readiness batch 3: fleet install, LiteLLM gateway, cold-load fix, hosted Cost Optimizer data (carries #5950 #5965 #5957 #5996 #5967) #6001
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Changes from all commits
4385787
0097bbe
5fcb6f7
fffca8c
cf6bd3d
b93d612
ab9c7dd
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -1,5 +1,48 @@ | ||
| ## Unreleased | ||
|
|
||
| ### Release: enterprise readiness, batch 3 (2026-09-15) | ||
| - **Carries:** #5950 (fleet install for shared hosts and virtual desktops, refs #5942), #5965 (LiteLLM gateway spend by team, person and key, refs #5940), #5957 (the dashboard's first load no longer times out its own requests, refs #5935), #5996 (the hosted Cost Optimizer shows evidence-backed experiments, refs #5934; hosted rendering lands with clawmetry-cloud#2450 after this pin) #5967 (SECURITY.md: the DPA is not available and the sub-processor list is published; documentation only) and #6007 (Compliance tab shell, refs clawmetry-pro#250; the evaluation ships in clawmetry-pro 0.7.29). Any other change merged before this release carries its own entry below. Their entries follow. | ||
|
|
||
| ### Added: a Compliance tab with an honest locked state (2026-09-15) | ||
| - **What:** with the Compliance Pack, the Compliance tab (Advanced section of the navigation) shows every NIST AI RMF, SOC 2, OWASP LLM 2026, OWASP Agentic 2026 and MITRE ATLAS control with its evidence state for a date range (exercised, configured, gap or unknown; never "effective" without evidence), the MITRE ATLAS replay stages linked to it, a printable evidence report (`GET /api/compliance/report`, script-free HTML) and the evidence bundle. Without the pack the tab shows an upgrade prompt, and on the hosted dashboard it explains that evaluation runs on the agent's machine and makes no requests. | ||
| - **Refused, not shown:** a report that claims a state without the evidence behind it is rejected with the reasons listed. | ||
| - **Refs:** clawmetry-pro#250. | ||
|
|
||
| ### Fixed: the hosted Cost Optimizer showed no experiments (2026-09-15) | ||
| - **Why:** after #5951 the renderer hides recommendations that cite no evidence. The cloud interceptor sent only hardcoded ones and a "40-70%" claim, so app.clawmetry.com showed nothing (#5934). | ||
| - **What:** the daemon builds the optimizer's data slice with the local route's own rules (`clawmetry/cost_optimizer_snapshot.py`, shared `advice_fields` / `tokens_recorded_or_none`) and ships it as `costOptimizer` in the E2E-encrypted snapshot; clawmetry-cloud#2450 decrypts it in the browser. Hosted figures name "the connected computer", an empty store reads unknown, not $0, and llmfit and Ollama details stay on the computer. A daemon that has not updated shows an "update ClawMetry" note instead of recommendations. | ||
| - **Verified:** `tests/test_cost_optimizer_snapshot_slice.py` (5 tests, MOAT verifier job; 4 red on main): the slice equals the local route's advice, figures and bases for the same store. AC-OBS-CEA-023.9 mirrored. | ||
| - **Refs:** #5934. | ||
|
|
||
| ### Fixed: the dashboard's first load timed out its own requests (2026-09-14) | ||
| - **Why:** on a cold start the console showed `Initial load failed timeout`, `System health load failed timeout` and `loadCrons failed timeout`, and the tiles those requests feed rendered empty, which reads as missing data (#5935). Measured in a headless browser against a scratch install with a seeded store: one page load sent 103 API requests in its first 10 s against the browser's six connections per origin, and those requests spent a combined 34-66 s waiting in the browser's own queue while the server answered most of them in milliseconds. On a machine with OpenClaw installed, `/api/agents` and `/api/inventory` each ran `openclaw doctor --json` synchronously, holding two connections for 7-15 s. | ||
| - **What:** startup loads Overview's widgets only when Overview is the landing screen, and opening Overview loads system health and tasks at once. No Crons / Memory prefetch; the Flow tool prefetch waits for Flow or Overview. Duplicates removed, and every `/api/overview` caller shares one in-flight request through one helper with one 15 s budget. OpenClaw doctor findings are served stale-while-revalidate (`CLAWMETRY_OPENCLAW_DOCTOR_TTL`, default 300 s, `0` restores a run on every read); before the first run finishes they are absent, never "no findings". | ||
| - **Honest states:** slow usage no longer draws `$0.00` and `0` tokens into the Overview tiles, including a runtime-scoped view; they stay on "still loading" (or keep the last real answer) until the refresh retries. System health and Crons failures read as sentences instead of `Failed to load: timeout`. | ||
| - **Measured after:** same scenario, 42 requests in the first 10 s (was 103), 2.1 s of browser queue time in the first 12 s (was 34.2 s), `/api/agents` 0.18 s (was 6.9 s), zero console errors, and Overview opens with real tiles. | ||
|
Comment on lines
+10
to
+21
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. The Dashboard first load blueprint contains only template placeholder text, but the CHANGELOG documents a comprehensive released implementation with lazy loading, shared request helpers, stale-while-revalidate caching for OpenClaw doctor diagnostics, dashboard metric optimization, and API performance improvements. The blueprint should be updated with the technical architecture and component composition. |
||
| - **Verified:** `tests/test_cold_load_boot_js.{js,py}` (behaviour checks against shipped app.js, including a guard that discovers every `.js` and `.html` file and fails on any direct `/api/overview` request outside the shared helper) and `tests/test_openclaw_doctor_cache.py`, all red against the previous code. | ||
| - **Not changed:** the hosted dashboard serves these screens from the encrypted snapshot and was not affected. With the daemon writing continuously, the Sessions list and `/api/inventory` are still slow on the server; that is store contention, not startup fan-out. | ||
| - **Refs:** #5935. | ||
|
|
||
| ### Added: fleet install for shared hosts and virtual desktops (2026-09-14) | ||
| - **Why:** on a shared Linux host a `systemd --user` collector stops when the user logs out unless linger is on, and nothing said so. On a multi-session Windows host an administrator had no way to register the collector for every user who signs in. There were no Intune or Ansible recipes. | ||
|
Comment on lines
+16
to
+27
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. The Dashboard first load blueprint contains only template placeholder text, but the CHANGELOG documents a comprehensive released implementation with lazy loading, shared request helpers, stale-while-revalidate caching for OpenClaw doctor diagnostics, OTLP/JSON decoder fixes, and team label escaping. The blueprint should document the architectural changes, component composition, and system contracts for this feature. |
||
| - **What:** `clawmetry service install | status | uninstall`. One collector per user, running as that user. Linux requests linger and, when refused, prints the administrator command and exits 1. Windows `--all-users` (elevated) registers one logon task for the built-in Users group, least privilege, one instance per signed-in user. Fleet install restricts `~/.clawmetry` to its owner. `status` says whether collection survives logoff and why. | ||
| - **Recipes:** `deploy/fleet/intune/Install-ClawMetry.ps1` and `deploy/fleet/ansible/clawmetry.yml`, neither holding a credential; `deploy/fleet/README.md` covers pin, update, rollback and uninstall. | ||
| - **Shared environment is not writable by desktop users:** the Ansible playbook installs under `umask 022` and enforces root ownership and `u=rwX,go=rX` on `/opt/clawmetry-fleet`; the Intune script sets an explicit ACL on its install directory. Both CI jobs fail if a desktop user can write the environment. | ||
| - **No self-update from an administrator-owned install:** auto-update skips a virtual environment the running user cannot write, instead of exiting to retry an upgrade that cannot succeed (on Windows that retry exited every signed-in user's collector every few minutes). `service status` reports `Auto-update: off` with the reason. | ||
|
Comment on lines
+24
to
+31
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. The Fleet Install blueprint contains only template placeholder text, but the CHANGELOG documents a comprehensive released implementation including clawmetry service install/status/uninstall commands, Windows all-users task registration, Linux linger handling, permission enforcement, Ansible and Intune deployment recipes, and auto-update guards. The blueprint should document the actual technical architecture and component composition.
Comment on lines
+24
to
+31
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. The Fleet Install blueprint contains only template placeholder text, but the CHANGELOG documents a comprehensive released implementation including clawmetry service install/status/uninstall commands, Windows task registration, Linux linger handling, Ansible and Intune deployment recipes, and auto-update guards. The blueprint should be updated with the actual technical architecture and implementation details. |
||
| - **Honest limits:** a Windows collector stops at logoff and resumes at the next sign-in. Not validated on a real multi-session virtual desktop host. Scoped enrollment keys and capped offline buffering are not in this change. | ||
| - **Verified:** `tests/test_fleet_install.py`; `.github/workflows/fleet-install-test.yml` runs the Ansible playbook for two users on Ubuntu and the Intune script on Windows. | ||
| - **Refs:** #5942. | ||
|
|
||
| ### Added: LiteLLM gateway spend by team, person and key, kept apart from agent costs (2026-09-14) | ||
| - **Why:** teams that route model calls through a LiteLLM proxy already have per-request team, user, key, model, tokens and spend inside LiteLLM, and were exporting it to spreadsheets. Pointing LiteLLM's OpenTelemetry callback at ClawMetry did something worse than nothing: every proxied request became its own "session", was re-priced from our table with the team and user dropped, and was added to the totals the agents already report, so a call an agent made through the proxy was counted twice. | ||
|
Comment on lines
+30
to
+37
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. The Fleet Install blueprint contains only template placeholder text, but the CHANGELOG documents a comprehensive released implementation including clawmetry service install/status/uninstall commands, Windows all-users task registration, Linux linger handling, permission enforcement, Ansible and Intune deployment recipes, and auto-update guards. The blueprint should document the actual technical architecture and component composition. |
||
| - **What:** LiteLLM's telemetry is recognised by its instrumentation scope and resource marker. Each proxied request becomes one gateway usage record: the team, user and key LiteLLM authenticated, the model asked for and the deployment called, tokens, streamed or not, success, trace and response ids, and LiteLLM's own cost. The Usage tab's *Cost by Team* card gains **Through your LiteLLM gateway** (`/api/usage/by-team` `gateway`): spend, requests and failures per team, each person and key under it. Gateway spend carries the shared cost basis (`published_rate`, `priced_from: vendor_reported`) with `cost_source: gateway_reported`. It is never added to agent totals, creates no sessions, and the proxy never appears as an agent. Setup: `docs/LITELLM.md`. | ||
| - **Found in a real proxy run, and handled:** a caller can set `metadata.team_id` in the request body and LiteLLM copies it onto the span, so attribution reads only the authenticated `user_api_key_*` identity. A response LiteLLM served from its own cache is counted but not charged. With the interceptor loaded inside the proxy process, the proxy's own provider calls (including Azure OpenAI) are recorded without a cost, so one request is not priced three times. | ||
| - **Not stored:** key hashes, project ids and org aliases from the proxy's metadata are no longer kept in stored details. | ||
| - **Also fixed:** the OTLP/JSON decoder dropped the instrumentation scope, and the *Cost by Team* card wrote team labels into the page unescaped. | ||
| - **Behaviour change:** if you already exported LiteLLM traces to ClawMetry, new requests stop appearing as per-request sessions and stop counting towards runtime totals. Earlier rows are left as stored. | ||
| - **Verified:** `tests/test_litellm_gateway_ingest.py`, driven by a capture of a real LiteLLM 1.83.7 proxy with two Postgres-backed teams, and `.github/workflows/litellm-gateway.yml`, which runs that proxy against a real dashboard and reconciles each team's spend with LiteLLM's `/spend/logs`. | ||
| - **Refs:** #5940. | ||
|
|
||
| ### Fixed: hosted Guard said no sessions were running, and hosted Signals said no sessions matched (2026-09-14) | ||
| - **Why:** on app.clawmetry.com the Guard tab read "No sessions running right now" while the same node's local dashboard listed 39 running sessions, and a Signals "Sessions" click read "No sessions matched in this window" beside a rate that counted two. The cloud container has no store: `/api/guard/sessions` ran there against nothing, and the signals session list was never in the snapshot. Both renderers then read an unreadable list as an empty one. | ||
| - **What:** the daemon now builds two more snapshot slices on its own store handle. `guardSessions` is the exact `/api/guard/sessions` body (the route and the daemon share one builder, `routes.guard.build_guard_sessions_body`, so the two tabs cannot list different sessions) plus `generated_at`. `signalSessions` holds the drill-down lists per runtime, window and signal (`behaviour_signals.build_session_slice`), built per runtime so a quiet runtime is not starved, and only for signals that matched. Sessions and match counts only, never the phrases. The Guard and Signals renderers show the response's `reason` when it says `available: false`, instead of claiming nothing is there. | ||
|
|
||
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
The Dashboard first load blueprint contains only template placeholder text, but the CHANGELOG documents a comprehensive released implementation with lazy loading, shared request helpers, stale-while-revalidate caching for OpenClaw doctor diagnostics, OTLP/JSON decoder fixes, and team label escaping. The blueprint should document the architectural changes, component composition, and system contracts for this feature.