[RELEASE] Enterprise readiness batch 3: fleet install, LiteLLM gateway, cold-load fix, hosted Cost Optimizer data (carries #5950 #5965 #5957 #5996 #5967) - #6001
Conversation
Publishes #5950 (fleet install for shared hosts and virtual desktops) and #5965 (LiteLLM gateway spend by team, person and key), with their CHANGELOG entries. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9
|
Both merged onto main after 0.12.879 and before this release branch was cut, so the release already contains their code. #5957 gets its CHANGELOG entry; #5967 is documentation only. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9
|
| - **Shared environment is not writable by desktop users:** the Ansible playbook installs under `umask 022` and enforces root ownership and `u=rwX,go=rX` on `/opt/clawmetry-fleet`; the Intune script sets an explicit ACL on its install directory. Both CI jobs fail if a desktop user can write the environment. | ||
| - **No self-update from an administrator-owned install:** auto-update skips a virtual environment the running user cannot write, instead of exiting to retry an upgrade that cannot succeed (on Windows that retry exited every signed-in user's collector every few minutes). `service status` reports `Auto-update: off` with the reason. | ||
| - **Honest limits:** a Windows collector stops at logoff and resumes at the next sign-in. Not validated on a real multi-session virtual desktop host. Scoped enrollment keys and capped offline buffering are not in this change. | ||
| - **Verified:** `tests/test_fleet_install.py`; `.github/workflows/fleet-install-test.yml` runs the Ansible playbook for two users on Ubuntu and the Intune script on Windows. | ||
| - **Refs:** #5942. | ||
|
|
||
| ### Added: LiteLLM gateway spend by team, person and key, kept apart from agent costs (2026-09-14) | ||
| - **Why:** teams that route model calls through a LiteLLM proxy already have per-request team, user, key, model, tokens and spend inside LiteLLM, and were exporting it to spreadsheets. Pointing LiteLLM's OpenTelemetry callback at ClawMetry did something worse than nothing: every proxied request became its own "session", was re-priced from our table with the team and user dropped, and was added to the totals the agents already report, so a call an agent made through the proxy was counted twice. |
There was a problem hiding this comment.
The Fleet Install blueprint contains only template placeholder text, but the CHANGELOG documents a comprehensive released implementation including clawmetry service install/status/uninstall commands, Windows all-users task registration, Linux linger handling, permission enforcement, Ansible and Intune deployment recipes, and auto-update guards. The blueprint should document the actual technical architecture and component composition.
|
|
||
| ### Fixed: the dashboard's first load timed out its own requests (2026-09-14) | ||
| - **Why:** on a cold start the console showed `Initial load failed timeout`, `System health load failed timeout` and `loadCrons failed timeout`, and the tiles those requests feed rendered empty, which reads as missing data (#5935). Measured in a headless browser against a scratch install with a seeded store: one page load sent 103 API requests in its first 10 s against the browser's six connections per origin, and those requests spent a combined 34-66 s waiting in the browser's own queue while the server answered most of them in milliseconds. On a machine with OpenClaw installed, `/api/agents` and `/api/inventory` each ran `openclaw doctor --json` synchronously, holding two connections for 7-15 s. | ||
| - **What:** startup loads Overview's widgets only when Overview is the landing screen, and opening Overview loads system health and tasks at once. No Crons / Memory prefetch; the Flow tool prefetch waits for Flow or Overview. Duplicates removed, and every `/api/overview` caller shares one in-flight request through one helper with one 15 s budget. OpenClaw doctor findings are served stale-while-revalidate (`CLAWMETRY_OPENCLAW_DOCTOR_TTL`, default 300 s, `0` restores a run on every read); before the first run finishes they are absent, never "no findings". | ||
| - **Honest states:** slow usage no longer draws `$0.00` and `0` tokens into the Overview tiles, including a runtime-scoped view; they stay on "still loading" (or keep the last real answer) until the refresh retries. System health and Crons failures read as sentences instead of `Failed to load: timeout`. | ||
| - **Measured after:** same scenario, 42 requests in the first 10 s (was 103), 2.1 s of browser queue time in the first 12 s (was 34.2 s), `/api/agents` 0.18 s (was 6.9 s), zero console errors, and Overview opens with real tiles. | ||
| - **Verified:** `tests/test_cold_load_boot_js.{js,py}` (behaviour checks against shipped app.js, including a guard that discovers every `.js` and `.html` file and fails on any direct `/api/overview` request outside the shared helper) and `tests/test_openclaw_doctor_cache.py`, all red against the previous code. | ||
| - **Not changed:** the hosted dashboard serves these screens from the encrypted snapshot and was not affected. With the daemon writing continuously, the Sessions list and `/api/inventory` are still slow on the server; that is store contention, not startup fan-out. | ||
| - **Refs:** #5935. | ||
|
|
||
| ### Added: fleet install for shared hosts and virtual desktops (2026-09-14) | ||
| - **Why:** on a shared Linux host a `systemd --user` collector stops when the user logs out unless linger is on, and nothing said so. On a multi-session Windows host an administrator had no way to register the collector for every user who signs in. There were no Intune or Ansible recipes. |
There was a problem hiding this comment.
The Dashboard first load blueprint contains only template placeholder text, but the CHANGELOG documents a comprehensive released implementation with lazy loading, shared request helpers, stale-while-revalidate caching for OpenClaw doctor diagnostics, OTLP/JSON decoder fixes, and team label escaping. The blueprint should document the architectural changes, component composition, and system contracts for this feature.
#5996 merged onto main before this release, so the release contains its snapshot slice. Adds its CHANGELOG entry; the hosted rendering ships with clawmetry-cloud#2450 once this version is pinned. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9
|
| - **Refs:** #5935. | ||
|
|
||
| ### Added: fleet install for shared hosts and virtual desktops (2026-09-14) | ||
| - **Why:** on a shared Linux host a `systemd --user` collector stops when the user logs out unless linger is on, and nothing said so. On a multi-session Windows host an administrator had no way to register the collector for every user who signs in. There were no Intune or Ansible recipes. | ||
| - **What:** `clawmetry service install | status | uninstall`. One collector per user, running as that user. Linux requests linger and, when refused, prints the administrator command and exits 1. Windows `--all-users` (elevated) registers one logon task for the built-in Users group, least privilege, one instance per signed-in user. Fleet install restricts `~/.clawmetry` to its owner. `status` says whether collection survives logoff and why. | ||
| - **Recipes:** `deploy/fleet/intune/Install-ClawMetry.ps1` and `deploy/fleet/ansible/clawmetry.yml`, neither holding a credential; `deploy/fleet/README.md` covers pin, update, rollback and uninstall. | ||
| - **Shared environment is not writable by desktop users:** the Ansible playbook installs under `umask 022` and enforces root ownership and `u=rwX,go=rX` on `/opt/clawmetry-fleet`; the Intune script sets an explicit ACL on its install directory. Both CI jobs fail if a desktop user can write the environment. | ||
| - **No self-update from an administrator-owned install:** auto-update skips a virtual environment the running user cannot write, instead of exiting to retry an upgrade that cannot succeed (on Windows that retry exited every signed-in user's collector every few minutes). `service status` reports `Auto-update: off` with the reason. |
There was a problem hiding this comment.
The Fleet Install blueprint contains only template placeholder text, but the CHANGELOG documents a comprehensive released implementation including clawmetry service install/status/uninstall commands, Windows all-users task registration, Linux linger handling, permission enforcement, Ansible and Intune deployment recipes, and auto-update guards. The blueprint should document the actual technical architecture and component composition.
|
|
||
| ### Fixed: the hosted Cost Optimizer showed no experiments (2026-09-15) | ||
| - **Why:** after #5951 the renderer hides recommendations that cite no evidence. The cloud interceptor sent only hardcoded ones and a "40-70%" claim, so app.clawmetry.com showed nothing (#5934). | ||
| - **What:** the daemon builds the optimizer's data slice with the local route's own rules (`clawmetry/cost_optimizer_snapshot.py`, shared `advice_fields` / `tokens_recorded_or_none`) and ships it as `costOptimizer` in the E2E-encrypted snapshot; clawmetry-cloud#2450 decrypts it in the browser. Hosted figures name "the connected computer", an empty store reads unknown, not $0, and llmfit and Ollama details stay on the computer. A daemon that has not updated shows an "update ClawMetry" note instead of recommendations. | ||
| - **Verified:** `tests/test_cost_optimizer_snapshot_slice.py` (5 tests, MOAT verifier job; 4 red on main): the slice equals the local route's advice, figures and bases for the same store. AC-OBS-CEA-023.9 mirrored. | ||
| - **Refs:** #5934. | ||
|
|
||
| ### Fixed: the dashboard's first load timed out its own requests (2026-09-14) | ||
| - **Why:** on a cold start the console showed `Initial load failed timeout`, `System health load failed timeout` and `loadCrons failed timeout`, and the tiles those requests feed rendered empty, which reads as missing data (#5935). Measured in a headless browser against a scratch install with a seeded store: one page load sent 103 API requests in its first 10 s against the browser's six connections per origin, and those requests spent a combined 34-66 s waiting in the browser's own queue while the server answered most of them in milliseconds. On a machine with OpenClaw installed, `/api/agents` and `/api/inventory` each ran `openclaw doctor --json` synchronously, holding two connections for 7-15 s. | ||
| - **What:** startup loads Overview's widgets only when Overview is the landing screen, and opening Overview loads system health and tasks at once. No Crons / Memory prefetch; the Flow tool prefetch waits for Flow or Overview. Duplicates removed, and every `/api/overview` caller shares one in-flight request through one helper with one 15 s budget. OpenClaw doctor findings are served stale-while-revalidate (`CLAWMETRY_OPENCLAW_DOCTOR_TTL`, default 300 s, `0` restores a run on every read); before the first run finishes they are absent, never "no findings". | ||
| - **Honest states:** slow usage no longer draws `$0.00` and `0` tokens into the Overview tiles, including a runtime-scoped view; they stay on "still loading" (or keep the last real answer) until the refresh retries. System health and Crons failures read as sentences instead of `Failed to load: timeout`. | ||
| - **Measured after:** same scenario, 42 requests in the first 10 s (was 103), 2.1 s of browser queue time in the first 12 s (was 34.2 s), `/api/agents` 0.18 s (was 6.9 s), zero console errors, and Overview opens with real tiles. |
There was a problem hiding this comment.
The Dashboard first load blueprint contains only template placeholder text, but the CHANGELOG documents a comprehensive released implementation with lazy loading, shared request helpers, stale-while-revalidate caching for OpenClaw doctor diagnostics, OTLP/JSON decoder fixes, and team label escaping. The blueprint should document the architectural changes, component composition, and system contracts for this feature.
Main's module map fell behind the source tree after the merges carried by this release, so Syntax & Lint failed on every open PR. Regenerated with scripts/gen_module_map.py; no source change. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9
|
| - **Refs:** #5935. | ||
|
|
||
| ### Added: fleet install for shared hosts and virtual desktops (2026-09-14) | ||
| - **Why:** on a shared Linux host a `systemd --user` collector stops when the user logs out unless linger is on, and nothing said so. On a multi-session Windows host an administrator had no way to register the collector for every user who signs in. There were no Intune or Ansible recipes. | ||
| - **What:** `clawmetry service install | status | uninstall`. One collector per user, running as that user. Linux requests linger and, when refused, prints the administrator command and exits 1. Windows `--all-users` (elevated) registers one logon task for the built-in Users group, least privilege, one instance per signed-in user. Fleet install restricts `~/.clawmetry` to its owner. `status` says whether collection survives logoff and why. | ||
| - **Recipes:** `deploy/fleet/intune/Install-ClawMetry.ps1` and `deploy/fleet/ansible/clawmetry.yml`, neither holding a credential; `deploy/fleet/README.md` covers pin, update, rollback and uninstall. | ||
| - **Shared environment is not writable by desktop users:** the Ansible playbook installs under `umask 022` and enforces root ownership and `u=rwX,go=rX` on `/opt/clawmetry-fleet`; the Intune script sets an explicit ACL on its install directory. Both CI jobs fail if a desktop user can write the environment. | ||
| - **No self-update from an administrator-owned install:** auto-update skips a virtual environment the running user cannot write, instead of exiting to retry an upgrade that cannot succeed (on Windows that retry exited every signed-in user's collector every few minutes). `service status` reports `Auto-update: off` with the reason. |
There was a problem hiding this comment.
The Fleet Install blueprint contains only template placeholder text, but the CHANGELOG documents a comprehensive released implementation including clawmetry service install/status/uninstall commands, Windows task registration, Linux linger handling, Ansible and Intune deployment recipes, and auto-update guards. The blueprint should be updated with the actual technical architecture and implementation details.
|
|
||
| ### Fixed: the hosted Cost Optimizer showed no experiments (2026-09-15) | ||
| - **Why:** after #5951 the renderer hides recommendations that cite no evidence. The cloud interceptor sent only hardcoded ones and a "40-70%" claim, so app.clawmetry.com showed nothing (#5934). | ||
| - **What:** the daemon builds the optimizer's data slice with the local route's own rules (`clawmetry/cost_optimizer_snapshot.py`, shared `advice_fields` / `tokens_recorded_or_none`) and ships it as `costOptimizer` in the E2E-encrypted snapshot; clawmetry-cloud#2450 decrypts it in the browser. Hosted figures name "the connected computer", an empty store reads unknown, not $0, and llmfit and Ollama details stay on the computer. A daemon that has not updated shows an "update ClawMetry" note instead of recommendations. | ||
| - **Verified:** `tests/test_cost_optimizer_snapshot_slice.py` (5 tests, MOAT verifier job; 4 red on main): the slice equals the local route's advice, figures and bases for the same store. AC-OBS-CEA-023.9 mirrored. | ||
| - **Refs:** #5934. | ||
|
|
||
| ### Fixed: the dashboard's first load timed out its own requests (2026-09-14) | ||
| - **Why:** on a cold start the console showed `Initial load failed timeout`, `System health load failed timeout` and `loadCrons failed timeout`, and the tiles those requests feed rendered empty, which reads as missing data (#5935). Measured in a headless browser against a scratch install with a seeded store: one page load sent 103 API requests in its first 10 s against the browser's six connections per origin, and those requests spent a combined 34-66 s waiting in the browser's own queue while the server answered most of them in milliseconds. On a machine with OpenClaw installed, `/api/agents` and `/api/inventory` each ran `openclaw doctor --json` synchronously, holding two connections for 7-15 s. | ||
| - **What:** startup loads Overview's widgets only when Overview is the landing screen, and opening Overview loads system health and tasks at once. No Crons / Memory prefetch; the Flow tool prefetch waits for Flow or Overview. Duplicates removed, and every `/api/overview` caller shares one in-flight request through one helper with one 15 s budget. OpenClaw doctor findings are served stale-while-revalidate (`CLAWMETRY_OPENCLAW_DOCTOR_TTL`, default 300 s, `0` restores a run on every read); before the first run finishes they are absent, never "no findings". | ||
| - **Honest states:** slow usage no longer draws `$0.00` and `0` tokens into the Overview tiles, including a runtime-scoped view; they stay on "still loading" (or keep the last real answer) until the refresh retries. System health and Crons failures read as sentences instead of `Failed to load: timeout`. | ||
| - **Measured after:** same scenario, 42 requests in the first 10 s (was 103), 2.1 s of browser queue time in the first 12 s (was 34.2 s), `/api/agents` 0.18 s (was 6.9 s), zero console errors, and Overview opens with real tiles. |
There was a problem hiding this comment.
The Dashboard first load blueprint contains only template placeholder text, but the CHANGELOG documents a comprehensive released implementation with lazy loading, shared request helpers, stale-while-revalidate caching for OpenClaw doctor diagnostics, dashboard metric optimization, and API performance improvements. The blueprint should be updated with the technical architecture and component composition.
Main's module map is behind the source tree, which failed Syntax & Lint on this PR. Same one-line regeneration as release PR #6001. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9
…d blueprints Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9
✅ Drift Bot (ClawMetry): no drift detectedDrift Bot analyzed the changed files against this project's blueprints and requirements and found no drift. |
|
Human handoff needed to complete this release. All CI is green on this PR. When you are ready to publish to PyPI:
This session cannot trigger PyPI publish or deploy Cloud Run — those require production secrets not available in the sandbox. Generated by Claude Code |
The Compliance tab shell merged onto main before this release, so the release contains it. Adds its CHANGELOG entry. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9
✅ Drift Bot (ClawMetry): no drift detectedDrift Bot analyzed the changed files against this project's blueprints and requirements and found no drift. |
This PR triggers the publish. It changes only
CHANGELOG.md, adding entries for two merged PRs whose fixers moved their entries out to stop conflicts between parallel PRs.What is being released
clawmetry service install|status|uninstall: per-user collectors on shared Linux hosts and multi-session Windows, Intune and Ansible recipes (refs Fleet install: Windows service, Linux linger and silent install for VDI / Azure Virtual Desktop #5942). This unblocks landing docs: Slack emoji-react approvals design proposal #836 (rollout guide).If more OSS PRs merge before this one, the coordinator adds their CHANGELOG entries here before it merges.
After this publishes
clawmetry serviceincli.pyandgateway_litellmin the package.No-PRD: CHANGELOG entry only; the product records are cited on each carried feature PR.
🤖 Generated with Claude Code
https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9