Skip to content

Cost basis labels: published-rate usage value, subscription value and metered usage kept apart - #5975

Merged
vivekchand merged 4 commits into
mainfrom
feat/cost-basis-labels-5937
Sep 14, 2026
Merged

vivekchand merged 4 commits into
mainfrom
feat/cost-basis-labels-5937

Conversation

@vivekchand

@vivekchand vivekchand commented Sep 14, 2026 •

Copy link
Copy Markdown
Owner

Refs #5937

Factory requirement (REQ-OBS-CEA-025, written before the code): https://factory.8090.ai/project/b415065f-ab2f-4f53-8864-0c009fd098cb/requirements/950d4687-45cc-45a8-9d58-0c0d82fd5d9b
Child of Cost and Efficiency Analytics. Seven criteria are mirrored into docs/acceptance_criteria.json, and each is cited by a test.

Merge order: branched off main, not stacked. #5959 (price book, #5936) is a conceptual dependency only. This PR defines the "contract rate" label and the evidence it needs (a rate version), and nothing produces that label yet. The two PRs share no files except CHANGELOG.md, docs/acceptance_criteria.json and docs/ac_coverage_baseline.json, which are additive; on a conflict, keep both sides and regenerate the baseline. No new HTTP route, so no cloud_route_policy entry or cloud PR is needed.

Review fixes (79c7121)

The review found AC-OBS-CEA-025.1 marked met while two live screens still printed unlabelled costs.

  • Usage tab session cost chart and table. loadUsage() kept cbd.top10 and dropped the provenance. The top10[].cost_usd entry now travels with the rows. The canvas bars get a caption ("Bar values and the Cost column: published rates"), and the Cost column heading gets the badge. An unknown cost reads "not available", not $0.0000.

  • Flow brain panel per-call list. Each call now carries a numeric cost_usd: null when the call had tokens but no price, so it reads "not available". The list is labelled once, under calls[].cost_usd, with a "Cost per call: published rates" heading. The colour thresholds are now checked largest first, so red is reachable.

  • AC-OBS-CEA-025.1 narrowed in the Factory requirement (tracked edit) and in the mirror, to exactly the figures that carry a label:

    • Overview spend tile.
    • Usage period cards, Cost Breakdown table and coverage banner.
    • Both Top Sessions by Cost tables and the session cost chart.
    • Brain cost card and per-call list.
    • Snapshot cost slice.

    The Usage tab's other cost cards and the Overview hero chip are now an explicit non-goal and are listed under Remaining. They are no longer claimed.

Why

Users ask whether a cost figure is an estimate or a bill, and the dashboard could not say.

  • Labels describe arithmetic, not money. Nearly every figure is recorded tokens priced at a provider's published rate, or the runtime's own per-call cost, which runtimes also compute from published rates. The shared badge said "derived". That is true of the arithmetic but says nothing about whether the number was invoiced.
  • Zero-bill claims. With a subscription detected, the Usage tab said "Your actual out-of-pocket spend is $0" and "$0 billed to you". The info icons said "actual incremental cost is $0". The Overview hero said "free on your plan". ClawMetry cannot see a plan fee, included allowance or overage, so each of those is a claim about a bill it has never seen.
  • Unlabelled figures. The Flow brain panel's cost was a pre-formatted string with no basis.

What

Vocabulary. clawmetry/cost_basis.py (new, short module) adds a financial basis that rides beside the existing provenance basis inside the same entry. The provenance wire shape is unchanged: consumers that ignore the new fields keep working.

  • The bases:
    • published_rate: usage value at published rates. This covers a runtime-reported cost too, whose source is named in rate_source.
    • contract: expected contract spend. It needs a recorded rate_version.
    • allocated_actual: allocated actual spend. It needs a recorded ledger_ref.
    • unknown.
  • A bill-shaped label without its evidence becomes an unknown entry with the reason. stamp() then nulls the figure, so an unevidenced "actual spend" cannot reach a screen as a number. Nothing produces contract or allocated_actual today.
  • Billing route is subscription, metered or unknown.
    • Subscription-covered value is labelled "value, not an extra bill".
    • The remainder is metered only when a metered runtime was actually detected, otherwise "not detected".
    • plan_fee_usd is None, labelled unknown, never 0.0.

Surfaces.

Surface Change
/api/usage (routes/usage.py) Every cost entry carries the financial basis and rate source. The coverage split carries its billing routes, and routing savings are labelled as a published-rate counterfactual.
Hosted snapshot (clawmetry/sync.py) The same labels on dailyUsage and on both branches of the spending triple (live, and the stale state fallback).
/api/sessions/cost-breakdown (routes/sessions.py) Every figure is labelled, on both the fast and the transcript paths. Rows priced at the blended per-token rate make the session figures estimated, with a count.
Usage session cost chart and table (app.js loadUsage / renderSessionCostChart) The breakdown's basis is kept: a caption above the canvas bars, a badge on the Cost column, and "not available" for an unknown cost.
Flow brain panel (routes/components.py) New numeric today_cost_usd with provenance, next to the legacy string. Calls that carried tokens but no price make the total a labelled floor. If nothing could be priced it is "not available", not $0.00. The per-call list carries cost_usd under a calls[].cost_usd entry, rendered as "Cost per call: published rates".
Badge (clawmetry/static/js/provenance.js) Shows the financial words ("published rates"). It takes keyboard focus, and its aria-label carries the explanation. On :focus-visible one fixed tip on <body> shows the same text a mouse user gets from the title tooltip.
Usage banner, cards, info icon (clawmetry/static/js/app.js) Included-in-plan value is shown apart from metered usage, or usage whose route was not detected. Each card has its basis badge. No $0-bill wording.
Overview Info-icon and hero copy fixed: the hero says "included in your plan, not an extra bill" instead of "free on your plan".
Inventory The subscription chip tooltip no longer says "Usage adds $0 extra".
dashboard.css Financial-basis badge casing and the focus tip.

Verification

Regression guard, red before green. tests/test_cost_basis_labels.py has 18 tests.

  • How it tests: it walks the real payload builders, and it runs the shipped renderBillingCoverageBanner, _planLabel, renderSessionCostChart, loadBrainData and cmProv.badge out of app.js and provenance.js under node.
  • Red, original 15: copied onto origin/main at 658342f238 with only the new vocabulary module added, 12 of the 15 fail. The 3 that pass test the vocabulary alone.
  • Red, review-fix 3: the three tests added for the review fail on the previous head b2b438f5e6, checked in a separate detached worktree:
    • test_the_usage_session_cost_chart_and_table_show_their_basis
    • test_the_flow_brain_call_list_cost_is_labelled
    • test_the_flow_brain_call_list_renders_its_basis
  • Green: all 18 pass on this branch. 61 pass together with test_provenance.py and test_provenance_render_coverage.py.

A guard that never ran now runs. tests/test_provenance.py and tests/test_provenance_render_coverage.py were named in no workflow, and main is at 66 unbadged renders against their ceiling of 63. This PR:

  • wires both into the MOAT verifier job, with the new file;
  • converts dead fallbacks, reaching 62;
  • lowers UNBADGED_CEILING to 62, which is downward only.

docs/ci_test_coverage_baseline.json tightened (unlisted 920 to 917).

Other gates: check_ac_coverage.py --check OK at 124/194 on the rebased branch. gen_module_map.py --check OK. check_py39_annotations.py OK. node --check passes on both JS files.

Real dashboard in a browser, first pass. I ran this branch's dashboard.py with a scratch HOME (a Claude Max OAuth marker, so a subscription is detected), a scratch seeded DuckDB and a random port, and drove it with Chrome DevTools.

  • Overview tile: reads $0.59 with a published rates badge.
    • With keyboard focus, the explanation sits 6px below the badge, fully in the viewport, not clipped.
    • The page contains no "free on your plan", "billed to you" or "out-of-pocket".
  • Usage banner: "Usage included in Claude Max 20x. The figures below are usage value at published rates ... this value is not an extra bill. ... ClawMetry cannot see your plan fee, included allowance or overages, so none of them are counted here."
  • Usage cards: "about $0.59 · published rates".
  • Flow brain panel: the cost card reads "$0.13 published rates".
  • Mixed and undetected banners are covered by the node-run tests rather than the browser, because a scratch HOME cannot fake a metered runtime and a subscription at once without editing the detector.

Real dashboard in a browser, review-fix pass. This time the scratch HOME was seeded with today's OpenClaw transcript: one priced call ($0.0042) and one call on a model the runtime did not price. I ran dashboard.py on a random port and checked it in Chrome.

  • Usage tab: the caption reads "Bar values and the Cost column: published rates". The table header reads "Session Tokens Cost published rates Model Date", and the row reads sess-live-1 2K $0.0042.
  • Flow brain panel: the list is headed "Cost per call: published rates", with rows $0.0042 and $0.0006 (the second is ClawMetry's price-table fallback). The badge has tabindex=0, and its tip gives the formula, the window ("one assistant call"), the rate source and the source.
  • Over HTTP: /api/component/brain returns calls[].cost_usd with cost_basis=published_rate, and /api/sessions/cost-breakdown returns top10[].cost_usd labelled "published rates".

Remaining (not in this PR; #5937 stays open)

  • Other Usage tab cost cards. Each is served by its own endpoint, none of which carries a financial basis yet, so none shows a basis:

    • Where the money goes
    • Cost Forecast
    • Cache Re-read Tax
    • cache hit rate card
    • Cost By Plugin / Skill
    • Cost Comparison
    • Spend Optimization
    • Skill Cost Leaderboard
    • Cost by Team

    These are a non-goal in REQ-OBS-CEA-025, and the render ratchet keeps their count from growing.

  • Overview hero chip. It shows a cost with no basis and an unclear window, and it now carries the "included in your plan, not an extra bill" suffix. Its figure also differed from the tile's today figure in the first browser pass ($12.58 vs $0.59). That predates this PR and was not investigated.

  • Sessions tab transcript chips. The per-turn and per-tool cost chips come from _buildReplayEvent over several local and hosted transcript sources, so labelling them needs a transcript payload change on both sides. The Sessions-list chips in loadSessions() write to #sessions-list, which no template renders, so they were left untouched rather than claimed.

  • Contract spend and actual spend. Needs the price book (Price book contract: negotiated rates, Azure OpenAI deployment aliases, effective dates #5959 / Pricing: custom price book for negotiated rates + Azure OpenAI deployment aliases #5936) plus a valuation engine, and an invoice or ledger ingest. The labels and their evidence rule exist; no figure uses them.

  • Hand-check against every source, and a cache and window-boundary audit. The refinement asks for cache tokens not counted twice, and for late usage kept consistent across local, snapshot and hosted views. Not done here.

  • Budgets (Attribution: project + user tags and per-project budgets (burn vs budget) #5941) and the Cost Optimizer (Cost Optimizer: 33s spinner on cold load, debug label and unlabelled figures #5934) choosing a basis. They render through the same formatter and inherit the labels; their thresholds are unchanged.

  • Hosted dashboard. It renders the OSS app.js, so the corrected copy and the snapshot's labels reach it with the next pin. The hosted Usage interceptor's billingCoverage carries no provenance, so the split figures render without a badge there. The same applies to hosted brain and cost-breakdown payloads that lack the new entries: they keep their legacy strings and invent no label. A cloud-side pass-through is a follow-up.

  • Badge tab stops (review note, not addressed). Every provenance badge, compact ones included, now takes tabindex=0, which can add many tab stops on dense tables.

  • Other unbadged money renders. 62 remain in app.js, held by the ratchet.

🤖 Generated with Claude Code

https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9

@8090-software-factory

Copy link
Copy Markdown

✅ Drift Bot (ClawMetry): no drift detected

Drift Bot analyzed the changed files against this project's blueprints and requirements and found no drift.

@github-actions

github-actions Bot commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor

Visual diff

Comparing c5c84c81a8e0 (head) against the PR base branch.

44 of 70 comparison(s) flagged (>1% pixel diff).

View Before After Diff
desktop overview before after diff · 0.01%
desktop flow before after diff · 0.08%
desktop brain ⚠️ before after diff · 100.00%
desktop usage ⚠️ before after diff · 100.00%
desktop crons before after diff · 0.01%
desktop memory ⚠️ before after diff · 4.63%
desktop security before after diff · 0.18%
desktop subagents before after diff · 0.02%
desktop transcripts ⚠️ before after diff · 100.00%
desktop logs ⚠️ before after diff · 7.28%
desktop skills before after diff · 0.01%
desktop models before after diff · 0.00%
desktop approvals before after diff · 0.26%
desktop alerts ⚠️ before after diff · 100.00%
desktop notifications before after diff · 0.47%
desktop limits ⚠️ before after diff · 100.00%
desktop history before after diff · 0.27%
desktop channels ⚠️ before after diff · 100.00%
desktop harness ⚠️ before after diff · 100.00%
desktop inventory ⚠️ before after diff · 100.00%
desktop nemoclaw ⚠️ before after diff · 5.83%
desktop guard ⚠️ before after diff · 100.00%
desktop signals ⚠️ before after diff · 3.07%
desktop policy before after diff · 0.42%
desktop selfevolve ⚠️ before after diff · 100.00%
desktop swimlane ⚠️ before after diff · 100.00%
desktop tool-catalog ⚠️ before after diff · 100.00%
desktop tracing ⚠️ before after diff · 100.00%
desktop turn-anatomy ⚠️ before after diff · 3.10%
desktop version-impact ⚠️ before after diff · 100.00%
desktop context-economics before after diff · 0.67%
desktop agents ⚠️ before after diff · 100.00%
desktop evals ⚠️ before after diff · 16.80%
desktop bench ⚠️ before after diff · 2.51%
desktop trail ⚠️ before after diff · 100.00%
mobile overview before after diff · 0.00%
mobile flow ⚠️ before after diff · 5.82%
mobile brain ⚠️ before after diff · 6.64%
mobile usage ⚠️ before after diff · 100.00%
mobile crons ⚠️ before after diff · 100.00%
mobile memory ⚠️ before after diff · 4.54%
mobile security ⚠️ before after diff · 100.00%
mobile subagents ⚠️ before after diff · 100.00%
mobile transcripts ⚠️ before after diff · 3.43%
mobile logs before after diff · 0.02%
mobile skills before after diff · 0.07%
mobile models ⚠️ before after diff · 1.94%
mobile approvals ⚠️ before after diff · 100.00%
mobile alerts ⚠️ before after diff · 3.57%
mobile notifications ⚠️ before after diff · 100.00%
mobile limits before after diff · 0.01%
mobile history ⚠️ before after diff · 2.24%
mobile channels ⚠️ before after diff · 1.61%
mobile harness before after diff · 0.00%
mobile inventory before after diff · 0.02%
mobile nemoclaw ⚠️ before after diff · 100.00%
mobile guard before after diff · 0.00%
mobile signals ⚠️ before after diff · 100.00%
mobile policy before after diff · 0.02%
mobile selfevolve ⚠️ before after diff · 100.00%
mobile swimlane before after diff · 0.01%
mobile tool-catalog ⚠️ before after diff · 100.00%
mobile tracing ⚠️ before after diff · 1.14%
mobile turn-anatomy ⚠️ before after diff · 100.00%
mobile version-impact ⚠️ before after diff · 3.18%
mobile context-economics before after diff · 0.02%
mobile agents before after diff · 0.01%
mobile evals before after diff · 0.01%
mobile bench before after diff · 0.00%
mobile trail before after diff · 0.01%

Folder: c5c84c81a8e0. Full PNGs also attached as a workflow artefact.

Generated by visual-diff bot. Pixel diffs >1% flagged; eyeball the table before merging. This check is non-blocking — fail = bot bug, not a code problem.

@vivekchand

Copy link
Copy Markdown
Owner Author

Coordinator review

Verdict: fix before merge. One blocking item. Most of the PR holds up: the vocabulary, the evidence rule, the removed zero-bill copy, the focus tip, and the tightened ratchets.

Reviewed head 5fa9794fa9, which is rebased on current origin/main (658342f238). CI is green (47 pass, Scorecard skipped); the PR is MERGEABLE and CLEAN. I reran tests/test_cost_basis_labels.py, test_provenance.py and test_provenance_render_coverage.py in a scratch venv: 58 passed.

Blocking

1. AC-OBS-CEA-025.1 is marked met, but two of its named surfaces still show cost figures with no visible basis.

The criterion says "Every cost figure on ... the per-session cost breakdown, the Flow brain panel ... shown as visible text beside the figure." Its evidence is API-level only (/api/sessions/cost-breakdown returns cost_basis). These renders are still unlabelled:

  • Usage tab, session cost chart and table.
    • loadUsage() keeps only window._sessionCostData = cbd.top10, so the payload's provenance is dropped before rendering.
    • renderSessionCostChart() then draws '$' + cost.toFixed(4) on the canvas bars and in the table's Cost column, with no badge.
    • Both elements are live: templates/tabs/usage.html:140-141.
  • Flow brain panel, per-call list.
    • loadBrainData() at app.js:26620-26628 prints c.cost, a pre-formatted "$0.0123" string from routes/components.py:1788/2016, with no basis.
    • Only the summary Cost card was labelled.

AC-025.1 also says every cost figure on the Usage tab is labelled. The PR itself notes that 62 renders in app.js still have no badge. Several are Usage-tab cards: cost comparison, cache analytics, forecast, by-team and spend optimisation.

Fix, either way:

  • (a) Carry cbd.provenance['top10[].cost_usd'] into renderSessionCostChart and put a cmProv.badge on the Cost column heading, as renderTopSessionsByCost already does. Label the brain call-list column the same way, with a test for each.
  • (b) Narrow AC-025.1, in the Factory requirement and in docs/acceptance_criteria.json, to the figures actually labelled: the Overview tile, the Usage summary cards and banner, the Top Sessions table, the brain summary card, and the snapshot cost slice. Then list the session cost chart, the brain call list and the other Usage cards under Remaining.

A criterion marked met while its text is false on a live screen is a hidden gap.

Non-blocking (list under Remaining or follow up)

  • Overview hero chip.
    • The 💸 $12.58 chip has no basis, although it is an Overview cost figure (the issue's done-when covers every Overview cost). The PR notes its mismatch with the tile, but not that it is unlabelled.
    • With the new "included in your plan, not an extra bill" suffix, a figure of unclear window now carries a financial claim. Worth resolving together.
  • Tab stops. cmProv.badge() now adds tabindex="0" to every provenance badge, including compact one-letter badges in dense tables and non-cost score badges, which could add dozens of tab stops per tab. Consider limiting focusability to cost badges or non-compact badges.
  • _setCost fallback. When cmProv is missing it prints a raw float with no $ (String(usage[key])). The path is unreachable because provenance.js loads first, but it is a worse fallback than before.
  • Visual diff. The 100% flags on usage and brain are only page-height changes (5111→5160 px, 905→906 px), not a regression.
  • Remaining items listed honestly: Sessions-tab transcript chips, contract and actual spend producers, the cache and window audit, budgets and optimizer basis, the hosted billingCoverage badge pass-through, and the unverified mixed-route banner in a live browser.

Gates checked

🤖 Generated with Claude Code

https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9

vivekchand and others added 3 commits September 14, 2026 08:43
… metered usage kept apart

Every cost figure now carries a financial basis beside its provenance
basis (usage value at published rates / expected contract spend /
allocated actual spend / not available). Contract and actual labels are
refused without the rate version or ledger reference they rest on.

The Usage coverage banner, cards and info text, the Overview hero and
the inventory chip no longer describe subscription-covered usage as a
$0 bill: it is shown as value, apart from metered usage, with an
undetected route labelled undetected and the plan fee reported as not
visible. The Sessions cost chips and the Flow brain panel cost render
through the shared badge, which is keyboard-focusable and shows its
explanation on focus.

Refs #5937 (REQ-OBS-CEA-025).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9
…hip edits

A CSS ::after tip on the badge was clipped to one line by the Overview
tile's overflow:hidden, and hiding it on scroll hid it the moment Tab
navigation scrolled the badge into view. provenance.js now shows one
fixed tip on <body> on :focus-visible and follows the badge on scroll.

loadSessions() writes to #sessions-list, which no template renders, so
its chip edits could not be verified and are reverted. The live
Sessions-tab transcript chips are listed as not in this change.

Refs #5937 (REQ-OBS-CEA-025).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9
Review of #5975 found AC-OBS-CEA-025.1 claimed on screens that still
printed unlabelled costs:

- loadUsage kept only cbd.top10 and dropped the cost breakdown's
  provenance, so the session cost chart and its table printed bare
  dollar figures. The entry now travels with the rows: a caption above
  the canvas bars and a badge on the Cost column, with an unknown cost
  reading "not available" instead of $0.0000.
- The Flow brain panel's per-call list printed a pre-formatted string.
  Each call now carries cost_usd (null when it had tokens but no price)
  under one calls[].cost_usd entry, rendered as "Cost per call: published
  rates". The colour thresholds are checked largest first, so red is
  reachable.

AC-OBS-CEA-025.1 is narrowed, in the Factory requirement and the mirror,
to the figures that are actually labelled. The other Usage cost cards and
the Overview hero chip are named as a non-goal instead of being claimed.

Three new tests run the shipped renderers under node and the payload
builder in Python; all three fail on the previous head.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9
@vivekchand
vivekchand force-pushed the feat/cost-basis-labels-5937 branch from 5fa9794 to 79c7121 Compare September 14, 2026 06:53
@8090-software-factory

Copy link
Copy Markdown

✅ Drift Bot (ClawMetry): no drift detected

Drift Bot analyzed the changed files against this project's blueprints and requirements and found no drift.

@vivekchand

Copy link
Copy Markdown
Owner Author

Review follow-up (head 79c7121). This addresses the blocking item on AC-OBS-CEA-025.1.

Option (a), for the two screens you named:

  • Usage session cost chart and table. loadUsage() now keeps cbd.provenance['top10[].cost_usd'], stored as window._sessionCostEntry. renderSessionCostChart() then renders:
    • a caption above the canvas: "Bar values and the Cost column: published rates";
    • a cmProv.badge on the Cost column heading;
    • each cost cell through cmProv.figure, so a null cost reads "not available", not $0.0000.
  • Flow brain per-call list.
    • _brain_call_costs() in routes/components.py gives every call a numeric cost_usd, which is null when the call had tokens but no price. The list is labelled once, as calls[].cost_usd. Both the DuckDB path and the transcript path go through it.
    • loadBrainData() renders "Cost per call: published rates" above the list, and each cell through cmProv.figure.
    • The colour thresholds are now checked largest first.

Option (b), for the rest. AC-025.1 is narrowed in the Factory requirement (tracked edit) and in docs/acceptance_criteria.json to the figures that actually carry a label. The other Usage cost cards (spend flow, forecast, cache re-read tax, cache hit rate, plugin/skill, comparison, spend optimization, skill leaderboard, team) and the Overview hero chip are now a named non-goal and are listed under Remaining in the PR body. They are no longer claimed.

Guards. There are 3 new tests in tests/test_cost_basis_labels.py. They run the shipped renderSessionCostChart and loadBrainData under node, plus the Python payload builder. All 3 fail on the previous head b2b438f5e6 (checked in a separate detached worktree), and all 18 pass on this head. The render ratchet stays at 62, and check_ac_coverage --check is OK.

Browser check. Scratch HOME seeded with today's OpenClaw transcript, random port, Chrome:

  • Usage tab: the caption and the Cost-column badge both read "published rates". The row reads $0.0042.
  • Brain panel: the list is headed "Cost per call: published rates". The badge takes keyboard focus (tabindex=0), and its tip gives the formula, the window, the rate source and the source.

Non-blocking notes:

  • Hero chip: now listed under Remaining.
  • Visual-diff flags: page height only, as you found.
  • tabindex on every badge: not changed. It is listed under Remaining as a follow-up.

Regenerate the generated inventory so the gen_module_map --check lint
guard passes on CI.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KUtV6jyUVMhBRWXjSSef9S
@8090-software-factory

Copy link
Copy Markdown

✅ Drift Bot (ClawMetry): no drift detected

Drift Bot analyzed the changed files against this project's blueprints and requirements and found no drift.

github-actions Bot pushed a commit that referenced this pull request Sep 14, 2026
@vivekchand

Copy link
Copy Markdown
Owner Author

Ready to merge. Every blocking review item is addressed (see the review follow-up comment above). CI is green, drift-bot and the product-record gate pass, and GitHub reports the PR as mergeable and CLEAN. I have not merged it.

Merge-after dependencies: none. The PR is branched off main and not stacked. #5959 (price book) is a conceptual dependency only. If #5959 merges first, the overlap is additive: CHANGELOG.md, docs/acceptance_criteria.json and docs/ac_coverage_baseline.json. Rebase, keep both sides, and regenerate the baseline with python3 scripts/check_ac_coverage.py --update-baseline.

Companion PRs: none. No new HTTP route, so there is no cloud_route_policy entry and no cloud PR.

Factory: the AC-OBS-CEA-025.1 narrowing and the new non-goal are pending tracked suggestions on https://factory.8090.ai/project/b415065f-ab2f-4f53-8864-0c009fd098cb/requirements/950d4687-45cc-45a8-9d58-0c0d82fd5d9b. Accept them so the requirement matches the mirror in docs/acceptance_criteria.json.

Post-merge / post-release verification (after the [RELEASE] PR publishes to PyPI):

  1. Install the release in a scratch venv under a scratch HOME, not the real daemon: python3 -m venv /tmp/v && /tmp/v/bin/pip install clawmetry==<released version>.
  2. Seed today's OpenClaw transcript in $HOME/.openclaw/agents/main/sessions/ with one priced assistant call (usage.cost.total), then start the dashboard on a free port.
  3. Check the API over HTTP:
    • GET /api/component/brain: every call has cost_usd, and provenance["calls[].cost_usd"].cost_basis == "published_rate".
    • GET /api/sessions/cost-breakdown: provenance["top10[].cost_usd"].cost_basis_label == "published rates".
    • GET /api/usage: covered_usd has billing route subscription, plan_fee_usd is null, and no entry claims contract or allocated_actual.
  4. Check the pages in a browser:
    • Usage tab: the session cost chart caption reads "Bar values and the Cost column: published rates", and the Cost column heading has the same badge.
    • Flow, Brain node: the list heading reads "Cost per call: published rates".
    • Overview: the spend tile has a "published rates" badge, and Tab-focusing it shows the formula and rate source below the badge.
    • Wording: no page says "free on your plan", "billed to you" or "out-of-pocket".
  5. After the cloud pin auto-bumps to the release, check the hosted dashboard:
    • Usage tab: the subscription copy says "not an extra bill", with no $0-bill wording.
    • Snapshot: dailyUsage.provenance entries carry cost_basis.
    • Known gap: the hosted billingCoverage split renders without a badge. It is listed as a follow-up, not a regression.

Remaining scope stays on #5937; see the Remaining section of the PR body.

🤖 Generated with Claude Code

https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9

@vivekchand
vivekchand merged commit c80cbff into main Sep 14, 2026
54 of 72 checks passed
vivekchand added a commit that referenced this pull request Sep 14, 2026
…s, effective dates

A local price book (~/.clawmetry/pricing.json, schema clawmetry.price_book/1)
with exact/prefix/pattern entries, effective_from/effective_to, rates or a
discount off the published rate, and Azure deployment aliases scoped to their
resource. POST /api/pricing/resolve says which entry applies to each usage
record and why; GET /api/pricing/book shows the validated book, rejected
entries and recorded content-addressed versions; POST /api/pricing/valuations
defines the contract a future engine answers (501 until one exists). The
interceptor now captures Azure OpenAI calls with their deployment and host.

Integrated with the cost basis labels (#5975): every figure this surface
returns carries the shared cost_basis vocabulary. A list figure is
published_rate, no amount is unknown, and a contract valuation goes through
cost_basis.label, so it cannot claim "contract rate" without its rate version.

Rebased onto current main as one commit. acceptance_criteria.json carries the
AC-OBS-CEA-024 block once (earlier merges had duplicated it and the
AC-GOV-FWM block). The CHANGELOG entry is left to the release.

Refs #5936.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9
vivekchand added a commit that referenced this pull request Sep 14, 2026
Attributes each session's spend to a project derived when read (operator
assignment, else the recorded git repository containing its cwd, else the
directory, else a visible Unassigned row), with per-project budgets on a
declared period, timezone, currency and basis, 50/80/100% alerts latched once
per budget period, and GET /api/usage/export?by=project.

A budget on a repository keeps counting its sessions after that repository is
assigned to a named project, and reports the named project under
reported_under. The usage-bucket cache key snaps to a 15-minute edge so
daemon ticks reuse it.

Rebased on main as one commit. Integrated with the cost basis labels (#5975):
/api/projects, /api/projects/budgets and /api/projects/budgets/alerts now stamp
every dollar figure as usage value at published rates through
clawmetry.cost_basis, the CSV carries a cost_basis column, and the alert text
and notice say "at published rates, not an invoice" instead of "estimated
spend". The CHANGELOG entry is left to the release.

Refs #5941. REQ-OBS-PRJ-001.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9
vivekchand added a commit that referenced this pull request Sep 14, 2026
… CHANGELOG entry

Builds on c1b7fed (which routed gatewayMoney through cmFigure with a
hardcoded client-side entry). The label now comes from the server, in the
one cost vocabulary (clawmetry/cost_basis.py): gateway_litellm stamps
provenance on the gateway object with basis measured, cost_basis
published_rate, and a rate_source naming LiteLLM (ClawMetry does not
re-price it). A null cost renders the unknown state "not reported" via a
cost_basis unknown entry, not a plain string. The column header carries
the badge; the exact reported amount stays in the tooltip, so a fraction
of a cent is not rounded away.

The gateway object's top-level `cost_basis: "gateway_reported"` collided
with that vocabulary and is renamed `cost_source` (the ledger row's name).
CI script, test and docs/LITELLM.md updated.

CHANGELOG.md restored to main's version (release notes are written at
release time).

Revert-proof: the new test assertions and the unbadged-render ratchet both
fail against the previous gateway_litellm.py / app.js and pass with this.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9
vivekchand added a commit that referenced this pull request Sep 14, 2026
Publishes the six merged changes listed in CHANGELOG.md: Cost Optimizer
honesty (#5951), cost basis labels (#5975), framework IDs on Guard findings
(#5952), ATLAS replay scorecard (#5961), PR provenance SARIF (#5974), and
the Copilot VS Code control refusal (#5954).


Claude-Session: https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
vivekchand added a commit that referenced this pull request Sep 14, 2026
…s, effective dates (#5959)

* Price book contract: negotiated rates, Azure OpenAI deployment aliases, effective dates

A local price book (~/.clawmetry/pricing.json, schema clawmetry.price_book/1)
with exact/prefix/pattern entries, effective_from/effective_to, rates or a
discount off the published rate, and Azure deployment aliases scoped to their
resource. POST /api/pricing/resolve says which entry applies to each usage
record and why; GET /api/pricing/book shows the validated book, rejected
entries and recorded content-addressed versions; POST /api/pricing/valuations
defines the contract a future engine answers (501 until one exists). The
interceptor now captures Azure OpenAI calls with their deployment and host.

Integrated with the cost basis labels (#5975): every figure this surface
returns carries the shared cost_basis vocabulary. A list figure is
published_rate, no amount is unknown, and a contract valuation goes through
cost_basis.label, so it cannot claim "contract rate" without its rate version.

Rebased onto current main as one commit. acceptance_criteria.json carries the
AC-OBS-CEA-024 block once (earlier merges had duplicated it and the
AC-GOV-FWM block). The CHANGELOG entry is left to the release.

Refs #5936.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9

* chore: regenerate MODULE_MAP.md (259 modules after price-book additions)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017cF1pNfkt6jLiKc8qF9Zro

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
vivekchand added a commit that referenced this pull request Sep 14, 2026
Attributes each session's spend to a project derived when read (operator
assignment, else the recorded git repository containing its cwd, else the
directory, else a visible Unassigned row), with per-project budgets on a
declared period, timezone, currency and basis, 50/80/100% alerts latched once
per budget period, and GET /api/usage/export?by=project.

A budget on a repository keeps counting its sessions after that repository is
assigned to a named project, and reports the named project under
reported_under. The usage-bucket cache key snaps to a 15-minute edge so
daemon ticks reuse it.

Rebased on main as one commit. Integrated with the cost basis labels (#5975):
/api/projects, /api/projects/budgets and /api/projects/budgets/alerts now stamp
every dollar figure as usage value at published rates through
clawmetry.cost_basis, the CSV carries a cost_basis column, and the alert text
and notice say "at published rates, not an invoice" instead of "estimated
spend". The CHANGELOG entry is left to the release.

Refs #5941. REQ-OBS-PRJ-001.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9
vivekchand added a commit that referenced this pull request Sep 14, 2026
… CHANGELOG entry

Builds on c1b7fed (which routed gatewayMoney through cmFigure with a
hardcoded client-side entry). The label now comes from the server, in the
one cost vocabulary (clawmetry/cost_basis.py): gateway_litellm stamps
provenance on the gateway object with basis measured, cost_basis
published_rate, and a rate_source naming LiteLLM (ClawMetry does not
re-price it). A null cost renders the unknown state "not reported" via a
cost_basis unknown entry, not a plain string. The column header carries
the badge; the exact reported amount stays in the tooltip, so a fraction
of a cent is not rounded away.

The gateway object's top-level `cost_basis: "gateway_reported"` collided
with that vocabulary and is renamed `cost_source` (the ledger row's name).
CI script, test and docs/LITELLM.md updated.

CHANGELOG.md restored to main's version (release notes are written at
release time).

Revert-proof: the new test assertions and the unbadged-render ratchet both
fail against the previous gateway_litellm.py / app.js and pass with this.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9
vivekchand added a commit that referenced this pull request Sep 15, 2026
Attributes each session's spend to a project derived when read (operator
assignment, else the recorded git repository containing its cwd, else the
directory, else a visible Unassigned row), with per-project budgets on a
declared period, timezone, currency and basis, 50/80/100% alerts latched once
per budget period, and GET /api/usage/export?by=project.

A budget on a repository keeps counting its sessions after that repository is
assigned to a named project, and reports the named project under
reported_under. The usage-bucket cache key snaps to a 15-minute edge so
daemon ticks reuse it.

Rebased on main as one commit. Integrated with the cost basis labels (#5975):
/api/projects, /api/projects/budgets and /api/projects/budgets/alerts now stamp
every dollar figure as usage value at published rates through
clawmetry.cost_basis, the CSV carries a cost_basis column, and the alert text
and notice say "at published rates, not an invoice" instead of "estimated
spend". The CHANGELOG entry is left to the release.

Refs #5941. REQ-OBS-PRJ-001.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9
vivekchand added a commit that referenced this pull request Sep 15, 2026
…t) (#5968)

* Project attribution and per-project budgets: burn against budget (#5941)

Attributes each session's spend to a project derived when read (operator
assignment, else the recorded git repository containing its cwd, else the
directory, else a visible Unassigned row), with per-project budgets on a
declared period, timezone, currency and basis, 50/80/100% alerts latched once
per budget period, and GET /api/usage/export?by=project.

A budget on a repository keeps counting its sessions after that repository is
assigned to a named project, and reports the named project under
reported_under. The usage-bucket cache key snaps to a 15-minute edge so
daemon ticks reuse it.

Rebased on main as one commit. Integrated with the cost basis labels (#5975):
/api/projects, /api/projects/budgets and /api/projects/budgets/alerts now stamp
every dollar figure as usage value at published rates through
clawmetry.cost_basis, the CSV carries a cost_basis column, and the alert text
and notice say "at published rates, not an invoice" instead of "estimated
spend". The CHANGELOG entry is left to the release.

Refs #5941. REQ-OBS-PRJ-001.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Jm9d7s4fN55hN3YzQo75o9

* chore: regenerate MODULE_MAP.md (259→260 modules) for project attribution budgets branch

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HkHeWgNjveVVQbtiEqkENY

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants