Skip to content

feat(lore-0311): five usage plans, and the dashboard states the key's own plan - #351

Open
adamkoot wants to merge 22 commits into
developfrom
feat/0311_paid-plans-and-dashboard-shows-the-keys-own-plan
Open

adamkoot wants to merge 22 commits into
developfrom
feat/0311_paid-plans-and-dashboard-shows-the-keys-own-plan

Conversation

@adamkoot

Copy link
Copy Markdown
Collaborator

Summary

Lore task 0311: five usage plans, and the dashboard shows each key's own plan.

  • Four paid plans in CDK. Basic, Analyst, Lite and Pro (variant B: quota CoinGecko ×10, rate ×0.6, burst 5× the rate) sit beside the free plan, which is unchanged. The synth diff against 1966975 shows the free plan, its key, its plan key and its SSM parameter byte-identical.

    Plan Quota / month Rate Burst
    free 100 000 1 req/s 5
    Basic 1 000 000 3 req/s 15
    Analyst 5 000 000 5 req/s 25
    Lite 20 000 000 10 req/s 50
    Pro 50 000 000 25 req/s 125
  • /api/usage reports the key's own plan. Gateway::plan_of lists GetUsagePlans?keyId= and keeps only plans on our API id and stage (PORTAL_API_ID_PARAM, PORTAL_API_STAGE). Usage is read on that plan. The response gains plan { tier, name, rate_limit_per_second, burst_limit, quota_limit, quota_period }, and limit is the plan's quota.

    • Other plans: a plan that is not one of the five reports as custom.
    • Distinct states: a key on no plan answers plan: null, and a plan without a quota answers null counters. Neither answer is zeros.
    • Periods: MONTH is the calendar month whatever the offset, because offset counts requests. DAY is the UTC day. WEEK and unknown periods are reported, not guessed.
    • Cache: the plan lives in the same 60 s cache entry.
  • A rework keeps the plan. The early delete of revoked keys is gone. The order is now create, attach, delete. The target plan is resolved before anything is deleted:

    1. the new key's own plan, if it already has one on our stage;
    2. else the newest revoked record on a non-free plan for our stage;
    3. else free.

    A 409, or a 400 naming the "same API Stage" from CreateUsagePlanKey, is read back with plan_of, so the key is never pushed to free. The once-per-period cap is unchanged.

  • Dashboard.

    • Plan pill: a second pill beside "Active" in the Rate Limit card: Free, Basic, Analyst, Lite, Pro or Custom.
    • Figures: per-second and per-minute (×60) rates from the plan. Quota and reset come from the plan too, with labels that follow the quota period.
    • Contact line: links to RUMBLEFISH_CONTACT, with copy per tier.
    • Before the plan is known: while the plan loads, or when loading fails, neither card shows a figure. The free plan's rate is no longer shown to a paid key.
  • IAM. The api-handler role's own policy gets exactly three grants: GET /usageplans, GET /usageplans/*/usage and POST /usageplans/*/keys. They replace ApiGateway's PortalAttachKeyToFreePlan, so they deploy together with the code that needs them.

  • Runbook. docs/runbooks/manual-api-key-tier.md covers:

    • upgrading a user, with guards against leaving a key on no plan;
    • recovery when a plan move fails;
    • a revoked key waiting for the next period;
    • the deploy order and the revert.

Rollout (not in this PR)

  1. Deploy Compute first, then ApiGateway. Read each cdk diff before deploying: the free plan must show no change.
  2. Deploy the portal bundle (make -C infra sync-portal-explorer) once Compute is live.
  3. To revert: move every key off the paid plans first, then deploy ApiGateway, then Compute.

Compute keeps exporting the role name for one release. The deployed ApiGateway policy still imports it, and CloudFormation refuses to drop an export that is in use. Remove the export in a later release.

After merge, on production (there is no dev environment), follow the checklist in the runbook:

  • move a test key free → Basic → Pro → free, and after each move check the gateway throttle and that the dashboard shows the new plan within 60 s;
  • rework a Basic key, and in the next period check that the new key is on Basic;
  • run the recovery path once. It also measures AWS's exact 400 message for a key already on another plan on the same stage.

Tests

Gate Result
cargo test --workspace 1247 passed, 0 failed, 250 ignored
cargo clippy -p prices-api --all-targets -D warnings / fmt --check clean
infra typecheck / lint / test 37/37
portal test 265 passed, 4 skipped
portal typecheck / lint, prettier clean
CDK synth diff vs 1966975 free plan unchanged; +4 paid plans; IAM exactly the three grants

Basic, Analyst, Lite and Pro (variant B figures) as per-env config
beside the untouched free plan; the synthesized free plan, its key,
plan key and SSM parameter are byte-identical to 1966975.

The portal attach policy keeps its construct id and name and now holds
exactly three sid'd grants: GET /usageplans, GET /usageplans/*/usage,
POST /usageplans/*/keys. The api-handler gets PORTAL_API_ID_PARAM and
PORTAL_API_STAGE so the backend can pick the plan on our stage.
Gateway::plan_of lists GetUsagePlans by key, keeps the plan on our API
stage (PORTAL_API_ID_PARAM + PORTAL_API_STAGE) and parses the tier from
pricing-api-<tier>-<stage>; anything else is custom. usage_of and the
attach now take the plan id, and PlanNotFound names the plan it got.

/api/usage gains `plan` and states the plan's quota as `limit`
(used + remaining is a warn-logged cross-check). A key on no plan for
our stage answers plan: null; a plan without a quota answers null
counters without calling GetUsage; MONTH is the calendar month whatever
the offset, DAY the UTC day, WEEK and unknown periods are reported, not
guessed. Each cached answer is re-checked against its own period rule.

The mock control plane learns GetUsagePlans, plan-aware attach and
usage reads.
The next issue after a rework used to delete the revoked keys first,
and their plan membership went with them: a paid user came back on
free. The early delete is gone. The target plan is resolved before
anything is deleted: the winner's own plan for our stage (no attach),
else the latest revoked key's plan, else free. Step 5 sweeps the
revoked keys only after the new key is attached, so the order is
create, attach, delete, and a crash between create and attach retries
onto the same previous plan. The once-per-period cap is unchanged.
The Rate Limit card states the plan /api/usage reports: a second pill
beside Active (Free, Basic, Analyst, Lite, Pro or Custom), the plan's
per-second rate and x60 per minute, "Unlimited" for a plan without a
throttle, and a stated no-plan state with no pill. /config's free
figure remains for the no-key state, the landing page and an
unanswered usage call. Monthly Usage takes its limit and reset from
the plan and states the no-plan, unlimited and unsupported-period
cases instead of zeros.

The contact line is now a link to RUMBLEFISH_CONTACT with copy per
tier.
The tier runbook now starts from the five CDK plans and walks an
operator through moving a user's key: exact-name key lookup, delete
then create the plan key (seconds of 403 between, value unchanged),
get-usage-plans --key-id to verify, the counter restarting at zero and
a rework keeping the plan. It adds the post-merge production checklist
and keeps the hand-made Custom plan procedure, which must stay on our
API stage to be reported as Custom. The deploy-prep grant list names
the three /usageplans grants.
The new handler calls GetUsagePlans on every sign-in and usage load,
but the grants sat in ApiGateway's policy, which deploys after
Compute: until then every sign-in failed and /api/usage answered 502.
The three grants (GET /usageplans, GET /usageplans/*/usage,
POST /usageplans/*/keys) now live on the api-handler role's own
policy, which the function depends on, so they land before the code.
ApiGateway's PortalAttachKeyToFreePlan policy is removed. Compute
keeps exporting the role name for one release, because the deployed
ApiGateway policy still imports it.
A rework read the plan off the latest revoked record only, so an
older sibling on a paid plan was ignored and the new key landed on
free. The target plan now comes from the newest revoked record on a
non-free plan for our stage, else free.

CreateUsagePlanKey on a key already on another plan for the same
stage answers 409 or a 400 naming the "same API Stage". Both now mean
"already on a plan": the plan is read back with plan_of, the key is
never put on free, and a plan that never shows ends in
?issue=failed. The mock control plane answers both and can lag.
The unit test compared a period rule with itself. An integration
test now puts a MONTH plan at offset 7 and checks that period_start
and the GetUsage window both start on the 1st and limit equals the
quota.
While /api/usage loaded or failed, a paid key was shown 1 req/s and
"Contact us about a paid plan", and the key card showed the free rate
beside a Rate Limit card with the real one. Both cards now take the
key's plan and state a loading or failed state instead of a figure.
Quota labels and the reset follow the quota period; the regenerate
strip keeps the monthly cap. An unknown tier falls back to Custom,
fractional rates print with two decimals at most, and the no-plan
card has its own contact line. The unused rateLimit prop is gone.
The runbook looks the current plan up instead of taking it by hand,
refuses an empty or unchanged id, and only creates the new plan key
after the delete succeeded. It adds a recovery for a refused create,
the procedure for a revoked key waiting for the next period, the
0311 deploy order (Compute, then ApiGateway), and a revert that moves
keys off the paid plans before ApiGateway deletes them.
With every key under a name disabled, current_key took the earliest
record. Two revoked records sit side by side when the listing lagged the
create after a rolled rework (nothing is swept on that path) and the new
key was then reworked in the same period. The usage card then read the
old key's counter, empty for the period, and the new key's traffic
vanished from the dashboard until the 1st.

The newest revocation now decides. The reveal and the revoke only read
the enabled flag and the revocation instant, so they are unchanged.
A key on no usage plan was cached as an ordinary usage answer, and a
successful issue evicted only "no key". The no-plan copy tells the user
that signing out and in again puts the key back on a plan; the sign-in
did attach it, but the dashboard kept saying "not on a usage plan" for
the rest of the TTL, or up to STALE_KEEP while the control plane
throttles.

The no-plan answer is now evicted and epoch-guarded like "no key".
… not at cold start

A burst of /v1 cold starts throttled SSM, and one failed read closed the
portal in that execution environment for its life. The api-handler cold
start now reads only the mTLS bundle. The portal's five sources load on
the first portal request that needs them (/config included), concurrently
under a 4 s budget, through one shared cell: a success is kept for the
environment, a failure answers that request only and the next retries.

- /config: enabled only when the flag is on and the load succeeded
- /key, /usage, /me: 503 on a failed load; login keeps its landings;
  the callback lands on ?signin=failed, and past SOURCES_ALLOWANCE
  (500 ms) on a retryable failure before the token exchange
- PortalLoadError::Keys names the variable of the read that failed
- serve.rs keeps its eager, expect()ing loaders
- tokio test-util enabled for prices-api tests only (paused clock)
…t closure

The portal no longer closes at cold start; a failed load answers one
request and the next retries. The portal-closed filter and alarm are
replaced 1-for-1 by prices-${env}-api-handler-portal-load-failed on the
new "portal sources failed to load" line (Prices/ApiHandler,
PortalSourcesLoadFailed). New logical ids, so CloudFormation replaces
both; DashboardAlarmCount is unchanged.

- guard test renamed; it now reads sources.rs for the line, main.rs for
  the subscriber, and pins the alarm's metric wiring and the 1024-char
  description limit
- ci.yml runs it when sources.rs or main.rs changes
- infra comments and runbooks describe the lazy load and its retry
The extension's own SSM retry outlasts our 2 s client timeout, so a
throttled read surfaced as "extension unreachable" with no retry on our
side. Every portal read through the extension (the OAuth secret, the
plan and API ids, the eligibility parameters) now goes through
with_retry: 3 attempts, full-jitter backoff capped at 100 ms then
300 ms, from getrandom. Only MtlsError::Fetch is retried; missing env,
parse and validation errors are not.

- the per-issuance eligibility read keeps PARAMETER_TIMEOUT (2 s) as the
  bound around the whole retry, so the callback arithmetic is unchanged
- prices_clickhouse::mtls, and the /v1 mTLS path, are untouched
- stale cold-start wording in the portal modules describes the lazy load
- callback on a failed load follows the action the round-trip claims
  (pending cookie, then state, read unverified to pick between two
  failure literals only): an issue press lands on ?issue=failed, which the
  signed-in dashboard renders, instead of ?signin=failed, which it does not
- login on a failed load lands on ?signin=failed / ?issue=failed, not the
  permanent "not yet available" card, and logs no second ERROR line
- a lazy load before the callback is paid out of the token exchange's own
  4 s timeout; the allowance rises to 2 s and the 1 s redirect margin is
  back, pinned in budget_arithmetic_fits_the_lambda
- tests that relied on the environment loader now supply their sources
…ed load

- retry an extension read only on unreachable/timeout, 5xx, 429 and a 400
  that names throttling; a 403/404/other 400, a missing field, a parse or
  client-build failure returns at once; the classified messages are pinned
  against mtls.rs, which stays untouched
- a failed load is remembered for LOAD_FAILURE_COOLDOWN (2 s): inside it a
  portal request answers unavailable without loading and without a second
  alarm line; queued waiters no longer load in turn
- the log assertions live in a binary of their own (portal_load_logs.rs)
- a malformed OAuth secret's error no longer echoes field values
- pin that a closed portal gets no loader, that a hung read gets one retry
  inside the load budget, and the observed retry schedule
- app_with_portal debug-asserts the config carries no portal sources
- guard test requires the prefix in the message literal, not a field
- alarm description, runbook and stack comments state the cooldown
- USAGE_TIMEOUT_MS 15 s -> 20 s: /usage can load the portal's sources
  first (4 s) before its 10 s deadline, so its 503 can arrive at 14 s
- KEY_TIMEOUT_MS keeps 20 s; its comment now cites the same arithmetic
Measure why the portal closed on 2026-09-24 while /v1 kept serving,
how much the shared ClickHouse holds per endpoint mix, and the
cold-start curve against SSM. Record the lazy portal load deployed
the same day: a replay of the 09-24 herd (69 cold starts) now makes
zero SSM reads and closes no portal.
The route check still required 0188's single GetUsage grant on the
free plan, which 0311 replaced with three grants on the api-handler
role: list plans by key, usage on the key's plan, attach a key to a
plan. The old check fails on the current templates.

Pin exactly that set in ComputeStack and none in ApiGatewayStack.
Count any allowed apigateway resource whose glob reaches a usage
plan (`::/*`, `::/usage*`), not only those spelling `/usageplans`,
and refuse NotAction/NotResource statements, which cannot be
enumerated.
History for the 09-24 deploy and alarm, and for 09-25: the herd test,
the lazy portal load, SSM high-throughput, the deploy, the herd replay
and the live portal check. Eight of nine criteria met; the rework
criterion stays open with unit tests only, its live half skipped by
Adam's decision. Decisions 8-10 added.
let service_error = e.into_service_error();
if service_error.is_conflict_exception() {
Ok(Attachment::OnPlan)
if service_error.is_conflict_exception()

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A 409 on the plan we asked for now fails if GetUsagePlans lags behind the attach (medium)

Before this PR, 409 ConflictException from CreateUsagePlanKey returned OnPlan straight away. Now it becomes AlreadyOnAPlan, and attach asks plan_of(key_id) which plan the key is on. If that listing does not yet show the plan, the attempt returns Settled::Retry.

A 409 here already proves the key is on plan_id, the plan this call asked for. The Attachment doc says so itself: "already exists in the usage plan, i.e. THIS plan". Only the 400 "same API Stage" refusal leaves the plan unknown.

How it fails, in a double-submit:

  1. Request A attaches the winner to the target plan.
  2. Request B runs resolve_target_plan. Its plan_of(winner) does not show the plan yet, so B attaches the same key to the same plan and gets 409.
  3. B's second plan_of is still behind, so B returns Retry.
  4. Every re-run takes the same path until MAX_ATTEMPTS runs out. The result is Reconciled::Lost, then IssueOutcome::Failed, and the visitor lands on ?issue=failed while their key actually works.

Before this PR, the same race settled on the first 409. The PR's own Settled::Retry doc treats this lag as a real possibility.

Suggested fix: send ConflictException back to Attachment::OnPlan. Keep the plan_of check only for the 400 same-stage refusal.

CreateUsagePlanKey answers 409 when the key is already on the plan it
was asked for, so the refusal names the plan. Reading it back with
GetUsagePlans, which can lag the attach, let a double-submit re-run
until the attempts ran out and land on ?issue=failed for a key that
works. Before 0311 a 409 settled at once, and it does again.

Only the 400 "same API Stage" refusal, which does not say which plan,
still asks plan_of. The 409 race test now lags GetUsagePlans for good
and must still end in ?issue=ok; it fails on the previous mapping.

Refs: PR #351 review
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants