Conversation
Basic, Analyst, Lite and Pro (variant B figures) as per-env config beside the untouched free plan; the synthesized free plan, its key, plan key and SSM parameter are byte-identical to 1966975. The portal attach policy keeps its construct id and name and now holds exactly three sid'd grants: GET /usageplans, GET /usageplans/*/usage, POST /usageplans/*/keys. The api-handler gets PORTAL_API_ID_PARAM and PORTAL_API_STAGE so the backend can pick the plan on our stage.
Gateway::plan_of lists GetUsagePlans by key, keeps the plan on our API stage (PORTAL_API_ID_PARAM + PORTAL_API_STAGE) and parses the tier from pricing-api-<tier>-<stage>; anything else is custom. usage_of and the attach now take the plan id, and PlanNotFound names the plan it got. /api/usage gains `plan` and states the plan's quota as `limit` (used + remaining is a warn-logged cross-check). A key on no plan for our stage answers plan: null; a plan without a quota answers null counters without calling GetUsage; MONTH is the calendar month whatever the offset, DAY the UTC day, WEEK and unknown periods are reported, not guessed. Each cached answer is re-checked against its own period rule. The mock control plane learns GetUsagePlans, plan-aware attach and usage reads.
The next issue after a rework used to delete the revoked keys first, and their plan membership went with them: a paid user came back on free. The early delete is gone. The target plan is resolved before anything is deleted: the winner's own plan for our stage (no attach), else the latest revoked key's plan, else free. Step 5 sweeps the revoked keys only after the new key is attached, so the order is create, attach, delete, and a crash between create and attach retries onto the same previous plan. The once-per-period cap is unchanged.
The Rate Limit card states the plan /api/usage reports: a second pill beside Active (Free, Basic, Analyst, Lite, Pro or Custom), the plan's per-second rate and x60 per minute, "Unlimited" for a plan without a throttle, and a stated no-plan state with no pill. /config's free figure remains for the no-key state, the landing page and an unanswered usage call. Monthly Usage takes its limit and reset from the plan and states the no-plan, unlimited and unsupported-period cases instead of zeros. The contact line is now a link to RUMBLEFISH_CONTACT with copy per tier.
The tier runbook now starts from the five CDK plans and walks an operator through moving a user's key: exact-name key lookup, delete then create the plan key (seconds of 403 between, value unchanged), get-usage-plans --key-id to verify, the counter restarting at zero and a rework keeping the plan. It adds the post-merge production checklist and keeps the hand-made Custom plan procedure, which must stay on our API stage to be reported as Custom. The deploy-prep grant list names the three /usageplans grants.
The new handler calls GetUsagePlans on every sign-in and usage load, but the grants sat in ApiGateway's policy, which deploys after Compute: until then every sign-in failed and /api/usage answered 502. The three grants (GET /usageplans, GET /usageplans/*/usage, POST /usageplans/*/keys) now live on the api-handler role's own policy, which the function depends on, so they land before the code. ApiGateway's PortalAttachKeyToFreePlan policy is removed. Compute keeps exporting the role name for one release, because the deployed ApiGateway policy still imports it.
A rework read the plan off the latest revoked record only, so an older sibling on a paid plan was ignored and the new key landed on free. The target plan now comes from the newest revoked record on a non-free plan for our stage, else free. CreateUsagePlanKey on a key already on another plan for the same stage answers 409 or a 400 naming the "same API Stage". Both now mean "already on a plan": the plan is read back with plan_of, the key is never put on free, and a plan that never shows ends in ?issue=failed. The mock control plane answers both and can lag.
The unit test compared a period rule with itself. An integration test now puts a MONTH plan at offset 7 and checks that period_start and the GetUsage window both start on the 1st and limit equals the quota.
While /api/usage loaded or failed, a paid key was shown 1 req/s and "Contact us about a paid plan", and the key card showed the free rate beside a Rate Limit card with the real one. Both cards now take the key's plan and state a loading or failed state instead of a figure. Quota labels and the reset follow the quota period; the regenerate strip keeps the monthly cap. An unknown tier falls back to Custom, fractional rates print with two decimals at most, and the no-plan card has its own contact line. The unused rateLimit prop is gone.
The runbook looks the current plan up instead of taking it by hand, refuses an empty or unchanged id, and only creates the new plan key after the delete succeeded. It adds a recovery for a refused create, the procedure for a revoked key waiting for the next period, the 0311 deploy order (Compute, then ApiGateway), and a revert that moves keys off the paid plans before ApiGateway deletes them.
With every key under a name disabled, current_key took the earliest record. Two revoked records sit side by side when the listing lagged the create after a rolled rework (nothing is swept on that path) and the new key was then reworked in the same period. The usage card then read the old key's counter, empty for the period, and the new key's traffic vanished from the dashboard until the 1st. The newest revocation now decides. The reveal and the revoke only read the enabled flag and the revocation instant, so they are unchanged.
A key on no usage plan was cached as an ordinary usage answer, and a successful issue evicted only "no key". The no-plan copy tells the user that signing out and in again puts the key back on a plan; the sign-in did attach it, but the dashboard kept saying "not on a usage plan" for the rest of the TTL, or up to STALE_KEEP while the control plane throttles. The no-plan answer is now evicted and epoch-guarded like "no key".
… not at cold start A burst of /v1 cold starts throttled SSM, and one failed read closed the portal in that execution environment for its life. The api-handler cold start now reads only the mTLS bundle. The portal's five sources load on the first portal request that needs them (/config included), concurrently under a 4 s budget, through one shared cell: a success is kept for the environment, a failure answers that request only and the next retries. - /config: enabled only when the flag is on and the load succeeded - /key, /usage, /me: 503 on a failed load; login keeps its landings; the callback lands on ?signin=failed, and past SOURCES_ALLOWANCE (500 ms) on a retryable failure before the token exchange - PortalLoadError::Keys names the variable of the read that failed - serve.rs keeps its eager, expect()ing loaders - tokio test-util enabled for prices-api tests only (paused clock)
…t closure
The portal no longer closes at cold start; a failed load answers one
request and the next retries. The portal-closed filter and alarm are
replaced 1-for-1 by prices-${env}-api-handler-portal-load-failed on the
new "portal sources failed to load" line (Prices/ApiHandler,
PortalSourcesLoadFailed). New logical ids, so CloudFormation replaces
both; DashboardAlarmCount is unchanged.
- guard test renamed; it now reads sources.rs for the line, main.rs for
the subscriber, and pins the alarm's metric wiring and the 1024-char
description limit
- ci.yml runs it when sources.rs or main.rs changes
- infra comments and runbooks describe the lazy load and its retry
The extension's own SSM retry outlasts our 2 s client timeout, so a throttled read surfaced as "extension unreachable" with no retry on our side. Every portal read through the extension (the OAuth secret, the plan and API ids, the eligibility parameters) now goes through with_retry: 3 attempts, full-jitter backoff capped at 100 ms then 300 ms, from getrandom. Only MtlsError::Fetch is retried; missing env, parse and validation errors are not. - the per-issuance eligibility read keeps PARAMETER_TIMEOUT (2 s) as the bound around the whole retry, so the callback arithmetic is unchanged - prices_clickhouse::mtls, and the /v1 mTLS path, are untouched - stale cold-start wording in the portal modules describes the lazy load
- callback on a failed load follows the action the round-trip claims (pending cookie, then state, read unverified to pick between two failure literals only): an issue press lands on ?issue=failed, which the signed-in dashboard renders, instead of ?signin=failed, which it does not - login on a failed load lands on ?signin=failed / ?issue=failed, not the permanent "not yet available" card, and logs no second ERROR line - a lazy load before the callback is paid out of the token exchange's own 4 s timeout; the allowance rises to 2 s and the 1 s redirect margin is back, pinned in budget_arithmetic_fits_the_lambda - tests that relied on the environment loader now supply their sources
…ed load - retry an extension read only on unreachable/timeout, 5xx, 429 and a 400 that names throttling; a 403/404/other 400, a missing field, a parse or client-build failure returns at once; the classified messages are pinned against mtls.rs, which stays untouched - a failed load is remembered for LOAD_FAILURE_COOLDOWN (2 s): inside it a portal request answers unavailable without loading and without a second alarm line; queued waiters no longer load in turn - the log assertions live in a binary of their own (portal_load_logs.rs) - a malformed OAuth secret's error no longer echoes field values - pin that a closed portal gets no loader, that a hung read gets one retry inside the load budget, and the observed retry schedule - app_with_portal debug-asserts the config carries no portal sources - guard test requires the prefix in the message literal, not a field - alarm description, runbook and stack comments state the cooldown
- USAGE_TIMEOUT_MS 15 s -> 20 s: /usage can load the portal's sources first (4 s) before its 10 s deadline, so its 503 can arrive at 14 s - KEY_TIMEOUT_MS keeps 20 s; its comment now cites the same arithmetic
Measure why the portal closed on 2026-09-24 while /v1 kept serving, how much the shared ClickHouse holds per endpoint mix, and the cold-start curve against SSM. Record the lazy portal load deployed the same day: a replay of the 09-24 herd (69 cold starts) now makes zero SSM reads and closes no portal.
The route check still required 0188's single GetUsage grant on the free plan, which 0311 replaced with three grants on the api-handler role: list plans by key, usage on the key's plan, attach a key to a plan. The old check fails on the current templates. Pin exactly that set in ComputeStack and none in ApiGatewayStack. Count any allowed apigateway resource whose glob reaches a usage plan (`::/*`, `::/usage*`), not only those spelling `/usageplans`, and refuse NotAction/NotResource statements, which cannot be enumerated.
History for the 09-24 deploy and alarm, and for 09-25: the herd test, the lazy portal load, SSM high-throughput, the deploy, the herd replay and the live portal check. Eight of nine criteria met; the rework criterion stays open with unit tests only, its live half skipped by Adam's decision. Decisions 8-10 added.
| let service_error = e.into_service_error(); | ||
| if service_error.is_conflict_exception() { | ||
| Ok(Attachment::OnPlan) | ||
| if service_error.is_conflict_exception() |
There was a problem hiding this comment.
A 409 on the plan we asked for now fails if GetUsagePlans lags behind the attach (medium)
Before this PR, 409 ConflictException from CreateUsagePlanKey returned OnPlan straight away. Now it becomes AlreadyOnAPlan, and attach asks plan_of(key_id) which plan the key is on. If that listing does not yet show the plan, the attempt returns Settled::Retry.
A 409 here already proves the key is on plan_id, the plan this call asked for. The Attachment doc says so itself: "already exists in the usage plan, i.e. THIS plan". Only the 400 "same API Stage" refusal leaves the plan unknown.
How it fails, in a double-submit:
- Request A attaches the winner to the target plan.
- Request B runs
resolve_target_plan. Itsplan_of(winner)does not show the plan yet, so B attaches the same key to the same plan and gets409. - B's second
plan_ofis still behind, so B returnsRetry. - Every re-run takes the same path until
MAX_ATTEMPTSruns out. The result isReconciled::Lost, thenIssueOutcome::Failed, and the visitor lands on?issue=failedwhile their key actually works.
Before this PR, the same race settled on the first 409. The PR's own Settled::Retry doc treats this lag as a real possibility.
Suggested fix: send ConflictException back to Attachment::OnPlan. Keep the plan_of check only for the 400 same-stage refusal.
CreateUsagePlanKey answers 409 when the key is already on the plan it was asked for, so the refusal names the plan. Reading it back with GetUsagePlans, which can lag the attach, let a double-submit re-run until the attempts ran out and land on ?issue=failed for a key that works. Before 0311 a 409 settled at once, and it does again. Only the 400 "same API Stage" refusal, which does not say which plan, still asks plan_of. The 409 race test now lags GetUsagePlans for good and must still end in ?issue=ok; it fails on the previous mapping. Refs: PR #351 review
Summary
Lore task 0311: five usage plans, and the dashboard shows each key's own plan.
Four paid plans in CDK. Basic, Analyst, Lite and Pro (variant B: quota CoinGecko ×10, rate ×0.6, burst 5× the rate) sit beside the free plan, which is unchanged. The synth diff against 1966975 shows the free plan, its key, its plan key and its SSM parameter byte-identical.
/api/usagereports the key's own plan.Gateway::plan_oflistsGetUsagePlans?keyId=and keeps only plans on our API id and stage (PORTAL_API_ID_PARAM,PORTAL_API_STAGE). Usage is read on that plan. The response gainsplan { tier, name, rate_limit_per_second, burst_limit, quota_limit, quota_period }, andlimitis the plan's quota.custom.plan: null, and a plan without a quota answers null counters. Neither answer is zeros.offset, becauseoffsetcounts requests. DAY is the UTC day. WEEK and unknown periods are reported, not guessed.A rework keeps the plan. The early delete of revoked keys is gone. The order is now create, attach, delete. The target plan is resolved before anything is deleted:
A 409, or a 400 naming the "same API Stage" from
CreateUsagePlanKey, is read back withplan_of, so the key is never pushed to free. The once-per-period cap is unchanged.Dashboard.
Free,Basic,Analyst,Lite,ProorCustom.RUMBLEFISH_CONTACT, with copy per tier.IAM. The api-handler role's own policy gets exactly three grants:
GET /usageplans,GET /usageplans/*/usageandPOST /usageplans/*/keys. They replace ApiGateway'sPortalAttachKeyToFreePlan, so they deploy together with the code that needs them.Runbook.
docs/runbooks/manual-api-key-tier.mdcovers:Rollout (not in this PR)
cdk diffbefore deploying: the free plan must show no change.make -C infra sync-portal-explorer) once Compute is live.Compute keeps exporting the role name for one release. The deployed ApiGateway policy still imports it, and CloudFormation refuses to drop an export that is in use. Remove the export in a later release.
After merge, on production (there is no dev environment), follow the checklist in the runbook:
Tests
cargo test --workspacecargo clippy -p prices-api --all-targets -D warnings/fmt --check