Skip to content

release: protocol 28 readiness, balance change on the account page, assets by holders - #453

Merged
karolko9 merged 102 commits into
masterfrom
develop
Sep 16, 2026
Merged

karolko9 merged 102 commits into
masterfrom
develop

Conversation

@karolko9

@karolko9 karolko9 commented Sep 8, 2026 •

Copy link
Copy Markdown
Collaborator

Range: production-2026.09.07-1 → origin/develop (c2f2075c). The 2026-09-14 production deploy covered the range through ba3575e3; Compute was deployed again on 2026-09-16 at 9b0f05b6, which adds task 0374 and the pnpm migration. Later develop commits include lore records and the SPA fix 87412cb6, which is part of this PR but is not deployed yet.

Production is ahead of master, but not at this PR's head. Because of the protocol-28 vote on 2026-09-16 17:00 UTC, the range through ba3575e3 was deployed on 2026-09-14 straight from develop, from an operator laptop with make deploy-production-compute, -web and -ingestion — no production-* tag was cut. On 2026-09-16 Compute alone was deployed the same way from develop at 9b0f05b6 (the PR #459 merge). The later SPA fix 87412cb6 has not been deployed. Merging this PR deploys nothing by itself.

What ships

Protocol 28 readiness — task 0548 (PR #456)

  • stellar-xdr 27 → 28. Without it the first protocol-28 ledger fails to decode and the ingest queue dead-letters (the task 0368 incident). The new ContractExecutable::ExternalRef and ScVal::ExecutableTag arms render as their own types, never as wasm.
  • Galexie 28.0.1 pinned from the mirrored ECR digest (sha256:1d511631…). A protocol-27 captive core stops exporting at the vote (the task 0367 incident).
  • CAP-85 externally managed executables. A contract may run code set in another contract's storage. The reference is stored on the contract (executable_owner_id, executable_tag) and the owner's target in contract_executable_refs (owner_id, tag) → wasm_hash, resolved at read time. The contract page shows an "Externally managed" chip, names the owner and tag, and labels the hash "resolved"; the self-upgrade chip stays off for such contracts, since the owner can change their code at any time.
  • No more placeholder rows in soroban_contracts for addresses that were only mentioned; the transaction headline reads the called contract from the archive block instead.
  • Two defects found in review and fixed before deploy: the live upgrade prefetch selected 8 of the struct's 10 columns and silently skipped every contract upgrade (now SELECT ?fields, fail-closed); repair-tier1 reset columns added after its hand-typed list (now * REPLACE, including the pre-existing lp_positions.closed_at_ledger loss).

Balance change on the account page — task 0540

  • The account transaction list shows the signed per-asset change for that account.
  • The index floor is gone: the backfill of the whole ingested range passed its completion gate, so balance_changes is always an array (empty = no change). API contract: optional/nullable → required.
  • A fungible movement whose asset has no assets row renders as plain text instead of linking to a page that answers 404.

Assets list ordered by holders — task 0547

  • /v1/assets browses by holder count instead of the storage key (XLM, USDC, AQUA first).
  • Breaking: the cursor changes to (holder_rank, id); an old cursor is rejected invalid_cursor, so a client paginating across the deploy restarts from page one.

Targeted re-parse carries the pool tables — task 0518

TargetedTables gains the pool families. See below — this lives in backfill-runner.

Router pool reserves from the pool's own storage — task 0374 (PR #459, deployed 2026-09-16)

  • Router-family (Aquarius-style) reserve rows come from each pool's own instance storage, in raw units, instead of the deployment plane's PoolData, which stores stable-pool reserves multiplied by PrecisionMul — six mixed-decimal stable pools had one leg 10× or 10^11× too large.
  • A pool whose reserve layout the parser cannot read now logs an error instead of silently keeping its last reserves.
  • Backfilled after the deploy: the history of the 8 affected pools (774 keys) was re-derived with the merged parser and written on 2026-09-16. The reconciliation at ledger 64,454,699 finds 769 of 770 pools equal to their own storage; the exception is a contract whose code is no longer a pool (task 0325).
  • New integration test pool_reserves_reconciliation in backfill-runner, run on demand with a client certificate.

JS workspace on pnpm — task 0554

  • npm → pnpm across the lockfile, CI workflows, git hooks and infra/Makefile; no runtime change. develop CI is green on it, and the 2026-09-16 Compute deploy ran from a commit that includes it.
  • deploy-production.yml has not run on pnpm yet — the first tag after this merge is its first run.

CODE:ISSUER on the Assets list — task 0534 (not yet deployed)

A fully qualified classic asset pasted into the Assets-list search now opens that asset directly instead of producing an empty filtered list. This is the post-deploy SPA commit 87412cb6.

Bookkeeping and lore-only changes

Lore records for 0210, 0325, 0374, 0460, 0486, 0503, 0512, 0520, 0530, 0538, 0540, 0541, 0542, 0548, 0549, 0550, 0551, 0552, 0553, 0554, 0555, 0556 and 0557; review record, post-vote rollback floor and pnpm build prerequisites in docs/deployment.md. f3d84785 (NFT piece ids on the account page) is reverted by c2f2075c; the pair ships no change.

What the Compute / web / ingestion deploy does not cover

  • SPA commit 87412cb6 landed after the 2026-09-14 web deploy at ba3575e3, and the 2026-09-16 deploy was Compute only; the Assets-list CODE:ISSUER fix needs the next web deploy.
  • backfill-runner is built by hand on the box, so task 0518 and the 0548 repair-tier1 fix reach production only at its next manual build. Its changes after ba3575e3 are tests only (task 0374).
  • The placeholder-row cleanup (ALTER TABLE soroban_contracts DELETE WHERE wasm_uploaded_at_ledger = 0) is an operator step, run after the new indexer is live — done on 2026-09-14.

Deploy order / prerequisites

  • Schema already on production: both soroban_contracts columns (DEFAULT NULL) and contract_executable_refs were created on 2026-09-11. Task 0374 changes no schema. Nothing to run first.
  • API before SPA (task 0540): Compute was deployed first, then the SPA.
  • Galexie after Compute was verified; the restart took 56 minutes to a caught-up tip, ledgers contiguous across the gap.
  • Task 0374 backfill: done on 2026-09-16, after its Compute deploy (writer first).
  • Rollback floor after the vote: no Compute build older than 840f2b58 (the stellar-xdr 28 bump), and no Galexie image older than 28.0.1. Since task 0374, no Compute build older than 9b0f05b6, or router stable-pool reserves are written multiplied by PrecisionMul again.
  • First tag on pnpm: watch the install and build steps of deploy-production.yml (task 0554).

Issues

…ot need

Rollout steps 2-4 shipped on tag production-2026.09.07-1 (098bef9), one
combined window with the pool adapters. L0 = 64 317 019, so the backfill
range is 50 457 424 .. 64 317 019.

The window ran with no indexer pause and no read downtime, reversing the
order the rollout note prescribed. Two findings drove it: the deployed API
read net_settled inside the transaction-list aggregate, so dropping the
column before the deploy would have answered 500 on every transaction list;
and the driver rejects a table column missing from the row struct only when
it has no default. Giving the four now-unwritten columns DEFAULT NULL is a
metadata-only ALTER, which made old and new writers simultaneously valid.
The DROPs move to after the backfills, which keeps a rollback free.

Also recorded against 0374 and the archived 0518: the pool write path is on
production, its backfills and read half are not.
The coverage gate the rollout runs on the full range was run first against
the ledgers the new indexer had already written, from L0 onward: 1 546 984
token events against 1 546 984 distinct edges, exact. Both sides counted by
their own key so unmerged duplicates cannot flatter either.

Worth having before the historical pass starts: a disagreement here would
have repeated itself across 13.9 M ledgers.
A full-range `--only` run now writes liquidity_pools, pool_state_changes and
pool_instance_state, so the pool families' registry and reserve history ride
the 0540 descent instead of owing a second one plus a ~40 GB targeted fetch
for the config family. The registry gains from it: the re-parse puts its rows
through the live two-stage gate, which an in-DB generator structurally cannot
do, and it fills legs on the 86% of classic pools that lack them (45 396 of
52 800, measured).

None of the three carries a Tier-1 MIN column, so repair-tier1 stays unowed.

Two costs the flag does carry, measured rather than assumed. The registry
takes one row per (pool, ledger) change, so a full-range run inserts on the
order of 326 M rows into a table that holds 53 k after collapse; they do
collapse, the table is 3 MiB, but the insert and merge work runs for days
beside the main stream. And a re-emitted row ties on last_updated_ledger with
the one the original ingest wrote; a tie resolves to the last inserted row at
merge time, so the backfill wins, but until that merge both rows are live and
a read picks between them arbitrarily. The tie window opens as soon as a
worker passes a pool's last-change ledger, not at the end of the run, so the
OPTIMIZE FINAL on liquidity_pools belongs on a cadence during the run, not
only after it. A test pins the tie direction and fails if it ever reverses.

Also corrects the config family's population in the runbook (14 to 20 across
six factory deployments) and derives the activity-ledger list from what the
registry pass writes rather than from a constant.

Refs #405
The targeted re-parse now writes liquidity_pools, so the full-range run emits
a registry row per changed ledger with legs filled — which is step 2, for all
52 800 classic pools. Coverage is total: no classic pool has its last change
below the ingest floor (oldest 50 458 737, measured). Step 1 is moot, since
the Rust job computes the surrogate natively and the hash-equivalence question
never has to be settled.

Records the step-2 acceptance check, which can only run while both
representations exist, and the trap inside it: assets stores no asset_type 2
at all — alphanum4 and alphanum12 collapse into type 1 — while the legacy pair
column keeps the raw XDR distinction. Compared without normalising that, the
check reports a 52% failure that is not real. Run early against the 7 509
classic pools the live writer had already migrated: 7 509 consistent, 0
mismatched.
Adds the `Balance change` column to the account detail page: the signed
per-asset net for the account whose page it is, read from the per-transfer
`asset_transfers` index that shipped with the write half.

Account-relative by construction, which is why the field lives on
`AccountTransactionItem` alone. The retired `net_settled` aggregate carried
neither direction nor an account, so an inbound and an outbound transfer
rendered identically; a balance change with nobody to ask the question is
what made it mislead most on the global list.

Three states the wire keeps apart, because an absent measurement must never
render as a real value: `null` = below the index floor, not measured;
`[]` = measured, no token moved; non-empty = these assets moved. The floor
is a const, so a transaction below it is never even queried.

Assets come back in chain order, never ranked: they carry different decimals
and different prices, and no price exists anywhere here, so ordering by
amount would compare quantities that are not comparable.

A non-fungible movement has no amount by nature and is counted in pieces, one
entry per piece with its own token id, so every NFT is listed and linked
separately. The id comes from `nft_ownership` joined on the new owner, and is
used only when the set size matches the number of pieces the account moved --
otherwise the entry collapses without an id rather than name a piece we would
be guessing at.

Fees are not included: `asset_transfers` carries token movements only, and
40.2% of transactions are fee-bumps whose payer is the envelope's fee source,
which no column stores. The `Fee` column beside it carries them.

Read path measured on production before exposure (0243/0386 were both
read-shape outages): 20 ms / 13 824 rows for the signed sum, 44 ms / 268 k
for the asset identity beside it.

Refs #405
…view

Task 0540's `Balance change` column links a non-fungible movement, and
landing on `/nfts?contract={C…}` was rejected on sight — a list with a
56-character contract id typed into its search box is not where someone
clicking an NFT expects to be. 0540 now names the piece and links to it
wherever the ownership rows allow, so the column reaches this view only in
the cases it cannot: a set it could not prove, or a collection still
quarantined in `nft_ownership_pending`.

That narrows the exposure without fixing the view, and ties this task's
second half to its first: the same pipeline gaps that make the link
half-dead are what push a movement into the anonymous destination.
…ge key

The browse order is the `assets` primary key — alphabetical inside each type,
which tells a reader nothing. It was never chosen as a browse order; it fell
out of choosing a cheap keyset when task 0364 made the read two-phase, and no
column in the table is sortable, so it is the only order anyone ever sees.

`balance_aggregates` already holds a holder count per asset, and driving the
page off it measures 11 ms / 448 k rows — cheaper than the asset-identity read
the account page already pays. Filters still owe the join to `assets`, at
174 ms.

Records the dedup trap the first measurement hit: joining `assets` without
`LIMIT 1 BY id` returns native XLM six times, the table holding 984 k physical
rows for ~452 k assets.
…-xdr 27→28

Pubnet votes to protocol 28 on 2026-09-16 17:00 UTC. Two independent
failures land at that instant if untouched: the digest-pinned Galexie
image still ships a proto-27 captive core, and the workspace still pins
stellar-xdr 27, which cannot decode the new CAP-83 / CAP-85 union arms.

Same pair as 0367/0368 on the 26→27 upgrade, both of which were reactive
(16 h silent ingestion stall, then a 7.5k-message DLQ). This one is
planned ahead of the vote, which 0367's own future-work list called for.

Refs #367 #368
…e key

The browse order was the `assets` primary key, so the list read alphabetically
inside each type. That was never chosen as an order — it fell out of picking a
cheap keyset when task 0364 made the read two-phase — and since no column here
is sortable it is the only order a reader ever sees. The list now opens on XLM,
USDC and AQUA instead of on six codeless Soroban contracts.

`balance_aggregates` is joined LEFT: 4 222 assets have no aggregate row and an
inner join would delete them from the list. Their rank folds to -1, below a
measured zero, because "no data" and "no holders" are different statements.

Cost, measured on the statement the API issues, before shipping: 61 ms /
1.02 M rows per page, against a near-free key-prefix walk before — ordering on
a joined column reads all of `assets`. If the read quota ever binds, the
cheaper shape is a driver over `balance_aggregates` itself (measured 11 ms),
not a return to the alphabet.

Task 0485's requirement survives on different grounds: XLM still opens the
list, now because it has the most holders rather than because an empty asset
code sorts first. Its test is rewritten to say so.

`AssetRow.issuer_id` goes with the old cursor — it was read only to mint it.

BREAKING CHANGE: the `/v1/assets` cursor changes from the identity 4-tuple to
(holder_rank, id). A keyset cursor is the sort key, so it could not stay: a
comparator over columns the walk no longer orders by skips and repeats rows.
A cursor minted under the old order no longer deserializes and is rejected
`invalid_cursor` (ADR 0008 fail-clean); clients paginating across the deploy
restart from page one.
Raised on review: ordering by holder count means the keyset resumes on a value
that changes by itself, which the old identity 4-tuple could not. A count that
crosses the cursor's value between two pages takes its row with it — upward it
is missed, downward it is seen twice.

Measured rather than argued. `balance_aggregates` refreshes every two minutes
(full recompute, atomic exchange), so one query is always consistent and two
pages minutes apart may straddle snapshots. Sampling the top 60 rows twice, six
minutes apart: five assets changed their holder count, zero positions
reordered. Near the top the gaps are enormous, so a small drift moves nothing;
the tail is where it would bite (335 735 of 452 599 assets hold 1-3 holders)
and nobody pages that far.

Not a new property for this system either: `/accounts` already walks
`last_seen_ledger`, which moves on every transaction an account makes.

Records the options weighed and why accepting won over bucketing the key,
pinning an aggregate version, slowing the refresh, or OFFSET — and names what
would reopen it: a report of rows going missing while paging, not the theory.
Deployed 2026-09-08. Verified from the surface that changed rather than from a
row count: account GCKBNEKI shows all three states at once — `+1 NFT #44 XLEND`
on ledger 64 320 740 (whose other leg is the -4 681 USDC deposit), a MEASURED
`0` on a Manage Sell Offer, and signed amounts in US grouping.

A `Clawback` renders as an outflow (-1 436.3560918 ICE). That verb had never
been exercised against live rows before the deploy — the live window since L0
carried none.

Seven criteria stay open and all belong to the write half: the historical
backfill and its coverage gates.
…oduction

All seven acceptance criteria met, status completed, moved to archive.

The live list reads XLM 9 950 045, TXT 804 336, USDC 684 236, AQUA 129 397 —
twenty rows strictly descending, native appearing once.

The deploy also produced the decision record's own evidence, unplanned.
Against a measurement taken an hour earlier: XLM +69, USDC +51, SHX +7,
DOGET -1. Five counters moved, one downward, and the order did not change —
which is the property the "accept the mutable key" decision rests on, now
observed rather than argued.
`BalanceChange.asset` carried a bespoke token's contract StrKey whenever
`soroban_contracts` had a row for it — which it has for every deployed
contract. But `/assets/{id}` hydrates the key `(3, '', 0, surrogate)` out of
`assets`, so a token nobody registered there answers 404. Two different
questions were being read as one, and the cell drew a live-looking link to a
page that does not exist.

A FUNGIBLE movement whose asset has no `assets` row now emits an EMPTY
identity, which the cell already renders as plain text. A NON-FUNGIBLE entry
keeps its StrKey: its destination is the NFT pages, which are keyed on the
contract and answer for collections `assets` has never heard of — the reason
this read joins that table LEFT in the first place.

Measured on production: 4 of 51 421 fungible assets in a historical partition,
0 of 8 643 in the live one. Nothing renders wrong today; it would have
surfaced as the backfill lowers the index floor. Verified against production
rows that the two sampled unregistered contracts resolve
`resolves_on_asset_page = false` while USDC resolves true.

Found by a systematic sweep for places where two parts of the system answer
the same question differently; the rest of that sweep's findings are recorded
in tasks 0542 and 0376.
A systematic sweep for contradictions — places where two parts of the system
answer the same question differently on real data — showed that this task's
original framing had the symptom, not the disease. The shapes are not the
problem; the problem is that nothing owns the answer.

"Is this non-fungible?" is decided in four places with four rules, two of them
returning OPPOSITE verdicts on an i128 payload. "Which address kind counts?" is
decided in four more, one of which is the same rule re-spelled inline next to
its own helper. `nft.rs` shares no parsing with the path the other two decoders
use, which is why it credits a mint's admin where they credit its recipient.

Records the measured exposure of each: 136 movements where the decoder and the
classifier disagree, 339 of 1 089 NFT owners that are contracts the API cannot
render, 88 tokens whose ownership order inside a ledger is undecidable, and two
contradictions that are real in code with zero exposure today. Also records the
five axes that came out CONSISTENT, so nobody re-derives them.

The task now owns one definition of a token movement in `domain`, one rule for
which address kinds count, and the nft.rs admin-shape fix with the historical
re-derivation it implies. Absorbs 0409, which is the same disease in another
table. 0512, 0392, 0424, 0376 and 0486 keep their own outcomes and read this
task's vocabulary.

Also re-measures 0376's contract-owner gap from the current tables.
The schema recommends `LIMIT 1 BY id` as the cheap dedup for resolving a
surrogate to its StrKey, and that is exact: every version of a row carries an
identical `contract_id` (task 0344). Nothing said the rule stops there.

It does. The stub rows this table documents are real rows whose
`contract_type` / `wasm_hash` / `deployer_id` are NULL, so "pick any one
version" can pick a stub. Measured on production: the largest NFT collection
holds three rows — one complete with `contract_type = 2` and two stubs — where
`FINAL` answers 2 and `LIMIT 1 BY id` answers NULL.

No current reader is wrong; the contracts endpoint uses FINAL. Recorded so the
next one is not, and because a classifier verdict that reads as "unknown"
depending on the dedup is exactly the kind of disagreement task 0542 exists to
remove.
…ster

Two corrections to yesterday's contradiction sweep.

The enum is `Token = 0, Other = 1, Nft = 2, Fungible = 3`. The sweep read `1`
as Fungible, so it reported "the decoder says non-fungible, the classifier says
Fungible" where the truth is "the classifier says Other" — it recognised
nothing rather than disagreeing. Weaker as a contradiction, stronger as
evidence for 0512. Re-measured with the right values: 410 non-fungible
movements in 12 collections the classifier calls Nft, 136 in 3 it calls Other,
and asset-row coverage for Fungible contracts is 4 423 of 4 423, complete.

The second correction matters more. 0540 stopped the cell linking to an asset
page that answers 404, and stopped there. Tracing why the row is missing: the
rule that creates an `assets` row for a bespoke token fires on the classifier's
`Fungible` verdict and is complete for it. The four dead links are contracts
the classifier calls `Other` — each of which emits `{amount}` transfers, so the
chain has already shown they are fungible tokens. The registry keys on a guess
at the WASM's function names and never consults what the contract did.

So the fix shipped is a display-level mitigation of a classification gap, and
is recorded here as one. The fundamental rule belongs to this task: register an
asset from the evidence, not from a name match. It is self-healing for history
once the backfill lands, and it retires the `resolves_on_asset_page` flag that
exists only to route around the gap.
`resolves_on_asset_page` reads like a considered rule. It is not one: it routes
around a gap it does not fix.

An `assets` row for a bespoke token is created from the classifier's `Fungible`
verdict, and that verdict is a guess at the WASM's function names. A contract
that demonstrably moves fungible amounts but whose code is named unusually is
filed as `Other`, never gets a row, and so has no asset page — which is why the
link had to be withheld at all. Registering an asset from the EVIDENCE (it
moved an amount) rather than from the name retires the gap and this flag with
it; task 0542 owns that.

Comment only. Recorded here so the next reader does not inherit a half-measure
as the intended shape.
A worktree whose node_modules is a symlink to the main checkout resolves
every npm workspace package to MAIN's source tree: npm materialises
node_modules/@rumblefish/* as symlinks relative to node_modules' own
location, so ../../libs/ui lands in main. The worktree then typechecks
against whatever branch main is parked on, producing TypeScript errors in
files the branch never touched and blocking every commit, lore-only ones
included. tsconfig.base.json declares no paths mapping, so node_modules is
the only resolution route and nothing else catches it.

The worktree-hooks skill recommended that symlink, which is why the failure
kept recurring. It now recommends a real directory and says why.

tools/scripts/worktree-node-modules.sh provisions one from main's via APFS
copy-on-write (cp -Rpc): measured 20 s and 30 MB against npm ci's minutes
and 1 GB, with no shared blocks to write through. It replaces an existing
symlink, falls back to npm ci when the lockfile differs from main's, and
verifies that a workspace package resolves inside the worktree before
returning.
The claim rested on two 500 k-ledger windows. Counted across every partition
asset_transfers holds, 27 movements in 2 collections are an i128 token id
stored as an amount: both collections are in nft_ownership, the recorded
amounts are their sequential piece numbers, and neither has a single
amount IS NULL row. The column would print a piece number as a quantity.
It is masked only because those collections have no assets row either, so
the link flag refuses the link — the right answer for the wrong reason.
The "four Other contracts" was measured on one partition and conflated two
defects. Across every backfilled partition: 8 contracts, 45 movements, 45
transactions out of ~1.5 bn — and the count is a moving target, reading 36
an hour earlier as more partitions landed.

27 of the 45 are i128 token ids misread as amounts by two Nft-classified
collections, which step 6 owns; 18 are genuine fungible tokens from six
Other-classified contracts, which the registry-by-name gap owns. Of those
six, two are unarguably fungible and four emit only 0 and 1, so the evidence
rule alone cannot classify them — they need step 6's event-spec evidence
too. Resolving the ids first keeps those collections out of the registry as
fungible candidates.
The entry-state numbers here were taken from the quarantine queue, which can
say how deep the residual is but not how much of it MOVES. Measured from the
other end while working on task 0374: `asset_transfers`, full table, no
sampling — 1,352,496,561 rows, 143,782 distinct assets, of which 28 have no
`assets` row at all, together 563 rows.

`assets` only gets a row for a contract classified `Fungible`, so an orphan
there is exactly a contract this task's discriminator has not reached. Split by
movement shape rather than by verdict, which is what the queue cannot show:
15 emit only non-fungible movements, 13 emit only movements carrying an amount,
none are mixed. Three of the non-fungible emitters are absent from the NFT
registry — the launchpad-template class seen from the traffic side, the two
busiest carrying 88 and 44 rows.

Nothing is invisible: every one of the 28 resolves to an address through
`soroban_contracts`, and the value read LEFT-joins `assets` deliberately, so
those movements render with an address instead of a code rather than
disappearing. The residual costs a name, not a row.

No task filed — this is this task's own subject from a second angle. Recorded
so the discriminator's acceptance can be checked against traffic (does the
orphan count fall?) and not only against queue depth.
…h no sender

First reject measurement from the 0540 historical pass rather than a live
window: over ~2.08 M ledgers, 1 108 rejects, every one of them
unrecognised_topics, zero in the other four causes. 532 per million against
the ~300 per million measured on a recent window — same order, era not
regression.

The whole count is one contract family, 21 emitters. Their event carries the
token id in a topic and the recipient in the data, with no sender at all —
while the contract's own interface declares transfer(from, to, token_id). So
the sender is in the invocation and absent from the event, and no topic-shape
rule can complete the edge from the event alone. That is a question for the
one definition of a token movement, not a missing row in a shape table.

Also records why the classifier answers Other on them (none exposes a
discriminator function name; token_uri appears only as a mint parameter, which
a substring test misreads), that they are absent from nfts and nearly absent
from nfts_pending, and that emitter_not_sac has stayed at 0 across the history
swept so far — the first historical evidence for the gate 0540 added.
VALUE_FLOW_FLOOR_LEDGER ships at the deploy ledger and renders "not indexed"
below it, because an empty cell there would read as "nothing moved". Records
what lowering it is gated on, that it moves in three places at once (the
constant, the test that pins it, the frontend gate), and that it stops for good
at the ingest floor.

Also records why an intermediate drop is worth taking before the whole range
lands: the workers advance from the bottom of their own ranges, so the newest
ledgers arrive last, and a separate worker over the most recent window fills
that slice in hours. The overlap it creates costs nothing — the write is
idempotent under the row key, proven on a deliberate re-run.
The 27 misread movements are two decoders reading the same bytes and
disagreeing, not one decoder failing. The event is a CAP-67 [mint, to] with a
bare i128 payload; nft.rs writes it to nft_ownership.token_id and
asset_transfers writes it as an amount. Joined on (contract, ledger) all 27
pairs carry the same number under two meanings — each table internally
consistent, the contradiction visible only across them. That is the case this
task's one-definition rule exists to make unrepresentable.

It also bounds 0512's tier-4 SCVal discriminator, whose "no overlap on the
scalar types" was measured on transfer events. These are mint events, the only
signature either collection emits, and they are Nft-verdict with a bare i128.
The tier holds as stated but its evidence needs re-measuring per verb — SEP-50
specifies no mint, so that is where the ambiguity sits.

Bounded: 0 of 136 Nft-verdict contracts have an assets row, so no larger
population is hidden behind the anti-join.
The earlier figure enumerated partitions from a system.parts snapshot and
anti-joined them one at a time, so the three partitions the backfill filled
while the measurement ran were never visited. A full-table anti-join instead:
73 contracts move value with no assets row, 48 of them emitting amounts across
118 movements, 25 emitting only non-fungible ones, none mixed.

The undercount was entirely in the Other half — 6 of 46 contracts, 18 of 91
movements. The Nft half was exact at 2 contracts and 27 movements, because
those collections are confined to partitions the stale snapshot happened to
include.

Records the method error as the durable part: against a table being written, a
partition list captured up front is stale before the query ends. Anti-join the
table, not a remembered list of its parts. Every count here is a mid-backfill
snapshot and rises as partitions land, so cite the method, never the number.
A fourth backfill worker filled 64 128 000 .. 64 317 019 out of order — the
archive partition boundary below the fourteen-day mark, ahead of the three
workers walking up from the ingest floor. The newest ledgers are the ones an
account page is most often asked about and they would otherwise have arrived
last, so the floor moves to the start of that window and the column stops
saying "not indexed" for the most recent two weeks.

Verified before the change that the upper edge leaves no hole: the ledgers
table is continuous above the deploy ledger (28 161 rows, none missing) and no
ledger above it carries a token event without its edges, so the worker's range
meets the live block at exactly one overlapping ledger.

The constant's doc now states the rule the value has to obey — it may only move
to a range the backfill has provably covered, because lowering it ahead of the
data turns "not indexed" into a measured zero, which is the one failure it
exists to prevent. Two places move together, the constant and its pinning test;
the frontend holds no threshold of its own and renders on a null from the API.
…s it

Every consumer in this task guesses an event's meaning from its label because
nothing tells it what the event means. That absence is not real: a contract
declares its own events in its WASM, and the XDR library this repo already
depends on carries the type — name, which symbols lead the topics, and per
parameter its type and whether it sits in the topics or the data.

That last field is the question this task keeps running into. It is declared by
the contract, stored in code addressed by its own hash, so it is immutable for
that version and matches the build that emitted the event. It is not a claim
like a topic symbol; it is the published interface.

contract.rs reads the spec section and keeps FunctionV0 only, so every event
declaration on the chain passes through this repo and is dropped on one line.
The ingestion half already exists. This reframes the steps rather than adding
one: the shape inventory becomes declared-versus-emitted, the i128 ambiguity
resolves from the declaration, and the missing-sender family is readable as
soon as its declaration says where the sender sits.

Also records the same disease on the trace view, which colours a row from the
first topic symbol with no emitter check, and a structural hole beside it — the
tree is rebuilt from fn_call labels without separating host diagnostics from
contract events, so a frame could be injected. Measured at zero across
309 355 024 contract events in one partition.
… transaction

A fresh-eyes review of a real multi-operation claimable-balance transaction,
read against the raw events rather than the page.

The data is right and the presentation misleads in four places. The escrow
sentence item 16 shipped names the claimants, who receive nothing at creation
time, while the actual recipient — a new ledger entry with its own address —
is named nowhere on the card. Amount and asset are printed twice per
operation with nothing saying they are one fact from two sources. A canonical
asset id matches no strkey pattern, so it stays plain text where every address
beside it is a copyable link; item 11 recorded the cause and deferred the
resolution to 0456, and this is its presentation half. And a card gives no
sense of scale: one of eighty-five identical operations renders alone while
the account page sums the transaction, with nothing reconciling the two
numbers.

Extends this task rather than filing a new one: it already owns the
transaction-detail debts and item 16 is the same sentence.
The figures recorded here from the 0374 side were taken with a sound method —
full-table anti-join, no sampling, no partition list — against a table that was
growing underneath it. Task 0542 asked the same question hours later and read 73
contracts / 1,026 movements against the 28 / 563 recorded above.

Nothing about the method was wrong; the omission was saying that the number
moves. 0542 states the rule and this correction carries it here: cite the
method, never the number. A count against `asset_transfers` rises until the
backfill finishes, and a figure without its timestamp reads as a steady state.

The shape is what this task needs and it survives unchanged: orphans are
contracts the discriminator has not reached, they never mix fungible with
non-fungible movements, and every one still resolves to an address — the
residual costs a name, not a row. Live figures are pointed at 0542, which owns
that measurement.
Pubnet votes to protocol 28 on 2026-09-16 17:00 UTC. stellar-xdr 27 cannot
decode the new union arms, and deserialize_batch fails a whole batch rather
than one ledger — the shape of the protocol-27 incident, where the doorbell
redelivered until it dead-lettered. Done before the vote this time.

The compiler named four break sites, not the three the task predicted:

  scval.rs        ScVal::ExecutableTag        (SCV_EXECUTABLE_TAG = 22)
  scval.rs        ContractExecutable::ExternalRef
  operation.rs    ContractExecutable::ExternalRef
  invocation.rs   ContractExecutable::ExternalRef

The fourth comes from an arm no CAP write-up mentions; it was found by diffing
the .x sources between v27.0 and v28.0. Nothing else in the workspace broke.

External refs render as their own type, never as wasm with a borrowed or
zeroed hash — a consumer has to be able to tell a contract that carries its
own code from one that points at another contract's.

A decode gate covers all of it, including a real protocol-28 ledger pulled from
testnet (which voted three weeks ahead of pubnet) rather than a hand-built
value. Testnet carried no empty-tx-set ledger to capture, so CAP-83 is covered
by a constructed round-trip instead.

Also records a gap the compiler cannot see: an executable_update event whose
new executable is an external ref yields no hash, so the stored wasm_hash stays
at the pre-upgrade value. Pinned by a test and documented on the function; the
fix needs a decision on what the column should hold for a fleet member.

Contains a pure move as well — the inline test modules of operation.rs (859
lines) and invocation.rs (1035) went to sibling files, per the project rule and
the form event.rs already uses. Read it with --color-moved=zebra to separate
the move from the seven lines of real change.
karolko9 and others added 28 commits September 15, 2026 09:59
Decision C′ needed proof that pool storage holds raw reserves in every
contract version pools have run, not only today's. Of 58 versions, 55 were
decoded from a raw ledger where reserves changed and all wrote them to the
pool's own instance; the other 3 saw no reserve change. Three key layouts
exist in all history, and across 23 samples of the mixed-decimal pools
storage is raw while the plane is scaled, with no exception.

Records the parser rules this implies and the rollout: parser first, then a
list pass over the 726 ledgers of the eight affected pools.

Refs #405
Router-family reserve rows were copied from the deployment plane's PoolData
entry. That entry is the input to swap quotes: a stable pool writes its
reserves multiplied by PrecisionMul there, so the six non-empty stable pools
whose tokens differ in decimals stored one leg 10x or 10^11x too large.

Reserves now come from the pool's own instance storage, where the pool is
the ledger-authenticated owner and the units are raw. Measured before the
switch: storage equals the plane row for 508 of 514 pools, three key layouts
cover all 58 code versions pools have run, and storage holds raw units in
every sample of the affected pools.

A row stages only when an instance write moved the reserves against its
pre-image, or created the pool, because reward and config calls rewrite the
instance too. The plane stays a cross-check and logs a warning when it stops
equalling storage times PrecisionMul.

Re-decoding 874 real ledgers through the new code: every row of every other
pool is identical to production, the affected pools' rows equal the old
value divided by the multiplier and match decoded chain storage 23 of 23,
and the only rows no longer produced repeat the pool's previous reserves.
A list-driven re-decode test produces the history of the affected pools for
the backfill described in docs/backfills.md.

Refs #405
The reserve-source change is up for review as PR #459. Records what changed,
the before/after re-decode of 874 real ledgers (every other pool's rows
identical to production, the affected pools' rows equal to decoded chain
storage), the two local test failures that reproduce on clean develop, and
the writer-first rollout with the backfill payload's checksum.

Refs #405
Promote 0210 and scope its first stage to claimable balances. The stock
comes from ClaimableBalanceEntry changes in a dedicated table shaped like
balances, seeded from a checkpoint and joined to the snapshot
reconciliation (ADR 0057). This replaces the 2026-09-14 plan of seed plus
net B edges from asset_transfers, which cannot correct a dropped edge.

Amend ADR 0056 with the boundary rule: a holding leaves balances only
under mass churn and a never-reused id. Claimable balances qualify;
classic pool reserves go into balances in the LP stage.
Scope, pre-implementation findings and the decisions taken: strict
linker, per-worktree pnpm install replacing the node_modules symlink,
npx call sites moved to pnpm exec, pnpm pinned to 10.34.5 with hash.
Worktrees shared the main checkout's node_modules through a symlink, so
workspace packages resolved to whatever branch the main checkout sat on.
pnpm links each checkout from one store, giving every worktree its own
tree cheaply.

- pnpm 10.34.5 pinned with its hash in packageManager; workspace moved to
  pnpm-workspace.yaml with a build-script allow-list; lockfile imported,
  resolved versions preserved (web build and cdk synth byte-identical)
- CI installs via pnpm/action-setup pinned to a commit SHA; the
  deploy-production node_modules cache is replaced by the store cache
- npx call sites in hooks, Makefile, codegen, e2e and docs use pnpm
- post-checkout runs pnpm install instead of symlinking; the clone script
  is removed
- husky hooks hand over to the checkout's own script, since worktrees
  start the main checkout's copy

Task 0532 is absorbed and archived as superseded.
Record the npm vs pnpm CI step timings and archive the task. Remove the
worktree-hooks skill: after the migration its only durable content was
the rule never to bypass hooks, now stated in CLAUDE.md.
Docs-only pushes pay the full CI and pre-push clippy cost, the TS job
checks every project uncached, and dev builds carry full dependency
debug info. Records measurements and why a shared target dir is rejected.
A pool whose code is replaced stops being a pool while its registry row and
last reserves keep presenting it as one. The earlier decision derived that at
read time, which makes every reader re-apply the rule. The fact changes only
when the code does, so it is decided once, at the upgrade, and stored: the
detection point, the append-only table, the ClickHouse-only history backfill
and the read side are written out.
The research question — what separates "nothing to fetch" from "the fetch
fails" — was answered by measurement: every empty row was re-fetched. Four
defects share this path and ship together: the enrichment producer ignores the
NFT verdict staging already made, a failure is stored as "nothing to fetch"
and never retried, two URI shapes cannot be fetched at all, and the
token-metadata warning fires on contracts that are not tokens.

The file rename rode along with the previous commit.
A code version that renames its reserve keys would stop a pool's snapshots
with no trace: the instance parses, carries no reserve key we know, and no row
is staged. Two refusals now log an error, the way every other decoder refusal
does — a router instance with no known reserve key while its plane row still
shows reserves, and a pair instance carrying half a reserve pair.

The plane cross-check goes: on 1,565 ledgers it fired 17 times and was wrong
every time — 13 mixed-decimal stable pools whose instance predates the
PrecisionMul key, 4 concentrated pools, which do write plane rows after all.
The plane is kept only to notice the unreadable layout above.
Reads each registered pool's reserve entries with getLedgerEntries, decodes
them from XDR here rather than through the parser under test, bounds
ClickHouse at the ledger the RPC answered at, and fails on any pool whose
newest row differs. Skips without a client certificate, so cargo test stays
offline; it finds the certificate where api --bin local does.

First production run: 762 of 769 pools equal. The seven are the six
mixed-decimal stable pools awaiting the backfill and the pool whose code was
replaced (task 0325). The list pass now prints refused writes, the way a
backfill should.
Both sides appended to the 0374 task file: develop carries the per-version
measurement and the PR write-up, this branch the wide differential and today's
two decisions. Kept in time order, nothing dropped. Two sentences from the
develop side were corrected where today's commits made them stale: the plane
cross-check is gone, and the post-backfill check is now the reconciliation
test.
The worktree's node_modules was a symlink to the main checkout, so the husky
hooks ran nothing here and the prettier pass never happened. Installed the
worktree's own dependencies; the formatter now runs on commit.
The contract page says whether code can be replaced, not that it was, how
often or when. 486 of 4,450 bespoke token contracts have had their code
replaced, and the data is already ingested as executable_update events. The
task starts by settling what a version is, since our count and another
explorer's disagree for the same contract.
…rom-pool-storage

fix(0374): read router pool reserves from the pool's own storage
Compute deployed from the merge commit; the 8 mixed-decimal stable pools'
history re-derived with the merged parser, checked key by key against
production, loaded through a staging table and merged. Production now equals
the payload exactly, and 769 of 770 pools equal their own storage; the one
left is the pool whose code was replaced (task 0325).
Stellar's event identity, read from stellar-rpc's source, is (ledger, tx
application order, operation, event within operation) with sentinels for fee
events, and it covers every row soroban_events stores. The schema's two
reasons for our flat event_index do not hold. Decided to make the canonical
location the table's sort key, sourced from the existing side table. The task
0538 measurements are refreshed: transaction_id columns are 26% of the
database.
The identity-column research becomes the epic for moving every table
and list query from the surrogate transaction id to the canonical
position (ledger, application order, operation index, event position).

Records the two defects with one cause: storage spent on random
64-bit ids, and list pages ordered by hash instead of execution order
(11 list queries across 6 API modules). Adds the ordered steps,
constraints and programme acceptance criteria; the original
measurements stay below.
Exact join over every ingested ledger: every per-operation event has
its operation row, no operation row lacks an event, and the only
events without one are the native fee events at positions 0 and 1.
The side table carries byte-equal duplicates in one ledger range,
which the fill deduplicates on the key.
The canonical id of a fee event depends on its stage, which is not
stored. Position 0 is always the charge; position 1 is the refund,
settled after its own transaction before protocol 23 and after all
transactions from then on. Verified against live getEvents ids and
decoded archive meta across protocols 20 to 27.
This reverts the code part of commit f3d8478.

Replayed over every NFT movement in production asset_transfers, the
previous-owner rule named no piece the owner-group rule did not, and it
hid ids the owner-group rule showed whenever a token's earlier ownership
row is missing (2 account senders, 135 contract senders). The
several-senders-to-one-owner case it targeted does not occur today.

Task 0558 removes the inference by persisting the token id on
asset_transfers. The 0538 epic step recorded in that commit is kept.
@karolko9
karolko9 merged commit 597778b into master Sep 16, 2026
8 checks passed
karolko9 added a commit that referenced this pull request Sep 16, 2026
Verified findings from the review of release PR #453 that touch only
documentation:

- docs/backfills.md links the runnable Gate 7a query; the query in the
  0540 rollout note now counts under FINAL per partition (tested on
  partition 100) instead of a whole-partition uniqExact that exceeds the
  memory cap.
- docs/backups.md warns that a post-restore gap refill does not apply
  contract upgrades; 0548 records the measured sink gap and the shape of
  its fix.
- The XDR parsing overview no longer claims the pool parser reads
  PrecisionMul.
- ADR 0058 marks the original reserve source superseded, corrects the
  concentrated-pool plane claim and records both amendments.
- Stale plan text marked or corrected in 0548 (pnpm codegen command,
  CAP-85 decision closed), 0540 (rollout acceptance, rollout note, research
  note figures and row identity), 0325 (inconsistent counts), 0520 (history
  migration order) and 0530 (acceptance query now catches a missing native
  leg; same result on production today, 52,974 checked / 2 known).

Also closes 0534 after its production verification, records the first
pnpm tag deploy in 0554, and files 0559 for the five review findings that
change behaviour.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants