Conversation
scolear
force-pushed
the
feature/noise-sunset
branch
3 times, most recently
from
September 15, 2026 09:38
b2e1106 to
d279f56
Compare
Decman coordinated every multi-party workflow over a Noise P2P transport. Canton already carries partially signed topology proposals to every member and Daml contracts to every observer. This design replaces Noise with those native paths: a node party and registry contract per node, workflow proposals and acceptances in Daml, co-signing by transaction hash in the synchronizer store, signature rounds for prepared transactions, hash-pinned DAR uploads, and an operator-moved ACS snapshot.
The review found nine blocking defects: a Confirmation-only node party cannot submit commands, the bootstrap exemption would open an unauthenticated write on upgraded nodes, members never pinned the topology operation, an accepted proposal had no serial binding or expiry, acceptance provenance was unspecified, the kicked key had no trusted source, kick quorum ignored the stored namespace threshold, the ACS manifest signer was unbound, and registry publication ignored package vetting. This revision fixes each one and the should-fix items that followed from them.
…tion Decman nodes will coordinate through Canton instead of Noise. This commit adds the pieces every later change builds on: the Daml package decman-coordination-v1 (node registry, workflow proposals and acceptances, submission rounds, ACS manifests) with its test package and pinned release DAR; a node identity stored as a party_credentials row of kind 'node' with GET/PUT /node-identity; and the onledger Rust module with the Daml codec and client, the registry with heartbeat and peer health, proposal projection, synchronizer-store proposal discovery and co-signing by transaction hash, per-kind validation, the observer loop, and the engine API with stub drivers for every kind. Noise code is untouched here; the removal follows once the drivers are in place.
Pinned DAR uploads pass expected_main_package_id, which needs the sha256 of the main .dalf. The zip crate is already in the graph through utoipa-swagger-ui, so this adds no new code to the build.
Every kind driver needs the same key material: the dual-usage
{prefix}-key with its self-signed root NamespaceDelegation, the local
identity the validator compares against, and the member-key caches
that acceptances heal. Keeping them in one module lets the drivers
stay thin and lets kick attribution read a single trusted source.
The six workflow kinds now run on the on-ledger engine. A run starts as a WorkflowProposal; invitees accept on the ledger; topology changes are proposed into the synchronizer store and co-signed by transaction hash; contract deployments collect prepared-transaction signatures in SubmissionRounds; DAR distribution pins hashes and observes vetting; add-party replication publishes an AcsManifest and takes the snapshot through operator-driven export and import endpoints. With that in place the Noise transport has no caller. This removes the listener, the peer client and server, the heartbeat and ping loops, the Noise key file, the peer public keys and ports, the ten Noise CLI flags, and the hyper-noise, tokio-noise and secp256k1 dependencies. Migration 000021 fails any run that was in flight, renames the coordinator column to the participant id, drops the peer address columns, and adds the proposal columns the projection needs. Peer health is the registry heartbeat. The CLI, the integration harness, and the test crate compile against the new shapes; the coordination phases still carry TODO markers for their rewrite.
A member validates the kick party mapping only after the kick namespace change became effective, so the head owner set already excludes the kicked member while its signing key is still in the mapping. The dual-usage test compared the two sets for equality, so it failed, and the check fell through to the weaker legacy branch. A proposer could then drop a different member's signing key. The member now recognises the same party from either owner set.
The Noise transport was the only user of hyper in decman, and every `http::` path in the crate resolves through actix-web. CI runs cargo-machete, which fails on an unused direct dependency. Removing hyper does not clear the h2 0.3 advisory: actix-http still pulls it. Both rationales now say so.
The operator and developer documents still described a transport that no longer exists. They now describe what the code does: a node party per node, a registry contract that carries peer identity and liveness, workflow proposals and acceptances on the ledger, topology changes co-signed in the synchronizer store, signature rounds for contract deployment, hash-pinned DAR uploads, and an operator-moved ACS snapshot. The deployment guide gains a cutover runbook from 1.8.x and the manual rollback. The KMS guide explains that one key now signs both topology and ledger transactions for a party, and what that means for the key policy. The architecture document keeps a History section so a reader who meets a 1.x deployment knows what it was.
submit_and_wait returns before this node's own create is guaranteed visible to the next active-contract read, so an observer tick that read too early published a second DecmanNode. Peers then saw two entries for one node and read whichever they met first. The reader now retires the older entries instead of reporting them, and an update that names an entry the heartbeat has just replaced waits for the next tick rather than failing the registry stage.
A contracts run never finished. The coordinator waited in
WaitingForAcceptances with both invitees listed as missing, while each
member waited in SignSubmissions for a round the coordinator opens only
after the acceptances arrive. Neither side logged an error, because
neither side had one: both were waiting correctly.
POST /invitations/{cid}/accept records the operator's decision locally
and opens the peer run row. The ledger write was left to each kind's
member driver. Onboarding, add-party, kick and DARs made it. Contracts
and change-threshold never did, so their members published nothing and
both runs deadlocked.
Move the write into engine::drive, before the first member tick, and let
a kind opt out through KindDriver::member_publishes_own_acceptance. The
default is false, so a new kind is correct without knowing the rule.
Onboarding, add-party and kick opt out: their acceptance carries key
material a member step generates, and they still write it at that step.
DARs drops its own copy and uses the shared path, which the DARs phase
already covers end to end.
The write is idempotent twice over: against the tick snapshot, and
against a cid pinned on the run row. submit_and_wait can return before
the create reaches the next ACS read, and a second acceptance from one
acceptor fails the coordinator closed.
The acceptance deadlock left no trace in CI. run.sh exports its quiet preset to the node processes too, so every node recorded WARN and above and the whole coordination flow went unlogged. The uploaded node logs, the only evidence a CI failure leaves, said nothing at all. Give the nodes their own level through DECPM_NODE_RUST_LOG, defaulting to INFO with the on-ledger module at DEBUG. The runner output is unchanged. Log the step each run sits in once per tick, and log the member branch that waits for a SubmissionRound, which returned silently. Also stop the one-shot coordination DAR upload from reporting a false failure. spawn_supervised treats any return as death, so every node logged "may stay unvetted" at ERROR on a clean startup. spawn_once reports only a panic.
The coordinator waited for every invitee before it prepared a round. A contracts run needs party_signing_keys.threshold signatures, not one per invitee, so a single silent operator held up a party that had the quorum to act. That is the absent-owner hang contracts_quorum_completes covers: on a 2-of-3 party with only one member accepting, the run never finished. Advance once the counted acceptances plus the proposer reach the threshold. The proposer signs its own rounds, so it counts as one. An explicit Decline still fails the run. Silence and refusal are different answers, and an operator who says no should stop the run rather than be outvoted by a quorum that never heard from them.
Three phases still watched for 1.x evidence that the on-ledger model no longer produces. add_party_edge_cases reaped a declined run with a per-instance cancel, because a coordinator task used to keep idling after the decline. No task idles now: the decline fails the row and finishes the WorkflowProposal, so the run is already terminal and the cancel must refuse it. Assert the 409. generate_keys_idempotent and dismiss_failed_cleans_artifacts waited for a workflow_artifacts row. Only a contracts run writes artefacts now; every other kind keeps its progress on the run row. Both gate on the run's step instead, through a new db::workflow_run_step helper.
The add-party coordinator exported the party's ACS, published the AcsManifest, and then waited in AwaitReplication for a snapshot nobody moved. Decman used to relay those bytes between the nodes. It does not any more: the operator downloads the file from the exporting host and uploads it to the joiner (design D9), and the suite has to do the same. The phase now reads the manifest that names the joining participant, downloads the file it points at, uploads it, and checks that the received bytes hash to what the manifest states. Two byte-level HTTP helpers carry the gzip file, because every other fixture helper speaks JSON. A respawned node also kept the runner's quiet log level, so a node went dark after a chaos restart, which is where its own record matters most. It now takes the same DECPM_NODE_RUST_LOG default the shell harness uses.
A kick failed on a member with "its owner key has not yet been resolved". The member held no owner key for its peers, because it never wrote the party cache for that party at all. The member writes that cache in its Complete step. The proposer finishes the WorkflowProposal as soon as the party is effective, and a member that reads the outcome first completes its row through reconcile, which skips Complete. The member then joins the party and keeps no record of who owns which key, so the next kick on that party fails on that node. Write the cache when the member co-signs the party mapping instead. It has validated that membership by then and is about to sign it, so it records what it endorsed, and no completion race can get in front of it. The Complete step keeps its own call; the write is an upsert. G5 watched dec_party_identity, which only parties created before the Canton-native coordination hold. It now watches the dec_party_participant rows that carry a key, which is the material a later kick reads.
The add-party relay read whichever AcsManifest named the joiner and then always downloaded from P1. Every current host exports and publishes its own manifest, within a millisecond of the others, and two snapshots of one ACS are not byte-identical. Picking P2's manifest and P1's bytes made the import refuse the upload, which read as a broken export rather than a broken test. It passed or failed by whichever manifest sorted first. Choose a manifest whose exporter this suite can reach, then download from that host. That is the operator's rule too: carry the file from the host whose manifest the import names.
The snapshot-relay helper landed between the run function's doc comment and the function, so the two doc blocks merged. The phase lost its documentation, and clippy read the merged block as a doc list running into prose. Move the helper below the function it serves.
G7's gate moved to the run row, but its closing assertion still counted dec_party_identity rows, which only a party created before the Canton-native coordination holds. Count the keyed participant rows, the same material G5 watches. The old helper has no callers left.
G9 runs a Dars workflow beside an Onboarding, restarts the coordinator, and waits for both to finish. It accepted the Dars invitations but never uploaded the files, so the coordinator sat in AwaitVetting: DAR bytes do not travel between nodes any more, and the run finishes when the topology shows the packages vetted everywhere. Upload them on both members, the same step distribute_dars already performs.
The offline import re-validates every contract and fails on a package the target participant does not hold — but only after it disconnects from the synchronizer. That is the devnet failure of onboarding a node without its DARs, and the design asks for a check before the disconnect window opens. The check only ran inside the import. The import is the operator's step now, so a joiner that could never succeed sat in SyncAcs looking healthy and waited for a file it would have rejected. Run it in SyncAcs instead. The manifests name the packages the party's contracts need, so the joiner compares them against its own vetted set while it is still connected and nothing has been carried, and fails the run with the missing ids. An empty snapshot still takes the fast path, and a completed import is never undone by a late check.
The exporter read the party's contracts with no ledger credential, so the read failed authentication and every AcsManifest named no packages. The joiner's pre-disconnect check then had nothing to check, and a node without the DARs learned that only when the import failed, after the disconnect window was already open. Pass the party's own credential down to the read. The export helpers take the token, and the add-party engine supplies the one it already uses to prepare submissions. A node that holds no credential for the party still exports; its manifest names no packages, and it says so.
code-coverage.sh runs the suite with the LLVM instrumentation on, and each test process drops a .profraw file beside it. Nothing ignored them, so a `git add` after a coverage run sweeps in hundreds of them.
scolear
force-pushed
the
feature/noise-sunset
branch
from
September 15, 2026 13:28
05d0ccc to
3d7ccd0
Compare
Eleven phases carried a note saying they still described the 1.x transport and needed a rewrite before the suite could run green. The suite runs green now, and each of these phases asserts an outcome the on-ledger model delivers, so the note is stale. The one marker left in check_peer_dars asks for coverage that does not exist yet, which is a different thing.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this changes
Decman nodes no longer talk to each other. Every workflow that used the Noise
transport now runs through Canton itself, and the transport is deleted.
Canton already carries two things between the operators' nodes. The synchronizer
topology store broadcasts partially signed topology transactions to every member.
The ledger delivers a Daml contract to every participant that hosts a stakeholder.
This branch uses both, so decman needs no channel of its own.
A workflow now runs like this:
WorkflowProposalcontract that names the invitees.partially signed proposal.
by transaction hash.
is met.
Why
The Noise transport duplicated work Canton already does, and it cost real
operational weight: a second listen port per node, a static key pair per node, a
peer table of addresses and public keys, a pre-shared key per pair of operators, a
full peer mesh as a precondition of every workflow, and a bespoke wire protocol of
about fifty message codes. A node that could reach its own Canton participant
could still fail to start a workflow because a peer's port was unreachable.
Removing it also closes a gap in the old trust model. A peer used to receive the
proposal bytes from the coordinator and could not see the other members' keys. Now
every member reads the same proposal from its own synchronizer store and compares
it against the on-ledger acceptances of every other member.
Key decisions
Node identity. Each node gets one node party, hosted on its own participant
with Submission permission. It signs the registry entry, the proposals, the
acceptances, the signature rounds and the manifests. It is stored as a
party_credentialsrow of kindnode, andPUT /node-identitysets it.Discovery. Operators exchange one string per peer:
participant_id,node_party_id,name. No addresses, no ports, no keys. Each nodeverifies through the topology store that the node party really is hosted on the
claimed participant before it names that peer anywhere.
Peer health. Each node publishes a
DecmanNodecontract and heartbeats on it.A peer reads as Active, Stale, Unknown or Unvetted. There is no live probe any
more, so the peers table reports the age of the last heartbeat, not a round trip.
Keys. A new party gives each member one key with both the Namespace and the
Protocol usage. Its self-signed root
NamespaceDelegationpublishes the public keyin the topology store, so the proposer can build the party mapping without asking
anyone. Parties created before this change keep their two keys and keep working.
Contract deployment. The proposer prepares each transaction and opens a
SubmissionRound. Members re-hash the bytes, check the transaction acts as theparty they accepted, and sign. The proposer executes once the party's signing
threshold is met. Canton's preparation-time tolerance bounds the window, which is
hours, not minutes.
DAR distribution. The proposer pins each file by SHA-256 and main package id.
Each operator uploads the same file locally, and the run completes when the
topology store shows the package vetted everywhere. DAR bytes never move between
nodes.
ACS replication. Canton online party replication is still alpha and its codecs
are protocol-version
dev, so it is unusable here. Add-party keeps offlinereplication. A current host exports the snapshot and publishes an
AcsManifest;the joining operator downloads and uploads the file through two new endpoints. The
joiner verifies the manifest signer, its hosting, and the activation serial before
it disconnects anything.
Upgrade
This is a breaking change, so the crate version moves to 2.0.0.
Migration
000020adds the node-identity column and the proposal-decision table.Migration
000021fails any run still in flight, rewrites the coordinator columnto a participant id, drops the peer address columns, and adds the proposal columns.
A mixed network is safe in both directions. An old node refuses to start a
workflow when a peer's Noise listener is gone. A new node refuses until every
invitee has vetted the coordination package and published a registry entry. The
runbook is in
docs/DEPLOYMENT_GUIDE.md, and the design records the manualrollback.
How this was tested
Everything below runs green on this branch:
cargo fmt --all -- --checkcargo clippy --all-targets --all-features --no-deps -- -D warningscargo test --workspace --lib— about 300 new tests cover the Daml codec,the mapping builders, the per-kind validation rules, the acceptance-counting
predicates, the quorum arithmetic and the migrations
cargo test -p decman-cli— 77 testsdpm build --all, anddpm testindecman-coordination-test— 30 testsnpm test(41 tests),tsc -b,vite buildThe integration suite passes end to end against a live three-node Canton —
all 33 phases,
1 passed; 0 failed. That covers onboarding, DAR distribution,contract deployment (including a 100-contract run through the submission
rounds), deployment at quorum with one operator absent, cancel, kick,
add-party with the operator-carried ACS snapshot, change-threshold, the
external-party flows, the governance and reward phases, and the whole chaos
block: coordinator and peer restarts mid-flight, coordinator retry, dismiss
cascades, key-generation idempotency across a restart, concurrent kinds
through a restart, cross-node concurrent workflows, and the sibling cancel and
decline cases.
The bugs this shook out
Removing the transport was the easy half. Running the suite against real Canton
found six defects that only a live network shows.
A member never published its acceptance.
POST /invitations/{cid}/acceptrecords the operator's decision locally; each kind's driver made the ledger
write. Onboarding, add-party, kick and DARs did. Contracts and change-threshold
did not, so their coordinators waited in
WaitingForAcceptanceswhile everymember waited to sign. Neither side logged an error, because neither side had
one: both were waiting correctly.
engine::drivenow makes that write forevery kind, and a kind opts out only when its acceptance carries key material
one of its own steps generates. The default is correct, so a new kind cannot
repeat this.
Contracts waited for every invitee. A contracts run needs
party_signing_keys.thresholdsignatures, not one acceptance per invitee, soone silent operator held up a party that had the quorum to act. The coordinator
now advances once the counted acceptances plus itself reach the threshold. An
explicit
Declinestill stops the run: silence and refusal are differentanswers.
A member could join a party and keep no record of who owns which key. The
member wrote the party's key cache in its
Completestep, but the proposerfinishes the
WorkflowProposalas soon as the party is effective, and a memberthat reads that outcome first completes its row through
reconcile, skippingComplete. The next kick on that party then failed on that node. The membernow writes the cache when it co-signs the party mapping, which it has validated
and is about to sign.
A joiner without the party's packages waited instead of stopping. The
offline import re-validates every contract and fails on a missing package, but
only after it disconnects from the synchronizer — the devnet failure of
onboarding a node without its DARs. The check lived inside the import, and the
import is the operator's step now, so a joiner that could never succeed sat in
SyncAcslooking healthy. It now compares the manifests' package ids againstits own vetted set while still connected, before anything is carried.
No manifest ever named a package. The exporter read the party's contracts
with no ledger credential, so the read failed authentication and every
AcsManifestcarried an empty package list, which made the check above inert.The export helpers now take the party's own token, which the add-party engine
already holds. A node with no credential for the party still exports, and says
that its manifest names no packages.
Duplicate registry entries.
submit_and_waitreturns before the node's owncreate reaches the next ACS read, so a node could publish a second
DecmanNode.It now retires the duplicates and tolerates the stale-id race.
The suite also carried assumptions the transport removal invalidated: a
declined run used to need reaping, member progress used to land in
workflow_artifacts, DAR bytes and ACS snapshots used to travel between nodes,and a party's keys used to live in
dec_party_identity. Those phases nowassert the on-ledger behaviour instead.
Known gaps to close before merge
acs::export_to_responsere-exports when it finds no spool file. Two exportsof one ACS are not byte-identical, so those bytes match no published
manifest. The endpoint should publish a manifest for what it serves, or
refuse. Until then an operator must download from the host whose manifest the
import names, which is what the suite does.
behind. The dismiss handler should call
acs::cleanup_spool.check_peer_darshas no scenario asserting that/participants-statusreports a peer
Activeonce its registry heartbeat lands (design D3). Thatcoverage replaced a reachability scenario the transport used to carry, and it
is still missing.
dismiss_failed_cleans_artifactscovers the cleanup cascade against anonboarding run, which writes no artefact rows. A contracts-run variant would
exercise a populated table.
that spike before the mainnet window.
Reviewing this
The diff is large because the transport touched everything, so review it commit
by commit:
docs: add the Noise-sunset coordination designand its revision — thedesign, and the nine defects an adversarial review found in it
feat(onledger): add the Daml coordination package and the Rust foundation—the new package, the node identity, and the module everything else builds on
feat(onledger): add the shared party-key and member-key helpersfeat: coordinate every workflow through Canton and remove Noise— the sixworkflow engines, the deletion, and the migration
fix(onledger)commits — what the live runs foundtest(e2e)commits — the suite moved onto the on-ledger modeldocs/NOISE_SUNSET_DESIGN.mdis the specification the code follows, andcrates/decman/src/onledger/README.mddocuments the new module's API.🤖 Generated with Claude Code