Skip to content

Coordinate through Canton and remove the Noise transport - #451

Draft
scolear wants to merge 24 commits into
mainfrom
feature/noise-sunset
Draft

scolear wants to merge 24 commits into
mainfrom
feature/noise-sunset

Conversation

@scolear

@scolear scolear commented Sep 15, 2026 •

Copy link
Copy Markdown
Member

What this changes

Decman nodes no longer talk to each other. Every workflow that used the Noise
transport now runs through Canton itself, and the transport is deleted.

Canton already carries two things between the operators' nodes. The synchronizer
topology store broadcasts partially signed topology transactions to every member.
The ledger delivers a Daml contract to every participant that hosts a stakeholder.
This branch uses both, so decman needs no channel of its own.

A workflow now runs like this:

  1. The proposer creates a WorkflowProposal contract that names the invitees.
  2. Each invitee sees a card, and its operator accepts or declines on the ledger.
  3. The proposer writes the topology change into the synchronizer store as a
    partially signed proposal.
  4. Each member validates that proposal against what it accepted, then co-signs it
    by transaction hash.
  5. Canton merges the signatures. The change becomes effective when the threshold
    is met.

Why

The Noise transport duplicated work Canton already does, and it cost real
operational weight: a second listen port per node, a static key pair per node, a
peer table of addresses and public keys, a pre-shared key per pair of operators, a
full peer mesh as a precondition of every workflow, and a bespoke wire protocol of
about fifty message codes. A node that could reach its own Canton participant
could still fail to start a workflow because a peer's port was unreachable.

Removing it also closes a gap in the old trust model. A peer used to receive the
proposal bytes from the coordinator and could not see the other members' keys. Now
every member reads the same proposal from its own synchronizer store and compares
it against the on-ledger acceptances of every other member.

Key decisions

Node identity. Each node gets one node party, hosted on its own participant
with Submission permission. It signs the registry entry, the proposals, the
acceptances, the signature rounds and the manifests. It is stored as a
party_credentials row of kind node, and PUT /node-identity sets it.

Discovery. Operators exchange one string per peer:
participant_id,node_party_id,name. No addresses, no ports, no keys. Each node
verifies through the topology store that the node party really is hosted on the
claimed participant before it names that peer anywhere.

Peer health. Each node publishes a DecmanNode contract and heartbeats on it.
A peer reads as Active, Stale, Unknown or Unvetted. There is no live probe any
more, so the peers table reports the age of the last heartbeat, not a round trip.

Keys. A new party gives each member one key with both the Namespace and the
Protocol usage. Its self-signed root NamespaceDelegation publishes the public key
in the topology store, so the proposer can build the party mapping without asking
anyone. Parties created before this change keep their two keys and keep working.

Contract deployment. The proposer prepares each transaction and opens a
SubmissionRound. Members re-hash the bytes, check the transaction acts as the
party they accepted, and sign. The proposer executes once the party's signing
threshold is met. Canton's preparation-time tolerance bounds the window, which is
hours, not minutes.

DAR distribution. The proposer pins each file by SHA-256 and main package id.
Each operator uploads the same file locally, and the run completes when the
topology store shows the package vetted everywhere. DAR bytes never move between
nodes.

ACS replication. Canton online party replication is still alpha and its codecs
are protocol-version dev, so it is unusable here. Add-party keeps offline
replication. A current host exports the snapshot and publishes an AcsManifest;
the joining operator downloads and uploads the file through two new endpoints. The
joiner verifies the manifest signer, its hosting, and the activation serial before
it disconnects anything.

Upgrade

This is a breaking change, so the crate version moves to 2.0.0.

Migration 000020 adds the node-identity column and the proposal-decision table.
Migration 000021 fails any run still in flight, rewrites the coordinator column
to a participant id, drops the peer address columns, and adds the proposal columns.

A mixed network is safe in both directions. An old node refuses to start a
workflow when a peer's Noise listener is gone. A new node refuses until every
invitee has vetted the coordination package and published a registry entry. The
runbook is in docs/DEPLOYMENT_GUIDE.md, and the design records the manual
rollback.

How this was tested

Everything below runs green on this branch:

  • cargo fmt --all -- --check
  • cargo clippy --all-targets --all-features --no-deps -- -D warnings
  • cargo test --workspace --lib — about 300 new tests cover the Daml codec,
    the mapping builders, the per-kind validation rules, the acceptance-counting
    predicates, the quorum arithmetic and the migrations
  • cargo test -p decman-cli — 77 tests
  • dpm build --all, and dpm test in decman-coordination-test — 30 tests
  • every built DAR is byte-identical to its committed release
  • the frontend: npm test (41 tests), tsc -b, vite build

The integration suite passes end to end against a live three-node Canton —
all 33 phases, 1 passed; 0 failed. That covers onboarding, DAR distribution,
contract deployment (including a 100-contract run through the submission
rounds), deployment at quorum with one operator absent, cancel, kick,
add-party with the operator-carried ACS snapshot, change-threshold, the
external-party flows, the governance and reward phases, and the whole chaos
block: coordinator and peer restarts mid-flight, coordinator retry, dismiss
cascades, key-generation idempotency across a restart, concurrent kinds
through a restart, cross-node concurrent workflows, and the sibling cancel and
decline cases.

The bugs this shook out

Removing the transport was the easy half. Running the suite against real Canton
found six defects that only a live network shows.

A member never published its acceptance. POST /invitations/{cid}/accept
records the operator's decision locally; each kind's driver made the ledger
write. Onboarding, add-party, kick and DARs did. Contracts and change-threshold
did not, so their coordinators waited in WaitingForAcceptances while every
member waited to sign. Neither side logged an error, because neither side had
one: both were waiting correctly. engine::drive now makes that write for
every kind, and a kind opts out only when its acceptance carries key material
one of its own steps generates. The default is correct, so a new kind cannot
repeat this.

Contracts waited for every invitee. A contracts run needs
party_signing_keys.threshold signatures, not one acceptance per invitee, so
one silent operator held up a party that had the quorum to act. The coordinator
now advances once the counted acceptances plus itself reach the threshold. An
explicit Decline still stops the run: silence and refusal are different
answers.

A member could join a party and keep no record of who owns which key. The
member wrote the party's key cache in its Complete step, but the proposer
finishes the WorkflowProposal as soon as the party is effective, and a member
that reads that outcome first completes its row through reconcile, skipping
Complete. The next kick on that party then failed on that node. The member
now writes the cache when it co-signs the party mapping, which it has validated
and is about to sign.

A joiner without the party's packages waited instead of stopping. The
offline import re-validates every contract and fails on a missing package, but
only after it disconnects from the synchronizer — the devnet failure of
onboarding a node without its DARs. The check lived inside the import, and the
import is the operator's step now, so a joiner that could never succeed sat in
SyncAcs looking healthy. It now compares the manifests' package ids against
its own vetted set while still connected, before anything is carried.

No manifest ever named a package. The exporter read the party's contracts
with no ledger credential, so the read failed authentication and every
AcsManifest carried an empty package list, which made the check above inert.
The export helpers now take the party's own token, which the add-party engine
already holds. A node with no credential for the party still exports, and says
that its manifest names no packages.

Duplicate registry entries. submit_and_wait returns before the node's own
create reaches the next ACS read, so a node could publish a second DecmanNode.
It now retires the duplicates and tolerates the stale-id race.

The suite also carried assumptions the transport removal invalidated: a
declined run used to need reaping, member progress used to land in
workflow_artifacts, DAR bytes and ACS snapshots used to travel between nodes,
and a party's keys used to live in dec_party_identity. Those phases now
assert the on-ledger behaviour instead.

Known gaps to close before merge

  • A devnet run.
  • acs::export_to_response re-exports when it finds no spool file. Two exports
    of one ACS are not byte-identical, so those bytes match no published
    manifest. The endpoint should publish a manifest for what it serves, or
    refuse. Until then an operator must download from the host whose manifest the
    import names, which is what the suite does.
  • A member that completes through the reconcile path can leave its spool file
    behind. The dismiss handler should call acs::cleanup_spool.
  • check_peer_dars has no scenario asserting that /participants-status
    reports a peer Active once its registry heartbeat lands (design D3). That
    coverage replaced a reachability scenario the transport used to carry, and it
    is still missing.
  • dismiss_failed_cleans_artifacts covers the cleanup cascade against an
    onboarding run, which writes no artefact rows. A contracts-run variant would
    exercise a populated table.
  • A KMS node has never generated a key with both usages. The design asks for
    that spike before the mainnet window.

Reviewing this

The diff is large because the transport touched everything, so review it commit
by commit:

  • docs: add the Noise-sunset coordination design and its revision — the
    design, and the nine defects an adversarial review found in it
  • feat(onledger): add the Daml coordination package and the Rust foundation —
    the new package, the node identity, and the module everything else builds on
  • feat(onledger): add the shared party-key and member-key helpers
  • feat: coordinate every workflow through Canton and remove Noise — the six
    workflow engines, the deletion, and the migration
  • the fix(onledger) commits — what the live runs found
  • the test(e2e) commits — the suite moved onto the on-ledger model

docs/NOISE_SUNSET_DESIGN.md is the specification the code follows, and
crates/decman/src/onledger/README.md documents the new module's API.

🤖 Generated with Claude Code

@scolear
scolear force-pushed the feature/noise-sunset branch 3 times, most recently from b2e1106 to d279f56 Compare September 15, 2026 09:38
Decman coordinated every multi-party workflow over a Noise P2P
transport. Canton already carries partially signed topology
proposals to every member and Daml contracts to every observer.
This design replaces Noise with those native paths: a node party
and registry contract per node, workflow proposals and
acceptances in Daml, co-signing by transaction hash in the
synchronizer store, signature rounds for prepared transactions,
hash-pinned DAR uploads, and an operator-moved ACS snapshot.
The review found nine blocking defects: a Confirmation-only node
party cannot submit commands, the bootstrap exemption would open an
unauthenticated write on upgraded nodes, members never pinned the
topology operation, an accepted proposal had no serial binding or
expiry, acceptance provenance was unspecified, the kicked key had no
trusted source, kick quorum ignored the stored namespace threshold,
the ACS manifest signer was unbound, and registry publication ignored
package vetting. This revision fixes each one and the should-fix
items that followed from them.
…tion

Decman nodes will coordinate through Canton instead of Noise. This
commit adds the pieces every later change builds on: the Daml package
decman-coordination-v1 (node registry, workflow proposals and
acceptances, submission rounds, ACS manifests) with its test package
and pinned release DAR; a node identity stored as a party_credentials
row of kind 'node' with GET/PUT /node-identity; and the onledger Rust
module with the Daml codec and client, the registry with heartbeat
and peer health, proposal projection, synchronizer-store proposal
discovery and co-signing by transaction hash, per-kind validation, the
observer loop, and the engine API with stub drivers for every kind.

Noise code is untouched here; the removal follows once the drivers
are in place.
Pinned DAR uploads pass expected_main_package_id, which needs the
sha256 of the main .dalf. The zip crate is already in the graph
through utoipa-swagger-ui, so this adds no new code to the build.
Every kind driver needs the same key material: the dual-usage
{prefix}-key with its self-signed root NamespaceDelegation, the local
identity the validator compares against, and the member-key caches
that acceptances heal. Keeping them in one module lets the drivers
stay thin and lets kick attribution read a single trusted source.
The six workflow kinds now run on the on-ledger engine. A run starts
as a WorkflowProposal; invitees accept on the ledger; topology changes
are proposed into the synchronizer store and co-signed by transaction
hash; contract deployments collect prepared-transaction signatures in
SubmissionRounds; DAR distribution pins hashes and observes vetting;
add-party replication publishes an AcsManifest and takes the snapshot
through operator-driven export and import endpoints.

With that in place the Noise transport has no caller. This removes the
listener, the peer client and server, the heartbeat and ping loops,
the Noise key file, the peer public keys and ports, the ten Noise CLI
flags, and the hyper-noise, tokio-noise and secp256k1 dependencies.
Migration 000021 fails any run that was in flight, renames the
coordinator column to the participant id, drops the peer address
columns, and adds the proposal columns the projection needs. Peer
health is the registry heartbeat. The CLI, the integration harness,
and the test crate compile against the new shapes; the coordination
phases still carry TODO markers for their rewrite.
A member validates the kick party mapping only after the kick namespace
change became effective, so the head owner set already excludes the
kicked member while its signing key is still in the mapping. The
dual-usage test compared the two sets for equality, so it failed, and
the check fell through to the weaker legacy branch. A proposer could
then drop a different member's signing key.

The member now recognises the same party from either owner set.
The Noise transport was the only user of hyper in decman, and every
`http::` path in the crate resolves through actix-web. CI runs
cargo-machete, which fails on an unused direct dependency.

Removing hyper does not clear the h2 0.3 advisory: actix-http still
pulls it. Both rationales now say so.
The operator and developer documents still described a transport that
no longer exists. They now describe what the code does: a node party
per node, a registry contract that carries peer identity and
liveness, workflow proposals and acceptances on the ledger, topology
changes co-signed in the synchronizer store, signature rounds for
contract deployment, hash-pinned DAR uploads, and an operator-moved
ACS snapshot.

The deployment guide gains a cutover runbook from 1.8.x and the
manual rollback. The KMS guide explains that one key now signs both
topology and ledger transactions for a party, and what that means for
the key policy. The architecture document keeps a History section so
a reader who meets a 1.x deployment knows what it was.
submit_and_wait returns before this node's own create is guaranteed
visible to the next active-contract read, so an observer tick that
read too early published a second DecmanNode. Peers then saw two
entries for one node and read whichever they met first.

The reader now retires the older entries instead of reporting them,
and an update that names an entry the heartbeat has just replaced
waits for the next tick rather than failing the registry stage.
A contracts run never finished. The coordinator waited in
WaitingForAcceptances with both invitees listed as missing, while each
member waited in SignSubmissions for a round the coordinator opens only
after the acceptances arrive. Neither side logged an error, because
neither side had one: both were waiting correctly.

POST /invitations/{cid}/accept records the operator's decision locally
and opens the peer run row. The ledger write was left to each kind's
member driver. Onboarding, add-party, kick and DARs made it. Contracts
and change-threshold never did, so their members published nothing and
both runs deadlocked.

Move the write into engine::drive, before the first member tick, and let
a kind opt out through KindDriver::member_publishes_own_acceptance. The
default is false, so a new kind is correct without knowing the rule.
Onboarding, add-party and kick opt out: their acceptance carries key
material a member step generates, and they still write it at that step.
DARs drops its own copy and uses the shared path, which the DARs phase
already covers end to end.

The write is idempotent twice over: against the tick snapshot, and
against a cid pinned on the run row. submit_and_wait can return before
the create reaches the next ACS read, and a second acceptance from one
acceptor fails the coordinator closed.
The acceptance deadlock left no trace in CI. run.sh exports its quiet
preset to the node processes too, so every node recorded WARN and above
and the whole coordination flow went unlogged. The uploaded node logs,
the only evidence a CI failure leaves, said nothing at all.

Give the nodes their own level through DECPM_NODE_RUST_LOG, defaulting
to INFO with the on-ledger module at DEBUG. The runner output is
unchanged.

Log the step each run sits in once per tick, and log the member branch
that waits for a SubmissionRound, which returned silently.

Also stop the one-shot coordination DAR upload from reporting a false
failure. spawn_supervised treats any return as death, so every node
logged "may stay unvetted" at ERROR on a clean startup. spawn_once
reports only a panic.
The coordinator waited for every invitee before it prepared a round. A
contracts run needs party_signing_keys.threshold signatures, not one per
invitee, so a single silent operator held up a party that had the quorum
to act. That is the absent-owner hang contracts_quorum_completes covers:
on a 2-of-3 party with only one member accepting, the run never finished.

Advance once the counted acceptances plus the proposer reach the
threshold. The proposer signs its own rounds, so it counts as one.

An explicit Decline still fails the run. Silence and refusal are
different answers, and an operator who says no should stop the run
rather than be outvoted by a quorum that never heard from them.
Three phases still watched for 1.x evidence that the on-ledger model no
longer produces.

add_party_edge_cases reaped a declined run with a per-instance cancel,
because a coordinator task used to keep idling after the decline. No task
idles now: the decline fails the row and finishes the WorkflowProposal, so
the run is already terminal and the cancel must refuse it. Assert the 409.

generate_keys_idempotent and dismiss_failed_cleans_artifacts waited for a
workflow_artifacts row. Only a contracts run writes artefacts now; every
other kind keeps its progress on the run row. Both gate on the run's step
instead, through a new db::workflow_run_step helper.
The add-party coordinator exported the party's ACS, published the
AcsManifest, and then waited in AwaitReplication for a snapshot nobody
moved. Decman used to relay those bytes between the nodes. It does not
any more: the operator downloads the file from the exporting host and
uploads it to the joiner (design D9), and the suite has to do the same.

The phase now reads the manifest that names the joining participant,
downloads the file it points at, uploads it, and checks that the received
bytes hash to what the manifest states. Two byte-level HTTP helpers carry
the gzip file, because every other fixture helper speaks JSON.

A respawned node also kept the runner's quiet log level, so a node went
dark after a chaos restart, which is where its own record matters most.
It now takes the same DECPM_NODE_RUST_LOG default the shell harness uses.
A kick failed on a member with "its owner key has not yet been resolved".
The member held no owner key for its peers, because it never wrote the
party cache for that party at all.

The member writes that cache in its Complete step. The proposer finishes
the WorkflowProposal as soon as the party is effective, and a member that
reads the outcome first completes its row through reconcile, which skips
Complete. The member then joins the party and keeps no record of who owns
which key, so the next kick on that party fails on that node.

Write the cache when the member co-signs the party mapping instead. It has
validated that membership by then and is about to sign it, so it records
what it endorsed, and no completion race can get in front of it. The
Complete step keeps its own call; the write is an upsert.

G5 watched dec_party_identity, which only parties created before the
Canton-native coordination hold. It now watches the dec_party_participant
rows that carry a key, which is the material a later kick reads.
The add-party relay read whichever AcsManifest named the joiner and then
always downloaded from P1. Every current host exports and publishes its
own manifest, within a millisecond of the others, and two snapshots of one
ACS are not byte-identical. Picking P2's manifest and P1's bytes made the
import refuse the upload, which read as a broken export rather than a
broken test. It passed or failed by whichever manifest sorted first.

Choose a manifest whose exporter this suite can reach, then download from
that host. That is the operator's rule too: carry the file from the host
whose manifest the import names.
The snapshot-relay helper landed between the run function's doc comment
and the function, so the two doc blocks merged. The phase lost its
documentation, and clippy read the merged block as a doc list running into
prose. Move the helper below the function it serves.
G7's gate moved to the run row, but its closing assertion still counted
dec_party_identity rows, which only a party created before the
Canton-native coordination holds. Count the keyed participant rows, the
same material G5 watches. The old helper has no callers left.
G9 runs a Dars workflow beside an Onboarding, restarts the coordinator,
and waits for both to finish. It accepted the Dars invitations but never
uploaded the files, so the coordinator sat in AwaitVetting: DAR bytes do
not travel between nodes any more, and the run finishes when the topology
shows the packages vetted everywhere. Upload them on both members, the
same step distribute_dars already performs.
The offline import re-validates every contract and fails on a package the
target participant does not hold — but only after it disconnects from the
synchronizer. That is the devnet failure of onboarding a node without its
DARs, and the design asks for a check before the disconnect window opens.

The check only ran inside the import. The import is the operator's step
now, so a joiner that could never succeed sat in SyncAcs looking healthy
and waited for a file it would have rejected.

Run it in SyncAcs instead. The manifests name the packages the party's
contracts need, so the joiner compares them against its own vetted set
while it is still connected and nothing has been carried, and fails the
run with the missing ids. An empty snapshot still takes the fast path,
and a completed import is never undone by a late check.
The exporter read the party's contracts with no ledger credential, so the
read failed authentication and every AcsManifest named no packages. The
joiner's pre-disconnect check then had nothing to check, and a node
without the DARs learned that only when the import failed, after the
disconnect window was already open.

Pass the party's own credential down to the read. The export helpers take
the token, and the add-party engine supplies the one it already uses to
prepare submissions. A node that holds no credential for the party still
exports; its manifest names no packages, and it says so.
code-coverage.sh runs the suite with the LLVM instrumentation on, and each
test process drops a .profraw file beside it. Nothing ignored them, so a
`git add` after a coverage run sweeps in hundreds of them.
@scolear
scolear force-pushed the feature/noise-sunset branch from 05d0ccc to 3d7ccd0 Compare September 15, 2026 13:28
Eleven phases carried a note saying they still described the 1.x transport
and needed a rewrite before the suite could run green. The suite runs green
now, and each of these phases asserts an outcome the on-ledger model
delivers, so the note is stale. The one marker left in check_peer_dars asks
for coverage that does not exist yet, which is a different thing.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant