# Architecture This document describes the architecture of the Zama KMS: a key-management service for fully homomorphic encryption facilitated by [TFHE-rs](https://github.com/zama-ai/tfhe-rs). The system supports key generation, CRS generation and decryption in both a single-party centralized service or as an `n`-party threshold MPC cluster. Input and output happen through gRPC and is designed to be triggered and consumed by [FHEVM](https://github.com/zama-ai/fhevm). The underlying MPC protocol is maliciously secure and robust; see [Noah's Ark (eprint 2023/815)](https://eprint.iacr.org/2023/815) for the formal treatment. ## System context At the top level, an FHEVM deployment is composed of three subsystems: 1. A **host chain** (EVM L1) that stores ciphertexts on-chain. 2. An **FHEVM Gateway** that coordinates user requests. 3. The **KMS** (this repository) that holds FHE key material and performs key generation, public/user decryption, CRS generation, and reshare operations. The KMS exposes a gRPC API. In the threshold deployment the KMS is itself a cluster of `n` independent parties (typically 13 parties, threshold `t = 4`) that run an MPC protocol among themselves; each party runs the same binary with its own configuration and secret share. A single deployment mode is chosen at startup via the server configuration (centralized vs. threshold). The gRPC surface is shared between modes; a few RPCs (preprocessing, reshare) are only meaningful in threshold mode. The configuration of the set of servers is handled through MPC contexts, which are also managed by the FHEVM. Threshold KMS deployments using Nitro Enclave remote attestation reject both new and stored contexts whose PCR allowlist is empty; stored contexts that fail validation are skipped during startup. Non-enclave and mocked-enclave deployments permit an empty allowlist. The system supports automatic backup, facilitated either through AWS KMS, or through a custom threshold protocol where Custodians hold keys that can be used to help KMS nodes decrypt encrypted backups. The settings and administration for this is also managed through gRPC calls with the notion of Custodian contexts. ## Communication interfaces and trust model Read this section before you review code for security issues or judge a security report. The full text is in [docs/explanations/trust_model.md](../docs/explanations/trust_model.md). A KMS core listens on two separate gRPC interfaces: 1. **Core-to-core interface** (`[threshold]` section, default port 50001, crate [threshold-networking](../core/threshold-networking/)). A peer-to-peer network between the KMS cores of one deployment. It carries the MPC protocol messages. 2. **Service interface** (`[service]` section, default port 50100, `CoreServiceEndpoint` in [kms-service.v1.proto](../core/grpc/proto/kms-service.v1.proto)). The [KMS connector](https://github.com/zama-ai/fhevm/tree/main/kms-connector) calls it to start an operation and to fetch the result. It is an orchestrator channel: the connector says which operation to run, and the cores run the MPC protocol over the core-to-core interface. The **core-to-core interface** is guarded by mutual TLS. A node only accepts connections from the allowlisted set of peers in its peer list and MPC contexts. The receiver checks that the sender named in each message matches the Common Name of the peer certificate. In Nitro Enclave deployments (`tls.auto`), the verifier also checks the PCR values in the attestation document against `trusted_releases`, so a node only talks to peers that run an allowlisted release. PCR0 is the hash of the whole enclave image file, PCR1 the hash of the kernel and bootstrap ramdisk, and PCR2 the hash of the application root filesystem; all three must match one allowlisted entry, and when `eif_signing_cert` is configured PCR8 (hash of the image signing certificate) is checked against the certificate bundled in the peer's TLS certificate. Production deployments always run this interface with TLS enabled; a TLS-off configuration needs the `insecure` cargo feature and is used only for debugging and testing. Authenticated peers are still mutually distrusting MPC parties: up to `t` of them may be malicious, so the content of a peer message is adversarial input and the protocol code validates it. The **service interface** has no TLS, no authentication and no authorization in the code, and it does not verify the intent of a request. It trusts and accepts every message it receives. The deployment guarantees, at the infrastructure level, that exactly one KMS connector can reach this interface. That connector is operated by the same party that runs the KMS core, so the two trust each other by definition. The interface is never publicly reachable. Validation of a request is split across the stack. The [KMS connector](https://github.com/zama-ai/fhevm/tree/main/kms-connector) performs the ACL checks on ciphertext handles and only forwards events emitted by the gateway contracts. Input proofs, verified on smart contract level and by the coprocessor, ensure that a ciphertext is well formed before it reaches the chain. The KMS core verifies EIP-712 signatures on user decryption requests, authenticates peers, and validates protocol messages. Request IDs work as follows. The gateway contracts assign each ID and bind it to its ciphertexts, and the connector resends the same payload on a retry, so the core assumes that a known ID carries the same ciphertexts as before. A meta store per operation type in the core tracks every accepted ID; MPC session IDs derive from the request ID, so this also stops a second MPC session under a used session ID. Key generation, preprocessing, CRS generation and context or epoch management reject a known ID with `AlreadyExists`. `PublicDecrypt` and `UserDecrypt` do the same unless the earlier attempt failed, in which case they reset the entry and decrypt again (`add_or_redo_failed_in_meta_store`). `PublicDecryptSync` and `UserDecryptSync` attach to the existing entry and return its result. A repeated decryption of the same ciphertexts is not a finding. The meta store is in-memory only, so a reboot of the core forgets every known request and session ID. The KMS connector keeps the state of each request in its [persistent database](https://github.com/zama-ai/fhevm/tree/main/kms-connector/connector-db); its [kms-worker](https://github.com/zama-ai/fhevm/tree/main/kms-connector/crates/kms-worker) marks a request as sent and only polls for the result on a retry, so a request is not run more often than necessary across core reboots. Consequences for agents: - The threat model assumes that at most `t` of the `n` parties are malicious. An attack that needs more than `t` malicious parties is out of scope. Every attack that works with at most `t` malicious parties is in scope: bypassing TLS, attestation or sender binding on the core-to-core interface, or breaking the confidentiality of the key material or the correctness of a result. - Do not report missing authentication, authorization or rate limiting on the service interface, or any finding in which the connector itself is the attacker, as a vulnerability. Such a finding describes the design. - Values inside a request that originate from external clients and pass through the smart contracts and the connector unchanged are untrusted: ciphertexts and handles, user public encryption keys, EIP-712 signatures and domains, and parameter selectors such as the FHE parameter set or keyset configuration. The core must process them without a service outage (crash, stall, unbounded allocation) and without a confidentiality break. Such a finding is in scope even though the request arrives over the service interface. - Do not add authentication or authorization to the service interface unless your human asks for it. - Code that is not used in production is out of scope for security findings: the experimental BGV/BFV schemes in [core/threshold-bgv/](../core/threshold-bgv/), the benchmark and experiment harnesses in [core/experiments/](../core/experiments/), and any other code marked as experimental. - Every security finding you report must cite the commit hash or tag you analyzed and the file path and line numbers of every code location it relies on. Verify each pointer against the checked-out tree before you report it. ## Workspace layout The repository is a Cargo workspace. The members are declared in [Cargo.toml](../Cargo.toml). ### Core cryptography / MPC | Crate | Path | Responsibility | | ---------------------- | ----------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------- | | `threshold-algebra` | [core/threshold-algebra/](../core/threshold-algebra/) | Finite-field and group primitives used by the MPC protocols | | `threshold-execution` | [core/threshold-execution/](../core/threshold-execution/) | Threshold FHE protocol execution: DKG, preprocessing, online protocols | | `threshold-bgv` | [core/threshold-bgv/](../core/threshold-bgv/) | Experimental BGV/BFV schemes with distributed keygen and threshold decryption | | `threshold-networking` | [core/threshold-networking/](../core/threshold-networking/) | Inter-party gRPC transport and choreography | | `threshold-hashing` | [core/threshold-hashing/](../core/threshold-hashing/) | Hashing primitives used across the MPC stack | | `threshold-types` | [core/threshold-types/](../core/threshold-types/) | Shared types and constants | | `experiments` | [core/experiments/](../core/experiments/) | Benchmark and experiment harnesses (see [docs/guides/threshold-benchmark.md](../docs/guides/threshold-benchmark.md)) | ### Service layer | Crate | Path | Responsibility | | ---------------- | ----------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `kms` | [core/service/](../core/service/) | KMS service library and binaries — the packaging around the core crypto | | `kms-grpc` | [core/grpc/](../core/grpc/) | Protobuf definitions + generated types and client stubs | | `core-client` | [core-client/](../core-client/) | CLI client that drives the gRPC API | | `observability` | [observability/](../observability/) | OpenTelemetry / Prometheus wiring | | `vsocktun` | [vsocktun/](vsocktun/) | Multi-queue, offload-aware TUN-to-VSOCK relay used by Nitro enclave deployment scripts to preserve end-to-end peer TCP while bridging enclave IP traffic through the parent, including raw virtio-net TUN frames when both ends support offload metadata; the parent side also bootstraps the enclave-side tunnel CIDR, MTU, shard count, and rewritten resolver config over the same VSOCK control port | | `bc2wrap` | [bc2wrap/](../bc2wrap/) | Version-pinned `bincode` wrapper used for on-disk and on-wire encoding | | `error-utils` | [core/error-utils/](../core/error-utils/) | Shared error types and helpers | | `thread-handles` | [core/thread-handles/](../core/thread-handles/) | Rayon thread-pool management | Auxiliary tools live under [tools/](../tools/): `kms-health-check` is a gRPC health probe and `generate-test-material` produces reproducible crypto test vectors. Shared test fixtures and generic local file helpers are in [core/test-utils/](../core/test-utils/). The [backward-compatibility/](../backward-compatibility/) crate is a separate Cargo workspace — see [Backward compatibility](#backward-compatibility). ## The service crate (`core/service`) The service crate is the main surface area. Key subdirectories under [core/service/src/](../core/service/src/): - [engine/](../core/service/src/engine/) — RPC handlers and KMS state machines. Split into [centralized/](../core/service/src/engine/centralized/) and [threshold/](../core/service/src/engine/threshold/) submodules. Other notable files: [base.rs](../core/service/src/engine/base.rs), [context.rs](../core/service/src/engine/context.rs), [backup_operator.rs](../core/service/src/engine/backup_operator.rs), [keyset_configuration.rs](../core/service/src/engine/keyset_configuration.rs), [material_integrity.rs](../core/service/src/engine/material_integrity.rs) (digest primitives over raw stored bytes, depended on by both the storage layer and the startup checks), [public_material_sync.rs](../core/service/src/engine/public_material_sync.rs) (the digest-verified peer fetcher shared with resharing, and the boot-time repair of public storage built on it) and [storage_material_verification.rs](../core/service/src/engine/storage_material_verification.rs) (the read-only startup checks on top of both — see [Boot-time storage verification](#boot-time-storage-verification)), [validation_non_wasm.rs](../core/service/src/engine/validation_non_wasm.rs) and [validation_wasm.rs](../core/service/src/engine/validation_wasm.rs) (the validation logic is compiled for both native and WASM so that clients can verify user-decryption responses in the browser). - [vault/](../core/service/src/vault/) — pluggable storage for key material. Backends include AWS S3, local file, AWS KMS, and AWS Nitro Enclaves. Root keys and key-encryption logic live in [vault/keychain/](../core/service/src/vault/keychain/). - [backup/](../core/service/src/backup/) — custodian-based secret-sharing backup of long-term signing / root keys, used for disaster recovery. See [Backup and recovery](#backup-and-recovery) below. - [cryptography/](../core/service/src/cryptography/) — AES-GCM-SIV, signcryption, hybrid ML-KEM (post-quantum), MLKEM1024-P384 (a composite of post-quantum ML-KEM-1024 and classical P-384), and attestation (Nitro NSM + certificate chain verification). Custodian backup uses MLKEM1024-P384 for all three of its keypairs — the custodian's long-term key, the operator's ephemeral recovery key, and the operator's per-context backup vault key — selected in one place, `backup::BACKUP_PKE_SCHEME`. A new custodian context is rejected unless every custodian encryption key, and the operator's own backup key, uses that scheme (`InternalCustodianContext::new` / `validated_nodes`). On the signing side, every signature in the custodian-backup chain is a composite under `backup::BACKUP_SIGNING_SCHEMES` (ECDSA/secp256k1 and ML-DSA-87), and every signature in it must verify. Each party publishes a `VerfKeySet`, and a key set that does not cover these schemes is rejected (`ensure_backup_schemes`). User decryption accepts ML-KEM-512 only. Randomly generated MLKEM1024-P384 keypairs use a 256-bit-seeded CSPRNG. The custodian key derives directly from 256-bit mnemonic entropy. Signing lives under [cryptography/signing/](../core/service/src/cryptography/signing/): a scheme-tagged `Signature` plus one backend per scheme — ECDSA/secp256k1 (`ecdsa`, the legacy default and EIP-712 home), EdDSA/ed25519 (`eddsa`), and ML-DSA/FIPS-204 (`mldsa`) — behind the `SigningScheme` trait and the `unified_sign`/`unified_verify` entry points. The historic `cryptography::signatures` path is a re-export facade. A node persists two private objects: its ECDSA signing key (`PrivDataType::SigningKey`, the authoritative on-chain identity) and an independent, CSPRNG-generated `RootSigningSeed` (`PrivDataType::SigningSeed`), both under `SIGNING_KEY_ID`. The seed will eventually be the root of _every_ signing key of the node, ECDSA included: keys are derived on demand from the _seed_. To keep backward compatibility, and to avoid making nodes roll their ECDSA keys, an ECDSA key is also stored on its own, and the seed serves every non-ECDSA scheme. That is, if a stored ECDSA key exists, then the seed does _not_ derive the ECDSA material. The two halves come together in memory as `signing::identity::NodeSigningIdentity`, which `get_core_signing_identity` assembles and `BaseKmsStruct::signing_identity` hands out. `NodeSigningIdentity` is never persisted, and it is the only type with the multi-scheme `unified_sign_with` / `unified_verifying_key` methods: `PrivateSigKey` is the ECDSA leaf type, which client wallets and the WASM surface also use. An identity with no seed — a node that has not yet run `kms-gen-keys` — can only do ECDSA, and errors with `SigningError::MissingRootSeed` for anything else. On the client side, `Client::verify_result_signatures` checks a result's per-scheme `signatures` against the peers' published keys, which `Client::new_client` reads from `PubDataType::TypedVerfKey`, and rejects a result that omits a scheme the client asked for (`Client::signing_schemes`). Every scheme's public verification material — ECDSA's included — is stored under the handle `consts::signing_material_id(scheme)` gives, in the data types `key_setup::NON_LEGACY_VERF_MATERIAL_TYPES` names: `PubDataType::TypedVerfKey` holds the scheme's _own_ verification key type (`PublicSigKey`, `Ed25519VerfKey`, `MlDsaVerfKey
`), and `TypedVerfAddress` its
`address_text()` (`0x`-prefixed hex; for ECDSA the EIP-55 address).
ECDSA's material is _additionally_ written to the deprecated `key_setup::LEGACY_VERF_MATERIAL_TYPES`
(`PubDataType::VerfKey`/`VerfAddress`, a bare `PublicSigKey` and the same
address text) for existing external consumers; those two are scheduled for
removal and nothing new should read them. Both copies are validated against the
signing key when backfilling.
- [client/](../core/service/src/client/) and
[testing/](../core/service/src/testing/) — client-side helpers (including
local key-material utilities used by `core-client`) and test-only wiring.
- [bin/](../core/service/src/bin/) — entry points (see below).
### Task randomness
[`RngSource`](../core/service/src/engine/rng_source.rs) supplies task seeds from two parent
RNGs per KMS instance: a 128-bit-seeded `AesRng` and a 256-bit-seeded `ChaCha20Rng`. `AesRng`
([rng.rs](../core/threshold-types/src/rng.rs)) runs AES-128 in counter mode through
`tfhe-csprng` and wipes its key on drop. A fork never
carries more entropy than its parent. The wide path therefore needs its own parent, rather than a
wider fork of the narrow one. `BaseKmsStruct` instances and `SessionMaker` share the source
through `Arc`. Each task receives an owned RNG with a separate seed. After a fork, the parent
discards 256 bytes of output, so the seed does not stay in its buffer. Initialization seeds each
parent from an independent draw, which combines OS entropy with entropy from the configured
security module. Refresh also mixes output from the existing parents. Entropy failures return
errors and leave both parents unchanged. Refresh logs report success or failure without seed
values.
Threshold epoch creation refreshes once in `new_mpc_epoch`, before either the resharing
or PRSS session forks its RNG. This includes old-committee parties that skip PRSS initialization.
A successful refresh protects future task seeds once fresh entropy is unknown to the attacker.
Existing task RNGs remain unchanged. The source does not provide backtracking resistance
within a reseeding interval. Centralized services seed at construction; epoch refresh
applies to threshold services.
### Binaries
All under [core/service/src/bin/](../core/service/src/bin/):
- [kms-server.rs](../core/service/src/bin/kms-server.rs) — main service process.
- [kms-init.rs](../core/service/src/bin/kms-init.rs) — post-deployment cluster
initialization.
- [kms-gen-keys.rs](../core/service/src/bin/kms-gen-keys.rs) — generate the server
signing identity — the `RootSigningSeed` plus the ECDSA signing key, derived
from the seed on a fresh node and left untouched on an upgraded one — and, in
threshold mode, per-party self-signed CA certificates for mTLS. It is the
**only** thing that ever creates a seed. Also derives and persists every
scheme's public verification material: ECDSA's from the persisted signing key,
every other scheme's from the seed. Reads a keygen TOML with
`--config-file`; `[keygen] repopulate = true` backfills the per-scheme
verification material from the signing identity already in private storage
instead of generating keys (the same backfill runs automatically on server
start via `migration::migrate_public_verification_material`, which warns and
skips when the seed is absent), `[keygen] show_existing = true` prints the
existing signing-material handles and exits, and `[keygen] overwrite = true`
deletes the signing key and the seed together with the verification material
derived from them, since generating an identity alongside another identity's
derived material is rejected. Supports `mock_enclave` in config for local dev
when compiled with the `insecure` feature.
- [kms-custodian.rs](../core/service/src/bin/kms-custodian.rs) — custodian-side
tool for producing and recovering backup shares.
- [kms-gen-tls-certs.rs](../core/service/src/bin/kms-gen-tls-certs.rs) — TLS
certificate generation for inter-party mTLS.
## gRPC surface
Protobuf definitions live in [core/grpc/proto/](../core/grpc/proto/). The main
service definition is
[kms-service.v1.proto](../core/grpc/proto/kms-service.v1.proto); shared messages
are in [kms.v1.proto](../core/grpc/proto/kms.v1.proto); an insecure transport
variant is in
[kms-service-insecure.v1.proto](../core/grpc/proto/kms-service-insecure.v1.proto);
metastore status types in
[metastore-status.v1.proto](../core/grpc/proto/metastore-status.v1.proto).
The primary service is `CoreServiceEndpoint`. Its RPCs group into:
- **Key generation** — `KeyGenPreproc` / `KeyGenPreprocResult` (threshold
preprocessing), `KeyGen`, and `NewMpcEpoch` for key rotation. Multiple keyset
configurations are supported (standard, decompression-only, compressed
variants). Insecure key generation still requires an explicit preprocessing
ID for both the centralized and threshold cases.
Standard threshold keygen persists a dedicated OPRF LWE secret-key share in
each party's private key material and includes the corresponding OPRF server
key in the generated TFHE server key. Legacy private keysets that predate this
field are upgraded with the OPRF share absent; `UseExisting` keygen generates
and persists a fresh OPRF share for such legacy material before regenerating
public keys. When the parameter set carries transciphering parameters, keygen
additionally persists a _second_, independently sampled LWE secret-key share
and includes the matching transciphering server key; as for the OPRF key,
`UseExisting` keygen generates a fresh transciphering share when the existing
keyset has none. Key generation and CRS generation write persistent material
only after generation completes. An abort updates request state but does not
purge storage.
- **Decryption** — `PublicDecrypt` (returns plaintext) and `UserDecrypt`
(user-initiated, EIP-712 authenticated). `PublicDecryptSync` / `UserDecryptSync`
start a decryption and wait for its result in the same call, so the caller does
not need the `Get*DecryptionResult` round trip; a known `request_id` attaches to
the running or succeeded attempt, and redoes a failed one, just like the async
variants.
- **Noise-flooded user decryption** — Each party masks its partial decryption
to protect its key share before signcrypting the result for the user. The epoch's
PRSS setup and each ciphertext's session ID let that party derive the mask locally.
This path needs no network session. Public decryption and bit-decomposition user
decryption use network sessions.
- **CRS** — `CrsGen` for ZK-proof common reference strings.
- **Resharing** — `NewMpcEpoch` with `previous_epoch` set rotates parties /
refreshes secret shares as part of epoch creation; the outcome is fetched via
`GetEpochResult`. The `preproc_id` supplied per key in `previous_epoch` is
caller-controlled but ends up in the EIP-712 struct signed for the new epoch,
so before any resharing protocol runs each party checks it against the
preprocessing ID stored in that key's `KeyGenMetadata` and rejects a mismatch.
Each party also rejects a key unless it has a non-empty public-key digest and exactly one
non-empty server-key or compressed-keyset digest before role dispatch.
The resharing session matches the parties of the two contexts by MPC identity.
A party with the same MPC identity in both contexts is a `Both` party, and the
session connects to it at its set 2 URL. The URL is not compared. The request
fails if the contexts give one MPC identity two different signer addresses, or
one signer address two different MPC identities. A party without a listed
signer address matches by MPC identity alone. A node that changes its signing
key is a new party: it needs a new MPC identity, and it runs as a separate
core during the reshare.
What a missing keyset means depends on the party's `TwoSetsRole`: set 1 and
both sets must hold the key material, so failing to read it rejects the
request, whereas a pure set 2 party (a node joining the new context) never
held the key and logs a warning instead. When resharing legacy key material
that has no dedicated OPRF/transciphering secret-key share, the
OPRF/transciphering sub-protocol is skipped and the reshared private keyset
keeps that field absent. Which of these optional shares to reshare is decided
from the input keyset, and every party must agree. A storage failure during
resharing rolls the new epoch back on the party that fails. That party deletes
the key shares and the CRS metadata that its own resharing wrote under the new
epoch. The party deletes the epoch data and forgets the epoch only once the
epoch holds no key share and no CRS metadata. The `VerifiedCrsMaterial`
constructor checks the expected digest and deserializes the CRS from the same
bytes. A reshare retains the exact public bytes that it verifies, and restores
missing material during its locked storage phase. A failed reshare deletes only
public material that its storage phase created. It deletes that material only
after private cleanup succeeds. One lock serializes all reshare storage and
rollback on a party. Public material that existed before the storage phase
remains unchanged. A failed private deletion keeps the epoch and its public
material so cleanup can be retried. `DestroyMpcEpoch` erases a whole epoch
instead, and covers the material of every request. `DestroyMpcContext` takes
a stable snapshot of the context's registered epochs and erases their secret
shares before it forgets the context and removes its TLS trust-root references.
A trust root remains if another live context uses it. This order leaves no
usable key shares after the party set retires. The kms-connector is the source
of truth for which epochs belong to a context. In-memory lifecycle leases
serialize creation against destruction. `NewMpcEpoch` holds shared leases for
its target context and epoch through all PRSS, resharing, and persistence work.
A reshare also holds shared leases for its source context and epoch.
`DestroyMpcEpoch` and `DestroyMpcContext` require exclusive leases before
taking snapshots or deleting data. A conflicting destruction is refused with
`FailedPrecondition`, including while PRSS is still running and the new epoch
has not yet been registered in the session maker; callers retry once creation
has settled. MPC context updates serialize the existence check with storage and
cache or session updates. A failed deletion keeps the in-memory context if its
persistent entry remains, which permits a retry before or after restart.
The TLS verifier stores one trust root per context and MPC identity. Different
trust roots for one identity coexist while their contexts remain active. Each
root is evaluated with only its context's PCR allowlist, and a handshake
succeeds when one complete root and PCR check succeeds. The verifier rejects
registration for an active context ID instead of replacing its trust roots or
PCR values. Removing a context removes its roots, while another context's copy
of the same root remains trusted.
- **Session management** — creation, result retrieval, and cleanup for
long-running threshold sessions.
EIP-712 signature validation on user-decryption requests is shared between
the server and in-browser verifiers via the `validation_wasm` build.
## Deployment modes
Mode is selected in the server TOML config — a party runs in threshold mode
when the optional `[threshold]` section is present; see the sample
files in `core/service/config/` (`default_centralized.toml`,
`default_1.toml`..`default_4.toml`, and the compose-specific variants).
In both modes, when the same key or CRS ID has metadata under multiple epochs,
startup loads the metadata from the greatest epoch ID into the result meta
store. Epoch IDs are compared as big-endian integers.
### Centralized
A single `RealCentralizedKms` instance holds all key material. No MPC; keys
live in the configured vault backend. Preprocessing / reshare RPCs are not
applicable.
A centralized node keeps no epoch registry and stores no `EpochData`, so every
epoch-scoped entry it holds belongs to the default epoch. `KeyGen`, `CrsGen`,
`KeyGenPreproc`, `NewMpcEpoch`, `PublicDecrypt` and `UserDecrypt` reject any
other epoch ID with `InvalidArgument`. A request that omits the epoch ID still
falls back to the default epoch, so a caller that never sets the field is
unaffected.
### Threshold
`n` parties each run a `ThresholdKms` server. Each party holds a secret share
of the FHE secret key and participates in the MPC protocol for every
sensitive operation. Parties reach each other over gRPC via
`threshold-networking` (typically with mTLS using certs generated by
`kms-gen-tls-certs`). Preprocessing runs asynchronously and produces material
consumed by the online phase.
## Backup and recovery
Long-term private material held by a KMS node — signing keys, FHE secret-key
shares, custodian / MPC context state — is automatically backed up so that a
node whose local storage is lost can be rebuilt without reconstructing the
whole cluster. Secrets are wrapped into versioned `BackupCiphertext`s
(tagged by `RequestId` and `PrivDataType`) and written to the configured
backup vault, typically S3.
The payload-wrapping key is protected by one of two **keychains**, selected
in server config and unified behind `KeychainProxy`
([core/service/src/vault/keychain/](../core/service/src/vault/keychain/)):
- **`AwsKms`** — wrapping key is an AWS KMS CMK. Default and bootstrap path.
- **`SecretSharing`** — wrapping key is Shamir-shared across a set of
**custodians**, offline entities who each hold a key share plus a BIP39
seed phrase. `NewCustodianContext` requires the backup vault to be configured
with this keychain already, and the keychain can only encrypt once that call
has installed a context, so a node configured for it makes no backups until
its first context exists. New custodian contexts are rejected unless every custodian
encryption key is unique and no two custodians share any verification key, unless every
custodian encryption key uses `BACKUP_PKE_SCHEME`, and unless every custodian
publishes a verification key for every scheme in `BACKUP_SIGNING_SCHEMES`.
Every key in this path is MLKEM1024-P384 (`backup::BACKUP_PKE_SCHEME`), and the
custodian's is derived from 256 bits of seed-phrase entropy — a 24-word mnemonic —
so the phrase does not cap the scheme's security level. A vault written under an
older ML-KEM-512 context is not readable by a node holding a composite key, but
each ciphertext carries its own `pke_type`, so a vault spanning both schemes
decrypts as long as the matching key is installed. That remains true for
material already written; what is refused is *creating* a new context under a
weaker scheme.
Custodian workflows are driven through the
[kms-custodian](../core/service/src/bin/kms-custodian.rs) CLI and the
`NewCustodianContext` / `DestroyCustodianContext` / `CustodianRecoveryInit`
/ `CustodianBackupRecovery` RPCs defined in
[kms-service.v1.proto](../core/grpc/proto/kms-service.v1.proto).
A separate `RestoreFromBackup` RPC completes restoration on the node for the non-custodian AWS-KMS path.
The `RecoveryValidationMaterial` describing a custodian context — the custodian-signcrypted
shares of the backup decryption key, plus the commitments and the context itself — lives in the
**backup vault**, as the one object there that the keychain does not encrypt: it is what recovery
needs in order to reconstruct that very key, so encrypting it under the key would be circular. Its
integrity comes from the operator's composite signature under `BACKUP_SIGNING_SCHEMES`, checked at
startup once the signing key is available, together with a check — applied on every load — that the
object is stored under the context id its payload names. A node without a root seed cannot derive
the ML-DSA key, so it fails that startup check as soon as its vault holds any recovery material. The
material also embeds the operator's backup key set (`operator_verf_keys`), for recovery mode below.
It sits outside the `