# BYO — the whole system, in sequence

Every part of a BYO host's life, from an owner adding someone else's Spark on the dashboard to the host being handed back — and what keeps it serving in between. This document is the source of truth for how the BYO system behaves: where code, specs or other docs disagree with it, they are the ones to change.

1. **A**–**H** are the phases, from an owner adding a machine to the two ways it can end. In onboarding an operator appears once, to approve (**D**).
2. **I**, **J** and **M** are out of band, and none of them happens on an ordinary onboarding:
   1. **I** is the operator's way *in* — break glass — from the moment `operator-access` completes.
   2. **J** is how a host gets a newer agent than the one it was installed with.
   3. **M** is how a change to who the operators are reaches every host, rented ones included.
   4. **N** is how a host whose enrolment completed ends for good, so its machines can be enrolled into another host.
3. **K** is a variant, not a phase: what **A**–**D** become when the host is more than one machine — **K1** the pair (`spark-2`), **K2** the four-Spark switchless ring (`spark-4`).
4. **L** is the host after go-live: a new binary, either machine restarting, and a reboot of both.

Six areas, one continuous numbering across the whole document, so **step 44** means one thing and **E** means one phase.

> The Python agent in `platform/agent/`, and the hosts that run it, are out of scope. The BYO agent is `platform/byo-agent/`.

## The shape, and why

**Everything is provisioned before the box is ever touched.** Clicking "add a
GPU" allocates the relay slot, the address, the tunnel and the `hosts` row — not
the box enrolling, and not an operator afterwards. By the time the owner runs
the one-liner, the only thing missing is the machine itself.

That ordering is what makes the rest fall out:

1. **The agent enrols once and keeps going.** Its single call returns every
   credential, so it goes straight into preparing the host. No second start, no
   waiting, nothing on the box watching for anything.
2. **The operator appears once**, at the end, to approve a host they can read
   the whole preparation of — and that approval is the only thing between an
   enrolled machine and a paying customer.
   - An earlier draft had an operator provisioning each host by hand between
     enrolment and preparation. That put a person in the middle of every
     onboarding, which is precisely what BYO is supposed to remove.
3. **Each host gets its own relay, created for it here.** Same topology as the
   fleet has today — one relay per host — except that building it is an API call
   rather than a person following [new_spark.md](../new_spark.md) §1–§2.
   - The instance configures itself from `relay/gcp/bootstrap.sh` at first boot,
     so nobody SSHes into a relay to bring it up.
   - It boots while the owner is still reading the one-liner, so its start-up
     costs nothing anybody waits on.

## The six areas

Everything belongs to one of six places. The diagrams group them, because which
side of a boundary a step happens on is usually the point — a credential minted
in an edge function and a credential minted on the host are very different
facts.

Who belongs to each area — every actor's ID and label, and the verbs they use — is in the [BYO glossary](byo_glossary.md), and only there.

**1. The owner's Spark host** — the machine being onboarded, and the programs on it. The agent is two of them, because preparing needs root and serving must not have it.

**2. `app.enverge.ai`** — the Vercel deployment: the owner's and operator's
pages, the static installer assets, and the proxy that redirects to Blob.

**3. Supabase edge functions** — everything with authority. Nothing here runs
on a host.

**4. The Supabase database** — `host_enrollments`, `host_profiles`, `hosts`,
and `vms` from **E** onwards. Written by edge functions under the service role;
read by the pages under RLS. There is no relays table: a relay is named by its
host's slug, so the `hosts` row is its record, exactly as it is for the
hand-built fleet relays.

**5. Relays** — one per host, created during provisioning. Includes the GCP API
that builds the instance, reserves the address and points it at the relay. Same
topology and naming as the fleet runs today — `<slug>-relay` and
`<slug>-relay-ip`; the difference is that nobody builds it by hand.

**6. Cloudflare** — the operator tunnel and DNS.

And three people, none of whom belong to an area: the owner, the operator and the customer.

The customer never learns any of this. They land on a `spark-1` host and cannot
tell whose hardware it is, which is the tenancy decision spec63 makes
explicitly rather than by omission.
## A. Add a GPU — the backend provisions everything (1–16)

```mermaid
sequenceDiagram
  autonumber 1
  actor OWN as Owner
  box app.enverge.ai
    participant HOSTS as /host
  end
  box Supabase edge functions
    participant PRV as host-provision
    participant RLY as relay-provision
    participant ONB as admin-onboard-host
    participant HST as admin-host-action
  end
  box Supabase database
    participant DB as Postgres
  end
  box Relays
    participant RELAY as The host's relay
    participant GCP as GCP API
  end
  box Cloudflare
    participant CF as Cloudflare
  end

  Note over OWN,PRV: only a community_host (or superadmin) sees the button, and host-provision refuses anyone else
  OWN->>HOSTS: signed in, click "add a GPU"
  HOSTS->>PRV: POST — the caller's identity and the profile they picked
  PRV->>DB: read the existing slugs, mint the next in the profile's family
  PRV->>RLY: provision a relay for this slug
  RLY->>GCP: create the relay instance, reserve an IPv4, forward tcp/22
  RLY->>RELAY: bring up WireGuard, install the host's peer
  RLY-->>PRV: endpoint, the relay's WireGuard key, 10.8.0.2, public IPv4
  PRV->>ONB: mint the operator tunnel for this slug
  ONB->>CF: tunnel, ingress ssh://localhost:22, DNS route
  PRV->>HST: create {slug, gpu_type, agent_url, ssh_ipv4, active false}
  HST->>DB: insert hosts, active = false
  HST->>CF: wildcard A record for the slug
  HST->>DB: record the zone that record covers
  PRV->>DB: insert host_enrollments — owner, token hash, host_id, wg_conf
  PRV-->>HOSTS: the enrolment token, returned once
  HOSTS-->>OWN: curl app.enverge.ai/install-agent | sudo sh -s TOKEN, good for an hour
```

1. **Nothing here needs the machine.** BYO accepts one hardware configuration,
   so the relay slot, the address and the sizing are known before a box exists.
   The box contributes one thing: proof it really is that configuration, checked
   at step 28.
2. **Step 11 writes `active = false`**, and nothing writes `true` except the
   operator at the end.
3. **Relay provisioning is its own edge function.** It is the only part of this
   that talks to two systems we do not otherwise touch from here — the GCP API,
   and the relay itself.
   - `host-provision` stays an orchestrator, and replacing a dead relay later
     does not have to go through host onboarding to get one.
   - It is the slowest thing in the flow and the only one nobody waits on: the
     relay finishes booting while the owner is still finding a terminal.
4. **Provisioning ahead of time costs something.** An owner who clicks and never
   runs the command leaves a relay, an IPv4, a tunnel and an inactive `hosts`
   row allocated. Four things contain it:
   - only an owner granted `community_host` from the `/admin` users tab (or a
     superadmin) can click at all,
   - a cap of three outstanding allocations per owner,
   - a token that expires within the hour,
   - `host-provision`'s `reap`, which hands back everything a lapsed or rejected
     enrolment was holding. **Nothing calls `reap` on a schedule yet.**
5. **This has been run against real GCP**, not just written: address reserved,
   instance created, metadata read, bundle fetched, `wg0` up with the host's
   peer, `ssh_mapd` on `10.8.0.1:8754`, Caddy on `:80`/`:443`, and the
   `LOFISPARK_VMSSH` chain seeded with the IAP exemption and a blackhole
   default.
   - It found one bug: `accessConfigs[].natIP` takes the address literal, not a
     self-link, so every create was failing 400.
   - Still unproven is time-to-certificate, which needs the Cloudflare record
     and so a deploy.

## B. Install and enrol — one shot (17–30)

```mermaid
sequenceDiagram
  autonumber 17
  actor OWN as Owner
  box Owner's Spark host
    participant INS as Installer
    participant SD0 as systemd on node0
    participant PRP as Agent — preparing
  end
  box app.enverge.ai
    participant APP as the app
    participant BLB as Vercel Blob
  end
  box Supabase edge functions
    participant ENR as host-enroll
  end
  box Supabase database
    participant DB as Postgres
  end

  OWN->>INS: run the one-liner, token and all
  INS->>INS: refuse immediately if no token was passed
  INS->>APP: GET /install-agent, then the binary and units
  APP-->>INS: 302 to Blob for the binary
  INS->>BLB: fetch it
  INS->>INS: verify the checksum, or install nothing at all
  INS->>INS: write the token to /etc/enverge/enrol-token, 0600
  INS->>SD0: install, daemon-reload, enable, start
  SD0->>PRP: start
  PRP->>PRP: Ed25519 keypair, then fingerprint the hardware
  PRP->>ENR: POST {token, pubkey, fingerprint}
  ENR->>DB: redeem — lock the row, require unspent and unexpired,<br/>check the fingerprint against host_profiles
  ENR-->>PRP: 200 with every credential: host token, wg_conf,<br/>tunnel token, host profile, the signed operator keys
  PRP->>PRP: write credentials.json 0600, delete the spent token
```

As built: [the agent's side of enrolment, generated from its tests](../../../byo-agent/internal/enroll/testdata/sequences.md#1-first-enrolment-single-host). Scenario 1 enrols, 2–4 are refused, and 5 finds the host already enrolled.

1. **Step 28 is the single-use point.** The row is locked, so two machines
   running the same one-liner cannot both become this host.
   - Unknown, spent and expired are one answer: which of the three it was is not
     something a caller has earned.
   - Hardware that is not a GB10 Spark is rejected here, and the allocation it
     was holding is released.
2. **Step 29 is the whole payoff of provisioning first.** One call returns
   everything, so the agent has no reason to stop and nothing to come back for.
3. **The operator keys are read at redeem, not at provisioning.** Step 29 returns
   the signed list as it stands when the box enrols, and the public key its
   signature is checked against (**M**).
   - Enrolment is only the first delivery. Every change after it arrives
     through **M**.

## C. Prepare (31–38)

The agent runs the whole of [new_spark.md](../new_spark.md) §4 and §6 as a state
machine. Every step has the same shape, and the shape is the point:

```mermaid
sequenceDiagram
  autonumber 31
  box Owner's Spark host
    participant PRP as Agent — preparing
    participant N0 as node0
  end
  box Supabase edge functions
    participant REP as host-report
  end
  box Supabase database
    participant DB as Postgres
  end

  loop the 19 steps, in order, stopping at the first that will not complete
    PRP->>N0: check — is it already like this?
    alt already satisfied
      PRP->>REP: already-satisfied
    else needs applying
      PRP->>N0: apply
      PRP->>N0: check again — verify rather than trust the apply
      PRP->>REP: applied, or failed with the reason
    end
    REP->>DB: write hosts.prepare_state
  end
  PRP->>REP: self-check — machine-check every verification at once
  REP->>DB: the host is ready for review
```

As built: [the prepare run, generated from the agent's tests](../../../byo-agent/internal/prepare/testdata/sequences.md#1-every-path-a-step-can-take). Scenarios 1–3 take the runner down each of its paths, and 4 is the real runbook.

1. **Check and apply are separate**, for two reasons:
   - a re-run changes nothing on an already-prepared host, which is the property
     the whole machine rests on,
   - a **dry run** calls only `check`, so it reports what it would change while
     touching nothing.
2. **The re-check after apply is not belt and braces.** A step that reports
   success without its check passing is how a host arrives "prepared" and
   broken, and nothing later would notice.
3. **It stops at the first step that will not complete.** The orderings below
   are load-bearing, so a host that failed `lxd-install` must never reach
   `owner-lockout`.

### The 19 steps

| # | Step | Notes |
|---|---|---|
| 1 | `preflight` | Linux, Ubuntu, `uname -m` and `nvidia-smi` match the allocated profile. **Fail closed** |
| 2 | `docker-purge` | purge dockerd, `iptables -P FORWARD ACCEPT` |
| 3 | `swap-off` | `swapoff -a`, comment `/etc/fstab`, mask `swap.img.swap` |
| 4 | `earlyoom` | the profile's floor, verbatim ([spec28](../../specs/spec28.md), retuned [spec51](../../specs/spec51.md)) |
| 5 | `zfs-arc-cap` | `/etc/modprobe.d/zfs.conf`, 4 GiB |
| 6 | `headless` | `multi-user.target`, mask `gdm` |
| 7 | `watchdog` | `RuntimeWatchdogSec` + the boot-reset reporter ([spec46](../../specs/spec46.md)); **skipped** when the profile says `watchdog_sec: 0` (a guest with no watchdog device), **fails closed** when it wants one and `/dev/watchdog0` is missing |
| 8 | `operator-access` | `spark` user, every operator's key, NOPASSWD sudo, the tunnel minted back in **A**. `authorized_keys` is written as the whole list, so it holds exactly the operators and nobody else |
| 9 | `lxd-install` | apt `lxd`; **do not** rely on `lxd init --storage-backend` |
| 10 | `lxd-network` | `lxc network create lxdbr0`; **`lxd init` is not run, so nothing else makes it** |
| 11 | `firewall` | if `ufw` is active: allow in + route on `lxdbr0`, allow in on `wg0`, `DEFAULT_FORWARD_POLICY=ACCEPT`, reload. No ufw or inactive = satisfied |
| 12 | `storage-pool` | explicit `lxc storage create default zfs` |
| 13 | `nvidia-runtime` | `nvidia-container-toolkit-base` + `libnvidia-container-tools`, for `nvidia.runtime` injection |
| 14 | `gpu-accounting` | NVML per-process accounting. Driver state, so re-asserted |
| 15 | `profiles` | `lofispark-base` + `lofispark-gpu` |
| 16 | `base-image` | import the prebuilt image, else build it |
| 17 | `relay-peer` | apply the WireGuard peer collected in **B**; **collision-check first** |
| 18 | `owner-lockout` | gated on `hosts.lockout_owner`, **ships false** |
| 19 | `self-check` | machine-check every verification, and report |

**Orderings that are landmines, not preferences:**

Named rather than numbered, because a step inserted in the middle renumbers
every landmine below it and the numbers are what a reader trusts.

- **`operator-access` before `owner-lockout`, and verified.** Prove Enverge's
  access before removing the owner's. The reverse strands a box neither party
  can administer.
- **`docker-purge` before `base-image`.** Docker's `FORWARD` DROP breaks
  `lxdbr0` NAT, so the image build loses egress to the NVIDIA repo.
- **`zfs-arc-cap` before `lxd-install`.** `modprobe.d` binds at module load, and
  installing LXD loads `zfs` even for a non-zfs pool.
- **`swap-off` is what arms `earlyoom`.** It fires only when `MemAvailable`
  *and* free swap are low — an AND. A GPU runaway is unswappable, so leaving
  swap on disarms the floor entirely.
- **`lxd-network` before `profiles`.** `lofispark-base` attaches eth0 to
  `lxdbr0`, and `lxd init` — which would have made both the pool and the bridge
  — is deliberately not run. Creating the pool explicitly without the bridge
  gives a host with LXD up, a healthy pool, and no way to attach a NIC; the
  failure then surfaces at `profiles`, several steps after the omission.
- **`firewall` before `base-image` and `relay-peer`.** A default-deny `ufw`
  (cloud images and DGX OS alike) drops the bridge's DHCP, forwarded container
  egress and the relay's packets, and none of the three failures names the
  firewall: the image build dies resolving `archive.ubuntu.com` and the control
  plane just hangs. `DEFAULT_FORWARD_POLICY` has to change in the file —
  `docker-purge`'s `iptables -P FORWARD ACCEPT` lasts only until the next
  `ufw reload`.
- **`base-image` builds from a script the agent carries.** On the fleet it
  arrives with `make sparkN-copy`; a BYO host has nobody to run that, so the
  builder and its login banner are embedded in the binary. It is run with the
  profile's `cuda_repo_arch` — the script defaults to `sbsa`, which is arm64's
  NVIDIA repo and wrong everywhere else.
- **`storage-pool` must create the pool explicitly.** A `dir` pool costs ~69 s
  on every VM create against 0.6 s on zfs, on a host that looks perfectly
  healthy.
- **`base-image` must end with the alias.** The lifecycle launches by alias; an
  image without one fails every create.
- **`relay-peer` must check for a subnet collision first.** `10.8.0.0/24` is a
  common home-router default, and bringing `wg0` up into a colliding range makes
  the host unreachable the instant it succeeds.

## D. Approve (39–40)

```mermaid
sequenceDiagram
  autonumber 39
  actor OPS as Operator
  box app.enverge.ai
    participant ADM as /admin
  end
  box Supabase database
    participant DB as Postgres
  end

  OPS->>ADM: read the prepare report — every step, and how it went
  ADM->>DB: active = true, with the pool it should serve
```

1. **This is the only human decision in the flow**, and the only thing between
   an enrolled machine and a paying customer on the shared `spark-1` tier.
2. **`gpu_type` is not decided here.** It was fixed when the allocation was made
   — BYO accepts one hardware configuration, and the box proved it was that at
   enrolment. Activation is the single flag, plus the placement chosen with it.
3. **The backend refuses to activate a host whose preparation did not finish.**
   - The rule used to live in the render condition of an "approve" button on the
     enrolments table, so the hosts table — a second control issuing the same
     update — walked straight past it: a host that stopped at `storage-pool`
     could be activated by anyone who used the other one.
   - A gate only one of two paths honours is not a gate, so it moved into
     `admin-host-action` and the duplicate button is gone.

## E. Create an instance (41–74)

Five callers, three shapes. What separates them is **how the host is chosen**,
because that decides when the row can be written and where the SSH hostname
comes from.

| # | Trigger | Who calls | Drawn in |
|---|---|---|---|
| 1 | the customer asks | the proxy (**PXY**) | **E1** — the picker chooses, and the customer is waiting |
| 2 | capacity frees and the queue drains | `drain-waitlist` | **E2** — no host until one accepts |
| 3 | a paid pre-book window opens | `drain-prebook` | **E3** — the booking pinned its host when it was made |
| 4 | an operator places one on a dedicated host | `admin-vm-action` | **E2**'s shape, host supplied instead of rotated |
| 5 | an operator launches a booking by hand | `admin-prebook-action` | **E3**'s shape, one booking by id |

The two operator paths are the manual twins of the two automatic ones — same
steps, with the cron's gate replaced by a superadmin and the selection replaced
by whatever the operator named. That is why they are not drawn separately:

1. **4 is 2 done by hand.** A dedicated owner often lands in the general queue,
   because the picker cannot reach their box; the operator places the VM there
   and closes the queued row the drain would have closed.
2. **5 is 3 done by hand**, targeted by booking id, and with a `force` that
   revives a booking torn down inside its own window.

**The row is a precondition for dispatch, in all five.** On a host whose agent
writes none, a create sent without one produces a container the database has no
record of: unbilled, invisible to the customer, and holding capacity the picker
cannot see. Every caller now refuses rather than dispatches, keyed on the same
`pubkey` marker — a fleet host is exempt, because its agent inserts the row a
moment later and refusing there would break a path that works.

To the customer every one of these is identical to any other host, which is the
point of BYO joining the `spark-1` pool rather than a tier of its own. They are
not identical in who writes the row: see **G**, where that difference is the
whole story.

### E1. From the dashboard (41–52)

A customer's create, once the host is `active`.

```mermaid
sequenceDiagram
  autonumber 41
  actor CUS as Customer
  box app.enverge.ai
    participant PXY as the proxy
  end
  box Supabase database
    participant DB as Postgres
  end
  box Relays
    participant RELAY as The host's relay
  end
  box Owner's Spark host
    participant SRV as Agent — serving
  end

  CUS->>PXY: POST /api/vms + Bearer JWT
  PXY->>DB: billing, then can_launch_gpu, then pick_hosts_for_gpu
  DB-->>PXY: this host — active, no live VM, least recently used
  PXY->>DB: rpc create_vm — stamps created_at, mints the id and hostname
  PXY->>RELAY: POST /vms + the host's agent_token
  RELAY->>SRV: over WireGuard to 10.8.0.2:8080
  SRV->>SRV: launch the head, wait for sshd, install the customer's key, push the banner
  SRV->>SRV: check GPU health and start the startup script — advisory, never fatal
  SRV->>RELAY: forward SSH to the instance, at the port it published
  SRV-->>PXY: 201 with the port
  PXY->>DB: complete the row
  PXY-->>CUS: the instance, and where to SSH
```

As built: [the agent's side of a create, generated from its tests](../../../byo-agent/internal/vm/testdata/sequences.md#1-create-single-host). Scenarios 1–2 create an instance, and 3–5 refuse and revert.

- **Step 44 stamps `created_at` before the host is contacted**, so billing runs
  from when capacity was committed rather than from whenever a host answered. A
  create that is never answered is marked `creation_failed_at` and bills nothing.
- **Step 49 is the agent programming the relay, and it had to be.** An earlier
  draft had the backend do it, to keep the last secret off the host. It cannot:
  `ssh_mapd` listens on `10.8.0.1:8754`, inside the tunnel, and Caddy proxies
  only `:8080` — there is no route from Vercel. Publishing a second endpoint on
  every relay to create one would be more attack surface than the thing it
  protects.
  The token is the **relay's own**, minted per relay, and on a per-host relay it
  can forward one host's `:22` to one port on that same host — which the agent
  decides anyway, being what publishes the port. It grants nothing a compromised
  host does not already have. **A shared relay would break that**, since one peer
  could forward another's `:22` elsewhere, so a pool must move this back to the backend
  first.
- **It is fatal, unlike the GPU check.** An instance whose `:22` goes nowhere is
  not an instance. The fleet's Python agent logs a map failure and carries on,
  which is how a host ends up healthy with a blackholed `:22` and nobody
  noticing.
- **Steps 47–48 split fatal from advisory.** An instance the customer cannot
  reach is not an instance; a GPU check or a startup script that fails is
  reported, and the instance is created anyway.

### E2. From the queue, when capacity frees up (53–64)

The queue is the only path that creates a VM with nobody present, acting on an
intent that may be days old — which is why billing is re-checked live rather
than trusted, and why the host is not known when the row is written.

```mermaid
sequenceDiagram
  autonumber 53
  actor CUS as Customer
  box app.enverge.ai
    participant PXY as the proxy
  end
  box Supabase edge functions
    participant DRN as drain-waitlist
  end
  box Supabase database
    participant DB as Postgres
  end
  box Owner's Spark host
    participant SRV as Agent — serving
  end

  CUS->>PXY: DELETE an instance — capacity for this gpu_type frees
  PXY->>DRN: POST {gpu_type}, fire and forget
  DRN->>DB: pick_hosts_for_gpu — every idle host, rotation order
  DRN->>DB: the queued row the first candidate's order chooses
  DRN->>DRN: Stripe status for that user — active or trialing only
  DRN->>DB: rpc create_vm — no host yet, so no hostname yet
  loop each candidate, until one takes it
    note over DRN,SRV: no row, and this host writes none — skip it, do not bench it
    DRN->>DRN: mint token.zone for THIS candidate, and write nothing
    DRN->>SRV: POST /vms + the name, over that host's relay
    SRV-->>DRN: 201 — readable or not — or a refusal that benches this host
  end
  DRN->>DB: only now — host_id, ssh_key_id, and the name
  DRN->>DB: close the waitlist row as fulfilled
  DRN->>DB: email the queued customer that it is ready
```

- **Step 58 inserts the row with no host, and that is the point.**
  `pick_hosts_for_gpu` excludes any host holding a live `vms` row, so a
  `host_id` written before a launch succeeded would make this row occupy the
  candidate that just refused it — benching a working host against every
  concurrent drain until the row is marked failed.
- **Step 59 mints per attempt, not once.** A name under the previous candidate's
  zone resolves to that host's relay, and the customer would land nowhere. It is
  minted in the function rather than by the database: a random label joined to a
  column, where TypeScript is type-checked, deployed by CLI and unit-tested.
- **Step 62 is the first write, and it is the whole difference from the
  dashboard path.** Every other create knows its host at insert, so `create_vm`
  stamps the host and the name in one statement. This one cannot, so it stamps
  them once a host has actually taken the VM.
- **The name is written only if one was minted.** A host whose own agent mints
  hostnames wrote that column back at step 60, and an empty string over it would
  take the address away from an instance already running.
- **A candidate has three outcomes, not two.**
  - **Refused** — benched, and the loop moves on. This is what stops one
    drifted host, agent-busy while the database read it as idle, stalling the
    whole queue.
  - **Skipped** — step 58 could not write the row and this host's agent writes
    none, so a container here would be one nothing can account for. Not benched:
    the host did nothing wrong, and the fleet hosts behind it are still tried.
    A fleet host is never skipped, because its agent inserts the row itself.
  - **Accepted** — including an answer that will not parse. A 2xx means the
    container exists, and the id, name and hostname were all *sent*, so the
    values are known without reading them back.
- **An unreadable 2xx must not be scored as a refusal.** It was: the parse sat
  inside the catch that scores attempts, so a successful launch benched a working
  host and sent the next candidate to build a second container for the same queue
  row. The guard that would normally catch that reads a row the *agent* writes,
  so on a BYO host it finds nothing and the loop carries on.

### E3. From a paid pre-booking, when the window opens (65–74)

A booking buys a named box for a date range. The window opening is the trigger,
so nobody is present — and the host was decided when the booking was made, not
when it launches.

```mermaid
sequenceDiagram
  autonumber 65
  participant CRON as Daily cron
  box Supabase edge functions
    participant PBK as drain-prebook
    participant NOT as notify-prebook
  end
  box Supabase database
    participant DB as Postgres
  end
  box Owner's Spark host
    participant SRV as Agent — serving
  end

  CRON->>PBK: tick
  PBK->>DB: earliest paid, open booking whose window covers today
  opt it has launched before
    PBK->>DB: is that VM still live? a soft delete keeps the id
  end
  PBK->>DB: the booker's customer key, from their account
  PBK->>DB: rpc create_vm on the host the booking pinned
  PBK->>SRV: POST /vms + the name, as the booker
  SRV-->>PBK: 201 with the port
  PBK->>DB: stamp launched_vm_id — the booking stays open
  PBK->>NOT: the launch email
  NOT->>DB: notified_at
```

1. **The gate is `paid_at` alone**, and no live Stripe check. Payment is
   collected off-band and recorded by ops, which is the deliberate difference
   from **E2** — and the reason a stale `paid_at` is a known risk rather than a
   caught one.
2. **The host is pinned on the booking, not picked.** A booking without a pin is
   a legacy row and falls back to the pool pick.
   - An unavailable pinned host leaves the booking open for the next tick and an
     operator to reassign. A paid window is not re-placed on a different box
     just because the intended one is offline.
3. **Step 67 is a self-heal, not an idempotence check.** VMs are soft-deleted,
   so `launched_vm_id` never nulls — without reading the row's liveness, a
   customer who deleted their box mid-window would be stranded until `end_date`.
   A non-live row falls through and the booking relaunches.
4. **The booking stays open after launching.** `expire-prebook` closes it at
   `end_date` and tears the VM down, which is trigger 5 of **G**.
5. **The customer key is the booker's own**, resolved from `ssh_keys` by `user_sub`
   rather than a per-booking id, so it can only ever be a key they hold. No key
   on the account leaves the booking open for an operator to follow up.

## F. Customer SSH (75–79)

No proxy, no agent, no HTTP. L4 forwarding only.

```mermaid
sequenceDiagram
  autonumber 75
  actor CUS as Customer
  box Cloudflare
    participant CF as Cloudflare
  end
  box Relays
    participant RELAY as The host's relay
  end
  box Owner's Spark host
    participant HEAD as Head container
  end

  CUS->>CF: token.slug.ssh.enverge.dev
  CF-->>CUS: the relay's IPv4 — a wildcard per host
  CUS->>RELAY: TCP :22
  RELAY->>HEAD: DNAT via LOFISPARK_VMSSH, over WireGuard, MSS clamped
  HEAD-->>CUS: sshd — banner, customer key auth
```

1. **Standard port 22, no client software.** The whole reason customer SSH was
   never moved onto the operator tunnel.
2. **Until the map has been written, `:22` is a blackhole.** Not an error —
   nothing is listening for that address yet. The agent:
   - writes it on create,
   - clears it on delete and stop,
   - re-asserts it when it starts, because a relay that restarted came back with
     an empty chain, and a host that rebooted has a container up the relay knows
     nothing about.
3. **The owner's box needs no inbound anything.** It dialled out; every address
   here is ours.

## G. Delete an instance (80–85)

Six callers, one shape. They differ only in what decided the instance should go.
Three of them (3–5) run on a schedule with nobody present, which is why the
shape has to be right without anyone reading the result.

| # | Trigger | Who calls |
|---|---|---|
| 1 | the customer asks | `DELETE /api/vms/{id}` through the proxy (**PXY**) |
| 2 | an operator, from `/admin` | `admin-vm-action` |
| 3 | a payment probe fails mid-dunning | `probe-payment` |
| 4 | dunning runs out | `teardown-vms`, the gated reaper |
| 5 | a pre-book window ends | `expire-prebook` |
| 6 | an operator ends a pre-booking early | `admin-prebook-action` |

```mermaid
sequenceDiagram
  autonumber 80
  participant DEL as Whichever decided
  box Supabase database
    participant DB as Postgres
  end
  box Owner's Spark host
    participant SRV as Agent — serving
  end
  box Supabase edge functions
    participant DRN as drain-waitlist
  end

  DEL->>DB: delete_requested_at — before the host is contacted
  DEL->>SRV: DELETE /vms/{id}, as the owner, over the host's relay
  alt 2xx, or 404 — the host has no such container
    SRV-->>DEL: deleted
    DEL->>DB: deleted_at — billing stops here
    DEL->>DRN: the freed gpu_type, so the queue moves
  else any other answer, or none at all
    SRV-->>DEL: nothing is written, and the row stays live
  end
```

As built: [the agent's side of a delete, generated from its tests](../../../byo-agent/internal/vm/testdata/sequences.md#6-delete-pair). Scenario 6, and 7 for an instance that is already gone.

1. **Two facts, deliberately separate.**
   - `delete_requested_at` — the backend accepted a delete. Stops nothing.
   - `deleted_at` — the container is *confirmed* gone. Billing stops here.
2. **The row is not closed optimistically**, because the database decides
   routing. `pick_hosts_for_gpu` excludes any host holding a live `vms` row, so
   a row closed before its container dies leaves a host that *looks* free:
   launches keep routing to an agent that refuses them, and every user behind it
   stalls.
   - So the bias runs the other way. A delete that does not complete keeps
     billing and keeps its host marked occupied, which is the honest state — the
     capacity really is still held.
   - Over-billing is refundable; lost utilization is not, and a poisoned queue
     costs every waiting customer rather than one.
3. **That trade is only defensible because the stuck case is visible.** Step 80
   is what `notify-stuck-deletes` scans for, and it also sizes the refund: the
   customer gave the instance up when they asked, not when an operator noticed.
4. **A delete ends when its caller stops waiting.** The backend waits 90s, well past the ~30s a pair takes; cut off earlier, the agent leaves the instance gone but node1's rails bare and the relay still forwarding SSH to the dead port. A relay write abandoned that way is not recorded as a relay failure. Either way, every heartbeat reads the relay's SSH map (`GET /map`) and writes it again if it does not match what the host runs, so `relay` health is the relay as it is this minute, and a stale or cleared target is put right within a beat.
5. **A 404 confirms, it does not fail.** The host has no such container — the
   orphan case, where it went away and the agent's own write never landed. The
   caller asked for it gone and it is gone, so the row closes rather than
   becoming one nothing can ever clear.
6. **Step 83 must precede step 84.** `pick_hosts_for_gpu` excludes a host
   holding a live row, so draining before the row closes hides the host that
   just freed up from the customer waiting for it.
7. **Every caller owns both edges, and none may delegate them.**
   - The fleet's Python agent PATCHes `deleted_at` from the host, so a caller
     that wrote nothing still looked correct there.
   - A BYO host's Go agent writes nothing to the database at all, so the same
     caller leaves the row live forever — billing, holding its host, and with
     step 80 unwritten, not even stuck. Just invisible.
   - `admin-vm-action` was the last caller doing this. It is the same asymmetry
     that produced the empty `ssh_hostname` and the orphaned container: backend
     code assuming the agent is a database writer.
8. **Both writes are guarded** — `delete_requested_at=is.null` and
   `deleted_at=is.null` — so on a host whose agent *does* write, the two compose:
   whichever lands first wins and the other is a no-op. Two independent paths to
   the same value beat either alone.

## H. Revoke, and uninstall (86–95)

Two different endings. **Revoke** cuts a host off and keeps it configured;
**uninstall** hands the machine back.

```mermaid
sequenceDiagram
  autonumber 86
  actor OPS as Operator
  box app.enverge.ai
    participant ADM as /admin
  end
  box Supabase edge functions
    participant PRV as host-provision
  end
  box Supabase database
    participant DB as Postgres
  end
  box Owner's Spark host
    participant SRV as Agent — serving
  end
  box Relays
    participant RELAY as The host's relay
  end
  box Cloudflare
    participant CF as Cloudflare
  end

  OPS->>ADM: revoke, or uninstall
  ADM->>DB: active = false — no new placements
  opt a tenant is still on it
    ADM->>SRV: delete the instance, and wait for it to confirm
  end
  opt uninstall only
    ADM->>SRV: restore the owner, then remove the operator keys and units
    SRV-->>ADM: done — and this is the last we hear from it
  end
  ADM->>PRV: release everything provisioned
  PRV->>DB: clear token_hash — nothing the agent presents matches
  PRV->>RELAY: destroy the relay and its address
  PRV->>CF: remove the tunnel and the DNS records
  PRV->>DB: close the host row
```

1. **Order is the whole design.** De-pool first, drain second (**G**), and only
   then remove access — the reverse strands a tenant on a host nobody can reach.
2. **The host confirms before backend state goes.** Once the relay is gone there
   is no path to that machine, so anything still needed from it must be asked
   for first.
3. **Restore the owner before removing ourselves.** The mirror of prepare's
   `operator-access` before `owner-lockout`, for the same reason.
4. **Revoke leaves the machine configured and locked.** That is why it is not
   offboarding: an owner who simply wants to stop hosting needs uninstall, and
   "reinstall your machine" is a poor answer for a self-serve product.

## I. Break glass — reaching the host itself (out of band)

Everything above is the platform driving a machine nobody is sitting in front
of. This is the other thing: a person, on a host, because something is wrong.

1. **It is deliberately never on the happy path.**
   [spec63](../../specs/spec63.md) splits **knowing** from **acting**: every
   diagnostic is what the agent reported and what `/admin` renders, because "SSH
   in and look" does not survive a thousand strangers' boxes.
2. **Acting on one is what this is for** — push a binary, restart a unit,
   unwedge a prepare that stopped.
3. **There are two ways in**, and the second exists because the first can fail.

```mermaid
sequenceDiagram
  autonumber 96
  actor OPS as Operator
  box Cloudflare
    participant CF as Cloudflare
  end
  box Relays
    participant RELAY as The host's relay
  end
  box Owner's Spark host
    participant CFD as cloudflared
    participant SSHD as sshd on node0
    participant N0 as node0
  end

  note over OPS,N0: 1. The tunnel — the ordinary way
  OPS->>CF: ssh -o ProxyCommand='cloudflared access ssh<br/>--hostname slug.enverge.dev' spark@…
  CF->>CFD: down the connection the host dialled out
  CFD->>SSHD: localhost:22
  SSHD-->>OPS: auth as spark, with the operator's own key, NOPASSWD sudo
  OPS->>N0: act — push a binary, restart a unit

  note over OPS,N0: 2. The relay — when the tunnel is gone
  OPS->>RELAY: gcloud compute ssh slug-relay --tunnel-through-iap
  RELAY->>SSHD: ssh spark@10.8.0.2, over WireGuard
  SSHD-->>OPS: same user, same operator key
```

### Why there are two

1. **The tunnel is the one that scales, and the one that can vanish.**
   - It rides `cloudflared`, a unit the `operator-access` step installed with a
     connector token minted backend-side.
   - It needs nothing inbound on the owner's network, which is why it works
     behind a home router.
   - It needs Cloudflare, the host's outbound internet, and that unit alive.
     Lose any of the three and the door is shut.
2. **The relay is the one we own.**
   - A GCP VM on our project, so IAP reaches it whatever the host is doing, and
     WireGuard puts it one hop from `10.8.0.2`.
   - That hop matters more than it looks: the relay DNATs *inbound* public `:22`
     into the tenant's container through `LOFISPARK_VMSSH`, but that chain is on
     `PREROUTING`, so a connection **originating on the relay** is not caught by
     it — `ssh spark@10.8.0.2` from the relay lands on the *host's* sshd, not
     the customer's container.
3. **The two fail independently.** The tunnel dies with Cloudflare, the host's
   uplink, or one systemd unit; the relay path dies only if WireGuard is down —
   which is the same as the host being off the network entirely, at which point
   there is nothing to reach.

### What the operator gets, and why that is the posture

`spark`, with their own operator key and NOPASSWD sudo — so, root. Not a concession: it *is*
sole administrative control, the same asymmetry that lets `owner-lockout` remove
the owner's login and never ours. A BYO host has exactly one administrator and
it is us.

1. **One account, one operator key per operator.** Every operator logs in as `spark`;
   what tells them apart is the operator key. sshd logs the fingerprint of the key it
   accepted, and `operator_keys` records each key's fingerprint against the
   operator's user sub, so a login is attributable without a second account.
2. **Leaving is one action.** Removing an operator's key reaches every host
   through **M**, rented or not, so no host keeps a door open for someone who
   has gone.

### What has to be true first

| Depends on | Why it can be missing |
|---|---|
| `operator-access` completed **and verified** | It is verified precisely because `owner-lockout` runs after it. An unverified tunnel plus a completed lockout is a box neither party can administer |
| The host has outbound internet | Both paths are outbound-dialled. A host off the network is unreachable by design, and that is what "we detect rather than prevent" means |
| The operator's key is on the host | Delivered at enrolment, then by **M**. An operator key added while the host was offline arrives when an operator sends again |
| `relay-peer` applied, for path 2 | Before it, the only way in is the tunnel |
| The relay still exists | Releasing an allocation destroys it. **Order matters in H**: drain the tenant and take what you need from the host *before* the relay goes, because after that there is no path to that machine at all |

### When neither works

1. **Nothing.** That is the honest answer, and it is the design: a host that
   has been unplugged, reinstalled or taken back looks exactly like one that is
   merely offline, so **we trust the absence of a heartbeat rather than the
   presence of a good signal**. The response is to de-pool and revoke, not to
   try harder to get in.
2. **There is a third path while `hosts.lockout_owner` ships false** — the
   owner's own login. That is why it ships false: until uninstall exists, their
   login is the recovery path if this flow wedges a box.

## J. Upgrading the agent (out of band)

1. **Publishing is not delivery.** CI builds and publishes on every merge, and
   that changes nothing on any host.
2. **Nothing on a BYO host ever fetches a binary** — no timer, no polling, no
   update endpoint. The agent's steady state is to answer the backend and
   otherwise sit still.
3. **The one thing that moves unasked is `enverge-agent-rollback`**, and it
   moves *backwards* — once, when serving will not stay up: it consumes
   `.previous` and will not restore the build that is already running (**L1**).

So an upgrade is an operator action, over the same tunnel as **I**.

```mermaid
sequenceDiagram
  autonumber 104
  participant CI as GitHub Actions
  actor OPS as Operator
  box app.enverge.ai
    participant APP as the app
    participant BLB as Vercel Blob
  end
  box Cloudflare
    participant CF as Cloudflare
  end
  box Owner's Spark host
    participant INS as Installer
    participant SD0 as systemd on node0
    participant SRV as Agent — serving
  end

  note over CI,BLB: on merge to main
  CI->>BLB: publish both arches at the stable path, overwriting
  CI->>APP: request the path an installer follows — fail the release if it 503s
  note over CI,SRV: nothing on any host has changed

  OPS->>CF: make byo-upgrade BYO_HOST=slug
  CF->>INS: run the installer, detached, with no token
  INS->>APP: GET the binary, its checksum, and the units
  APP-->>INS: 302 to Blob
  INS->>BLB: fetch, and verify — or install nothing at all
  INS->>INS: keep the outgoing binary as .previous
  INS->>SD0: write the units, daemon-reload, restart serving — preparation is not started
  SD0->>SRV: serve, as spark, on the new binary — L1 from here
```

1. **GitHub Actions belongs to none of the six areas.** It publishes, and
   publishing is where its involvement ends — which is the point of the section.
2. **The release verifies the path an owner's installer follows, not the
   upload.** `put()` returning a URL says the object exists; it says nothing
   about whether `app.enverge.ai/agent/…` resolves.
   - That gap once let `AGENT_BLOB_BASE_URL` sit blank in production — present
     in the dashboard, so it read as configured — while every install one-liner
     was broken.
3. **No token, and that matters more than it sounds.** The installer asks for
   one only when the machine is not enrolled; an enrolled host reuses its
   credentials.
   - The alternative would be an operator obtaining a fresh token, and the only
     way to get one is to add a GPU — which provisions a *second* host, relay
     and tunnel.
4. **Detached, deliberately.** A dropped SSH session must not land between
   writing the binary and writing the units.
5. **Preparation starting serving is one line, and it went missing.**
   - The ordering between the two units is declared on the serving one —
     `Requires=`/`After=` the prepare unit — and that direction only pulls
     preparation in when serving starts, never the reverse.
   - So the installer's "restart prepare" started nothing else: spark-c05
     prepared itself, exited 0, and served nothing until it was started by hand.
   - The prepare unit's `Wants=enverge-agent.service` closes it. Earlier hosts
     hid it because they had been rebooted since installing.
6. **Restarting serving is the installer's job, not the `Wants=`'s.** On an
   upgrade the serving unit is already running, so `Wants=` is satisfied and
   systemd starts nothing — the host would keep answering from the binary the
   install just replaced.
   - An upgrade restarts serving; the preparation it pulls in exits at once.
   - A first install starts preparation only; its `Wants=` starts serving once credentials exist.

7. **An upgrade does not re-check the host.** Preparation does not run after go-live (**L**).
   - The done list is still scoped to the build that wrote it. Steps are recorded by name, and a name says nothing about what the step now does: `lxd-install` gained `usermod -aG lxd` while keeping its name, so hosts that had completed it skipped the new work and reported "already satisfied".
   - That is what makes the deliberate run in **L** re-check every step on a newer binary, rather than trust names an older one wrote.
8. **An upgrade does not reach containers that already exist.** The login banner
   is pushed at create, so a running instance keeps what it was given.
   Recreating it is the only way to change what a customer sees.
9. **One command covers every BYO host**, unlike the per-host targets the fleet
   carries.
   - Those exist because the fleet's agent is copied out of a working tree; a
     BYO host fetches the published binary itself, so the only per-host input is
     the slug.
   - The corollary: upgrading several hosts is a loop an operator runs by hand —
     fine at two, and the first thing to hurt as the fleet grows.

## K. A host of more than one node (118–134)

One `hosts` row, one relay, one tier slot, one enrolment, one token, one
command on node0, one prepare run, one heartbeat. Agentless satellites are
reached from node0 — **nothing is ever run on them by anyone.** One table,
`host_nodes`, says which machines a host is made of. **K1** is two machines;
**K2** is four on a switchless ring. Everything from **E** onwards is
unchanged either way.

### K1. Two nodes — spark-2

A `spark-2` host is two Sparks cabled together over ConnectX-7 and sold as one
unit. The second machine is reached from the first.

```mermaid
sequenceDiagram
  autonumber 118
  actor OWN as Owner
  box app.enverge.ai
    participant HOSTS as /host
  end
  box Supabase edge functions
    participant PRV as host-provision
    participant ENR as host-enroll
    participant REP as host-report
  end
  box Supabase database
    participant DB as Postgres
  end
  box Owner's Spark — node0
    participant INS as Installer
    participant PRP as Agent — preparing
    participant SRV as Agent — serving
  end
  box Owner's Spark — node1
    participant N1 as node1
  end
  actor OPS as Operator

  OWN->>HOSTS: add a GPU — "2× DGX Spark", offered to staff only
  HOSTS->>PRV: POST profile gb10-spark-2
  PRV->>DB: one hosts row, one relay, one tunnel, one enrolment — A, unchanged
  PRV-->>OWN: one install command, for node0 — --peer --nodes=2
  OWN->>INS: run it on node0
  INS->>OWN: asks for node1's login on the terminal — it never leaves the LAN
  INS->>INS: write peer-login and peer-nodes=2; peer-prep — find node1 on the cable (LAN mDNS+ARP only if the rail is silent), pin its QSFP link-local
  PRP->>PRP: pin every RDMA port to IPv6 link-local eui64 in NetworkManager (required, DGX OS often leaves addr_gen_mode=none)
  PRP->>N1: detect — multicast, the owner key then password, count machines, pin NM link-local on node1, fingerprint ports. No prepare or peer-access
  PRP->>ENR: POST {token, pubkey, fingerprint} — node0, its rails, and node1 under peers
  ENR->>DB: the profile wants a pair — at least one cable, no more than the planned rails, one machine per cable, a GB10 that is not node0. Anything else is rejected, token spent. Accepted, a host_nodes row per machine
  ENR-->>PRP: every credential, and a host_profile carrying node_count and a subnet per rail
  PRP->>N1: peer-access — create spark with the operators' keys and node0's, delete the login. Nothing is decided here
  loop node1's steps, over the rail, inside node0's own run
    PRP->>N1: check, apply, check again — the same step code through a remote exec
    PRP->>REP: the step, with index 1 — written to node1's own host_nodes row
  end
  PRP->>N1: rail-verify — traffic on every rail, both directions, node0 driving both ends
  SRV->>REP: heartbeat — node0's rails and how many machines answer on each. node1's worker is checked in L3
  OPS->>DB: approve — D, unchanged
```

1. **This is [new_spark_dual.md](../new_spark_dual.md) with the operator
   removed.** The fleet's pairs are node0, a full platform host, and node1, a
   machine with no agent, which a person prepares over SSH and node0's agent drives
   over SSH at create and delete. Node1 keeps that role exactly — no agent,
   no relay, no tunnel, no row — and the person's part moves into node0's
   prepare.
   - NVIDIA's own [stacked Sparks guide](https://build.nvidia.com/spark/connect-two-sparks/stacked-sparks)
     is the same work with a person on each box: matching users,
     passwordless SSH both ways, static addresses on the QSFP interfaces.
     Steps 125 and 129–131 are that guide, run from one side.
   - An earlier draft enrolled node1 in its own right, so that each box
     proved itself. It cost a second profile that was not a hardware profile,
     a second prepare track folded into one column, a heartbeat with nowhere
     to land, and a second connector on node0's tunnel. And it bought
     nothing: the lifecycle that follows needs node0 driving node1 over SSH
     anyway, and "this is a pair" is rail traffic between two GB10s whichever
     box reports it.
2. **Step 123 is the one thing that has to be asked for.** Node0 needs root
   on a box nobody has touched, and the login is the owner's to give. It is
   collected on node0's terminal rather than in the dashboard because that
   way it never reaches the backend: root-only on node0, read at steps 125
   and 129, then deleted. From then on node1 admits node0's peer key and
   the operator keys. `peer-prep` then activates node1's ConnectX-7
   ports and pins link-local eui64 before prepare starts, because
   NetworkManager often leaves them disconnected while the kernel link is up
   — no `fe80::`, and the cable probe at 125 would see silence forever.
   **It asks the cable first.** A rail is point-to-point, so whatever answers
   the multicast on it is node1 by construction; node0 logs in, groups the
   two addresses a QSFP shows into the one machine behind them, and prepares
   it at its management address — reconfiguring node1's rails through the
   rail being spoken over is a way to cut the session. The **management LAN**
   (mDNS `_ssh._tcp` and ARP on the uplink — not Tailscale) is the fallback,
   for exactly the case above: a node1 whose rails cannot yet answer. A rack
   of eleven Sparks on one LAN, five accepting the same login, is what that
   ordering is for — the scan can only report the ambiguity, and the cable
   has already answered. `ENVERGE_PEER_ADDR` names node1 outright and skips
   both.
3. **Step 125 finds node1 by the cable.** A rail is a point-to-point link, so
   a multicast ping to the all-nodes address on each RDMA port returns
   node1's link-local addresses before any static address exists. The probe
   ignores node0's own addresses, so a cable looped between its two ports
   finds nothing. A ConnectX-7 QSFP is two PCIe functions on one wire, so one
   cable shows two of node1's addresses; the probe logs in and counts
   machines, not addresses. Login tries the peer user's existing SSH private
   keys, then the password from the install prompt. A port whose neighbours
   are all one other Spark is a link; a port with several machines is the
   owner's LAN — a Spark's uplink can sit in a QSFP port — and is left alone.
   When every address on a port refuses the login, that is named as a login
   failure, not as a LAN. The links are numbered as the rails of the plan
   exactly once — here, the first time node0 looks — and the numbering is
   recorded on node0 on the spot, so no later run re-derives it from the
   order ports come in. It then asks which of node1's ports holds each
   address it saw. That is how a crossed cable becomes a recorded fact on
   both rows rather than a failed verification.
   - **Link-local persistence is required and automatic — not a manual
     step.** DGX OS often leaves ConnectX-7 ports at `addr_gen_mode=none`,
     so they have carrier but no `fe80::` and the multicast probe sees
     silence; NetworkManager's default also rewrites a live sysctl on
     reconnect. Before probing, the agent pins each of node0's RDMA ports'
     NM connection to `ipv6.method=link-local` and `addr-gen-mode=eui64`,
     forces the live sysctl, and bounces the iface if an address is still
     missing. After a successful login it does the same on node1's rails
     over sudo. Later, `rail-network` (130) keeps eui64 on the planned rails
     via netplan. There is no owner/operator runbook for this.
   - The enrolment has not been accepted yet, so 125 still must not prepare
     the box or install peer-access. The only allowed pre-enrol write is
     that NM link-local pin (and the matching live sysctl), so the cable
     stays reachable; a rejected peer is otherwise left as found.
   - If node1 cannot be reached or the login is wrong, the agent does not
     enrol. It retries in-process every minute while the token lives, re-reading the login, so the token stays unspent.
4. **Step 127 is the one place a lone Spark is turned away, and it is early
   on purpose.** Redeem is where turning away is cheap: the token is spent,
   the row parks as `rejected`, `reap` releases the relay, the address, the
   tunnel and the host row, and the owner is still at the terminal to see
   it. Anything found later can only stop a prepare run and leave the
   allocation held. So every piece of evidence that can exist before
   enrolment is judged here, once: at least one cable and no more than the
   plan has rails for, one machine on every cable, a peer that is the profile's
   hardware and not node0 itself, and every machine carrying its own
   `node_index` — 0 for the enroller, each other one once. Rows are written
   at the index a machine carries, never at its place in the list. `nvidia-smi` on either box reports `GB10 / 1 / aarch64`,
   byte-identical to a single, which is why the dropdown alone was never
   evidence.
   - What was found becomes rows: one `host_nodes` row per machine, node0 at
     index 0 and the peer at index 1, each with its own fingerprint —
     hardware identity — and its own `rails`, the observed state of each
     port, with the rail each cable carries and the plan's address for that
     node on it stamped in. Every BYO host gets them, so a single Spark is
     one row and the node count is a count. `unique (hardware_id)` is the
     "not node0 itself" rule as a constraint, and it also stops one box
     being claimed by two hosts. It keys on the GPUs rather than
     `/etc/machine-id` because that file clones with the image it was
     captured in; `machine_id` stays on the row as reported hardware. An
     agent older than the field sends none and its `machine_id` is written
     there instead, which is the namespace its own row was already in.
   - The profile is the plan and the node row is the fact. `constants.rails`
     says a subnet and an MTU per rail, and a node's address is the subnet's
     tenth host plus its index; `host_nodes.rails` says what these two
     machines look like, port by port, and is refreshed at 130–133 as they
     change.
5. **Step 132 is what redeem cannot know.** Whether the pair *works* needs
   both machines prepared and addressed, so the traffic test is late by
   nature. It proves the same claim step 127 accepted, with the rails
   configured. Both are reported by the owner's own hardware, so neither is
   more trustworthy than the other — what differs is the cost of being
   wrong, which is lowest at 127 and highest here. Approval (**D**) follows
   it and stays the last thing in the flow.
6. **Step 129 decides nothing.** `peer-access` creates the `spark` user, installs the peer key and the operator keys, and deletes the login. Its check still refuses an SSH host key that is not the pinned one. After go-live it runs only in the deliberate run (**L**).
7. **Steps 130–131 are C, run on node1 from node0.** Every step already runs
   through an `Exec`; a remote one over the rail runs the same check, apply
   and re-check on node1 and reports each with `index: 1`, which
   `host-report` writes to node1's own `host_nodes` row rather than to the
   host's `prepare_state`. The block is node1's subset of the list in the
   same order, so every landmine ordering holds there too, and the deliberate
   run in **L** re-checks node1 as it re-checks node0. Node0's track
   keeps its shape and its meaning; node0's `self-check`, after the block,
   asserts node1, which is why steps 39–40 and both activation gates are
   untouched.
8. **Step 133 does not watch node1 over the rails.** The heartbeat refreshes node0's own `host_nodes.rails` — per port, how many machines answered — which reads zero on both ports while an instance holds the rails, so it cannot tell a dead node1 from a rented pair.
   1. That is why the node1 probe was dropped: it reported every rented pair `degraded` while every subsystem was green.
   2. There is no `host_status` row for node1, and `host_nodes.seen_at` is no longer written.
   3. node1's worker is checked over the management network: **L3**.
9. **Break glass (I) reaches node1 through node0.** Node1's `spark` user
   holds the operators' keys, so the tunnel or the relay lands on node0 and one
   more hop lands on node1. Node1 has no tunnel of its own.
10. **A single Spark walks the same steps and none of this fires.** Its
    install command carries no pair marker, so 123–124 never happen; with no
    login file, 125 never happens; its profile has no `node_count` above one,
    so 127 judges nothing and 129–132 are skipped, reported as skipped the way
    `watchdog` is on a board without one. The rails still appear in its
    fingerprint, which is a few more fields and nothing else.
    - A single Spark presented as a pair with its own login typed in as
      node1's is caught twice: the probe ignores node0's own addresses, and
      redeem refuses a peer whose identity is node0's. That identity is
      `hardware_id` — the machine's GPU UUIDs, digested — not
      `/etc/machine-id`: systemd writes that file once, and only when it is
      empty, so an OS image captured with it populated hands every unit the
      same value and two genuinely different Sparks present as one machine
      that found itself. The agent makes the same comparison before it
      enrols, so a real self-pair costs a retry rather than the token.
    - A single Spark whose rails have carrier because they are on the
      owner's switch is fine against the single profile, where nothing
      checks rails, and fails the one-neighbour rule against the pair
      profile, which is the intended answer.
11. **The customer's instance is two containers**: the head on node0 and a worker on node1, named `<head>-w1`. node0's agent creates and deletes them together over the management network — not over the rail SSH step 129 established, because the dual profile moves the rail netdevs into the containers the moment one starts.
    1. At create, any worker already on node1 is removed first, the head mints a per-instance launcher key, and the worker trusts that key alone. A worker that cannot start fails the create: one GPU sold as two is not an instance. A create refused after the worker started — the customer key or the relay failing — reverts the worker along with the head, and restores both machines' rail addressing.
    2. At delete, the worker goes with the head, and both machines' rail addressing is restored.
    3. Keeping the two running together — through start and stop, and after either machine restarts — is **L**.

### What K writes, and where

One new table, `host_nodes`, one row per machine of a host. Everything else
is a table that existed, written at the steps it was always written at, with
the meaning it always had. The database is written at 120, 127, 131 and 133,
plus the approval at 134.

| Step | Reads | Writes |
|---|---|---|
| 118 | `host_profiles.staff_only`, to decide whether the picker shows the row | — |
| 119 | `staff_only` for the gate, `constants.node_count` for the command | — |
| 120 | — | `hosts` row; `host_enrollments` row with `host_id`, `profile_id`, `token_hash`, `wg_conf`. Same as a single |
| 121 | — | Nothing stored. The response carries `node_count` and `--peer --nodes=N` on the command (`N` from the profile) |
| 123–124 | — | `/etc/enverge/peer-login` and `peer-nodes` on node0; `peer-prep` finds node1 on the rail (falling back to the management LAN only when the rail is silent) and activates its RDMA NM connections + link-local eui64 (writes on node1 before enrol, so the cable is visible) |
| 125 | `peer-login`; node0's RDMA NM connections and sysctls; node1's `nvidia-smi` (including GPU uuids), `uname`, `/etc/os-release`, `/etc/machine-id`, sysfs ports, and which port holds each address | On node0: NM `ipv6.method=link-local` + `addr-gen-mode=eui64` on every RDMA port (and live `addr_gen_mode=0`); `peer.json` with the `links` the moment their numbering is assigned. On node1, after a successful login only: the same NM link-local pin on its rails (via sudo). No prepare, no `spark` user, no keys. In memory the fingerprint gains `node_index` 0, `machine_id`, `hardware_id`, `rails[].neighbours_discovered`, `rails[].index`, and `peers[]` each with its own `node_index` |
| 126 | — | Nothing yet. The request body is the fingerprint from 125 |
| 127 | `host_enrollments` by token hash for `profile_id`; `host_profiles` for `gpu_model`, `arch`, `gpu_count`, `constants.node_count`, `constants.rails` | Refused: `host_enrollments.state = rejected`, `token_hash = null`, `fingerprint`, `pubkey`, `source_ip`, `redeemed_at`. Accepted: the same columns with `state = ready` and node0's identity alone, `hosts.pubkey` and `hosts.profile_id`, then a `host_nodes` row per machine — identity in `fingerprint`, what detect saw of each port in `rails`, with the rail `index` each cable carries and the plan's address for that node stamped in |
| 128 | `hosts` for the token; `constants` merged with `arch` and `gpu_model`; the tunnel session; `operator_keys`; `wg_conf` | `host_enrollments.wg_conf = null`; `credentials.json` on node0, carrying `node_count`, the `rails` plan and the operator keys |
| 129 | `peer-login`; every RDMA port, for discovery; node1's `ip -6 addr`, for which port each cable lands on | On node1: the `spark` user, `authorized_keys` with node0's peer key and every operator key, `sudoers.d/enverge-spark`. On node0: `peer.key`, `peer.json` with `addr`, `host_key` and the `links` — index, node0's port, node1's port; `peer-login` deleted |
| 130 | `peer.json` for the links; the plan's subnet per rail and node1's index, for node1's addresses | Whatever that step writes on node1: netplan on node1's ends of the links (including `ipv6-address-generation: eui64`, which keeps the 125 pin on the planned rails), earlyoom, LXD, the base image |
| 131 | — | `host_nodes.prepare_state` on the row at index 1, replaced with `{step, state, attempt, error, final, at}`. Node0's own steps keep writing `hosts.prepare_state` as before. A `rail-network` transition, from either machine, also carries that machine's ports as observed after it ran — rail `index`, MTU, `active_mtu`, the address on the interface — merged into that node's `host_nodes.rails` |
| 132 | the links; the plan's subnets and both indexes | `host_nodes.rails[port].verified_at` on both rows for every rail that carried traffic. On failure, `hosts.prepare_state` with `step: rail-verify`, `state: failed` |
| 133 | node0's own ports | `host_status` for node0; node0's `host_nodes.rails` refreshed — MTU, `active_mtu`, `neighbours_discovered` per port — so a cable that comes loose after approval shows on the port that owns it. Nothing for node1 (**L3**) |
| 134 | `host_enrollments.state` must be `ready`; `hosts.prepare_state` must be `self-check` applied or already-satisfied | `host_enrollments.state = enrolled`, then `hosts.active = true` with a pool |

Node1 is written at three steps: 125 (NM link-local pin only), 129, and 130.
Step 125 is otherwise a look that makes the refusal at 127 possible without
preparing the box — the pin is the one pre-enrol exception, applied by the
agent on both nodes so the owner never has to.

### The shapes

**`host_nodes`**, the new table. One row per machine per host, for every
BYO host, so a single Spark is one row. Identity and state are separate
columns, because one never changes and the other is what an operator reads.
The index is explicit: what the agent, the backend and the plan agree on.

```sql
create table host_nodes (
  host_id       text not null references hosts(id) on delete cascade,
  index         int  not null,          -- 0 is node0: the relay, the tunnel, the agent
  machine_id    text,                   -- /etc/machine-id, as reported: hardware detail, not identity
  hardware_id   text not null,          -- sha256 over the sorted GPU uuids: what tells identical boxes apart
  fingerprint   jsonb not null,         -- hardware identity: gpu, memory, arch, os, driver
  rails         jsonb not null default '{}',  -- the ports as they are, keyed by interface
  prepare_state jsonb,                  -- that machine's latest {step, state, attempt, error, final, at}
  seen_at       timestamptz,            -- no longer written: node1 is not probed (133)
  created_at    timestamptz not null default now(),
  primary key (host_id, index),
  unique (hardware_id)
);
```

**`host_profiles.constants`**, the plan, written once. `staff_only` is a
column, being policy. The GB10 constants unchanged plus the two things a
pair adds: how many machines, and a rail plan — each rail named by an
explicit `index`, never by its place in the list, with a subnet and an MTU.
The subnets and hosts are the ones NVIDIA's stacked-Sparks guide assigns:
a node's address on a rail is the subnet's tenth host plus the node's own
index — node 0 is `.10`, node 1 is `.11` — and no interface is named,
because which port carries which rail is the cable's decision.

```json
{
  "node_count": 2,
  "rails": [
    { "index": 0, "subnet": "192.168.100.0/24",   "mtu": 9000 },
    { "index": 1, "subnet": "192.168.101.0/24", "mtu": 9000 }
  ]
}
```

**`host_nodes.rails`**, the fact beside the plan, keyed by this machine's
own interface. Written at 127 from what detect saw; the rail `index` and the
address stamped onto the ports that carry a cable; refreshed at 131 after
`rail-network` ran on that machine; `verified_at` stamped at 132;
`neighbours_discovered` kept current at 133 for node0. A port nobody is on
appears with `neighbours_discovered: 0` and no `index`. Here node1's second
cable landed on its other port — recorded, not a fault.

```json
{
  "enp1s0f0np0":   { "index": 0, "device": "rocep1s0f0",   "mtu": 9000, "active_mtu": 4096,
                     "address": "192.168.100.11/24",   "neighbours_discovered": 1, "verified_at": "2026-09-08T14:20:03Z" },
  "enP2p1s0f1np1": { "index": 1, "device": "roceP2p1s0f1", "mtu": 9000, "active_mtu": 4096,
                     "address": "192.168.101.11/24", "neighbours_discovered": 1, "verified_at": "2026-09-08T14:20:03Z" },
  "enP2p1s0f0np0": { "device": "roceP2p1s0f0", "mtu": 1500, "neighbours_discovered": 0 }
}
```

**The enrolment request at 126**, what the agent sends. This is a wire
shape, not a stored one: `host-enroll` splits it into the enrolment's own
`fingerprint` (node0's identity, without `rails` or `peers`) and the
`host_nodes` rows, identity into `fingerprint` and what detect saw of each
port into `rails`.

```json
{
  "gpu_model": "GB10", "gpu_count": 1, "gpu_mem_mib": 122880, "arch": "aarch64",
  "os_id": "ubuntu", "os_version": "24.04", "driver_version": "580.65.06",
  "machine_id": "3f2a…", "hardware_id": "b41c…", "node_index": 0,
  "rails": [
    { "device": "rocep1s0f0",   "iface": "enp1s0f0np0",   "mtu": 1500, "active_mtu": 1024, "neighbours_discovered": 1, "index": 0 },
    { "device": "roceP2p1s0f0", "iface": "enP2p1s0f0np0", "mtu": 1500, "active_mtu": 1024, "neighbours_discovered": 1, "index": 1 },
    { "device": "rocep1s0f1",   "iface": "enp1s0f1np1",   "mtu": 1500, "neighbours_discovered": 0 }
  ],
  "peers": [
    { "gpu_model": "GB10", "gpu_count": 1, "arch": "aarch64", "machine_id": "9c17…", "hardware_id": "7e08…", "os_id": "ubuntu", "node_index": 1,
      "rails": [ { "iface": "enp1s0f0np0", "index": 0, "neighbours_discovered": 1 }, { "iface": "enP2p1s0f1np1", "index": 1, "neighbours_discovered": 1 } ] }
  ]
}
```

A refused pair lands whole on `host_enrollments.fingerprint` under
`state = rejected`, which is how the enrolments page shows what turned up.

**`hosts.prepare_state`**, node0's track, unchanged in shape and meaning.
**`host_nodes.prepare_state`** on the row at index 1 has the same shape
and is node1's:

```json
{ "step": "base-image", "state": "applied", "attempt": 1, "error": null, "final": false, "at": "2026-09-08T14:02:11Z" }
```

**`host_status`**, node0's row, written at 133. No column for node1: half a pair shows as `lxd = degraded` (**L3**).

**Node0's files**, under `/etc/enverge`, root-only, never in the database:

- `peer-login` — node1's username and password from 124, deleted at 129
- `peer.key` — node0's peer key, Ed25519
- `peer.json` — the links from the moment they were numbered, then the pinned node1: `{"addr": "fe80::…%enp1s0f0np0", "host_key": "ssh-ed25519 …", "peer_index": 1, "links": [{"index": 0, "local": "enp1s0f0np0", "remote": "enp1s0f0np0", "neighbour_addr": "fe80::…%enp1s0f0np0"}, {"index": 1, "local": "enP2p1s0f0np0", "remote": "enP2p1s0f1np1", "neighbour_addr": "fe80::…%enP2p1s0f0np0"}]}`

The password never leaves node0, and the backend never learns node1's
link-local address. What it knows about node1 is its `host_nodes` row:
hardware, the state of its ports, progress, and when it was last seen.

### K2. Four nodes — switchless ring

A `spark-4` host is four Sparks on a switchless CX7 ring, sold as one unit. Everything in **K** holds; what differs is how many satellites there are, how node0 finds the one with no cable to it, and how many workers an instance carries.

```mermaid
sequenceDiagram
  autonumber 118
  actor OWN as Owner
  box app.enverge.ai
    participant HOSTS as /host
  end
  box Supabase edge functions
    participant PRV as host-provision
    participant ENR as host-enroll
    participant REP as host-report
  end
  box Supabase database
    participant DB as Postgres
  end
  box Owner's Spark — node0
    participant INS as Installer
    participant PRP as Agent — preparing
    participant SRV as Agent — serving
  end
  box Owner's Sparks — satellites
    participant N1 as node1
    participant N2 as node2
    participant N3 as node3
  end
  actor OPS as Operator

  OWN->>HOSTS: add a GPU — "4× DGX Spark", offered to staff only
  HOSTS->>PRV: POST profile gb10-spark-4
  PRV->>DB: one hosts row, one relay, one tunnel, one enrolment — A, unchanged
  PRV-->>OWN: one install command — --peer --nodes=4
  OWN->>INS: run it on node0
  INS->>OWN: asks once for a shared login that works on every satellite
  INS->>INS: write peer-login and peer-nodes=4; peer-prep hardens every satellite found on the cables
  PRP->>PRP: pin every RDMA port to IPv6 link-local eui64 in NetworkManager
  PRP->>N1: detect — a cabled neighbour, grouped by hardware_id
  PRP->>N2: detect — the other cabled neighbour
  PRP->>N3: detect — the diagonal, one hop through a neighbour's cable; pin NM link-local. No prepare
  PRP->>ENR: POST fingerprint — node0 plus peers[] with node_index 1..3
  ENR->>DB: judgePair wants node_count-1 peers; host_nodes row per machine
  ENR-->>PRP: credentials and host_profile with node_count 4 and four rail edges
  PRP->>N3: peer-access on every satellite — spark user, keys, peer.json entry; the diagonal over mgmt; then forget peer-login
  loop each satellite k=1..3, nodek/… over rail or mgmt
    PRP->>N1: check, apply, check again — same step code through a remote exec
    PRP->>REP: the step, with index k — that host_nodes row
  end
  PRP->>N1: rail-verify — every cable, from both of its ends
  SRV->>REP: heartbeat — node0's rails; workers checked over mgmt in L3
  OPS->>DB: approve — D, unchanged
```

1. **The command carries `--nodes=N`.** `host-provision` puts `--peer --nodes=${node_count}` on the command whenever `node_count > 1`. The installer writes `/etc/enverge/peer-nodes`, so detect and peer-prep know the target before enrol returns the profile.
2. **One shared login for every satellite.** Asked once on node0's terminal, into the same `peer-login` as a pair. Wrong password: re-run the one-liner.
3. **The satellites are the machines on the cables.** node0 finds its cabled neighbours, logs in to each, and asks what answers on *their* rails; the diagonal is reached one hop through either of them, until nothing new answers. Its management address and SSH host key are read from itself over the cable, and later access over the LAN is held to that key. The management LAN (mDNS + ARP, not Tailscale) is scanned only to wake a machine whose cabled ports have carrier and no link-local, and the cables are walked again: a machine that only shares the LAN is never a satellite, however the password is set. More machines on the cables than the profile has satellites is refused.
4. **A satellite is its `hardware_id`, as in K.** Machines on the cables are grouped and deduplicated by GPU identity — never `/etc/machine-id`, which a ring flashed from one image shares on all four boxes. `node_index` 1..N−1 is assigned by sorting `hardware_id` and pinned in `peer.json`, so later runs do not reshuffle.
5. **`peer.json` lists every satellite:** `peers: [{index, addr?, iface?, host_key, mgmt_addr?, links}, …]`, with rail `addr`/`iface` only where there is a cable from node0. A satellite's `links` are its own cables, each carrying `peer` — the node at the other end — so the diagonal's two cables are recorded on it even though node0 has no end on either. A pair-shaped file (flat fields, one peer) reads too.
6. **Prepare runs `node{k}/…` for each satellite** with K1's node1 step code. The diagonal is reached over the management network. `rail-verify` sends traffic across every cable from both of its ends, the two between the satellites and the diagonal included.
7. **An instance is four containers:** the head on node0 and `<head>-w1`…`-w3` on the satellites, created and deleted together over the management network. A worker that will not start fails the create, and every worker already launched is reverted.
   - Each container gets its node's two cables from `lofispark-dual`, rendered per node. On a ring every node also forwards IP and carries a route to each rail it is not on, through the neighbour that is — so the diagonal reaches the head, and the head the diagonal, as the switchless recipe does it. The head's `/etc/enverge/cluster.md` names every node, its rail IPs, and which one has no cable to it.
   - The profile's description carries a digest of its content; `profiles` rewrites a profile that does not match what this node renders now (a machine re-enrolled from another setup), except while an instance holds the rails.
8. **The profile plan** (`gb10-spark-4`) is `node_count: 4` and four rails, **one per cable** (`192.168.100–103/24`), as the switchless recipe the ring is built from. Each cable is addressed on the PCI-domain-0 port at both ends (`enp1s0f*`), never on its `enP2` twin, and cables are numbered from the cable map by (lower node, higher node). A port that no longer carries a rail loses its rail address, so a re-numbering never leaves the same address on two ports of one wire. Redeem allows fewer cables than rails and refuses more.

## L. After go-live — upgrades and reboots (135–158)

Everything above ends with a host serving. This is what keeps it serving afterwards, when nobody is onboarding anything: a new binary, either machine restarting, and a reboot of both.

One constraint bounds all of it: node0 is the gateway and satellites have no agent ([ARCHITECTURE.md](../../../ARCHITECTURE.md) — *Dual-host clusters*). Whatever happens on a satellite is done by node0's agent over the management network, or by host configuration node0's preparation laid down there — never by a program of ours running on the satellite. On a pair that satellite is node1 (**L3**); on a ring it is each of node1..node3 the same way.

**J** could treat the agent as one program. Here its three units are the point: `PRP`, `SRV` and `RB` in the [glossary](byo_glossary.md).

### The rules

1. **A reboot runs no preparation.** Serving starts at boot on its own. The last act of a completed preparation — handing the credentials to the serving user — is the record: later starts find it and exit at once, whatever the binary, so serving's `Requires=` is satisfied without a run, and a host that never completed preparation still cannot serve.
2. **The rollback fires when serving fails to start.** It judges the binary. A preparation step failing on the host's state — a cable, a lent rail, node1 — is not a bad binary.
3. **Serving restores the instance on both machines when it starts** — LXD, the head, then the worker on node1 over the management network — and only then points the relay.
4. **The pair is kept in lockstep**: the worker exists and runs exactly when the head does (*Lockstep*, below).
5. **"An instance holds the rails" is asked of both machines** before preparation touches a live host, node1 over the management network. When node1 cannot be reached, node0's answer is enough.
6. **Preparation does not run after go-live.** Not on a boot, not on an upgrade: an upgrade swaps the binary and restarts serving (L1). A change to host configuration is applied deliberately, on an idle host (*Host configuration after go-live*, below).

### L1. A new binary (135–141)

```mermaid
sequenceDiagram
  autonumber 135
  actor OPS as Operator
  box Owner's Spark — node0
    participant INS as Installer
    participant SD0 as systemd on node0
    participant SRV as Agent — serving
    participant RB as Rollback
  end
  box Supabase edge functions
    participant REP as host-report
  end

  OPS->>INS: make byo-upgrade — J, up to keeping the outgoing binary as .previous
  INS->>SD0: restart serving on the new binary
  SD0->>SRV: start — L2 from here, on the new binary
  alt serving comes up
    SRV->>REP: heartbeat, from the new binary
  else serving fails to start
    SD0->>RB: OnFailure= on the serving unit
    RB->>RB: live binary to .failed, .previous over it — once
    RB->>SD0: start serving on the previous binary
  end
```

1. **Serving is down only while it restarts.** The tenant's containers keep running and their SSH keeps working. What pauses is the VM API and the heartbeat.
2. **The rollback runs once, and never to the build already running.** A previous binary that fails too leaves the host down and alerted `agent_unreachable`, rather than flipping between two broken builds.
3. **No preparation runs.** The new binary serves the host the previous one prepared. What a newer agent would change about the host waits for the deliberate run below.

### L2. node0 restarts (142–147)

```mermaid
sequenceDiagram
  autonumber 142
  box Owner's Spark — node0
    participant SD0 as systemd on node0
    participant SRV as Agent — serving
    participant LXD0 as LXD on node0
    participant HEAD as Head container
  end
  box Owner's Spark — node1
    participant LXD1 as LXD on node1
  end
  box Relays
    participant RELAY as The host's relay
  end
  box Supabase edge functions
    participant REP as host-report
  end

  SD0->>SRV: start at boot — no preparation
  SRV->>LXD0: connect — this starts LXD if boot activation did not
  SRV->>HEAD: start it if it should be running, and wait until it is
  SRV->>LXD1: over the management network — start the worker if it should be running
  SRV->>RELAY: forward SSH to the head, now that it runs
  SRV->>REP: heartbeat
```

1. **Serving does not rely on LXD starting itself.** Both containers carried `boot.autostart: "true"` from `lofispark-base` at the time, yet LXD's boot activation started nothing on either machine of spark-21 across three boots (LXD 5.21.7). LXD starts when something first calls it. Serving waits for `snap.lxd.activate`, which opens LXD's socket to the `lxd` group.
2. **What ran before a reboot runs after it, and nothing else.** With `boot.autostart` unset, LXD restores each instance's last state, and the agent starts the head only when that state is running.
   - Older hosts carry `boot.autostart: "true"` — *always* start — until `make byo-prepare` removes it. spark-15 to spark-23 had it removed by hand on 2026-09-14.
3. **Head, then worker, then relay.** SSH is forwarded once, at start. Forwarding it before the head runs would send `:22` nowhere until the next restart.
4. **A node1 that cannot be reached does not stop node0 serving.** The heartbeat keeps asking (L3).

### L3. node1 restarts (148–152)

```mermaid
sequenceDiagram
  autonumber 148
  box Owner's Spark — node0
    participant SRV as Agent — serving
  end
  box Owner's Spark — node1
    participant LXD1 as LXD on node1
    participant W1 as Worker container
  end
  box Supabase edge functions
    participant REP as host-report
  end

  SRV->>LXD1: every heartbeat, over the management network — is the worker running while the head is?
  alt node1 answers, and the worker is stopped
    SRV->>LXD1: start the worker
    LXD1->>W1: start — it takes node1's rails back
    SRV->>REP: heartbeat — the pair is whole again
  else node1 cannot be reached, or the worker will not start
    SRV->>REP: heartbeat — degraded, half a pair
  end
```

1. **The check rides the heartbeat**, the one schedule serving keeps. It asks node1's LXD over the management network — never the rail, which is inside the containers.
2. **A restarted worker needs nothing else.** Its rail addresses, `/etc/nccl.conf` and the head's launcher key live inside the container and survive a restart. On spark-21 both containers came back with their rails addressed, RDMA active at MTU 4096, and the head reaching the worker as `user`.
3. **Half a pair is `degraded`, and de-pooling stays an operator's decision.** The signal K.8 once took from the rail is taken from the container, which reads the same during a rental as outside one. It lands on `lxd`, since `host_status` has no column for a pair.
4. **Both machines at once** — a power cut — is L2 and then L3. Serving restores what it reaches at start, and the heartbeat finishes the job when node1 answers.

### L4. A reboot of the pair (153–158)

```mermaid
sequenceDiagram
  autonumber 153
  actor CUS as Customer
  box app.enverge.ai
    participant PXY as the proxy
  end
  box Owner's Spark — node0
    participant SRV as Agent — serving
    participant SD0 as systemd on node0
  end
  box Owner's Spark — node1
    participant SD1 as systemd on node1
  end

  note over CUS,PXY: or an operator, from /admin through admin-vm-action
  CUS->>PXY: reboot
  PXY->>SRV: POST /vms/{id}/host/reboot, as the owner
  alt node1 answers
    SRV->>SD1: over the management network — reboot in 2 s
  else node1 cannot be reached
    SRV->>SRV: audited — the reboot is half done
  end
  SRV->>SD0: reboot in 2 s
  SRV->>PXY: 202
```

As built: [the agent's side of a reboot, generated from its tests](../../../byo-agent/internal/vm/testdata/sequences.md#8-reboot-pair). Scenario 8, with node1 answering.

1. **node1 first.** It has no agent, so once node0 is down nothing is left to tell it anything.
2. **Both reboots are deferred**, so each command returns before its machine goes down, and the caller hears a 202 rather than a dropped connection.
3. **An unreachable node1 does not block the reboot.** Refusing would leave the customer with no reset at all. node0 reboots, and the heartbeat reports half a pair until node1 answers (L3).
4. **The customer and an operator press the same button.** The dashboard's reboot and `/admin`'s are one call, so a pair is reset the same way whoever resets it.
5. **Recovery is L2 and then L3**, as for a power cut. Nothing about the instance is rebuilt: both containers come back as they were.

### Lockstep

The rule is one sentence: **the worker exists and runs exactly when the head does.** Nothing else needs keeping in step — the rail addresses, `/etc/nccl.conf` and the launcher key live inside the containers.

| Event | The worker | Built |
|---|---|---|
| Create | launched after the head, trusting its launcher key; failing to start fails the create, and a create refused later reverts it along with the head | yes |
| Delete | removed with the head; both machines' rail addressing restored | yes |
| Start, stop, restart | follows the head; a start that fails is repaired from the heartbeat | yes |
| node0 restarts | started by serving, after the head (L2) | yes |
| node1 restarts | started from the heartbeat (L3) | yes |
| Host reboot | node1 rebooted first, over the management network; back through L2 and L3 (L4) | yes |
| A worker with no head | removed | at create only |

### Host configuration after go-live

1. **Nothing runs preparation on its own once a host is live** — not a boot (rule 1), not an upgrade (L1).
2. **A change to host configuration is applied deliberately.** When a newer agent changes what a step does — `lxd-install` gaining `usermod -aG lxd` under the same name was the first — an operator runs preparation on the host over the operator tunnel (**I**): `make byo-prepare BYO_HOST=…`.
3. **Only on an idle host.** Rule 5 is how the run knows: no running instance holding the rails on either machine, or on node0 when node1 cannot be reached. A rented pair is left alone until it is not.
4. **It is C again, in full.** The done list is scoped to the build that wrote it, so a newer binary re-checks every step, and the pair's steps prove the cable once more — possible only because nothing holds it.
5. **Its failure is reported, not rolled back.** It shows in `/admin`, leaves serving running, and does not fire the rollback, which judges serving only (rule 2).
6. **The operator keys are not host configuration.** They change whenever an operator joins or leaves, not with the agent, and they reach live hosts through **M** rather than by preparing again.

## M. Changing who the operators are (out of band) (159–177)

The operator keys are the one thing on a host that changes on Enverge's side
rather than the owner's or the agent's. An operator joins, an operator leaves,
and every host has to agree — including the ones a customer is renting, which
preparing again (**L**) would not touch until they are idle.

```mermaid
sequenceDiagram
  autonumber 159
  actor OPS as Operator
  box app.enverge.ai
    participant ADM as /admin
  end
  box Supabase edge functions
    participant OPK as admin-operator-keys
  end
  box Supabase database
    participant DB as Postgres
  end
  box Owner's Spark host
    participant SRV as Agent — serving
  end

  OPS->>ADM: add their own key, or remove one
  ADM->>OPK: the change, as a superadmin
  OPK->>DB: apply it to the newest list, signed or waiting, refuse an empty result, and reserve the next version
  OPK-->>ADM: the exact bytes to sign — the whole list and its version
  ADM-->>OPS: the bytes, to sign on their own machine
  OPS->>OPS: check the list the save command prints, then sign on their<br/>own machine with the signing key, copied from the password manager for this one signature
  OPS->>ADM: the signature
  ADM->>OPK: the signature
  OPK->>OPK: check the signature against the public key every host holds
  OPK->>DB: write operator_keys and the signed list
  loop every enrolled host that is not released
    OPK->>SRV: set the operator keys — the signed list
    SRV->>SRV: check the signature and that the version is not older,<br/>then write spark's authorized_keys on node0, and on node1 through node0
    SRV-->>OPK: the version it now holds on each machine, or why it refused
    OPK->>DB: record the answer on the host — no answer is recorded too
  end
  OPK-->>ADM: which hosts hold the new version, and why the rest do not
  ADM-->>OPS: the hosts not updated, each with its reason
  opt the operator chooses to, once the reason is dealt with
    OPS->>ADM: send again, to the hosts not updated
    ADM->>OPK: send again
    OPK->>SRV: set the operator keys — the same signed list, answered as in 169–172
  end
```

1. **The list is sent whole, never as a change.** Adding and removing are the
   same call, and a push that is missed is repaired by the next one rather than
   replayed.
2. **An empty list is refused twice.** `admin-operator-keys` will not write a
   change that leaves no operator, and the agent will not apply a list with no
   keys. Either would leave a box nobody can administer — the same reason
   `operator-access` is verified before `owner-lockout` (**C**).
3. **A rented host is changed too, and that is safe.** Setting the keys writes
   `spark`'s login file and nothing else: not the rails, not a container, not a
   unit. That is why it is not a preparation step, which on a live host waits
   for idle (**L**).
4. **Serving can do it without root.** `authorized_keys` belongs to `spark`,
   the user serving already runs as, and node1's belongs to the `spark` user
   node0 already reaches it as. The serving unit may write `spark`'s `.ssh`
   and nothing else in `/home`.
5. **The agent keeps the list where preparation reads it.** Preparing again
   (**L**) writes `authorized_keys` from what the agent holds, so each push also
   updates that copy — the signed list and its version, as received. Otherwise a later run would put back the list from
   enrolment, and with it any operator removed since.
6. **A failed push fails, and says so.** Nothing retries on its own. `/admin`
   lists every host that was not updated with the reason, and the attempt stops
   there until an operator sends again.
   - Sending again re-sends the stored signed list, so it needs no signature
     and no password manager — one click, once the host is reachable.
   - An offline host is out of date but not open: with it off the network,
     nobody logs in to it. What matters is the time between it coming back and
     someone sending again, and `/admin` is where that shows.
7. **Only a signed list is applied, so the host token is not enough.**
   - **Why.** `hosts.agent_token` is readable from the database and used by the
     proxy, and the agent's API answers it from the internet. Until now it could
     create and delete containers. If it could also set the login keys, it would
     be root on the host, and a leak of the database or the proxy would be root
     on every BYO host.
   - **How.** A person signs every list, on their own machine. The host
     receives the matching public key once, at enrolment (step 29), and checks
     every list against it. No call changes that public key.
   - **Replays.** The version is inside what is signed, and the agent refuses a
     version older than the one it holds. A token holder who kept an old
     signed list cannot use it to bring back an operator removed since. The
     same version is taken again only with the same keys, which is what lets
     sending again land after an answer was lost.
   - **Where the signing key lives.** In the operators' password manager. To
     sign, an operator copies its private half from the web vault into a
     throwaway SSH agent on their own machine, for one signature: the sign
     command clears the clipboard once the agent has it and stops the agent
     when it is done, so it never reaches the disk. No server holds it and no
     Enverge page sees it, so a leak of the database, the edge-function
     secrets or the proxy signs nothing: changing the keys takes the password
     manager and a superadmin session together.
   - **What that exposes.** For the seconds between copying and signing, the
     private half is in the operator's clipboard, and so within reach of
     anything that reads it there — a clipboard manager, or another device
     sharing the clipboard. That is the price of signing from the web vault.
8. **What is signed is checked where it is signed.** `/admin` is code served by
   Vercel, so a tampered page could offer an attacker's list for signing. The
   signing command prints the operators and key fingerprints from the bytes
   themselves, and the operator reads that before approving — what they check
   is exactly what they sign, whatever the page showed.
9. **Only a change needs a signature.** Enrolment (step 29) and sending again
   use the latest signed list as stored, so neither needs the signing key. What cannot happen is any change to who gets in
   without someone signing it — which is the point.
10. **Every operator's changes add up in one draft.** There is one list: every
    operator's keys, one per machine, in one version. A change is one key added
    or removed, applied by `admin-operator-keys` to the newest list — the one
    waiting for a signature, if there is one — so Tudor adding his key and Breno
    adding his before either signs end in one version holding both, and one
    signature publishes it. Two changes landing at once cannot overwrite each
    other: the second is applied again on top of the first. Only the newest
    version can be signed, and it carries every change before it.
    - A key added belongs to whoever adds it, signed in as themselves: nobody
      grants access in someone else's name.
11. **A host enrolled before signed lists trusts no signing key.** It refuses
    every list and keeps the single operator key it enrolled with, until an operator
    tells it which signing key to trust over the tunnel (**I**), then sends again. That
    is the only way a host's trusted signing key is ever set after enrolment.
12. **A proposal nobody will sign is withdrawn, for everyone.** Withdrawing
    deletes it, and any older unsigned proposal it replaced, so none of them
    comes back as the one waiting. Only an unsigned proposal can be withdrawn:
    nothing unsigned ever reached a host, so there is nothing to undo.
13. **An operator who leaves takes the signing key with them.** The signing
    key is one Bitwarden entry shared by every operator, and signing from the
    web vault puts its private half through each signer's clipboard, so anyone
    who has signed may hold a copy. Removing their operator key is not enough:
    they could sign a list that adds it back. The signing key is changed too —
    trusted on every host by hand (**I**), then a new list signed with it.

## N. Retire a host (out of band) (178–190)

A host that finished enrolling ends here, for good: its relay, tunnel and DNS go, and nothing it was issued still works. The row stays, and so do its instances and its nodes, so the host's history is still readable. Its machines are free to be enrolled again, as a different setup. This is not **H**: **H** revokes an enrolment; **N** retires a host whose enrolment completed.

```mermaid
sequenceDiagram
  autonumber 178
  actor OPS as Operator
  box app.enverge.ai
    participant ADM as /admin
  end
  box Supabase edge functions
    participant HST as admin-host-action
    participant ONB as admin-onboard-host
    participant RLY as relay-provision
  end
  box Supabase database
    participant DB as Postgres
  end
  box Relays
    participant GCP as GCP API
  end
  box Cloudflare
    participant CF as Cloudflare
  end

  OPS->>ADM: retire a host
  ADM->>HST: retire a host
  HST->>DB: refuse if it never enrolled, its enrolment is still open, it is active or dedicated, has a live instance or a pre-booking to come, or is already retired
  HST->>ONB: revoke the operator tunnel, by slug
  ONB->>CF: delete the tunnel and its hostname
  HST->>RLY: release the relay, by slug
  RLY->>GCP: delete the relay instance, then its address
  RLY->>CF: delete the relay's record
  HST->>CF: delete the customer SSH wildcard, and the lab record if there is one
  HST->>DB: release the enrolled enrolment, and clear every enrolment's token and WireGuard config
  HST->>DB: clear the host's agent and SSH-map tokens
  HST->>DB: retired_at on the host and on every one of its nodes
  HST-->>ADM: retired — its machines can be enrolled into another host
```

1. **Only a host that is already out of service.** Taking it out of placement (`active = false`) and draining its instances (**G**) are the operator's decisions, made first. Retire refuses a host that is active or dedicated, has a live instance, or has a pre-booking whose window has not ended, because an inactive host can still be serving a dedicated tenant or be owed to a paid reservation. A failed create is not a live instance.
2. **Outside first, the stamp last.** Every teardown call is the one the enrolment release already makes, and each is idempotent. If one fails part way, nothing is stamped, and running retire again carries on from where it stopped. `retired_at` on a host means everything it held is gone.
3. **The row is never deleted.** `vms.host_id` keeps a host's instances and billing history pointing at it, and the row keeps its slug reserved. A slug that is freed while a teardown is still running is how a new host once lost its relay to the old host's release.
4. **The customer SSH wildcard is retire's own job.** Relay release removes the relay's record, and revoke removes the tunnel's hostname. Only a hard delete removes `*.<slug>.ssh`, so a retired host that left it up would keep a name pointing at an address GCP can hand to someone else.
5. **Only a host whose enrolment completed.** A host that never redeemed (no `pubkey`) was not provisioned by an enrolment, and one whose enrolment is still `pending` or `ready` is an enrolment to revoke (**H**), not a host to retire. So when retire runs, the only enrolment it closes is the `enrolled` one: released, keeping its `host_id` as the record of the setup, with every enrolment's token and WireGuard config cleared. Nothing is left open for the expiry sweep, so the sweep never tries to delete a retired host.
6. **Nothing the host was issued still works.** Its `token_hash` and SSH-map token are cleared, so an agent left running on the box can no longer report or fetch anything. `agent_token` stays: it is what the backend presents to the agent, not something the host holds against us. A retired host cannot be reactivated, edited or deleted, and a check on the table holds it inactive and undedicated for every writer, so no placement path can reach it.
7. **Its machines are free, and their history stays.** A node's `hardware_id` is unique only among nodes that are not retired. The box can join a new host at once, and the retired rows still say which host it was in, and when.
8. **Nothing on the box is touched.** Retire works on backend state alone. The agent, cloudflared, WireGuard and our keys on the machines are removed by the owner or overwritten by the next enrolment. Handing a machine back is **H**'s uninstall.

## What has to exist for this to work

| Missing | Phase | Why it matters here |
|---|---|---|
| A **schedule** for `host-provision` `reap` | A | The sweep is built; nothing calls it. Until something does, every abandoned click leaks a relay VM and a reserved address |
| **Rotating the signing key** | M | No call to the agent changes the key a host trusts, so a new signing key is trusted host by hand over the tunnel, as **J** upgrades are. Losing the password-manager entry before that is done freezes every host's operator keys as they last were |
| **Uninstall** | H | Revoke is `active = false` plus the release path, which exists. Handing the machine *back* does not. It is what gates ever setting `hosts.lockout_owner`: until it exists, a locked host is one we cannot return |

## Failure legs worth knowing

| Phase | Step | Failure | What happens |
|---|---|---|---|
| A | 5 | the relay will not build | The owner is told to try later; what was allocated is released |
| B | 18 | no token passed | Nothing downloaded; points the owner at the dashboard |
| B | 22 | checksum missing or wrong | Installs nothing, and says so |
| B | 28 | token unknown, spent or expired | One answer for all three. The owner adds the GPU again |
| B | 28 | hardware is not a GB10 Spark | Rejected, and the allocation is released |
| C | 35 | a step will not complete | The run stops there. Visible in `/admin`, and the host never reaches approval |
| D | 39 | the report shows a GPU that never initialised | The operator does not approve it |
| E | 61 | every idle host refuses the launch | The queue row stays open and the `vms` row is marked `creation_failed_at`, so the next trigger retries the same customer |
| E | 60 | a host accepts but its answer cannot be read, and no id was minted | The loop stops rather than rotating: another attempt risks a second container for one queue row. Reported, and left for the next drain to resolve by owner and name |
| E | 62 | the launch worked but this write is lost | The instance is live and reachable — the agent published the port — but its row is missing the host and the customer key. Logged rather than swallowed, because nothing downstream would notice |
| G | 81 | the agent never answers | Nothing is closed. The row keeps billing and keeps its host, which is honest — the capacity is still held — and `delete_requested_at` is what surfaces it |
| G | 83 | the container is gone but this write is lost | The row outlives the container: still billing, still holding its host, and invisible to every sweep. What the fleet's agent hid, and a BYO host cannot |
| K | 125 | node1 does not answer, or the login is wrong | No enrolment, token unspent. The agent retries every 60 s; the owner re-runs the installer to enter the login again |
| K | 127 | no carrier on any rail | Rejected, token spent, the host released — a lone Spark cannot take a `spark-2` allocation |
| K | 127 | more cables than the plan has rails | Rejected: the extra cable has no rail to take |
| K | 127 | the peer is not a GB10 Spark, or is node0 itself | Rejected, and the hardware that turned up is on the enrolment, as it is for node0 |
| K | 132 | the rails carry no traffic once configured | `rail-verify` fails and retries every 60 s, visible in `/admin`; approval never opens |
| L | 139 | serving fails to start on a new binary | Rolled back to the previous binary, once. If that fails too, the host stays down and is alerted `agent_unreachable` |
| L | 145 | node1 cannot be reached when node0 starts | node0 serves; the heartbeat starts the worker when node1 answers (L3) |
| L | 152 | node1 stays unreachable | `degraded`, half a pair. De-pooling stays an operator's decision |
| L | 156 | node1 cannot be reached for a reboot | node0 reboots anyway, audited as half done; the heartbeat starts the worker when node1 answers (L3) |
| M | 167 | the signature fails the check, or the version is no longer the next one | Nothing is written and nothing is pushed. The operator starts again |
| M | 169 | the host does not answer | Recorded as not updated and listed in `/admin` with the reason. It stays out of date until an operator sends again (175–177) |
| M | 170 | no signing key trusted, the signature fails the check, or the version is older | Refused, and the keys the host holds stay as they are. The answer says why, and it shows in `/admin` |
| M | 170 | node1 cannot be reached | node0 is updated and the host answers with node1's version still old, so the host is listed as not updated until an operator sends again. Until then a removed operator's key still opens node1, which has no tunnel of its own: reaching it takes node0, which no longer admits them, or the owner's own network |
| N | 180 | the host never enrolled, its enrolment is still open, it is active or dedicated, has a live instance or a pre-booking to come, or is already retired | Refused with the reason, and nothing changes. An open enrolment is revoked instead (**H**) |
| N | 184 | the relay instance is still deleting | Nothing is stamped. The address stays reserved until the instance is gone, and retire is run again |
| N | 186 | the customer SSH wildcard cannot be deleted | Nothing is stamped, and retire is run again. The relay and tunnel are already gone, which is safe to repeat |

## Other ideas explored

Set aside, with the reason — because "why didn't you just…" gets asked twice.

1. **An acceptance probe before activation.** The backend would launch a scratch
   VM and run `nvidia-smi` itself, since the GPU claim is billing-relevant.
   - The scratch VM is launched by the agent, on the host, and its output
     returns through the host: against a hostile host it restates the claim
     rather than checking it.
   - A *broken* host is caught by the step-37 self-check anyway.
   - Still worth having one day, and distinct from this: the backend has never
     used the customer's own path (**F**) before a customer does.
2. **A shared relay pool.** Removes the slowest step from **A** — but nobody
   waits on that step, since the relay boots while the owner finds a terminal.
   - In exchange: capacity to watch, slots to reclaim, a stall when it is empty,
     and one fewer isolation boundary.
   - Revisit if per-host relays prove expensive at rest rather than slow to
     build.
3. **The box enrolling anonymously, printing a claim code.** Needed a
   pending-and-unowned state, an expiry sweep, an alphabet a human could
   transcribe, and an enrolment endpoint anyone could write rows to — all to
   answer a question the current flow never asks, because the owner is signed in
   before the machine is touched.
   - Dropping it is a reversal worth naming: it had no secret in the owner's
     command line, where one lands in shell history and in whatever gets pasted
     into a support thread.
   - Provisioning first makes that trade a bad one. The token buys only an
     allocation that already carries the owner's name, is single-use, and dies
     within the hour — so what survives in history is a spent secret.
4. **The host token authorising key changes, with no signature.** One call,
   no password manager, no manual step.
   - The token is readable from the database, used by the proxy, and answered
     by the agent from the internet. A call that sets the login keys would make
     it root on the host, and a leak of either holder root on every BYO host.
   - Signing moves that authority to a signing key no server holds (**M**).
5. **Signing on the server, or in the browser.** `admin-operator-keys` signing
   with the signing key as a Supabase secret, or pasted into `/admin`, would keep the
   change to one click.
   - A Supabase secret is readable by every edge function in the project.
   - `/admin` is code Vercel serves, so a tampered page could take a pasted signing key
     the next time someone uses it.
   - Signing on the operator's own machine, with the key loaded from the
     password manager for one signature, keeps it off every server and every
     Enverge page.
6. **Delivering key changes by preparing again.** It already exists, and
   `operator-access` already writes `authorized_keys`.
   - Preparing again waits for an idle host (**L**), so a rented host would
     keep a removed operator's key for as long as the rental lasts.
   - It is a loop over every host run by hand, where **M** is one push.
7. **Retrying a failed push from the heartbeat.** It would bring an offline
   host up to date without anyone acting.
   - A change would reach a host at a moment nobody chose, for a reason
     `/admin` does not show.
   - An offline host cannot be logged into, so there is little to protect in
     the meantime: `/admin` lists it, and sending again is one click.
8. **One Unix account per operator, instead of a shared `spark`.** Logins would
   be attributable by user name.
   - The agent manages one service user on node0 and node1. An account per
     operator is a home directory and a sudoers file per operator per machine,
     created and removed on every host with nothing to clean up after them.
   - sshd logs the fingerprint of the key it accepted, and `operator_keys`
     maps that to the operator, so logins are attributable already.