BYO — the whole system, in sequence

publicv1
2h ago
3 views0 comments0 reviews78 min read
raw .md ↗

Every part of a BYO host's life, from an owner adding someone else's Spark on the dashboard to the host being handed back — and what keeps it serving in between. This document is the source of truth for how the BYO system behaves: where code, specs or other docs disagree with it, they are the ones to change.

  1. A–H are the phases, from an owner adding a machine to the two ways it can end. In onboarding an operator appears once, to approve (D).
  2. I, J and M are out of band, and none of them happens on an ordinary onboarding:
    1. I is the operator's way in — break glass — from the moment operator-access completes.
    2. J is how a host gets a newer agent than the one it was installed with.
    3. M is how a change to who the operators are reaches every host, rented ones included.
    4. N is how a host whose enrolment completed ends for good, so its machines can be enrolled into another host.
  3. K is a variant, not a phase: what A–D become when the host is more than one machine — K1 the pair (spark-2), K2 the four-Spark switchless ring (spark-4).
  4. L is the host after go-live: a new binary, either machine restarting, and a reboot of both.

Six areas, one continuous numbering across the whole document, so step 44 means one thing and E means one phase.

The Python agent in platform/agent/, and the hosts that run it, are out of scope. The BYO agent is platform/byo-agent/.

Everything is provisioned before the box is ever touched. Clicking "add a GPU" allocates the relay slot, the address, the tunnel and the hosts row — not the box enrolling, and not an operator afterwards. By the time the owner runs the one-liner, the only thing missing is the machine itself.

That ordering is what makes the rest fall out:

  1. The agent enrols once and keeps going. Its single call returns every credential, so it goes straight into preparing the host. No second start, no waiting, nothing on the box watching for anything.
  2. The operator appears once, at the end, to approve a host they can read the whole preparation of — and that approval is the only thing between an enrolled machine and a paying customer.
    • An earlier draft had an operator provisioning each host by hand between enrolment and preparation. That put a person in the middle of every onboarding, which is precisely what BYO is supposed to remove.
  3. Each host gets its own relay, created for it here. Same topology as the fleet has today — one relay per host — except that building it is an API call rather than a person following new_spark.md §1–§2.
    • The instance configures itself from relay/gcp/bootstrap.sh at first boot, so nobody SSHes into a relay to bring it up.
    • It boots while the owner is still reading the one-liner, so its start-up costs nothing anybody waits on.

Everything belongs to one of six places. The diagrams group them, because which side of a boundary a step happens on is usually the point — a credential minted in an edge function and a credential minted on the host are very different facts.

Who belongs to each area — every actor's ID and label, and the verbs they use — is in the BYO glossary, and only there.

1. The owner's Spark host — the machine being onboarded, and the programs on it. The agent is two of them, because preparing needs root and serving must not have it.

2. app.enverge.ai — the Vercel deployment: the owner's and operator's pages, the static installer assets, and the proxy that redirects to Blob.

3. Supabase edge functions — everything with authority. Nothing here runs on a host.

4. The Supabase database — host_enrollments, host_profiles, hosts, and vms from E onwards. Written by edge functions under the service role; read by the pages under RLS. There is no relays table: a relay is named by its host's slug, so the hosts row is its record, exactly as it is for the hand-built fleet relays.

5. Relays — one per host, created during provisioning. Includes the GCP API that builds the instance, reserves the address and points it at the relay. Same topology and naming as the fleet runs today — <slug>-relay and <slug>-relay-ip; the difference is that nobody builds it by hand.

6. Cloudflare — the operator tunnel and DNS.

And three people, none of whom belong to an area: the owner, the operator and the customer.

The customer never learns any of this. They land on a spark-1 host and cannot tell whose hardware it is, which is the tenancy decision spec63 makes explicitly rather than by omission.

  1. Nothing here needs the machine. BYO accepts one hardware configuration, so the relay slot, the address and the sizing are known before a box exists. The box contributes one thing: proof it really is that configuration, checked at step 28.
  2. Step 11 writes active = false, and nothing writes true except the operator at the end.
  3. Relay provisioning is its own edge function. It is the only part of this that talks to two systems we do not otherwise touch from here — the GCP API, and the relay itself.
    • host-provision stays an orchestrator, and replacing a dead relay later does not have to go through host onboarding to get one.
    • It is the slowest thing in the flow and the only one nobody waits on: the relay finishes booting while the owner is still finding a terminal.
  4. Provisioning ahead of time costs something. An owner who clicks and never runs the command leaves a relay, an IPv4, a tunnel and an inactive hosts row allocated. Four things contain it:
    • only an owner granted community_host from the /admin users tab (or a superadmin) can click at all,
    • a cap of three outstanding allocations per owner,
    • a token that expires within the hour,
    • host-provision's reap, which hands back everything a lapsed or rejected enrolment was holding. Nothing calls reap on a schedule yet.
  5. This has been run against real GCP, not just written: address reserved, instance created, metadata read, bundle fetched, wg0 up with the host's peer, ssh_mapd on 10.8.0.1:8754, Caddy on :80/:443, and the LOFISPARK_VMSSH chain seeded with the IAP exemption and a blackhole default.
    • It found one bug: accessConfigs[].natIP takes the address literal, not a self-link, so every create was failing 400.
    • Still unproven is time-to-certificate, which needs the Cloudflare record and so a deploy.

As built: the agent's side of enrolment, generated from its tests. Scenario 1 enrols, 2–4 are refused, and 5 finds the host already enrolled.

  1. Step 28 is the single-use point. The row is locked, so two machines running the same one-liner cannot both become this host.
    • Unknown, spent and expired are one answer: which of the three it was is not something a caller has earned.
    • Hardware that is not a GB10 Spark is rejected here, and the allocation it was holding is released.
  2. Step 29 is the whole payoff of provisioning first. One call returns everything, so the agent has no reason to stop and nothing to come back for.
  3. The operator keys are read at redeem, not at provisioning. Step 29 returns the signed list as it stands when the box enrols, and the public key its signature is checked against (M).
    • Enrolment is only the first delivery. Every change after it arrives through M.

The agent runs the whole of new_spark.md §4 and §6 as a state machine. Every step has the same shape, and the shape is the point:

As built: the prepare run, generated from the agent's tests. Scenarios 1–3 take the runner down each of its paths, and 4 is the real runbook.

  1. Check and apply are separate, for two reasons:
    • a re-run changes nothing on an already-prepared host, which is the property the whole machine rests on,
    • a dry run calls only check, so it reports what it would change while touching nothing.
  2. The re-check after apply is not belt and braces. A step that reports success without its check passing is how a host arrives "prepared" and broken, and nothing later would notice.
  3. It stops at the first step that will not complete. The orderings below are load-bearing, so a host that failed lxd-install must never reach owner-lockout.

#The 19 steps

#StepNotes
1preflightLinux, Ubuntu, uname -m and nvidia-smi match the allocated profile. Fail closed
2docker-purgepurge dockerd, iptables -P FORWARD ACCEPT
3swap-offswapoff -a, comment /etc/fstab, mask swap.img.swap
4earlyoomthe profile's floor, verbatim (spec28, retuned spec51)
5zfs-arc-cap/etc/modprobe.d/zfs.conf, 4 GiB
6headlessmulti-user.target, mask gdm
7watchdogRuntimeWatchdogSec + the boot-reset reporter (spec46); skipped when the profile says watchdog_sec: 0 (a guest with no watchdog device), fails closed when it wants one and /dev/watchdog0 is missing
8operator-accessspark user, every operator's key, NOPASSWD sudo, the tunnel minted back in A. authorized_keys is written as the whole list, so it holds exactly the operators and nobody else
9lxd-installapt lxd; do not rely on lxd init --storage-backend
10lxd-networklxc network create lxdbr0; lxd init is not run, so nothing else makes it
11firewallif ufw is active: allow in + route on lxdbr0, allow in on wg0, DEFAULT_FORWARD_POLICY=ACCEPT, reload. No ufw or inactive = satisfied
12storage-poolexplicit lxc storage create default zfs
13nvidia-runtimenvidia-container-toolkit-base + libnvidia-container-tools, for nvidia.runtime injection
14gpu-accountingNVML per-process accounting. Driver state, so re-asserted
15profileslofispark-base + lofispark-gpu
16base-imageimport the prebuilt image, else build it
17relay-peerapply the WireGuard peer collected in B; collision-check first
18owner-lockoutgated on hosts.lockout_owner, ships false
19self-checkmachine-check every verification, and report

Orderings that are landmines, not preferences:

Named rather than numbered, because a step inserted in the middle renumbers every landmine below it and the numbers are what a reader trusts.

  • operator-access before owner-lockout, and verified. Prove Enverge's access before removing the owner's. The reverse strands a box neither party can administer.
  • docker-purge before base-image. Docker's FORWARD DROP breaks lxdbr0 NAT, so the image build loses egress to the NVIDIA repo.
  • zfs-arc-cap before lxd-install. modprobe.d binds at module load, and installing LXD loads zfs even for a non-zfs pool.
  • swap-off is what arms earlyoom. It fires only when MemAvailable and free swap are low — an AND. A GPU runaway is unswappable, so leaving swap on disarms the floor entirely.
  • lxd-network before profiles. lofispark-base attaches eth0 to lxdbr0, and lxd init — which would have made both the pool and the bridge — is deliberately not run. Creating the pool explicitly without the bridge gives a host with LXD up, a healthy pool, and no way to attach a NIC; the failure then surfaces at profiles, several steps after the omission.
  • firewall before base-image and relay-peer. A default-deny ufw (cloud images and DGX OS alike) drops the bridge's DHCP, forwarded container egress and the relay's packets, and none of the three failures names the firewall: the image build dies resolving archive.ubuntu.com and the control plane just hangs. DEFAULT_FORWARD_POLICY has to change in the file — docker-purge's iptables -P FORWARD ACCEPT lasts only until the next ufw reload.
  • base-image builds from a script the agent carries. On the fleet it arrives with make sparkN-copy; a BYO host has nobody to run that, so the builder and its login banner are embedded in the binary. It is run with the profile's cuda_repo_arch — the script defaults to sbsa, which is arm64's NVIDIA repo and wrong everywhere else.
  • storage-pool must create the pool explicitly. A dir pool costs ~69 s on every VM create against 0.6 s on zfs, on a host that looks perfectly healthy.
  • base-image must end with the alias. The lifecycle launches by alias; an image without one fails every create.
  • relay-peer must check for a subnet collision first. 10.8.0.0/24 is a common home-router default, and bringing wg0 up into a colliding range makes the host unreachable the instant it succeeds.
  1. This is the only human decision in the flow, and the only thing between an enrolled machine and a paying customer on the shared spark-1 tier.
  2. gpu_type is not decided here. It was fixed when the allocation was made — BYO accepts one hardware configuration, and the box proved it was that at enrolment. Activation is the single flag, plus the placement chosen with it.
  3. The backend refuses to activate a host whose preparation did not finish.
    • The rule used to live in the render condition of an "approve" button on the enrolments table, so the hosts table — a second control issuing the same update — walked straight past it: a host that stopped at storage-pool could be activated by anyone who used the other one.
    • A gate only one of two paths honours is not a gate, so it moved into admin-host-action and the duplicate button is gone.

Five callers, three shapes. What separates them is how the host is chosen, because that decides when the row can be written and where the SSH hostname comes from.

#TriggerWho callsDrawn in
1the customer asksthe proxy (PXY)E1 — the picker chooses, and the customer is waiting
2capacity frees and the queue drainsdrain-waitlistE2 — no host until one accepts
3a paid pre-book window opensdrain-prebookE3 — the booking pinned its host when it was made
4an operator places one on a dedicated hostadmin-vm-actionE2's shape, host supplied instead of rotated
5an operator launches a booking by handadmin-prebook-actionE3's shape, one booking by id

The two operator paths are the manual twins of the two automatic ones — same steps, with the cron's gate replaced by a superadmin and the selection replaced by whatever the operator named. That is why they are not drawn separately:

  1. 4 is 2 done by hand. A dedicated owner often lands in the general queue, because the picker cannot reach their box; the operator places the VM there and closes the queued row the drain would have closed.
  2. 5 is 3 done by hand, targeted by booking id, and with a force that revives a booking torn down inside its own window.

The row is a precondition for dispatch, in all five. On a host whose agent writes none, a create sent without one produces a container the database has no record of: unbilled, invisible to the customer, and holding capacity the picker cannot see. Every caller now refuses rather than dispatches, keyed on the same pubkey marker — a fleet host is exempt, because its agent inserts the row a moment later and refusing there would break a path that works.

To the customer every one of these is identical to any other host, which is the point of BYO joining the spark-1 pool rather than a tier of its own. They are not identical in who writes the row: see G, where that difference is the whole story.

#E1. From the dashboard (41–52)

A customer's create, once the host is active.

As built: the agent's side of a create, generated from its tests. Scenarios 1–2 create an instance, and 3–5 refuse and revert.

  • Step 44 stamps created_at before the host is contacted, so billing runs from when capacity was committed rather than from whenever a host answered. A create that is never answered is marked creation_failed_at and bills nothing.
  • Step 49 is the agent programming the relay, and it had to be. An earlier draft had the backend do it, to keep the last secret off the host. It cannot: ssh_mapd listens on 10.8.0.1:8754, inside the tunnel, and Caddy proxies only :8080 — there is no route from Vercel. Publishing a second endpoint on every relay to create one would be more attack surface than the thing it protects. The token is the relay's own, minted per relay, and on a per-host relay it can forward one host's :22 to one port on that same host — which the agent decides anyway, being what publishes the port. It grants nothing a compromised host does not already have. A shared relay would break that, since one peer could forward another's :22 elsewhere, so a pool must move this back to the backend first.
  • It is fatal, unlike the GPU check. An instance whose :22 goes nowhere is not an instance. The fleet's Python agent logs a map failure and carries on, which is how a host ends up healthy with a blackholed :22 and nobody noticing.
  • Steps 47–48 split fatal from advisory. An instance the customer cannot reach is not an instance; a GPU check or a startup script that fails is reported, and the instance is created anyway.

#E2. From the queue, when capacity frees up (53–64)

The queue is the only path that creates a VM with nobody present, acting on an intent that may be days old — which is why billing is re-checked live rather than trusted, and why the host is not known when the row is written.

  • Step 58 inserts the row with no host, and that is the point. pick_hosts_for_gpu excludes any host holding a live vms row, so a host_id written before a launch succeeded would make this row occupy the candidate that just refused it — benching a working host against every concurrent drain until the row is marked failed.
  • Step 59 mints per attempt, not once. A name under the previous candidate's zone resolves to that host's relay, and the customer would land nowhere. It is minted in the function rather than by the database: a random label joined to a column, where TypeScript is type-checked, deployed by CLI and unit-tested.
  • Step 62 is the first write, and it is the whole difference from the dashboard path. Every other create knows its host at insert, so create_vm stamps the host and the name in one statement. This one cannot, so it stamps them once a host has actually taken the VM.
  • The name is written only if one was minted. A host whose own agent mints hostnames wrote that column back at step 60, and an empty string over it would take the address away from an instance already running.
  • A candidate has three outcomes, not two.
    • Refused — benched, and the loop moves on. This is what stops one drifted host, agent-busy while the database read it as idle, stalling the whole queue.
    • Skipped — step 58 could not write the row and this host's agent writes none, so a container here would be one nothing can account for. Not benched: the host did nothing wrong, and the fleet hosts behind it are still tried. A fleet host is never skipped, because its agent inserts the row itself.
    • Accepted — including an answer that will not parse. A 2xx means the container exists, and the id, name and hostname were all sent, so the values are known without reading them back.
  • An unreadable 2xx must not be scored as a refusal. It was: the parse sat inside the catch that scores attempts, so a successful launch benched a working host and sent the next candidate to build a second container for the same queue row. The guard that would normally catch that reads a row the agent writes, so on a BYO host it finds nothing and the loop carries on.

#E3. From a paid pre-booking, when the window opens (65–74)

A booking buys a named box for a date range. The window opening is the trigger, so nobody is present — and the host was decided when the booking was made, not when it launches.

  1. The gate is paid_at alone, and no live Stripe check. Payment is collected off-band and recorded by ops, which is the deliberate difference from E2 — and the reason a stale paid_at is a known risk rather than a caught one.
  2. The host is pinned on the booking, not picked. A booking without a pin is a legacy row and falls back to the pool pick.
    • An unavailable pinned host leaves the booking open for the next tick and an operator to reassign. A paid window is not re-placed on a different box just because the intended one is offline.
  3. Step 67 is a self-heal, not an idempotence check. VMs are soft-deleted, so launched_vm_id never nulls — without reading the row's liveness, a customer who deleted their box mid-window would be stranded until end_date. A non-live row falls through and the booking relaunches.
  4. The booking stays open after launching. expire-prebook closes it at end_date and tears the VM down, which is trigger 5 of G.
  5. The customer key is the booker's own, resolved from ssh_keys by user_sub rather than a per-booking id, so it can only ever be a key they hold. No key on the account leaves the booking open for an operator to follow up.

No proxy, no agent, no HTTP. L4 forwarding only.

  1. Standard port 22, no client software. The whole reason customer SSH was never moved onto the operator tunnel.
  2. Until the map has been written, :22 is a blackhole. Not an error — nothing is listening for that address yet. The agent:
    • writes it on create,
    • clears it on delete and stop,
    • re-asserts it when it starts, because a relay that restarted came back with an empty chain, and a host that rebooted has a container up the relay knows nothing about.
  3. The owner's box needs no inbound anything. It dialled out; every address here is ours.

Six callers, one shape. They differ only in what decided the instance should go. Three of them (3–5) run on a schedule with nobody present, which is why the shape has to be right without anyone reading the result.

#TriggerWho calls
1the customer asksDELETE /api/vms/{id} through the proxy (PXY)
2an operator, from /adminadmin-vm-action
3a payment probe fails mid-dunningprobe-payment
4dunning runs outteardown-vms, the gated reaper
5a pre-book window endsexpire-prebook
6an operator ends a pre-booking earlyadmin-prebook-action

As built: the agent's side of a delete, generated from its tests. Scenario 6, and 7 for an instance that is already gone.

  1. Two facts, deliberately separate.
    • delete_requested_at — the backend accepted a delete. Stops nothing.
    • deleted_at — the container is confirmed gone. Billing stops here.
  2. The row is not closed optimistically, because the database decides routing. pick_hosts_for_gpu excludes any host holding a live vms row, so a row closed before its container dies leaves a host that looks free: launches keep routing to an agent that refuses them, and every user behind it stalls.
    • So the bias runs the other way. A delete that does not complete keeps billing and keeps its host marked occupied, which is the honest state — the capacity really is still held.
    • Over-billing is refundable; lost utilization is not, and a poisoned queue costs every waiting customer rather than one.
  3. That trade is only defensible because the stuck case is visible. Step 80 is what notify-stuck-deletes scans for, and it also sizes the refund: the customer gave the instance up when they asked, not when an operator noticed.
  4. A delete ends when its caller stops waiting. The backend waits 90s, well past the ~30s a pair takes; cut off earlier, the agent leaves the instance gone but node1's rails bare and the relay still forwarding SSH to the dead port. A relay write abandoned that way is not recorded as a relay failure. Either way, every heartbeat reads the relay's SSH map (GET /map) and writes it again if it does not match what the host runs, so relay health is the relay as it is this minute, and a stale or cleared target is put right within a beat.
  5. A 404 confirms, it does not fail. The host has no such container — the orphan case, where it went away and the agent's own write never landed. The caller asked for it gone and it is gone, so the row closes rather than becoming one nothing can ever clear.
  6. Step 83 must precede step 84. pick_hosts_for_gpu excludes a host holding a live row, so draining before the row closes hides the host that just freed up from the customer waiting for it.
  7. Every caller owns both edges, and none may delegate them.
    • The fleet's Python agent PATCHes deleted_at from the host, so a caller that wrote nothing still looked correct there.
    • A BYO host's Go agent writes nothing to the database at all, so the same caller leaves the row live forever — billing, holding its host, and with step 80 unwritten, not even stuck. Just invisible.
    • admin-vm-action was the last caller doing this. It is the same asymmetry that produced the empty ssh_hostname and the orphaned container: backend code assuming the agent is a database writer.
  8. Both writes are guarded — delete_requested_at=is.null and deleted_at=is.null — so on a host whose agent does write, the two compose: whichever lands first wins and the other is a no-op. Two independent paths to the same value beat either alone.

Two different endings. Revoke cuts a host off and keeps it configured; uninstall hands the machine back.

  1. Order is the whole design. De-pool first, drain second (G), and only then remove access — the reverse strands a tenant on a host nobody can reach.
  2. The host confirms before backend state goes. Once the relay is gone there is no path to that machine, so anything still needed from it must be asked for first.
  3. Restore the owner before removing ourselves. The mirror of prepare's operator-access before owner-lockout, for the same reason.
  4. Revoke leaves the machine configured and locked. That is why it is not offboarding: an owner who simply wants to stop hosting needs uninstall, and "reinstall your machine" is a poor answer for a self-serve product.

Everything above is the platform driving a machine nobody is sitting in front of. This is the other thing: a person, on a host, because something is wrong.

  1. It is deliberately never on the happy path. spec63 splits knowing from acting: every diagnostic is what the agent reported and what /admin renders, because "SSH in and look" does not survive a thousand strangers' boxes.
  2. Acting on one is what this is for — push a binary, restart a unit, unwedge a prepare that stopped.
  3. There are two ways in, and the second exists because the first can fail.

#Why there are two

  1. The tunnel is the one that scales, and the one that can vanish.
    • It rides cloudflared, a unit the operator-access step installed with a connector token minted backend-side.
    • It needs nothing inbound on the owner's network, which is why it works behind a home router.
    • It needs Cloudflare, the host's outbound internet, and that unit alive. Lose any of the three and the door is shut.
  2. The relay is the one we own.
    • A GCP VM on our project, so IAP reaches it whatever the host is doing, and WireGuard puts it one hop from 10.8.0.2.
    • That hop matters more than it looks: the relay DNATs inbound public :22 into the tenant's container through LOFISPARK_VMSSH, but that chain is on PREROUTING, so a connection originating on the relay is not caught by it — ssh [email protected] from the relay lands on the host's sshd, not the customer's container.
  3. The two fail independently. The tunnel dies with Cloudflare, the host's uplink, or one systemd unit; the relay path dies only if WireGuard is down — which is the same as the host being off the network entirely, at which point there is nothing to reach.

#What the operator gets, and why that is the posture

spark, with their own operator key and NOPASSWD sudo — so, root. Not a concession: it is sole administrative control, the same asymmetry that lets owner-lockout remove the owner's login and never ours. A BYO host has exactly one administrator and it is us.

  1. One account, one operator key per operator. Every operator logs in as spark; what tells them apart is the operator key. sshd logs the fingerprint of the key it accepted, and operator_keys records each key's fingerprint against the operator's user sub, so a login is attributable without a second account.
  2. Leaving is one action. Removing an operator's key reaches every host through M, rented or not, so no host keeps a door open for someone who has gone.

#What has to be true first

Depends onWhy it can be missing
operator-access completed and verifiedIt is verified precisely because owner-lockout runs after it. An unverified tunnel plus a completed lockout is a box neither party can administer
The host has outbound internetBoth paths are outbound-dialled. A host off the network is unreachable by design, and that is what "we detect rather than prevent" means
The operator's key is on the hostDelivered at enrolment, then by M. An operator key added while the host was offline arrives when an operator sends again
relay-peer applied, for path 2Before it, the only way in is the tunnel
The relay still existsReleasing an allocation destroys it. Order matters in H: drain the tenant and take what you need from the host before the relay goes, because after that there is no path to that machine at all

#When neither works

  1. Nothing. That is the honest answer, and it is the design: a host that has been unplugged, reinstalled or taken back looks exactly like one that is merely offline, so we trust the absence of a heartbeat rather than the presence of a good signal. The response is to de-pool and revoke, not to try harder to get in.
  2. There is a third path while hosts.lockout_owner ships false — the owner's own login. That is why it ships false: until uninstall exists, their login is the recovery path if this flow wedges a box.
  1. Publishing is not delivery. CI builds and publishes on every merge, and that changes nothing on any host.
  2. Nothing on a BYO host ever fetches a binary — no timer, no polling, no update endpoint. The agent's steady state is to answer the backend and otherwise sit still.
  3. The one thing that moves unasked is enverge-agent-rollback, and it moves backwards — once, when serving will not stay up: it consumes .previous and will not restore the build that is already running (L1).

So an upgrade is an operator action, over the same tunnel as I.

  1. GitHub Actions belongs to none of the six areas. It publishes, and publishing is where its involvement ends — which is the point of the section.

  2. The release verifies the path an owner's installer follows, not the upload. put() returning a URL says the object exists; it says nothing about whether app.enverge.ai/agent/… resolves.

    • That gap once let AGENT_BLOB_BASE_URL sit blank in production — present in the dashboard, so it read as configured — while every install one-liner was broken.
  3. No token, and that matters more than it sounds. The installer asks for one only when the machine is not enrolled; an enrolled host reuses its credentials.

    • The alternative would be an operator obtaining a fresh token, and the only way to get one is to add a GPU — which provisions a second host, relay and tunnel.
  4. Detached, deliberately. A dropped SSH session must not land between writing the binary and writing the units.

  5. Preparation starting serving is one line, and it went missing.

    • The ordering between the two units is declared on the serving one — Requires=/After= the prepare unit — and that direction only pulls preparation in when serving starts, never the reverse.
    • So the installer's "restart prepare" started nothing else: spark-c05 prepared itself, exited 0, and served nothing until it was started by hand.
    • The prepare unit's Wants=enverge-agent.service closes it. Earlier hosts hid it because they had been rebooted since installing.
  6. Restarting serving is the installer's job, not the Wants='s. On an upgrade the serving unit is already running, so Wants= is satisfied and systemd starts nothing — the host would keep answering from the binary the install just replaced.

    • An upgrade restarts serving; the preparation it pulls in exits at once.
    • A first install starts preparation only; its Wants= starts serving once credentials exist.
  7. An upgrade does not re-check the host. Preparation does not run after go-live (L).

    • The done list is still scoped to the build that wrote it. Steps are recorded by name, and a name says nothing about what the step now does: lxd-install gained usermod -aG lxd while keeping its name, so hosts that had completed it skipped the new work and reported "already satisfied".
    • That is what makes the deliberate run in L re-check every step on a newer binary, rather than trust names an older one wrote.
  8. An upgrade does not reach containers that already exist. The login banner is pushed at create, so a running instance keeps what it was given. Recreating it is the only way to change what a customer sees.

  9. One command covers every BYO host, unlike the per-host targets the fleet carries.

    • Those exist because the fleet's agent is copied out of a working tree; a BYO host fetches the published binary itself, so the only per-host input is the slug.
    • The corollary: upgrading several hosts is a loop an operator runs by hand — fine at two, and the first thing to hurt as the fleet grows.

One hosts row, one relay, one tier slot, one enrolment, one token, one command on node0, one prepare run, one heartbeat. Agentless satellites are reached from node0 — nothing is ever run on them by anyone. One table, host_nodes, says which machines a host is made of. K1 is two machines; K2 is four on a switchless ring. Everything from E onwards is unchanged either way.

#K1. Two nodes — spark-2

A spark-2 host is two Sparks cabled together over ConnectX-7 and sold as one unit. The second machine is reached from the first.

  1. This is new_spark_dual.md with the operator removed. The fleet's pairs are node0, a full platform host, and node1, a machine with no agent, which a person prepares over SSH and node0's agent drives over SSH at create and delete. Node1 keeps that role exactly — no agent, no relay, no tunnel, no row — and the person's part moves into node0's prepare.
    • NVIDIA's own stacked Sparks guide is the same work with a person on each box: matching users, passwordless SSH both ways, static addresses on the QSFP interfaces. Steps 125 and 129–131 are that guide, run from one side.
    • An earlier draft enrolled node1 in its own right, so that each box proved itself. It cost a second profile that was not a hardware profile, a second prepare track folded into one column, a heartbeat with nowhere to land, and a second connector on node0's tunnel. And it bought nothing: the lifecycle that follows needs node0 driving node1 over SSH anyway, and "this is a pair" is rail traffic between two GB10s whichever box reports it.
  2. Step 123 is the one thing that has to be asked for. Node0 needs root on a box nobody has touched, and the login is the owner's to give. It is collected on node0's terminal rather than in the dashboard because that way it never reaches the backend: root-only on node0, read at steps 125 and 129, then deleted. From then on node1 admits node0's peer key and the operator keys. peer-prep then activates node1's ConnectX-7 ports and pins link-local eui64 before prepare starts, because NetworkManager often leaves them disconnected while the kernel link is up — no fe80::, and the cable probe at 125 would see silence forever. It asks the cable first. A rail is point-to-point, so whatever answers the multicast on it is node1 by construction; node0 logs in, groups the two addresses a QSFP shows into the one machine behind them, and prepares it at its management address — reconfiguring node1's rails through the rail being spoken over is a way to cut the session. The management LAN (mDNS _ssh._tcp and ARP on the uplink — not Tailscale) is the fallback, for exactly the case above: a node1 whose rails cannot yet answer. A rack of eleven Sparks on one LAN, five accepting the same login, is what that ordering is for — the scan can only report the ambiguity, and the cable has already answered. ENVERGE_PEER_ADDR names node1 outright and skips both.
  3. Step 125 finds node1 by the cable. A rail is a point-to-point link, so a multicast ping to the all-nodes address on each RDMA port returns node1's link-local addresses before any static address exists. The probe ignores node0's own addresses, so a cable looped between its two ports finds nothing. A ConnectX-7 QSFP is two PCIe functions on one wire, so one cable shows two of node1's addresses; the probe logs in and counts machines, not addresses. Login tries the peer user's existing SSH private keys, then the password from the install prompt. A port whose neighbours are all one other Spark is a link; a port with several machines is the owner's LAN — a Spark's uplink can sit in a QSFP port — and is left alone. When every address on a port refuses the login, that is named as a login failure, not as a LAN. The links are numbered as the rails of the plan exactly once — here, the first time node0 looks — and the numbering is recorded on node0 on the spot, so no later run re-derives it from the order ports come in. It then asks which of node1's ports holds each address it saw. That is how a crossed cable becomes a recorded fact on both rows rather than a failed verification.
    • Link-local persistence is required and automatic — not a manual step. DGX OS often leaves ConnectX-7 ports at addr_gen_mode=none, so they have carrier but no fe80:: and the multicast probe sees silence; NetworkManager's default also rewrites a live sysctl on reconnect. Before probing, the agent pins each of node0's RDMA ports' NM connection to ipv6.method=link-local and addr-gen-mode=eui64, forces the live sysctl, and bounces the iface if an address is still missing. After a successful login it does the same on node1's rails over sudo. Later, rail-network (130) keeps eui64 on the planned rails via netplan. There is no owner/operator runbook for this.
    • The enrolment has not been accepted yet, so 125 still must not prepare the box or install peer-access. The only allowed pre-enrol write is that NM link-local pin (and the matching live sysctl), so the cable stays reachable; a rejected peer is otherwise left as found.
    • If node1 cannot be reached or the login is wrong, the agent does not enrol. It retries in-process every minute while the token lives, re-reading the login, so the token stays unspent.
  4. Step 127 is the one place a lone Spark is turned away, and it is early on purpose. Redeem is where turning away is cheap: the token is spent, the row parks as rejected, reap releases the relay, the address, the tunnel and the host row, and the owner is still at the terminal to see it. Anything found later can only stop a prepare run and leave the allocation held. So every piece of evidence that can exist before enrolment is judged here, once: at least one cable and no more than the plan has rails for, one machine on every cable, a peer that is the profile's hardware and not node0 itself, and every machine carrying its own node_index — 0 for the enroller, each other one once. Rows are written at the index a machine carries, never at its place in the list. nvidia-smi on either box reports GB10 / 1 / aarch64, byte-identical to a single, which is why the dropdown alone was never evidence.
    • What was found becomes rows: one host_nodes row per machine, node0 at index 0 and the peer at index 1, each with its own fingerprint — hardware identity — and its own rails, the observed state of each port, with the rail each cable carries and the plan's address for that node on it stamped in. Every BYO host gets them, so a single Spark is one row and the node count is a count. unique (hardware_id) is the "not node0 itself" rule as a constraint, and it also stops one box being claimed by two hosts. It keys on the GPUs rather than /etc/machine-id because that file clones with the image it was captured in; machine_id stays on the row as reported hardware. An agent older than the field sends none and its machine_id is written there instead, which is the namespace its own row was already in.
    • The profile is the plan and the node row is the fact. constants.rails says a subnet and an MTU per rail, and a node's address is the subnet's tenth host plus its index; host_nodes.rails says what these two machines look like, port by port, and is refreshed at 130–133 as they change.
  5. Step 132 is what redeem cannot know. Whether the pair works needs both machines prepared and addressed, so the traffic test is late by nature. It proves the same claim step 127 accepted, with the rails configured. Both are reported by the owner's own hardware, so neither is more trustworthy than the other — what differs is the cost of being wrong, which is lowest at 127 and highest here. Approval (D) follows it and stays the last thing in the flow.
  6. Step 129 decides nothing. peer-access creates the spark user, installs the peer key and the operator keys, and deletes the login. Its check still refuses an SSH host key that is not the pinned one. After go-live it runs only in the deliberate run (L).
  7. Steps 130–131 are C, run on node1 from node0. Every step already runs through an Exec; a remote one over the rail runs the same check, apply and re-check on node1 and reports each with index: 1, which host-report writes to node1's own host_nodes row rather than to the host's prepare_state. The block is node1's subset of the list in the same order, so every landmine ordering holds there too, and the deliberate run in L re-checks node1 as it re-checks node0. Node0's track keeps its shape and its meaning; node0's self-check, after the block, asserts node1, which is why steps 39–40 and both activation gates are untouched.
  8. Step 133 does not watch node1 over the rails. The heartbeat refreshes node0's own host_nodes.rails — per port, how many machines answered — which reads zero on both ports while an instance holds the rails, so it cannot tell a dead node1 from a rented pair.
    1. That is why the node1 probe was dropped: it reported every rented pair degraded while every subsystem was green.
    2. There is no host_status row for node1, and host_nodes.seen_at is no longer written.
    3. node1's worker is checked over the management network: L3.
  9. Break glass (I) reaches node1 through node0. Node1's spark user holds the operators' keys, so the tunnel or the relay lands on node0 and one more hop lands on node1. Node1 has no tunnel of its own.
  10. A single Spark walks the same steps and none of this fires. Its install command carries no pair marker, so 123–124 never happen; with no login file, 125 never happens; its profile has no node_count above one, so 127 judges nothing and 129–132 are skipped, reported as skipped the way watchdog is on a board without one. The rails still appear in its fingerprint, which is a few more fields and nothing else.
    • A single Spark presented as a pair with its own login typed in as node1's is caught twice: the probe ignores node0's own addresses, and redeem refuses a peer whose identity is node0's. That identity is hardware_id — the machine's GPU UUIDs, digested — not /etc/machine-id: systemd writes that file once, and only when it is empty, so an OS image captured with it populated hands every unit the same value and two genuinely different Sparks present as one machine that found itself. The agent makes the same comparison before it enrols, so a real self-pair costs a retry rather than the token.
    • A single Spark whose rails have carrier because they are on the owner's switch is fine against the single profile, where nothing checks rails, and fails the one-neighbour rule against the pair profile, which is the intended answer.
  11. The customer's instance is two containers: the head on node0 and a worker on node1, named <head>-w1. node0's agent creates and deletes them together over the management network — not over the rail SSH step 129 established, because the dual profile moves the rail netdevs into the containers the moment one starts.
    1. At create, any worker already on node1 is removed first, the head mints a per-instance launcher key, and the worker trusts that key alone. A worker that cannot start fails the create: one GPU sold as two is not an instance. A create refused after the worker started — the customer key or the relay failing — reverts the worker along with the head, and restores both machines' rail addressing.
    2. At delete, the worker goes with the head, and both machines' rail addressing is restored.
    3. Keeping the two running together — through start and stop, and after either machine restarts — is L.

#What K writes, and where

One new table, host_nodes, one row per machine of a host. Everything else is a table that existed, written at the steps it was always written at, with the meaning it always had. The database is written at 120, 127, 131 and 133, plus the approval at 134.

StepReadsWrites
118host_profiles.staff_only, to decide whether the picker shows the row—
119staff_only for the gate, constants.node_count for the command—
120—hosts row; host_enrollments row with host_id, profile_id, token_hash, wg_conf. Same as a single
121—Nothing stored. The response carries node_count and --peer --nodes=N on the command (N from the profile)
123–124—/etc/enverge/peer-login and peer-nodes on node0; peer-prep finds node1 on the rail (falling back to the management LAN only when the rail is silent) and activates its RDMA NM connections + link-local eui64 (writes on node1 before enrol, so the cable is visible)
125peer-login; node0's RDMA NM connections and sysctls; node1's nvidia-smi (including GPU uuids), uname, /etc/os-release, /etc/machine-id, sysfs ports, and which port holds each addressOn node0: NM ipv6.method=link-local + addr-gen-mode=eui64 on every RDMA port (and live addr_gen_mode=0); peer.json with the links the moment their numbering is assigned. On node1, after a successful login only: the same NM link-local pin on its rails (via sudo). No prepare, no spark user, no keys. In memory the fingerprint gains node_index 0, machine_id, hardware_id, rails[].neighbours_discovered, rails[].index, and peers[] each with its own node_index
126—Nothing yet. The request body is the fingerprint from 125
127host_enrollments by token hash for profile_id; host_profiles for gpu_model, arch, gpu_count, constants.node_count, constants.railsRefused: host_enrollments.state = rejected, token_hash = null, fingerprint, pubkey, source_ip, redeemed_at. Accepted: the same columns with state = ready and node0's identity alone, hosts.pubkey and hosts.profile_id, then a host_nodes row per machine — identity in fingerprint, what detect saw of each port in rails, with the rail index each cable carries and the plan's address for that node stamped in
128hosts for the token; constants merged with arch and gpu_model; the tunnel session; operator_keys; wg_confhost_enrollments.wg_conf = null; credentials.json on node0, carrying node_count, the rails plan and the operator keys
129peer-login; every RDMA port, for discovery; node1's ip -6 addr, for which port each cable lands onOn node1: the spark user, authorized_keys with node0's peer key and every operator key, sudoers.d/enverge-spark. On node0: peer.key, peer.json with addr, host_key and the links — index, node0's port, node1's port; peer-login deleted
130peer.json for the links; the plan's subnet per rail and node1's index, for node1's addressesWhatever that step writes on node1: netplan on node1's ends of the links (including ipv6-address-generation: eui64, which keeps the 125 pin on the planned rails), earlyoom, LXD, the base image
131—host_nodes.prepare_state on the row at index 1, replaced with {step, state, attempt, error, final, at}. Node0's own steps keep writing hosts.prepare_state as before. A rail-network transition, from either machine, also carries that machine's ports as observed after it ran — rail index, MTU, active_mtu, the address on the interface — merged into that node's host_nodes.rails
132the links; the plan's subnets and both indexeshost_nodes.rails[port].verified_at on both rows for every rail that carried traffic. On failure, hosts.prepare_state with step: rail-verify, state: failed
133node0's own portshost_status for node0; node0's host_nodes.rails refreshed — MTU, active_mtu, neighbours_discovered per port — so a cable that comes loose after approval shows on the port that owns it. Nothing for node1 (L3)
134host_enrollments.state must be ready; hosts.prepare_state must be self-check applied or already-satisfiedhost_enrollments.state = enrolled, then hosts.active = true with a pool

Node1 is written at three steps: 125 (NM link-local pin only), 129, and 130. Step 125 is otherwise a look that makes the refusal at 127 possible without preparing the box — the pin is the one pre-enrol exception, applied by the agent on both nodes so the owner never has to.

#The shapes

host_nodes, the new table. One row per machine per host, for every BYO host, so a single Spark is one row. Identity and state are separate columns, because one never changes and the other is what an operator reads. The index is explicit: what the agent, the backend and the plan agree on.

create table host_nodes (
  host_id       text not null references hosts(id) on delete cascade,
  index         int  not null,          -- 0 is node0: the relay, the tunnel, the agent
  machine_id    text,                   -- /etc/machine-id, as reported: hardware detail, not identity
  hardware_id   text not null,          -- sha256 over the sorted GPU uuids: what tells identical boxes apart
  fingerprint   jsonb not null,         -- hardware identity: gpu, memory, arch, os, driver
  rails         jsonb not null default '{}',  -- the ports as they are, keyed by interface
  prepare_state jsonb,                  -- that machine's latest {step, state, attempt, error, final, at}
  seen_at       timestamptz,            -- no longer written: node1 is not probed (133)
  created_at    timestamptz not null default now(),
  primary key (host_id, index),
  unique (hardware_id)
);

host_profiles.constants, the plan, written once. staff_only is a column, being policy. The GB10 constants unchanged plus the two things a pair adds: how many machines, and a rail plan — each rail named by an explicit index, never by its place in the list, with a subnet and an MTU. The subnets and hosts are the ones NVIDIA's stacked-Sparks guide assigns: a node's address on a rail is the subnet's tenth host plus the node's own index — node 0 is .10, node 1 is .11 — and no interface is named, because which port carries which rail is the cable's decision.

{
  "node_count": 2,
  "rails": [
    { "index": 0, "subnet": "192.168.100.0/24",   "mtu": 9000 },
    { "index": 1, "subnet": "192.168.101.0/24", "mtu": 9000 }
  ]
}

host_nodes.rails, the fact beside the plan, keyed by this machine's own interface. Written at 127 from what detect saw; the rail index and the address stamped onto the ports that carry a cable; refreshed at 131 after rail-network ran on that machine; verified_at stamped at 132; neighbours_discovered kept current at 133 for node0. A port nobody is on appears with neighbours_discovered: 0 and no index. Here node1's second cable landed on its other port — recorded, not a fault.

{
  "enp1s0f0np0":   { "index": 0, "device": "rocep1s0f0",   "mtu": 9000, "active_mtu": 4096,
                     "address": "192.168.100.11/24",   "neighbours_discovered": 1, "verified_at": "2026-09-08T14:20:03Z" },
  "enP2p1s0f1np1": { "index": 1, "device": "roceP2p1s0f1", "mtu": 9000, "active_mtu": 4096,
                     "address": "192.168.101.11/24", "neighbours_discovered": 1, "verified_at": "2026-09-08T14:20:03Z" },
  "enP2p1s0f0np0": { "device": "roceP2p1s0f0", "mtu": 1500, "neighbours_discovered": 0 }
}

The enrolment request at 126, what the agent sends. This is a wire shape, not a stored one: host-enroll splits it into the enrolment's own fingerprint (node0's identity, without rails or peers) and the host_nodes rows, identity into fingerprint and what detect saw of each port into rails.

{
  "gpu_model": "GB10", "gpu_count": 1, "gpu_mem_mib": 122880, "arch": "aarch64",
  "os_id": "ubuntu", "os_version": "24.04", "driver_version": "580.65.06",
  "machine_id": "3f2a…", "hardware_id": "b41c…", "node_index": 0,
  "rails": [
    { "device": "rocep1s0f0",   "iface": "enp1s0f0np0",   "mtu": 1500, "active_mtu": 1024, "neighbours_discovered": 1, "index": 0 },
    { "device": "roceP2p1s0f0", "iface": "enP2p1s0f0np0", "mtu": 1500, "active_mtu": 1024, "neighbours_discovered": 1, "index": 1 },
    { "device": "rocep1s0f1",   "iface": "enp1s0f1np1",   "mtu": 1500, "neighbours_discovered": 0 }
  ],
  "peers": [
    { "gpu_model": "GB10", "gpu_count": 1, "arch": "aarch64", "machine_id": "9c17…", "hardware_id": "7e08…", "os_id": "ubuntu", "node_index": 1,
      "rails": [ { "iface": "enp1s0f0np0", "index": 0, "neighbours_discovered": 1 }, { "iface": "enP2p1s0f1np1", "index": 1, "neighbours_discovered": 1 } ] }
  ]
}

A refused pair lands whole on host_enrollments.fingerprint under state = rejected, which is how the enrolments page shows what turned up.

hosts.prepare_state, node0's track, unchanged in shape and meaning. host_nodes.prepare_state on the row at index 1 has the same shape and is node1's:

{ "step": "base-image", "state": "applied", "attempt": 1, "error": null, "final": false, "at": "2026-09-08T14:02:11Z" }

host_status, node0's row, written at 133. No column for node1: half a pair shows as lxd = degraded (L3).

Node0's files, under /etc/enverge, root-only, never in the database:

  • peer-login — node1's username and password from 124, deleted at 129
  • peer.key — node0's peer key, Ed25519
  • peer.json — the links from the moment they were numbered, then the pinned node1: {"addr": "fe80::…%enp1s0f0np0", "host_key": "ssh-ed25519 …", "peer_index": 1, "links": [{"index": 0, "local": "enp1s0f0np0", "remote": "enp1s0f0np0", "neighbour_addr": "fe80::…%enp1s0f0np0"}, {"index": 1, "local": "enP2p1s0f0np0", "remote": "enP2p1s0f1np1", "neighbour_addr": "fe80::…%enP2p1s0f0np0"}]}

The password never leaves node0, and the backend never learns node1's link-local address. What it knows about node1 is its host_nodes row: hardware, the state of its ports, progress, and when it was last seen.

#K2. Four nodes — switchless ring

A spark-4 host is four Sparks on a switchless CX7 ring, sold as one unit. Everything in K holds; what differs is how many satellites there are, how node0 finds the one with no cable to it, and how many workers an instance carries.

  1. The command carries --nodes=N. host-provision puts --peer --nodes=${node_count} on the command whenever node_count > 1. The installer writes /etc/enverge/peer-nodes, so detect and peer-prep know the target before enrol returns the profile.
  2. One shared login for every satellite. Asked once on node0's terminal, into the same peer-login as a pair. Wrong password: re-run the one-liner.
  3. The satellites are the machines on the cables. node0 finds its cabled neighbours, logs in to each, and asks what answers on their rails; the diagonal is reached one hop through either of them, until nothing new answers. Its management address and SSH host key are read from itself over the cable, and later access over the LAN is held to that key. The management LAN (mDNS + ARP, not Tailscale) is scanned only to wake a machine whose cabled ports have carrier and no link-local, and the cables are walked again: a machine that only shares the LAN is never a satellite, however the password is set. More machines on the cables than the profile has satellites is refused.
  4. A satellite is its hardware_id, as in K. Machines on the cables are grouped and deduplicated by GPU identity — never /etc/machine-id, which a ring flashed from one image shares on all four boxes. node_index 1..N−1 is assigned by sorting hardware_id and pinned in peer.json, so later runs do not reshuffle.
  5. peer.json lists every satellite: peers: [{index, addr?, iface?, host_key, mgmt_addr?, links}, …], with rail addr/iface only where there is a cable from node0. A satellite's links are its own cables, each carrying peer — the node at the other end — so the diagonal's two cables are recorded on it even though node0 has no end on either. A pair-shaped file (flat fields, one peer) reads too.
  6. Prepare runs node{k}/… for each satellite with K1's node1 step code. The diagonal is reached over the management network. rail-verify sends traffic across every cable from both of its ends, the two between the satellites and the diagonal included.
  7. An instance is four containers: the head on node0 and <head>-w1…-w3 on the satellites, created and deleted together over the management network. A worker that will not start fails the create, and every worker already launched is reverted.
    • Each container gets its node's two cables from lofispark-dual, rendered per node. On a ring every node also forwards IP and carries a route to each rail it is not on, through the neighbour that is — so the diagonal reaches the head, and the head the diagonal, as the switchless recipe does it. The head's /etc/enverge/cluster.md names every node, its rail IPs, and which one has no cable to it.
    • The profile's description carries a digest of its content; profiles rewrites a profile that does not match what this node renders now (a machine re-enrolled from another setup), except while an instance holds the rails.
  8. The profile plan (gb10-spark-4) is node_count: 4 and four rails, one per cable (192.168.100–103/24), as the switchless recipe the ring is built from. Each cable is addressed on the PCI-domain-0 port at both ends (enp1s0f*), never on its enP2 twin, and cables are numbered from the cable map by (lower node, higher node). A port that no longer carries a rail loses its rail address, so a re-numbering never leaves the same address on two ports of one wire. Redeem allows fewer cables than rails and refuses more.

Everything above ends with a host serving. This is what keeps it serving afterwards, when nobody is onboarding anything: a new binary, either machine restarting, and a reboot of both.

One constraint bounds all of it: node0 is the gateway and satellites have no agent (ARCHITECTURE.md — Dual-host clusters). Whatever happens on a satellite is done by node0's agent over the management network, or by host configuration node0's preparation laid down there — never by a program of ours running on the satellite. On a pair that satellite is node1 (L3); on a ring it is each of node1..node3 the same way.

J could treat the agent as one program. Here its three units are the point: PRP, SRV and RB in the glossary.

#The rules

  1. A reboot runs no preparation. Serving starts at boot on its own. The last act of a completed preparation — handing the credentials to the serving user — is the record: later starts find it and exit at once, whatever the binary, so serving's Requires= is satisfied without a run, and a host that never completed preparation still cannot serve.
  2. The rollback fires when serving fails to start. It judges the binary. A preparation step failing on the host's state — a cable, a lent rail, node1 — is not a bad binary.
  3. Serving restores the instance on both machines when it starts — LXD, the head, then the worker on node1 over the management network — and only then points the relay.
  4. The pair is kept in lockstep: the worker exists and runs exactly when the head does (Lockstep, below).
  5. "An instance holds the rails" is asked of both machines before preparation touches a live host, node1 over the management network. When node1 cannot be reached, node0's answer is enough.
  6. Preparation does not run after go-live. Not on a boot, not on an upgrade: an upgrade swaps the binary and restarts serving (L1). A change to host configuration is applied deliberately, on an idle host (Host configuration after go-live, below).

#L1. A new binary (135–141)

  1. Serving is down only while it restarts. The tenant's containers keep running and their SSH keeps working. What pauses is the VM API and the heartbeat.
  2. The rollback runs once, and never to the build already running. A previous binary that fails too leaves the host down and alerted agent_unreachable, rather than flipping between two broken builds.
  3. No preparation runs. The new binary serves the host the previous one prepared. What a newer agent would change about the host waits for the deliberate run below.

#L2. node0 restarts (142–147)

  1. Serving does not rely on LXD starting itself. Both containers carried boot.autostart: "true" from lofispark-base at the time, yet LXD's boot activation started nothing on either machine of spark-21 across three boots (LXD 5.21.7). LXD starts when something first calls it. Serving waits for snap.lxd.activate, which opens LXD's socket to the lxd group.
  2. What ran before a reboot runs after it, and nothing else. With boot.autostart unset, LXD restores each instance's last state, and the agent starts the head only when that state is running.
    • Older hosts carry boot.autostart: "true" — always start — until make byo-prepare removes it. spark-15 to spark-23 had it removed by hand on 2026-09-14.
  3. Head, then worker, then relay. SSH is forwarded once, at start. Forwarding it before the head runs would send :22 nowhere until the next restart.
  4. A node1 that cannot be reached does not stop node0 serving. The heartbeat keeps asking (L3).

#L3. node1 restarts (148–152)

  1. The check rides the heartbeat, the one schedule serving keeps. It asks node1's LXD over the management network — never the rail, which is inside the containers.
  2. A restarted worker needs nothing else. Its rail addresses, /etc/nccl.conf and the head's launcher key live inside the container and survive a restart. On spark-21 both containers came back with their rails addressed, RDMA active at MTU 4096, and the head reaching the worker as user.
  3. Half a pair is degraded, and de-pooling stays an operator's decision. The signal K.8 once took from the rail is taken from the container, which reads the same during a rental as outside one. It lands on lxd, since host_status has no column for a pair.
  4. Both machines at once — a power cut — is L2 and then L3. Serving restores what it reaches at start, and the heartbeat finishes the job when node1 answers.

#L4. A reboot of the pair (153–158)

As built: the agent's side of a reboot, generated from its tests. Scenario 8, with node1 answering.

  1. node1 first. It has no agent, so once node0 is down nothing is left to tell it anything.
  2. Both reboots are deferred, so each command returns before its machine goes down, and the caller hears a 202 rather than a dropped connection.
  3. An unreachable node1 does not block the reboot. Refusing would leave the customer with no reset at all. node0 reboots, and the heartbeat reports half a pair until node1 answers (L3).
  4. The customer and an operator press the same button. The dashboard's reboot and /admin's are one call, so a pair is reset the same way whoever resets it.
  5. Recovery is L2 and then L3, as for a power cut. Nothing about the instance is rebuilt: both containers come back as they were.

#Lockstep

The rule is one sentence: the worker exists and runs exactly when the head does. Nothing else needs keeping in step — the rail addresses, /etc/nccl.conf and the launcher key live inside the containers.

EventThe workerBuilt
Createlaunched after the head, trusting its launcher key; failing to start fails the create, and a create refused later reverts it along with the headyes
Deleteremoved with the head; both machines' rail addressing restoredyes
Start, stop, restartfollows the head; a start that fails is repaired from the heartbeatyes
node0 restartsstarted by serving, after the head (L2)yes
node1 restartsstarted from the heartbeat (L3)yes
Host rebootnode1 rebooted first, over the management network; back through L2 and L3 (L4)yes
A worker with no headremovedat create only

#Host configuration after go-live

  1. Nothing runs preparation on its own once a host is live — not a boot (rule 1), not an upgrade (L1).
  2. A change to host configuration is applied deliberately. When a newer agent changes what a step does — lxd-install gaining usermod -aG lxd under the same name was the first — an operator runs preparation on the host over the operator tunnel (I): make byo-prepare BYO_HOST=….
  3. Only on an idle host. Rule 5 is how the run knows: no running instance holding the rails on either machine, or on node0 when node1 cannot be reached. A rented pair is left alone until it is not.
  4. It is C again, in full. The done list is scoped to the build that wrote it, so a newer binary re-checks every step, and the pair's steps prove the cable once more — possible only because nothing holds it.
  5. Its failure is reported, not rolled back. It shows in /admin, leaves serving running, and does not fire the rollback, which judges serving only (rule 2).
  6. The operator keys are not host configuration. They change whenever an operator joins or leaves, not with the agent, and they reach live hosts through M rather than by preparing again.

The operator keys are the one thing on a host that changes on Enverge's side rather than the owner's or the agent's. An operator joins, an operator leaves, and every host has to agree — including the ones a customer is renting, which preparing again (L) would not touch until they are idle.

  1. The list is sent whole, never as a change. Adding and removing are the same call, and a push that is missed is repaired by the next one rather than replayed.
  2. An empty list is refused twice. admin-operator-keys will not write a change that leaves no operator, and the agent will not apply a list with no keys. Either would leave a box nobody can administer — the same reason operator-access is verified before owner-lockout (C).
  3. A rented host is changed too, and that is safe. Setting the keys writes spark's login file and nothing else: not the rails, not a container, not a unit. That is why it is not a preparation step, which on a live host waits for idle (L).
  4. Serving can do it without root. authorized_keys belongs to spark, the user serving already runs as, and node1's belongs to the spark user node0 already reaches it as. The serving unit may write spark's .ssh and nothing else in /home.
  5. The agent keeps the list where preparation reads it. Preparing again (L) writes authorized_keys from what the agent holds, so each push also updates that copy — the signed list and its version, as received. Otherwise a later run would put back the list from enrolment, and with it any operator removed since.
  6. A failed push fails, and says so. Nothing retries on its own. /admin lists every host that was not updated with the reason, and the attempt stops there until an operator sends again.
    • Sending again re-sends the stored signed list, so it needs no signature and no password manager — one click, once the host is reachable.
    • An offline host is out of date but not open: with it off the network, nobody logs in to it. What matters is the time between it coming back and someone sending again, and /admin is where that shows.
  7. Only a signed list is applied, so the host token is not enough.
    • Why. hosts.agent_token is readable from the database and used by the proxy, and the agent's API answers it from the internet. Until now it could create and delete containers. If it could also set the login keys, it would be root on the host, and a leak of the database or the proxy would be root on every BYO host.
    • How. A person signs every list, on their own machine. The host receives the matching public key once, at enrolment (step 29), and checks every list against it. No call changes that public key.
    • Replays. The version is inside what is signed, and the agent refuses a version older than the one it holds. A token holder who kept an old signed list cannot use it to bring back an operator removed since. The same version is taken again only with the same keys, which is what lets sending again land after an answer was lost.
    • Where the signing key lives. In the operators' password manager. To sign, an operator copies its private half from the web vault into a throwaway SSH agent on their own machine, for one signature: the sign command clears the clipboard once the agent has it and stops the agent when it is done, so it never reaches the disk. No server holds it and no Enverge page sees it, so a leak of the database, the edge-function secrets or the proxy signs nothing: changing the keys takes the password manager and a superadmin session together.
    • What that exposes. For the seconds between copying and signing, the private half is in the operator's clipboard, and so within reach of anything that reads it there — a clipboard manager, or another device sharing the clipboard. That is the price of signing from the web vault.
  8. What is signed is checked where it is signed. /admin is code served by Vercel, so a tampered page could offer an attacker's list for signing. The signing command prints the operators and key fingerprints from the bytes themselves, and the operator reads that before approving — what they check is exactly what they sign, whatever the page showed.
  9. Only a change needs a signature. Enrolment (step 29) and sending again use the latest signed list as stored, so neither needs the signing key. What cannot happen is any change to who gets in without someone signing it — which is the point.
  10. Every operator's changes add up in one draft. There is one list: every operator's keys, one per machine, in one version. A change is one key added or removed, applied by admin-operator-keys to the newest list — the one waiting for a signature, if there is one — so Tudor adding his key and Breno adding his before either signs end in one version holding both, and one signature publishes it. Two changes landing at once cannot overwrite each other: the second is applied again on top of the first. Only the newest version can be signed, and it carries every change before it.
    • A key added belongs to whoever adds it, signed in as themselves: nobody grants access in someone else's name.
  11. A host enrolled before signed lists trusts no signing key. It refuses every list and keeps the single operator key it enrolled with, until an operator tells it which signing key to trust over the tunnel (I), then sends again. That is the only way a host's trusted signing key is ever set after enrolment.
  12. A proposal nobody will sign is withdrawn, for everyone. Withdrawing deletes it, and any older unsigned proposal it replaced, so none of them comes back as the one waiting. Only an unsigned proposal can be withdrawn: nothing unsigned ever reached a host, so there is nothing to undo.
  13. An operator who leaves takes the signing key with them. The signing key is one Bitwarden entry shared by every operator, and signing from the web vault puts its private half through each signer's clipboard, so anyone who has signed may hold a copy. Removing their operator key is not enough: they could sign a list that adds it back. The signing key is changed too — trusted on every host by hand (I), then a new list signed with it.

A host that finished enrolling ends here, for good: its relay, tunnel and DNS go, and nothing it was issued still works. The row stays, and so do its instances and its nodes, so the host's history is still readable. Its machines are free to be enrolled again, as a different setup. This is not H: H revokes an enrolment; N retires a host whose enrolment completed.

  1. Only a host that is already out of service. Taking it out of placement (active = false) and draining its instances (G) are the operator's decisions, made first. Retire refuses a host that is active or dedicated, has a live instance, or has a pre-booking whose window has not ended, because an inactive host can still be serving a dedicated tenant or be owed to a paid reservation. A failed create is not a live instance.
  2. Outside first, the stamp last. Every teardown call is the one the enrolment release already makes, and each is idempotent. If one fails part way, nothing is stamped, and running retire again carries on from where it stopped. retired_at on a host means everything it held is gone.
  3. The row is never deleted. vms.host_id keeps a host's instances and billing history pointing at it, and the row keeps its slug reserved. A slug that is freed while a teardown is still running is how a new host once lost its relay to the old host's release.
  4. The customer SSH wildcard is retire's own job. Relay release removes the relay's record, and revoke removes the tunnel's hostname. Only a hard delete removes *.<slug>.ssh, so a retired host that left it up would keep a name pointing at an address GCP can hand to someone else.
  5. Only a host whose enrolment completed. A host that never redeemed (no pubkey) was not provisioned by an enrolment, and one whose enrolment is still pending or ready is an enrolment to revoke (H), not a host to retire. So when retire runs, the only enrolment it closes is the enrolled one: released, keeping its host_id as the record of the setup, with every enrolment's token and WireGuard config cleared. Nothing is left open for the expiry sweep, so the sweep never tries to delete a retired host.
  6. Nothing the host was issued still works. Its token_hash and SSH-map token are cleared, so an agent left running on the box can no longer report or fetch anything. agent_token stays: it is what the backend presents to the agent, not something the host holds against us. A retired host cannot be reactivated, edited or deleted, and a check on the table holds it inactive and undedicated for every writer, so no placement path can reach it.
  7. Its machines are free, and their history stays. A node's hardware_id is unique only among nodes that are not retired. The box can join a new host at once, and the retired rows still say which host it was in, and when.
  8. Nothing on the box is touched. Retire works on backend state alone. The agent, cloudflared, WireGuard and our keys on the machines are removed by the owner or overwritten by the next enrolment. Handing a machine back is H's uninstall.
MissingPhaseWhy it matters here
A schedule for host-provision reapAThe sweep is built; nothing calls it. Until something does, every abandoned click leaks a relay VM and a reserved address
Rotating the signing keyMNo call to the agent changes the key a host trusts, so a new signing key is trusted host by hand over the tunnel, as J upgrades are. Losing the password-manager entry before that is done freezes every host's operator keys as they last were
UninstallHRevoke is active = false plus the release path, which exists. Handing the machine back does not. It is what gates ever setting hosts.lockout_owner: until it exists, a locked host is one we cannot return
PhaseStepFailureWhat happens
A5the relay will not buildThe owner is told to try later; what was allocated is released
B18no token passedNothing downloaded; points the owner at the dashboard
B22checksum missing or wrongInstalls nothing, and says so
B28token unknown, spent or expiredOne answer for all three. The owner adds the GPU again
B28hardware is not a GB10 SparkRejected, and the allocation is released
C35a step will not completeThe run stops there. Visible in /admin, and the host never reaches approval
D39the report shows a GPU that never initialisedThe operator does not approve it
E61every idle host refuses the launchThe queue row stays open and the vms row is marked creation_failed_at, so the next trigger retries the same customer
E60a host accepts but its answer cannot be read, and no id was mintedThe loop stops rather than rotating: another attempt risks a second container for one queue row. Reported, and left for the next drain to resolve by owner and name
E62the launch worked but this write is lostThe instance is live and reachable — the agent published the port — but its row is missing the host and the customer key. Logged rather than swallowed, because nothing downstream would notice
G81the agent never answersNothing is closed. The row keeps billing and keeps its host, which is honest — the capacity is still held — and delete_requested_at is what surfaces it
G83the container is gone but this write is lostThe row outlives the container: still billing, still holding its host, and invisible to every sweep. What the fleet's agent hid, and a BYO host cannot
K125node1 does not answer, or the login is wrongNo enrolment, token unspent. The agent retries every 60 s; the owner re-runs the installer to enter the login again
K127no carrier on any railRejected, token spent, the host released — a lone Spark cannot take a spark-2 allocation
K127more cables than the plan has railsRejected: the extra cable has no rail to take
K127the peer is not a GB10 Spark, or is node0 itselfRejected, and the hardware that turned up is on the enrolment, as it is for node0
K132the rails carry no traffic once configuredrail-verify fails and retries every 60 s, visible in /admin; approval never opens
L139serving fails to start on a new binaryRolled back to the previous binary, once. If that fails too, the host stays down and is alerted agent_unreachable
L145node1 cannot be reached when node0 startsnode0 serves; the heartbeat starts the worker when node1 answers (L3)
L152node1 stays unreachabledegraded, half a pair. De-pooling stays an operator's decision
L156node1 cannot be reached for a rebootnode0 reboots anyway, audited as half done; the heartbeat starts the worker when node1 answers (L3)
M167the signature fails the check, or the version is no longer the next oneNothing is written and nothing is pushed. The operator starts again
M169the host does not answerRecorded as not updated and listed in /admin with the reason. It stays out of date until an operator sends again (175–177)
M170no signing key trusted, the signature fails the check, or the version is olderRefused, and the keys the host holds stay as they are. The answer says why, and it shows in /admin
M170node1 cannot be reachednode0 is updated and the host answers with node1's version still old, so the host is listed as not updated until an operator sends again. Until then a removed operator's key still opens node1, which has no tunnel of its own: reaching it takes node0, which no longer admits them, or the owner's own network
N180the host never enrolled, its enrolment is still open, it is active or dedicated, has a live instance or a pre-booking to come, or is already retiredRefused with the reason, and nothing changes. An open enrolment is revoked instead (H)
N184the relay instance is still deletingNothing is stamped. The address stays reserved until the instance is gone, and retire is run again
N186the customer SSH wildcard cannot be deletedNothing is stamped, and retire is run again. The relay and tunnel are already gone, which is safe to repeat

Set aside, with the reason — because "why didn't you just…" gets asked twice.

  1. An acceptance probe before activation. The backend would launch a scratch VM and run nvidia-smi itself, since the GPU claim is billing-relevant.
    • The scratch VM is launched by the agent, on the host, and its output returns through the host: against a hostile host it restates the claim rather than checking it.
    • A broken host is caught by the step-37 self-check anyway.
    • Still worth having one day, and distinct from this: the backend has never used the customer's own path (F) before a customer does.
  2. A shared relay pool. Removes the slowest step from A — but nobody waits on that step, since the relay boots while the owner finds a terminal.
    • In exchange: capacity to watch, slots to reclaim, a stall when it is empty, and one fewer isolation boundary.
    • Revisit if per-host relays prove expensive at rest rather than slow to build.
  3. The box enrolling anonymously, printing a claim code. Needed a pending-and-unowned state, an expiry sweep, an alphabet a human could transcribe, and an enrolment endpoint anyone could write rows to — all to answer a question the current flow never asks, because the owner is signed in before the machine is touched.
    • Dropping it is a reversal worth naming: it had no secret in the owner's command line, where one lands in shell history and in whatever gets pasted into a support thread.
    • Provisioning first makes that trade a bad one. The token buys only an allocation that already carries the owner's name, is single-use, and dies within the hour — so what survives in history is a spent secret.
  4. The host token authorising key changes, with no signature. One call, no password manager, no manual step.
    • The token is readable from the database, used by the proxy, and answered by the agent from the internet. A call that sets the login keys would make it root on the host, and a leak of either holder root on every BYO host.
    • Signing moves that authority to a signing key no server holds (M).
  5. Signing on the server, or in the browser. admin-operator-keys signing with the signing key as a Supabase secret, or pasted into /admin, would keep the change to one click.
    • A Supabase secret is readable by every edge function in the project.
    • /admin is code Vercel serves, so a tampered page could take a pasted signing key the next time someone uses it.
    • Signing on the operator's own machine, with the key loaded from the password manager for one signature, keeps it off every server and every Enverge page.
  6. Delivering key changes by preparing again. It already exists, and operator-access already writes authorized_keys.
    • Preparing again waits for an idle host (L), so a rented host would keep a removed operator's key for as long as the rental lasts.
    • It is a loop over every host run by hand, where M is one push.
  7. Retrying a failed push from the heartbeat. It would bring an offline host up to date without anyone acting.
    • A change would reach a host at a moment nobody chose, for a reason /admin does not show.
    • An offline host cannot be logged into, so there is little to protect in the meantime: /admin lists it, and sending again is one click.
  8. One Unix account per operator, instead of a shared spark. Logins would be attributable by user name.
    • The agent manages one service user on node0 and node1. An account per operator is a home directory and a sudoers file per operator per machine, created and removed on every host with nothing to clean up after them.
    • sshd logs the fingerprint of the key it accepted, and operator_keys maps that to the operator, so logins are attributable already.

comments (0)

reviews (0)