BYO — the whole system, in sequence
Every part of a BYO host's life, from an owner adding someone else's Spark on the dashboard to the host being handed back — and what keeps it serving in between. This document is the source of truth for how the BYO system behaves: where code, specs or other docs disagree with it, they are the ones to change.
- A–H are the phases, from an owner adding a machine to the two ways it can end. In onboarding an operator appears once, to approve (D).
- I, J and M are out of band, and none of them happens on an ordinary onboarding:
- I is the operator's way in — break glass — from the moment
operator-accesscompletes. - J is how a host gets a newer agent than the one it was installed with.
- M is how a change to who the operators are reaches every host, rented ones included.
- N is how a host whose enrolment completed ends for good, so its machines can be enrolled into another host.
- I is the operator's way in — break glass — from the moment
- K is a variant, not a phase: what A–D become when the host is more than one machine — K1 the pair (
spark-2), K2 the four-Spark switchless ring (spark-4). - L is the host after go-live: a new binary, either machine restarting, and a reboot of both.
Six areas, one continuous numbering across the whole document, so step 44 means one thing and E means one phase.
The Python agent in
platform/agent/, and the hosts that run it, are out of scope. The BYO agent isplatform/byo-agent/.
Everything is provisioned before the box is ever touched. Clicking "add a
GPU" allocates the relay slot, the address, the tunnel and the hosts row — not
the box enrolling, and not an operator afterwards. By the time the owner runs
the one-liner, the only thing missing is the machine itself.
That ordering is what makes the rest fall out:
- The agent enrols once and keeps going. Its single call returns every credential, so it goes straight into preparing the host. No second start, no waiting, nothing on the box watching for anything.
- The operator appears once, at the end, to approve a host they can read
the whole preparation of — and that approval is the only thing between an
enrolled machine and a paying customer.
- An earlier draft had an operator provisioning each host by hand between enrolment and preparation. That put a person in the middle of every onboarding, which is precisely what BYO is supposed to remove.
- Each host gets its own relay, created for it here. Same topology as the
fleet has today — one relay per host — except that building it is an API call
rather than a person following new_spark.md §1–§2.
- The instance configures itself from
relay/gcp/bootstrap.shat first boot, so nobody SSHes into a relay to bring it up. - It boots while the owner is still reading the one-liner, so its start-up costs nothing anybody waits on.
- The instance configures itself from
Everything belongs to one of six places. The diagrams group them, because which side of a boundary a step happens on is usually the point — a credential minted in an edge function and a credential minted on the host are very different facts.
Who belongs to each area — every actor's ID and label, and the verbs they use — is in the BYO glossary, and only there.
1. The owner's Spark host — the machine being onboarded, and the programs on it. The agent is two of them, because preparing needs root and serving must not have it.
2. app.enverge.ai — the Vercel deployment: the owner's and operator's
pages, the static installer assets, and the proxy that redirects to Blob.
3. Supabase edge functions — everything with authority. Nothing here runs on a host.
4. The Supabase database — host_enrollments, host_profiles, hosts,
and vms from E onwards. Written by edge functions under the service role;
read by the pages under RLS. There is no relays table: a relay is named by its
host's slug, so the hosts row is its record, exactly as it is for the
hand-built fleet relays.
5. Relays — one per host, created during provisioning. Includes the GCP API
that builds the instance, reserves the address and points it at the relay. Same
topology and naming as the fleet runs today — <slug>-relay and
<slug>-relay-ip; the difference is that nobody builds it by hand.
6. Cloudflare — the operator tunnel and DNS.
And three people, none of whom belong to an area: the owner, the operator and the customer.
The customer never learns any of this. They land on a spark-1 host and cannot
tell whose hardware it is, which is the tenancy decision spec63 makes
explicitly rather than by omission.
- Nothing here needs the machine. BYO accepts one hardware configuration, so the relay slot, the address and the sizing are known before a box exists. The box contributes one thing: proof it really is that configuration, checked at step 28.
- Step 11 writes
active = false, and nothing writestrueexcept the operator at the end. - Relay provisioning is its own edge function. It is the only part of this
that talks to two systems we do not otherwise touch from here — the GCP API,
and the relay itself.
host-provisionstays an orchestrator, and replacing a dead relay later does not have to go through host onboarding to get one.- It is the slowest thing in the flow and the only one nobody waits on: the relay finishes booting while the owner is still finding a terminal.
- Provisioning ahead of time costs something. An owner who clicks and never
runs the command leaves a relay, an IPv4, a tunnel and an inactive
hostsrow allocated. Four things contain it:- only an owner granted
community_hostfrom the/adminusers tab (or a superadmin) can click at all, - a cap of three outstanding allocations per owner,
- a token that expires within the hour,
host-provision'sreap, which hands back everything a lapsed or rejected enrolment was holding. Nothing callsreapon a schedule yet.
- only an owner granted
- This has been run against real GCP, not just written: address reserved,
instance created, metadata read, bundle fetched,
wg0up with the host's peer,ssh_mapdon10.8.0.1:8754, Caddy on:80/:443, and theLOFISPARK_VMSSHchain seeded with the IAP exemption and a blackhole default.- It found one bug:
accessConfigs[].natIPtakes the address literal, not a self-link, so every create was failing 400. - Still unproven is time-to-certificate, which needs the Cloudflare record and so a deploy.
- It found one bug:
As built: the agent's side of enrolment, generated from its tests. Scenario 1 enrols, 2–4 are refused, and 5 finds the host already enrolled.
- Step 28 is the single-use point. The row is locked, so two machines
running the same one-liner cannot both become this host.
- Unknown, spent and expired are one answer: which of the three it was is not something a caller has earned.
- Hardware that is not a GB10 Spark is rejected here, and the allocation it was holding is released.
- Step 29 is the whole payoff of provisioning first. One call returns everything, so the agent has no reason to stop and nothing to come back for.
- The operator keys are read at redeem, not at provisioning. Step 29 returns
the signed list as it stands when the box enrols, and the public key its
signature is checked against (M).
- Enrolment is only the first delivery. Every change after it arrives through M.
The agent runs the whole of new_spark.md §4 and §6 as a state machine. Every step has the same shape, and the shape is the point:
As built: the prepare run, generated from the agent's tests. Scenarios 1–3 take the runner down each of its paths, and 4 is the real runbook.
- Check and apply are separate, for two reasons:
- a re-run changes nothing on an already-prepared host, which is the property the whole machine rests on,
- a dry run calls only
check, so it reports what it would change while touching nothing.
- The re-check after apply is not belt and braces. A step that reports success without its check passing is how a host arrives "prepared" and broken, and nothing later would notice.
- It stops at the first step that will not complete. The orderings below
are load-bearing, so a host that failed
lxd-installmust never reachowner-lockout.
#The 19 steps
| # | Step | Notes |
|---|---|---|
| 1 | preflight | Linux, Ubuntu, uname -m and nvidia-smi match the allocated profile. Fail closed |
| 2 | docker-purge | purge dockerd, iptables -P FORWARD ACCEPT |
| 3 | swap-off | swapoff -a, comment /etc/fstab, mask swap.img.swap |
| 4 | earlyoom | the profile's floor, verbatim (spec28, retuned spec51) |
| 5 | zfs-arc-cap | /etc/modprobe.d/zfs.conf, 4 GiB |
| 6 | headless | multi-user.target, mask gdm |
| 7 | watchdog | RuntimeWatchdogSec + the boot-reset reporter (spec46); skipped when the profile says watchdog_sec: 0 (a guest with no watchdog device), fails closed when it wants one and /dev/watchdog0 is missing |
| 8 | operator-access | spark user, every operator's key, NOPASSWD sudo, the tunnel minted back in A. authorized_keys is written as the whole list, so it holds exactly the operators and nobody else |
| 9 | lxd-install | apt lxd; do not rely on lxd init --storage-backend |
| 10 | lxd-network | lxc network create lxdbr0; lxd init is not run, so nothing else makes it |
| 11 | firewall | if ufw is active: allow in + route on lxdbr0, allow in on wg0, DEFAULT_FORWARD_POLICY=ACCEPT, reload. No ufw or inactive = satisfied |
| 12 | storage-pool | explicit lxc storage create default zfs |
| 13 | nvidia-runtime | nvidia-container-toolkit-base + libnvidia-container-tools, for nvidia.runtime injection |
| 14 | gpu-accounting | NVML per-process accounting. Driver state, so re-asserted |
| 15 | profiles | lofispark-base + lofispark-gpu |
| 16 | base-image | import the prebuilt image, else build it |
| 17 | relay-peer | apply the WireGuard peer collected in B; collision-check first |
| 18 | owner-lockout | gated on hosts.lockout_owner, ships false |
| 19 | self-check | machine-check every verification, and report |
Orderings that are landmines, not preferences:
Named rather than numbered, because a step inserted in the middle renumbers every landmine below it and the numbers are what a reader trusts.
operator-accessbeforeowner-lockout, and verified. Prove Enverge's access before removing the owner's. The reverse strands a box neither party can administer.docker-purgebeforebase-image. Docker'sFORWARDDROP breakslxdbr0NAT, so the image build loses egress to the NVIDIA repo.zfs-arc-capbeforelxd-install.modprobe.dbinds at module load, and installing LXD loadszfseven for a non-zfs pool.swap-offis what armsearlyoom. It fires only whenMemAvailableand free swap are low — an AND. A GPU runaway is unswappable, so leaving swap on disarms the floor entirely.lxd-networkbeforeprofiles.lofispark-baseattaches eth0 tolxdbr0, andlxd init— which would have made both the pool and the bridge — is deliberately not run. Creating the pool explicitly without the bridge gives a host with LXD up, a healthy pool, and no way to attach a NIC; the failure then surfaces atprofiles, several steps after the omission.firewallbeforebase-imageandrelay-peer. A default-denyufw(cloud images and DGX OS alike) drops the bridge's DHCP, forwarded container egress and the relay's packets, and none of the three failures names the firewall: the image build dies resolvingarchive.ubuntu.comand the control plane just hangs.DEFAULT_FORWARD_POLICYhas to change in the file —docker-purge'siptables -P FORWARD ACCEPTlasts only until the nextufw reload.base-imagebuilds from a script the agent carries. On the fleet it arrives withmake sparkN-copy; a BYO host has nobody to run that, so the builder and its login banner are embedded in the binary. It is run with the profile'scuda_repo_arch— the script defaults tosbsa, which is arm64's NVIDIA repo and wrong everywhere else.storage-poolmust create the pool explicitly. Adirpool costs ~69 s on every VM create against 0.6 s on zfs, on a host that looks perfectly healthy.base-imagemust end with the alias. The lifecycle launches by alias; an image without one fails every create.relay-peermust check for a subnet collision first.10.8.0.0/24is a common home-router default, and bringingwg0up into a colliding range makes the host unreachable the instant it succeeds.
- This is the only human decision in the flow, and the only thing between
an enrolled machine and a paying customer on the shared
spark-1tier. gpu_typeis not decided here. It was fixed when the allocation was made — BYO accepts one hardware configuration, and the box proved it was that at enrolment. Activation is the single flag, plus the placement chosen with it.- The backend refuses to activate a host whose preparation did not finish.
- The rule used to live in the render condition of an "approve" button on the
enrolments table, so the hosts table — a second control issuing the same
update — walked straight past it: a host that stopped at
storage-poolcould be activated by anyone who used the other one. - A gate only one of two paths honours is not a gate, so it moved into
admin-host-actionand the duplicate button is gone.
- The rule used to live in the render condition of an "approve" button on the
enrolments table, so the hosts table — a second control issuing the same
update — walked straight past it: a host that stopped at
Five callers, three shapes. What separates them is how the host is chosen, because that decides when the row can be written and where the SSH hostname comes from.
| # | Trigger | Who calls | Drawn in |
|---|---|---|---|
| 1 | the customer asks | the proxy (PXY) | E1 — the picker chooses, and the customer is waiting |
| 2 | capacity frees and the queue drains | drain-waitlist | E2 — no host until one accepts |
| 3 | a paid pre-book window opens | drain-prebook | E3 — the booking pinned its host when it was made |
| 4 | an operator places one on a dedicated host | admin-vm-action | E2's shape, host supplied instead of rotated |
| 5 | an operator launches a booking by hand | admin-prebook-action | E3's shape, one booking by id |
The two operator paths are the manual twins of the two automatic ones — same steps, with the cron's gate replaced by a superadmin and the selection replaced by whatever the operator named. That is why they are not drawn separately:
- 4 is 2 done by hand. A dedicated owner often lands in the general queue, because the picker cannot reach their box; the operator places the VM there and closes the queued row the drain would have closed.
- 5 is 3 done by hand, targeted by booking id, and with a
forcethat revives a booking torn down inside its own window.
The row is a precondition for dispatch, in all five. On a host whose agent
writes none, a create sent without one produces a container the database has no
record of: unbilled, invisible to the customer, and holding capacity the picker
cannot see. Every caller now refuses rather than dispatches, keyed on the same
pubkey marker — a fleet host is exempt, because its agent inserts the row a
moment later and refusing there would break a path that works.
To the customer every one of these is identical to any other host, which is the
point of BYO joining the spark-1 pool rather than a tier of its own. They are
not identical in who writes the row: see G, where that difference is the
whole story.
#E1. From the dashboard (41–52)
A customer's create, once the host is active.
As built: the agent's side of a create, generated from its tests. Scenarios 1–2 create an instance, and 3–5 refuse and revert.
- Step 44 stamps
created_atbefore the host is contacted, so billing runs from when capacity was committed rather than from whenever a host answered. A create that is never answered is markedcreation_failed_atand bills nothing. - Step 49 is the agent programming the relay, and it had to be. An earlier
draft had the backend do it, to keep the last secret off the host. It cannot:
ssh_mapdlistens on10.8.0.1:8754, inside the tunnel, and Caddy proxies only:8080— there is no route from Vercel. Publishing a second endpoint on every relay to create one would be more attack surface than the thing it protects. The token is the relay's own, minted per relay, and on a per-host relay it can forward one host's:22to one port on that same host — which the agent decides anyway, being what publishes the port. It grants nothing a compromised host does not already have. A shared relay would break that, since one peer could forward another's:22elsewhere, so a pool must move this back to the backend first. - It is fatal, unlike the GPU check. An instance whose
:22goes nowhere is not an instance. The fleet's Python agent logs a map failure and carries on, which is how a host ends up healthy with a blackholed:22and nobody noticing. - Steps 47–48 split fatal from advisory. An instance the customer cannot reach is not an instance; a GPU check or a startup script that fails is reported, and the instance is created anyway.
#E2. From the queue, when capacity frees up (53–64)
The queue is the only path that creates a VM with nobody present, acting on an intent that may be days old — which is why billing is re-checked live rather than trusted, and why the host is not known when the row is written.
- Step 58 inserts the row with no host, and that is the point.
pick_hosts_for_gpuexcludes any host holding a livevmsrow, so ahost_idwritten before a launch succeeded would make this row occupy the candidate that just refused it — benching a working host against every concurrent drain until the row is marked failed. - Step 59 mints per attempt, not once. A name under the previous candidate's zone resolves to that host's relay, and the customer would land nowhere. It is minted in the function rather than by the database: a random label joined to a column, where TypeScript is type-checked, deployed by CLI and unit-tested.
- Step 62 is the first write, and it is the whole difference from the
dashboard path. Every other create knows its host at insert, so
create_vmstamps the host and the name in one statement. This one cannot, so it stamps them once a host has actually taken the VM. - The name is written only if one was minted. A host whose own agent mints hostnames wrote that column back at step 60, and an empty string over it would take the address away from an instance already running.
- A candidate has three outcomes, not two.
- Refused — benched, and the loop moves on. This is what stops one drifted host, agent-busy while the database read it as idle, stalling the whole queue.
- Skipped — step 58 could not write the row and this host's agent writes none, so a container here would be one nothing can account for. Not benched: the host did nothing wrong, and the fleet hosts behind it are still tried. A fleet host is never skipped, because its agent inserts the row itself.
- Accepted — including an answer that will not parse. A 2xx means the container exists, and the id, name and hostname were all sent, so the values are known without reading them back.
- An unreadable 2xx must not be scored as a refusal. It was: the parse sat inside the catch that scores attempts, so a successful launch benched a working host and sent the next candidate to build a second container for the same queue row. The guard that would normally catch that reads a row the agent writes, so on a BYO host it finds nothing and the loop carries on.
#E3. From a paid pre-booking, when the window opens (65–74)
A booking buys a named box for a date range. The window opening is the trigger, so nobody is present — and the host was decided when the booking was made, not when it launches.
- The gate is
paid_atalone, and no live Stripe check. Payment is collected off-band and recorded by ops, which is the deliberate difference from E2 — and the reason a stalepaid_atis a known risk rather than a caught one. - The host is pinned on the booking, not picked. A booking without a pin is
a legacy row and falls back to the pool pick.
- An unavailable pinned host leaves the booking open for the next tick and an operator to reassign. A paid window is not re-placed on a different box just because the intended one is offline.
- Step 67 is a self-heal, not an idempotence check. VMs are soft-deleted,
so
launched_vm_idnever nulls — without reading the row's liveness, a customer who deleted their box mid-window would be stranded untilend_date. A non-live row falls through and the booking relaunches. - The booking stays open after launching.
expire-prebookcloses it atend_dateand tears the VM down, which is trigger 5 of G. - The customer key is the booker's own, resolved from
ssh_keysbyuser_subrather than a per-booking id, so it can only ever be a key they hold. No key on the account leaves the booking open for an operator to follow up.
No proxy, no agent, no HTTP. L4 forwarding only.
- Standard port 22, no client software. The whole reason customer SSH was never moved onto the operator tunnel.
- Until the map has been written,
:22is a blackhole. Not an error — nothing is listening for that address yet. The agent:- writes it on create,
- clears it on delete and stop,
- re-asserts it when it starts, because a relay that restarted came back with an empty chain, and a host that rebooted has a container up the relay knows nothing about.
- The owner's box needs no inbound anything. It dialled out; every address here is ours.
Six callers, one shape. They differ only in what decided the instance should go. Three of them (3–5) run on a schedule with nobody present, which is why the shape has to be right without anyone reading the result.
| # | Trigger | Who calls |
|---|---|---|
| 1 | the customer asks | DELETE /api/vms/{id} through the proxy (PXY) |
| 2 | an operator, from /admin | admin-vm-action |
| 3 | a payment probe fails mid-dunning | probe-payment |
| 4 | dunning runs out | teardown-vms, the gated reaper |
| 5 | a pre-book window ends | expire-prebook |
| 6 | an operator ends a pre-booking early | admin-prebook-action |
As built: the agent's side of a delete, generated from its tests. Scenario 6, and 7 for an instance that is already gone.
- Two facts, deliberately separate.
delete_requested_at— the backend accepted a delete. Stops nothing.deleted_at— the container is confirmed gone. Billing stops here.
- The row is not closed optimistically, because the database decides
routing.
pick_hosts_for_gpuexcludes any host holding a livevmsrow, so a row closed before its container dies leaves a host that looks free: launches keep routing to an agent that refuses them, and every user behind it stalls.- So the bias runs the other way. A delete that does not complete keeps billing and keeps its host marked occupied, which is the honest state — the capacity really is still held.
- Over-billing is refundable; lost utilization is not, and a poisoned queue costs every waiting customer rather than one.
- That trade is only defensible because the stuck case is visible. Step 80
is what
notify-stuck-deletesscans for, and it also sizes the refund: the customer gave the instance up when they asked, not when an operator noticed. - A delete ends when its caller stops waiting. The backend waits 90s, well past the ~30s a pair takes; cut off earlier, the agent leaves the instance gone but node1's rails bare and the relay still forwarding SSH to the dead port. A relay write abandoned that way is not recorded as a relay failure. Either way, every heartbeat reads the relay's SSH map (
GET /map) and writes it again if it does not match what the host runs, sorelayhealth is the relay as it is this minute, and a stale or cleared target is put right within a beat. - A 404 confirms, it does not fail. The host has no such container — the orphan case, where it went away and the agent's own write never landed. The caller asked for it gone and it is gone, so the row closes rather than becoming one nothing can ever clear.
- Step 83 must precede step 84.
pick_hosts_for_gpuexcludes a host holding a live row, so draining before the row closes hides the host that just freed up from the customer waiting for it. - Every caller owns both edges, and none may delegate them.
- The fleet's Python agent PATCHes
deleted_atfrom the host, so a caller that wrote nothing still looked correct there. - A BYO host's Go agent writes nothing to the database at all, so the same caller leaves the row live forever — billing, holding its host, and with step 80 unwritten, not even stuck. Just invisible.
admin-vm-actionwas the last caller doing this. It is the same asymmetry that produced the emptyssh_hostnameand the orphaned container: backend code assuming the agent is a database writer.
- The fleet's Python agent PATCHes
- Both writes are guarded —
delete_requested_at=is.nullanddeleted_at=is.null— so on a host whose agent does write, the two compose: whichever lands first wins and the other is a no-op. Two independent paths to the same value beat either alone.
Two different endings. Revoke cuts a host off and keeps it configured; uninstall hands the machine back.
- Order is the whole design. De-pool first, drain second (G), and only then remove access — the reverse strands a tenant on a host nobody can reach.
- The host confirms before backend state goes. Once the relay is gone there is no path to that machine, so anything still needed from it must be asked for first.
- Restore the owner before removing ourselves. The mirror of prepare's
operator-accessbeforeowner-lockout, for the same reason. - Revoke leaves the machine configured and locked. That is why it is not offboarding: an owner who simply wants to stop hosting needs uninstall, and "reinstall your machine" is a poor answer for a self-serve product.
Everything above is the platform driving a machine nobody is sitting in front of. This is the other thing: a person, on a host, because something is wrong.
- It is deliberately never on the happy path.
spec63 splits knowing from acting: every
diagnostic is what the agent reported and what
/adminrenders, because "SSH in and look" does not survive a thousand strangers' boxes. - Acting on one is what this is for — push a binary, restart a unit, unwedge a prepare that stopped.
- There are two ways in, and the second exists because the first can fail.
#Why there are two
- The tunnel is the one that scales, and the one that can vanish.
- It rides
cloudflared, a unit theoperator-accessstep installed with a connector token minted backend-side. - It needs nothing inbound on the owner's network, which is why it works behind a home router.
- It needs Cloudflare, the host's outbound internet, and that unit alive. Lose any of the three and the door is shut.
- It rides
- The relay is the one we own.
- A GCP VM on our project, so IAP reaches it whatever the host is doing, and
WireGuard puts it one hop from
10.8.0.2. - That hop matters more than it looks: the relay DNATs inbound public
:22into the tenant's container throughLOFISPARK_VMSSH, but that chain is onPREROUTING, so a connection originating on the relay is not caught by it —ssh [email protected]from the relay lands on the host's sshd, not the customer's container.
- A GCP VM on our project, so IAP reaches it whatever the host is doing, and
WireGuard puts it one hop from
- The two fail independently. The tunnel dies with Cloudflare, the host's uplink, or one systemd unit; the relay path dies only if WireGuard is down — which is the same as the host being off the network entirely, at which point there is nothing to reach.
#What the operator gets, and why that is the posture
spark, with their own operator key and NOPASSWD sudo — so, root. Not a concession: it is
sole administrative control, the same asymmetry that lets owner-lockout remove
the owner's login and never ours. A BYO host has exactly one administrator and
it is us.
- One account, one operator key per operator. Every operator logs in as
spark; what tells them apart is the operator key. sshd logs the fingerprint of the key it accepted, andoperator_keysrecords each key's fingerprint against the operator's user sub, so a login is attributable without a second account. - Leaving is one action. Removing an operator's key reaches every host through M, rented or not, so no host keeps a door open for someone who has gone.
#What has to be true first
| Depends on | Why it can be missing |
|---|---|
operator-access completed and verified | It is verified precisely because owner-lockout runs after it. An unverified tunnel plus a completed lockout is a box neither party can administer |
| The host has outbound internet | Both paths are outbound-dialled. A host off the network is unreachable by design, and that is what "we detect rather than prevent" means |
| The operator's key is on the host | Delivered at enrolment, then by M. An operator key added while the host was offline arrives when an operator sends again |
relay-peer applied, for path 2 | Before it, the only way in is the tunnel |
| The relay still exists | Releasing an allocation destroys it. Order matters in H: drain the tenant and take what you need from the host before the relay goes, because after that there is no path to that machine at all |
#When neither works
- Nothing. That is the honest answer, and it is the design: a host that has been unplugged, reinstalled or taken back looks exactly like one that is merely offline, so we trust the absence of a heartbeat rather than the presence of a good signal. The response is to de-pool and revoke, not to try harder to get in.
- There is a third path while
hosts.lockout_ownerships false — the owner's own login. That is why it ships false: until uninstall exists, their login is the recovery path if this flow wedges a box.
- Publishing is not delivery. CI builds and publishes on every merge, and that changes nothing on any host.
- Nothing on a BYO host ever fetches a binary — no timer, no polling, no update endpoint. The agent's steady state is to answer the backend and otherwise sit still.
- The one thing that moves unasked is
enverge-agent-rollback, and it moves backwards — once, when serving will not stay up: it consumes.previousand will not restore the build that is already running (L1).
So an upgrade is an operator action, over the same tunnel as I.
-
GitHub Actions belongs to none of the six areas. It publishes, and publishing is where its involvement ends — which is the point of the section.
-
The release verifies the path an owner's installer follows, not the upload.
put()returning a URL says the object exists; it says nothing about whetherapp.enverge.ai/agent/…resolves.- That gap once let
AGENT_BLOB_BASE_URLsit blank in production — present in the dashboard, so it read as configured — while every install one-liner was broken.
- That gap once let
-
No token, and that matters more than it sounds. The installer asks for one only when the machine is not enrolled; an enrolled host reuses its credentials.
- The alternative would be an operator obtaining a fresh token, and the only way to get one is to add a GPU — which provisions a second host, relay and tunnel.
-
Detached, deliberately. A dropped SSH session must not land between writing the binary and writing the units.
-
Preparation starting serving is one line, and it went missing.
- The ordering between the two units is declared on the serving one —
Requires=/After=the prepare unit — and that direction only pulls preparation in when serving starts, never the reverse. - So the installer's "restart prepare" started nothing else: spark-c05 prepared itself, exited 0, and served nothing until it was started by hand.
- The prepare unit's
Wants=enverge-agent.servicecloses it. Earlier hosts hid it because they had been rebooted since installing.
- The ordering between the two units is declared on the serving one —
-
Restarting serving is the installer's job, not the
Wants='s. On an upgrade the serving unit is already running, soWants=is satisfied and systemd starts nothing — the host would keep answering from the binary the install just replaced.- An upgrade restarts serving; the preparation it pulls in exits at once.
- A first install starts preparation only; its
Wants=starts serving once credentials exist.
-
An upgrade does not re-check the host. Preparation does not run after go-live (L).
- The done list is still scoped to the build that wrote it. Steps are recorded by name, and a name says nothing about what the step now does:
lxd-installgainedusermod -aG lxdwhile keeping its name, so hosts that had completed it skipped the new work and reported "already satisfied". - That is what makes the deliberate run in L re-check every step on a newer binary, rather than trust names an older one wrote.
- The done list is still scoped to the build that wrote it. Steps are recorded by name, and a name says nothing about what the step now does:
-
An upgrade does not reach containers that already exist. The login banner is pushed at create, so a running instance keeps what it was given. Recreating it is the only way to change what a customer sees.
-
One command covers every BYO host, unlike the per-host targets the fleet carries.
- Those exist because the fleet's agent is copied out of a working tree; a BYO host fetches the published binary itself, so the only per-host input is the slug.
- The corollary: upgrading several hosts is a loop an operator runs by hand — fine at two, and the first thing to hurt as the fleet grows.
One hosts row, one relay, one tier slot, one enrolment, one token, one
command on node0, one prepare run, one heartbeat. Agentless satellites are
reached from node0 — nothing is ever run on them by anyone. One table,
host_nodes, says which machines a host is made of. K1 is two machines;
K2 is four on a switchless ring. Everything from E onwards is
unchanged either way.
#K1. Two nodes — spark-2
A spark-2 host is two Sparks cabled together over ConnectX-7 and sold as one
unit. The second machine is reached from the first.
- This is new_spark_dual.md with the operator
removed. The fleet's pairs are node0, a full platform host, and node1, a
machine with no agent, which a person prepares over SSH and node0's agent drives
over SSH at create and delete. Node1 keeps that role exactly — no agent,
no relay, no tunnel, no row — and the person's part moves into node0's
prepare.
- NVIDIA's own stacked Sparks guide is the same work with a person on each box: matching users, passwordless SSH both ways, static addresses on the QSFP interfaces. Steps 125 and 129–131 are that guide, run from one side.
- An earlier draft enrolled node1 in its own right, so that each box proved itself. It cost a second profile that was not a hardware profile, a second prepare track folded into one column, a heartbeat with nowhere to land, and a second connector on node0's tunnel. And it bought nothing: the lifecycle that follows needs node0 driving node1 over SSH anyway, and "this is a pair" is rail traffic between two GB10s whichever box reports it.
- Step 123 is the one thing that has to be asked for. Node0 needs root
on a box nobody has touched, and the login is the owner's to give. It is
collected on node0's terminal rather than in the dashboard because that
way it never reaches the backend: root-only on node0, read at steps 125
and 129, then deleted. From then on node1 admits node0's peer key and
the operator keys.
peer-prepthen activates node1's ConnectX-7 ports and pins link-local eui64 before prepare starts, because NetworkManager often leaves them disconnected while the kernel link is up — nofe80::, and the cable probe at 125 would see silence forever. It asks the cable first. A rail is point-to-point, so whatever answers the multicast on it is node1 by construction; node0 logs in, groups the two addresses a QSFP shows into the one machine behind them, and prepares it at its management address — reconfiguring node1's rails through the rail being spoken over is a way to cut the session. The management LAN (mDNS_ssh._tcpand ARP on the uplink — not Tailscale) is the fallback, for exactly the case above: a node1 whose rails cannot yet answer. A rack of eleven Sparks on one LAN, five accepting the same login, is what that ordering is for — the scan can only report the ambiguity, and the cable has already answered.ENVERGE_PEER_ADDRnames node1 outright and skips both. - Step 125 finds node1 by the cable. A rail is a point-to-point link, so
a multicast ping to the all-nodes address on each RDMA port returns
node1's link-local addresses before any static address exists. The probe
ignores node0's own addresses, so a cable looped between its two ports
finds nothing. A ConnectX-7 QSFP is two PCIe functions on one wire, so one
cable shows two of node1's addresses; the probe logs in and counts
machines, not addresses. Login tries the peer user's existing SSH private
keys, then the password from the install prompt. A port whose neighbours
are all one other Spark is a link; a port with several machines is the
owner's LAN — a Spark's uplink can sit in a QSFP port — and is left alone.
When every address on a port refuses the login, that is named as a login
failure, not as a LAN. The links are numbered as the rails of the plan
exactly once — here, the first time node0 looks — and the numbering is
recorded on node0 on the spot, so no later run re-derives it from the
order ports come in. It then asks which of node1's ports holds each
address it saw. That is how a crossed cable becomes a recorded fact on
both rows rather than a failed verification.
- Link-local persistence is required and automatic — not a manual
step. DGX OS often leaves ConnectX-7 ports at
addr_gen_mode=none, so they have carrier but nofe80::and the multicast probe sees silence; NetworkManager's default also rewrites a live sysctl on reconnect. Before probing, the agent pins each of node0's RDMA ports' NM connection toipv6.method=link-localandaddr-gen-mode=eui64, forces the live sysctl, and bounces the iface if an address is still missing. After a successful login it does the same on node1's rails over sudo. Later,rail-network(130) keeps eui64 on the planned rails via netplan. There is no owner/operator runbook for this. - The enrolment has not been accepted yet, so 125 still must not prepare the box or install peer-access. The only allowed pre-enrol write is that NM link-local pin (and the matching live sysctl), so the cable stays reachable; a rejected peer is otherwise left as found.
- If node1 cannot be reached or the login is wrong, the agent does not enrol. It retries in-process every minute while the token lives, re-reading the login, so the token stays unspent.
- Link-local persistence is required and automatic — not a manual
step. DGX OS often leaves ConnectX-7 ports at
- Step 127 is the one place a lone Spark is turned away, and it is early
on purpose. Redeem is where turning away is cheap: the token is spent,
the row parks as
rejected,reapreleases the relay, the address, the tunnel and the host row, and the owner is still at the terminal to see it. Anything found later can only stop a prepare run and leave the allocation held. So every piece of evidence that can exist before enrolment is judged here, once: at least one cable and no more than the plan has rails for, one machine on every cable, a peer that is the profile's hardware and not node0 itself, and every machine carrying its ownnode_index— 0 for the enroller, each other one once. Rows are written at the index a machine carries, never at its place in the list.nvidia-smion either box reportsGB10 / 1 / aarch64, byte-identical to a single, which is why the dropdown alone was never evidence.- What was found becomes rows: one
host_nodesrow per machine, node0 at index 0 and the peer at index 1, each with its own fingerprint — hardware identity — and its ownrails, the observed state of each port, with the rail each cable carries and the plan's address for that node on it stamped in. Every BYO host gets them, so a single Spark is one row and the node count is a count.unique (hardware_id)is the "not node0 itself" rule as a constraint, and it also stops one box being claimed by two hosts. It keys on the GPUs rather than/etc/machine-idbecause that file clones with the image it was captured in;machine_idstays on the row as reported hardware. An agent older than the field sends none and itsmachine_idis written there instead, which is the namespace its own row was already in. - The profile is the plan and the node row is the fact.
constants.railssays a subnet and an MTU per rail, and a node's address is the subnet's tenth host plus its index;host_nodes.railssays what these two machines look like, port by port, and is refreshed at 130–133 as they change.
- What was found becomes rows: one
- Step 132 is what redeem cannot know. Whether the pair works needs both machines prepared and addressed, so the traffic test is late by nature. It proves the same claim step 127 accepted, with the rails configured. Both are reported by the owner's own hardware, so neither is more trustworthy than the other — what differs is the cost of being wrong, which is lowest at 127 and highest here. Approval (D) follows it and stays the last thing in the flow.
- Step 129 decides nothing.
peer-accesscreates thesparkuser, installs the peer key and the operator keys, and deletes the login. Its check still refuses an SSH host key that is not the pinned one. After go-live it runs only in the deliberate run (L). - Steps 130–131 are C, run on node1 from node0. Every step already runs
through an
Exec; a remote one over the rail runs the same check, apply and re-check on node1 and reports each withindex: 1, whichhost-reportwrites to node1's ownhost_nodesrow rather than to the host'sprepare_state. The block is node1's subset of the list in the same order, so every landmine ordering holds there too, and the deliberate run in L re-checks node1 as it re-checks node0. Node0's track keeps its shape and its meaning; node0'sself-check, after the block, asserts node1, which is why steps 39–40 and both activation gates are untouched. - Step 133 does not watch node1 over the rails. The heartbeat refreshes node0's own
host_nodes.rails— per port, how many machines answered — which reads zero on both ports while an instance holds the rails, so it cannot tell a dead node1 from a rented pair.- That is why the node1 probe was dropped: it reported every rented pair
degradedwhile every subsystem was green. - There is no
host_statusrow for node1, andhost_nodes.seen_atis no longer written. - node1's worker is checked over the management network: L3.
- That is why the node1 probe was dropped: it reported every rented pair
- Break glass (I) reaches node1 through node0. Node1's
sparkuser holds the operators' keys, so the tunnel or the relay lands on node0 and one more hop lands on node1. Node1 has no tunnel of its own. - A single Spark walks the same steps and none of this fires. Its
install command carries no pair marker, so 123–124 never happen; with no
login file, 125 never happens; its profile has no
node_countabove one, so 127 judges nothing and 129–132 are skipped, reported as skipped the waywatchdogis on a board without one. The rails still appear in its fingerprint, which is a few more fields and nothing else.- A single Spark presented as a pair with its own login typed in as
node1's is caught twice: the probe ignores node0's own addresses, and
redeem refuses a peer whose identity is node0's. That identity is
hardware_id— the machine's GPU UUIDs, digested — not/etc/machine-id: systemd writes that file once, and only when it is empty, so an OS image captured with it populated hands every unit the same value and two genuinely different Sparks present as one machine that found itself. The agent makes the same comparison before it enrols, so a real self-pair costs a retry rather than the token. - A single Spark whose rails have carrier because they are on the owner's switch is fine against the single profile, where nothing checks rails, and fails the one-neighbour rule against the pair profile, which is the intended answer.
- A single Spark presented as a pair with its own login typed in as
node1's is caught twice: the probe ignores node0's own addresses, and
redeem refuses a peer whose identity is node0's. That identity is
- The customer's instance is two containers: the head on node0 and a worker on node1, named
<head>-w1. node0's agent creates and deletes them together over the management network — not over the rail SSH step 129 established, because the dual profile moves the rail netdevs into the containers the moment one starts.- At create, any worker already on node1 is removed first, the head mints a per-instance launcher key, and the worker trusts that key alone. A worker that cannot start fails the create: one GPU sold as two is not an instance. A create refused after the worker started — the customer key or the relay failing — reverts the worker along with the head, and restores both machines' rail addressing.
- At delete, the worker goes with the head, and both machines' rail addressing is restored.
- Keeping the two running together — through start and stop, and after either machine restarts — is L.
#What K writes, and where
One new table, host_nodes, one row per machine of a host. Everything else
is a table that existed, written at the steps it was always written at, with
the meaning it always had. The database is written at 120, 127, 131 and 133,
plus the approval at 134.
| Step | Reads | Writes |
|---|---|---|
| 118 | host_profiles.staff_only, to decide whether the picker shows the row | — |
| 119 | staff_only for the gate, constants.node_count for the command | — |
| 120 | — | hosts row; host_enrollments row with host_id, profile_id, token_hash, wg_conf. Same as a single |
| 121 | — | Nothing stored. The response carries node_count and --peer --nodes=N on the command (N from the profile) |
| 123–124 | — | /etc/enverge/peer-login and peer-nodes on node0; peer-prep finds node1 on the rail (falling back to the management LAN only when the rail is silent) and activates its RDMA NM connections + link-local eui64 (writes on node1 before enrol, so the cable is visible) |
| 125 | peer-login; node0's RDMA NM connections and sysctls; node1's nvidia-smi (including GPU uuids), uname, /etc/os-release, /etc/machine-id, sysfs ports, and which port holds each address | On node0: NM ipv6.method=link-local + addr-gen-mode=eui64 on every RDMA port (and live addr_gen_mode=0); peer.json with the links the moment their numbering is assigned. On node1, after a successful login only: the same NM link-local pin on its rails (via sudo). No prepare, no spark user, no keys. In memory the fingerprint gains node_index 0, machine_id, hardware_id, rails[].neighbours_discovered, rails[].index, and peers[] each with its own node_index |
| 126 | — | Nothing yet. The request body is the fingerprint from 125 |
| 127 | host_enrollments by token hash for profile_id; host_profiles for gpu_model, arch, gpu_count, constants.node_count, constants.rails | Refused: host_enrollments.state = rejected, token_hash = null, fingerprint, pubkey, source_ip, redeemed_at. Accepted: the same columns with state = ready and node0's identity alone, hosts.pubkey and hosts.profile_id, then a host_nodes row per machine — identity in fingerprint, what detect saw of each port in rails, with the rail index each cable carries and the plan's address for that node stamped in |
| 128 | hosts for the token; constants merged with arch and gpu_model; the tunnel session; operator_keys; wg_conf | host_enrollments.wg_conf = null; credentials.json on node0, carrying node_count, the rails plan and the operator keys |
| 129 | peer-login; every RDMA port, for discovery; node1's ip -6 addr, for which port each cable lands on | On node1: the spark user, authorized_keys with node0's peer key and every operator key, sudoers.d/enverge-spark. On node0: peer.key, peer.json with addr, host_key and the links — index, node0's port, node1's port; peer-login deleted |
| 130 | peer.json for the links; the plan's subnet per rail and node1's index, for node1's addresses | Whatever that step writes on node1: netplan on node1's ends of the links (including ipv6-address-generation: eui64, which keeps the 125 pin on the planned rails), earlyoom, LXD, the base image |
| 131 | — | host_nodes.prepare_state on the row at index 1, replaced with {step, state, attempt, error, final, at}. Node0's own steps keep writing hosts.prepare_state as before. A rail-network transition, from either machine, also carries that machine's ports as observed after it ran — rail index, MTU, active_mtu, the address on the interface — merged into that node's host_nodes.rails |
| 132 | the links; the plan's subnets and both indexes | host_nodes.rails[port].verified_at on both rows for every rail that carried traffic. On failure, hosts.prepare_state with step: rail-verify, state: failed |
| 133 | node0's own ports | host_status for node0; node0's host_nodes.rails refreshed — MTU, active_mtu, neighbours_discovered per port — so a cable that comes loose after approval shows on the port that owns it. Nothing for node1 (L3) |
| 134 | host_enrollments.state must be ready; hosts.prepare_state must be self-check applied or already-satisfied | host_enrollments.state = enrolled, then hosts.active = true with a pool |
Node1 is written at three steps: 125 (NM link-local pin only), 129, and 130. Step 125 is otherwise a look that makes the refusal at 127 possible without preparing the box — the pin is the one pre-enrol exception, applied by the agent on both nodes so the owner never has to.
#The shapes
host_nodes, the new table. One row per machine per host, for every
BYO host, so a single Spark is one row. Identity and state are separate
columns, because one never changes and the other is what an operator reads.
The index is explicit: what the agent, the backend and the plan agree on.
create table host_nodes (
host_id text not null references hosts(id) on delete cascade,
index int not null, -- 0 is node0: the relay, the tunnel, the agent
machine_id text, -- /etc/machine-id, as reported: hardware detail, not identity
hardware_id text not null, -- sha256 over the sorted GPU uuids: what tells identical boxes apart
fingerprint jsonb not null, -- hardware identity: gpu, memory, arch, os, driver
rails jsonb not null default '{}', -- the ports as they are, keyed by interface
prepare_state jsonb, -- that machine's latest {step, state, attempt, error, final, at}
seen_at timestamptz, -- no longer written: node1 is not probed (133)
created_at timestamptz not null default now(),
primary key (host_id, index),
unique (hardware_id)
);
host_profiles.constants, the plan, written once. staff_only is a
column, being policy. The GB10 constants unchanged plus the two things a
pair adds: how many machines, and a rail plan — each rail named by an
explicit index, never by its place in the list, with a subnet and an MTU.
The subnets and hosts are the ones NVIDIA's stacked-Sparks guide assigns:
a node's address on a rail is the subnet's tenth host plus the node's own
index — node 0 is .10, node 1 is .11 — and no interface is named,
because which port carries which rail is the cable's decision.
{
"node_count": 2,
"rails": [
{ "index": 0, "subnet": "192.168.100.0/24", "mtu": 9000 },
{ "index": 1, "subnet": "192.168.101.0/24", "mtu": 9000 }
]
}
host_nodes.rails, the fact beside the plan, keyed by this machine's
own interface. Written at 127 from what detect saw; the rail index and the
address stamped onto the ports that carry a cable; refreshed at 131 after
rail-network ran on that machine; verified_at stamped at 132;
neighbours_discovered kept current at 133 for node0. A port nobody is on
appears with neighbours_discovered: 0 and no index. Here node1's second
cable landed on its other port — recorded, not a fault.
{
"enp1s0f0np0": { "index": 0, "device": "rocep1s0f0", "mtu": 9000, "active_mtu": 4096,
"address": "192.168.100.11/24", "neighbours_discovered": 1, "verified_at": "2026-09-08T14:20:03Z" },
"enP2p1s0f1np1": { "index": 1, "device": "roceP2p1s0f1", "mtu": 9000, "active_mtu": 4096,
"address": "192.168.101.11/24", "neighbours_discovered": 1, "verified_at": "2026-09-08T14:20:03Z" },
"enP2p1s0f0np0": { "device": "roceP2p1s0f0", "mtu": 1500, "neighbours_discovered": 0 }
}
The enrolment request at 126, what the agent sends. This is a wire
shape, not a stored one: host-enroll splits it into the enrolment's own
fingerprint (node0's identity, without rails or peers) and the
host_nodes rows, identity into fingerprint and what detect saw of each
port into rails.
{
"gpu_model": "GB10", "gpu_count": 1, "gpu_mem_mib": 122880, "arch": "aarch64",
"os_id": "ubuntu", "os_version": "24.04", "driver_version": "580.65.06",
"machine_id": "3f2a…", "hardware_id": "b41c…", "node_index": 0,
"rails": [
{ "device": "rocep1s0f0", "iface": "enp1s0f0np0", "mtu": 1500, "active_mtu": 1024, "neighbours_discovered": 1, "index": 0 },
{ "device": "roceP2p1s0f0", "iface": "enP2p1s0f0np0", "mtu": 1500, "active_mtu": 1024, "neighbours_discovered": 1, "index": 1 },
{ "device": "rocep1s0f1", "iface": "enp1s0f1np1", "mtu": 1500, "neighbours_discovered": 0 }
],
"peers": [
{ "gpu_model": "GB10", "gpu_count": 1, "arch": "aarch64", "machine_id": "9c17…", "hardware_id": "7e08…", "os_id": "ubuntu", "node_index": 1,
"rails": [ { "iface": "enp1s0f0np0", "index": 0, "neighbours_discovered": 1 }, { "iface": "enP2p1s0f1np1", "index": 1, "neighbours_discovered": 1 } ] }
]
}
A refused pair lands whole on host_enrollments.fingerprint under
state = rejected, which is how the enrolments page shows what turned up.
hosts.prepare_state, node0's track, unchanged in shape and meaning.
host_nodes.prepare_state on the row at index 1 has the same shape
and is node1's:
{ "step": "base-image", "state": "applied", "attempt": 1, "error": null, "final": false, "at": "2026-09-08T14:02:11Z" }
host_status, node0's row, written at 133. No column for node1: half a pair shows as lxd = degraded (L3).
Node0's files, under /etc/enverge, root-only, never in the database:
peer-login— node1's username and password from 124, deleted at 129peer.key— node0's peer key, Ed25519peer.json— the links from the moment they were numbered, then the pinned node1:{"addr": "fe80::…%enp1s0f0np0", "host_key": "ssh-ed25519 …", "peer_index": 1, "links": [{"index": 0, "local": "enp1s0f0np0", "remote": "enp1s0f0np0", "neighbour_addr": "fe80::…%enp1s0f0np0"}, {"index": 1, "local": "enP2p1s0f0np0", "remote": "enP2p1s0f1np1", "neighbour_addr": "fe80::…%enP2p1s0f0np0"}]}
The password never leaves node0, and the backend never learns node1's
link-local address. What it knows about node1 is its host_nodes row:
hardware, the state of its ports, progress, and when it was last seen.
#K2. Four nodes — switchless ring
A spark-4 host is four Sparks on a switchless CX7 ring, sold as one unit. Everything in K holds; what differs is how many satellites there are, how node0 finds the one with no cable to it, and how many workers an instance carries.
- The command carries
--nodes=N.host-provisionputs--peer --nodes=${node_count}on the command whenevernode_count > 1. The installer writes/etc/enverge/peer-nodes, so detect and peer-prep know the target before enrol returns the profile. - One shared login for every satellite. Asked once on node0's terminal, into the same
peer-loginas a pair. Wrong password: re-run the one-liner. - The satellites are the machines on the cables. node0 finds its cabled neighbours, logs in to each, and asks what answers on their rails; the diagonal is reached one hop through either of them, until nothing new answers. Its management address and SSH host key are read from itself over the cable, and later access over the LAN is held to that key. The management LAN (mDNS + ARP, not Tailscale) is scanned only to wake a machine whose cabled ports have carrier and no link-local, and the cables are walked again: a machine that only shares the LAN is never a satellite, however the password is set. More machines on the cables than the profile has satellites is refused.
- A satellite is its
hardware_id, as in K. Machines on the cables are grouped and deduplicated by GPU identity — never/etc/machine-id, which a ring flashed from one image shares on all four boxes.node_index1..N−1 is assigned by sortinghardware_idand pinned inpeer.json, so later runs do not reshuffle. peer.jsonlists every satellite:peers: [{index, addr?, iface?, host_key, mgmt_addr?, links}, …], with railaddr/ifaceonly where there is a cable from node0. A satellite'slinksare its own cables, each carryingpeer— the node at the other end — so the diagonal's two cables are recorded on it even though node0 has no end on either. A pair-shaped file (flat fields, one peer) reads too.- Prepare runs
node{k}/…for each satellite with K1's node1 step code. The diagonal is reached over the management network.rail-verifysends traffic across every cable from both of its ends, the two between the satellites and the diagonal included. - An instance is four containers: the head on node0 and
<head>-w1…-w3on the satellites, created and deleted together over the management network. A worker that will not start fails the create, and every worker already launched is reverted.- Each container gets its node's two cables from
lofispark-dual, rendered per node. On a ring every node also forwards IP and carries a route to each rail it is not on, through the neighbour that is — so the diagonal reaches the head, and the head the diagonal, as the switchless recipe does it. The head's/etc/enverge/cluster.mdnames every node, its rail IPs, and which one has no cable to it. - The profile's description carries a digest of its content;
profilesrewrites a profile that does not match what this node renders now (a machine re-enrolled from another setup), except while an instance holds the rails.
- Each container gets its node's two cables from
- The profile plan (
gb10-spark-4) isnode_count: 4and four rails, one per cable (192.168.100–103/24), as the switchless recipe the ring is built from. Each cable is addressed on the PCI-domain-0 port at both ends (enp1s0f*), never on itsenP2twin, and cables are numbered from the cable map by (lower node, higher node). A port that no longer carries a rail loses its rail address, so a re-numbering never leaves the same address on two ports of one wire. Redeem allows fewer cables than rails and refuses more.
Everything above ends with a host serving. This is what keeps it serving afterwards, when nobody is onboarding anything: a new binary, either machine restarting, and a reboot of both.
One constraint bounds all of it: node0 is the gateway and satellites have no agent (ARCHITECTURE.md — Dual-host clusters). Whatever happens on a satellite is done by node0's agent over the management network, or by host configuration node0's preparation laid down there — never by a program of ours running on the satellite. On a pair that satellite is node1 (L3); on a ring it is each of node1..node3 the same way.
J could treat the agent as one program. Here its three units are the point: PRP, SRV and RB in the glossary.
#The rules
- A reboot runs no preparation. Serving starts at boot on its own. The last act of a completed preparation — handing the credentials to the serving user — is the record: later starts find it and exit at once, whatever the binary, so serving's
Requires=is satisfied without a run, and a host that never completed preparation still cannot serve. - The rollback fires when serving fails to start. It judges the binary. A preparation step failing on the host's state — a cable, a lent rail, node1 — is not a bad binary.
- Serving restores the instance on both machines when it starts — LXD, the head, then the worker on node1 over the management network — and only then points the relay.
- The pair is kept in lockstep: the worker exists and runs exactly when the head does (Lockstep, below).
- "An instance holds the rails" is asked of both machines before preparation touches a live host, node1 over the management network. When node1 cannot be reached, node0's answer is enough.
- Preparation does not run after go-live. Not on a boot, not on an upgrade: an upgrade swaps the binary and restarts serving (L1). A change to host configuration is applied deliberately, on an idle host (Host configuration after go-live, below).
#L1. A new binary (135–141)
- Serving is down only while it restarts. The tenant's containers keep running and their SSH keeps working. What pauses is the VM API and the heartbeat.
- The rollback runs once, and never to the build already running. A previous binary that fails too leaves the host down and alerted
agent_unreachable, rather than flipping between two broken builds. - No preparation runs. The new binary serves the host the previous one prepared. What a newer agent would change about the host waits for the deliberate run below.
#L2. node0 restarts (142–147)
- Serving does not rely on LXD starting itself. Both containers carried
boot.autostart: "true"fromlofispark-baseat the time, yet LXD's boot activation started nothing on either machine of spark-21 across three boots (LXD 5.21.7). LXD starts when something first calls it. Serving waits forsnap.lxd.activate, which opens LXD's socket to thelxdgroup. - What ran before a reboot runs after it, and nothing else. With
boot.autostartunset, LXD restores each instance's last state, and the agent starts the head only when that state is running.- Older hosts carry
boot.autostart: "true"— always start — untilmake byo-prepareremoves it. spark-15 to spark-23 had it removed by hand on 2026-09-14.
- Older hosts carry
- Head, then worker, then relay. SSH is forwarded once, at start. Forwarding it before the head runs would send
:22nowhere until the next restart. - A node1 that cannot be reached does not stop node0 serving. The heartbeat keeps asking (L3).
#L3. node1 restarts (148–152)
- The check rides the heartbeat, the one schedule serving keeps. It asks node1's LXD over the management network — never the rail, which is inside the containers.
- A restarted worker needs nothing else. Its rail addresses,
/etc/nccl.confand the head's launcher key live inside the container and survive a restart. On spark-21 both containers came back with their rails addressed, RDMA active at MTU 4096, and the head reaching the worker asuser. - Half a pair is
degraded, and de-pooling stays an operator's decision. The signal K.8 once took from the rail is taken from the container, which reads the same during a rental as outside one. It lands onlxd, sincehost_statushas no column for a pair. - Both machines at once — a power cut — is L2 and then L3. Serving restores what it reaches at start, and the heartbeat finishes the job when node1 answers.
#L4. A reboot of the pair (153–158)
As built: the agent's side of a reboot, generated from its tests. Scenario 8, with node1 answering.
- node1 first. It has no agent, so once node0 is down nothing is left to tell it anything.
- Both reboots are deferred, so each command returns before its machine goes down, and the caller hears a 202 rather than a dropped connection.
- An unreachable node1 does not block the reboot. Refusing would leave the customer with no reset at all. node0 reboots, and the heartbeat reports half a pair until node1 answers (L3).
- The customer and an operator press the same button. The dashboard's reboot and
/admin's are one call, so a pair is reset the same way whoever resets it. - Recovery is L2 and then L3, as for a power cut. Nothing about the instance is rebuilt: both containers come back as they were.
#Lockstep
The rule is one sentence: the worker exists and runs exactly when the head does. Nothing else needs keeping in step — the rail addresses, /etc/nccl.conf and the launcher key live inside the containers.
| Event | The worker | Built |
|---|---|---|
| Create | launched after the head, trusting its launcher key; failing to start fails the create, and a create refused later reverts it along with the head | yes |
| Delete | removed with the head; both machines' rail addressing restored | yes |
| Start, stop, restart | follows the head; a start that fails is repaired from the heartbeat | yes |
| node0 restarts | started by serving, after the head (L2) | yes |
| node1 restarts | started from the heartbeat (L3) | yes |
| Host reboot | node1 rebooted first, over the management network; back through L2 and L3 (L4) | yes |
| A worker with no head | removed | at create only |
#Host configuration after go-live
- Nothing runs preparation on its own once a host is live — not a boot (rule 1), not an upgrade (L1).
- A change to host configuration is applied deliberately. When a newer agent changes what a step does —
lxd-installgainingusermod -aG lxdunder the same name was the first — an operator runs preparation on the host over the operator tunnel (I):make byo-prepare BYO_HOST=…. - Only on an idle host. Rule 5 is how the run knows: no running instance holding the rails on either machine, or on node0 when node1 cannot be reached. A rented pair is left alone until it is not.
- It is C again, in full. The done list is scoped to the build that wrote it, so a newer binary re-checks every step, and the pair's steps prove the cable once more — possible only because nothing holds it.
- Its failure is reported, not rolled back. It shows in
/admin, leaves serving running, and does not fire the rollback, which judges serving only (rule 2). - The operator keys are not host configuration. They change whenever an operator joins or leaves, not with the agent, and they reach live hosts through M rather than by preparing again.
The operator keys are the one thing on a host that changes on Enverge's side rather than the owner's or the agent's. An operator joins, an operator leaves, and every host has to agree — including the ones a customer is renting, which preparing again (L) would not touch until they are idle.
- The list is sent whole, never as a change. Adding and removing are the same call, and a push that is missed is repaired by the next one rather than replayed.
- An empty list is refused twice.
admin-operator-keyswill not write a change that leaves no operator, and the agent will not apply a list with no keys. Either would leave a box nobody can administer — the same reasonoperator-accessis verified beforeowner-lockout(C). - A rented host is changed too, and that is safe. Setting the keys writes
spark's login file and nothing else: not the rails, not a container, not a unit. That is why it is not a preparation step, which on a live host waits for idle (L). - Serving can do it without root.
authorized_keysbelongs tospark, the user serving already runs as, and node1's belongs to thesparkuser node0 already reaches it as. The serving unit may writespark's.sshand nothing else in/home. - The agent keeps the list where preparation reads it. Preparing again
(L) writes
authorized_keysfrom what the agent holds, so each push also updates that copy — the signed list and its version, as received. Otherwise a later run would put back the list from enrolment, and with it any operator removed since. - A failed push fails, and says so. Nothing retries on its own.
/adminlists every host that was not updated with the reason, and the attempt stops there until an operator sends again.- Sending again re-sends the stored signed list, so it needs no signature and no password manager — one click, once the host is reachable.
- An offline host is out of date but not open: with it off the network,
nobody logs in to it. What matters is the time between it coming back and
someone sending again, and
/adminis where that shows.
- Only a signed list is applied, so the host token is not enough.
- Why.
hosts.agent_tokenis readable from the database and used by the proxy, and the agent's API answers it from the internet. Until now it could create and delete containers. If it could also set the login keys, it would be root on the host, and a leak of the database or the proxy would be root on every BYO host. - How. A person signs every list, on their own machine. The host receives the matching public key once, at enrolment (step 29), and checks every list against it. No call changes that public key.
- Replays. The version is inside what is signed, and the agent refuses a version older than the one it holds. A token holder who kept an old signed list cannot use it to bring back an operator removed since. The same version is taken again only with the same keys, which is what lets sending again land after an answer was lost.
- Where the signing key lives. In the operators' password manager. To sign, an operator copies its private half from the web vault into a throwaway SSH agent on their own machine, for one signature: the sign command clears the clipboard once the agent has it and stops the agent when it is done, so it never reaches the disk. No server holds it and no Enverge page sees it, so a leak of the database, the edge-function secrets or the proxy signs nothing: changing the keys takes the password manager and a superadmin session together.
- What that exposes. For the seconds between copying and signing, the private half is in the operator's clipboard, and so within reach of anything that reads it there — a clipboard manager, or another device sharing the clipboard. That is the price of signing from the web vault.
- Why.
- What is signed is checked where it is signed.
/adminis code served by Vercel, so a tampered page could offer an attacker's list for signing. The signing command prints the operators and key fingerprints from the bytes themselves, and the operator reads that before approving — what they check is exactly what they sign, whatever the page showed. - Only a change needs a signature. Enrolment (step 29) and sending again use the latest signed list as stored, so neither needs the signing key. What cannot happen is any change to who gets in without someone signing it — which is the point.
- Every operator's changes add up in one draft. There is one list: every
operator's keys, one per machine, in one version. A change is one key added
or removed, applied by
admin-operator-keysto the newest list — the one waiting for a signature, if there is one — so Tudor adding his key and Breno adding his before either signs end in one version holding both, and one signature publishes it. Two changes landing at once cannot overwrite each other: the second is applied again on top of the first. Only the newest version can be signed, and it carries every change before it.- A key added belongs to whoever adds it, signed in as themselves: nobody grants access in someone else's name.
- A host enrolled before signed lists trusts no signing key. It refuses every list and keeps the single operator key it enrolled with, until an operator tells it which signing key to trust over the tunnel (I), then sends again. That is the only way a host's trusted signing key is ever set after enrolment.
- A proposal nobody will sign is withdrawn, for everyone. Withdrawing deletes it, and any older unsigned proposal it replaced, so none of them comes back as the one waiting. Only an unsigned proposal can be withdrawn: nothing unsigned ever reached a host, so there is nothing to undo.
- An operator who leaves takes the signing key with them. The signing key is one Bitwarden entry shared by every operator, and signing from the web vault puts its private half through each signer's clipboard, so anyone who has signed may hold a copy. Removing their operator key is not enough: they could sign a list that adds it back. The signing key is changed too — trusted on every host by hand (I), then a new list signed with it.
A host that finished enrolling ends here, for good: its relay, tunnel and DNS go, and nothing it was issued still works. The row stays, and so do its instances and its nodes, so the host's history is still readable. Its machines are free to be enrolled again, as a different setup. This is not H: H revokes an enrolment; N retires a host whose enrolment completed.
- Only a host that is already out of service. Taking it out of placement (
active = false) and draining its instances (G) are the operator's decisions, made first. Retire refuses a host that is active or dedicated, has a live instance, or has a pre-booking whose window has not ended, because an inactive host can still be serving a dedicated tenant or be owed to a paid reservation. A failed create is not a live instance. - Outside first, the stamp last. Every teardown call is the one the enrolment release already makes, and each is idempotent. If one fails part way, nothing is stamped, and running retire again carries on from where it stopped.
retired_aton a host means everything it held is gone. - The row is never deleted.
vms.host_idkeeps a host's instances and billing history pointing at it, and the row keeps its slug reserved. A slug that is freed while a teardown is still running is how a new host once lost its relay to the old host's release. - The customer SSH wildcard is retire's own job. Relay release removes the relay's record, and revoke removes the tunnel's hostname. Only a hard delete removes
*.<slug>.ssh, so a retired host that left it up would keep a name pointing at an address GCP can hand to someone else. - Only a host whose enrolment completed. A host that never redeemed (no
pubkey) was not provisioned by an enrolment, and one whose enrolment is stillpendingorreadyis an enrolment to revoke (H), not a host to retire. So when retire runs, the only enrolment it closes is theenrolledone: released, keeping itshost_idas the record of the setup, with every enrolment's token and WireGuard config cleared. Nothing is left open for the expiry sweep, so the sweep never tries to delete a retired host. - Nothing the host was issued still works. Its
token_hashand SSH-map token are cleared, so an agent left running on the box can no longer report or fetch anything.agent_tokenstays: it is what the backend presents to the agent, not something the host holds against us. A retired host cannot be reactivated, edited or deleted, and a check on the table holds it inactive and undedicated for every writer, so no placement path can reach it. - Its machines are free, and their history stays. A node's
hardware_idis unique only among nodes that are not retired. The box can join a new host at once, and the retired rows still say which host it was in, and when. - Nothing on the box is touched. Retire works on backend state alone. The agent, cloudflared, WireGuard and our keys on the machines are removed by the owner or overwritten by the next enrolment. Handing a machine back is H's uninstall.
| Missing | Phase | Why it matters here |
|---|---|---|
A schedule for host-provision reap | A | The sweep is built; nothing calls it. Until something does, every abandoned click leaks a relay VM and a reserved address |
| Rotating the signing key | M | No call to the agent changes the key a host trusts, so a new signing key is trusted host by hand over the tunnel, as J upgrades are. Losing the password-manager entry before that is done freezes every host's operator keys as they last were |
| Uninstall | H | Revoke is active = false plus the release path, which exists. Handing the machine back does not. It is what gates ever setting hosts.lockout_owner: until it exists, a locked host is one we cannot return |
| Phase | Step | Failure | What happens |
|---|---|---|---|
| A | 5 | the relay will not build | The owner is told to try later; what was allocated is released |
| B | 18 | no token passed | Nothing downloaded; points the owner at the dashboard |
| B | 22 | checksum missing or wrong | Installs nothing, and says so |
| B | 28 | token unknown, spent or expired | One answer for all three. The owner adds the GPU again |
| B | 28 | hardware is not a GB10 Spark | Rejected, and the allocation is released |
| C | 35 | a step will not complete | The run stops there. Visible in /admin, and the host never reaches approval |
| D | 39 | the report shows a GPU that never initialised | The operator does not approve it |
| E | 61 | every idle host refuses the launch | The queue row stays open and the vms row is marked creation_failed_at, so the next trigger retries the same customer |
| E | 60 | a host accepts but its answer cannot be read, and no id was minted | The loop stops rather than rotating: another attempt risks a second container for one queue row. Reported, and left for the next drain to resolve by owner and name |
| E | 62 | the launch worked but this write is lost | The instance is live and reachable — the agent published the port — but its row is missing the host and the customer key. Logged rather than swallowed, because nothing downstream would notice |
| G | 81 | the agent never answers | Nothing is closed. The row keeps billing and keeps its host, which is honest — the capacity is still held — and delete_requested_at is what surfaces it |
| G | 83 | the container is gone but this write is lost | The row outlives the container: still billing, still holding its host, and invisible to every sweep. What the fleet's agent hid, and a BYO host cannot |
| K | 125 | node1 does not answer, or the login is wrong | No enrolment, token unspent. The agent retries every 60 s; the owner re-runs the installer to enter the login again |
| K | 127 | no carrier on any rail | Rejected, token spent, the host released — a lone Spark cannot take a spark-2 allocation |
| K | 127 | more cables than the plan has rails | Rejected: the extra cable has no rail to take |
| K | 127 | the peer is not a GB10 Spark, or is node0 itself | Rejected, and the hardware that turned up is on the enrolment, as it is for node0 |
| K | 132 | the rails carry no traffic once configured | rail-verify fails and retries every 60 s, visible in /admin; approval never opens |
| L | 139 | serving fails to start on a new binary | Rolled back to the previous binary, once. If that fails too, the host stays down and is alerted agent_unreachable |
| L | 145 | node1 cannot be reached when node0 starts | node0 serves; the heartbeat starts the worker when node1 answers (L3) |
| L | 152 | node1 stays unreachable | degraded, half a pair. De-pooling stays an operator's decision |
| L | 156 | node1 cannot be reached for a reboot | node0 reboots anyway, audited as half done; the heartbeat starts the worker when node1 answers (L3) |
| M | 167 | the signature fails the check, or the version is no longer the next one | Nothing is written and nothing is pushed. The operator starts again |
| M | 169 | the host does not answer | Recorded as not updated and listed in /admin with the reason. It stays out of date until an operator sends again (175–177) |
| M | 170 | no signing key trusted, the signature fails the check, or the version is older | Refused, and the keys the host holds stay as they are. The answer says why, and it shows in /admin |
| M | 170 | node1 cannot be reached | node0 is updated and the host answers with node1's version still old, so the host is listed as not updated until an operator sends again. Until then a removed operator's key still opens node1, which has no tunnel of its own: reaching it takes node0, which no longer admits them, or the owner's own network |
| N | 180 | the host never enrolled, its enrolment is still open, it is active or dedicated, has a live instance or a pre-booking to come, or is already retired | Refused with the reason, and nothing changes. An open enrolment is revoked instead (H) |
| N | 184 | the relay instance is still deleting | Nothing is stamped. The address stays reserved until the instance is gone, and retire is run again |
| N | 186 | the customer SSH wildcard cannot be deleted | Nothing is stamped, and retire is run again. The relay and tunnel are already gone, which is safe to repeat |
Set aside, with the reason — because "why didn't you just…" gets asked twice.
- An acceptance probe before activation. The backend would launch a scratch
VM and run
nvidia-smiitself, since the GPU claim is billing-relevant.- The scratch VM is launched by the agent, on the host, and its output returns through the host: against a hostile host it restates the claim rather than checking it.
- A broken host is caught by the step-37 self-check anyway.
- Still worth having one day, and distinct from this: the backend has never used the customer's own path (F) before a customer does.
- A shared relay pool. Removes the slowest step from A — but nobody
waits on that step, since the relay boots while the owner finds a terminal.
- In exchange: capacity to watch, slots to reclaim, a stall when it is empty, and one fewer isolation boundary.
- Revisit if per-host relays prove expensive at rest rather than slow to build.
- The box enrolling anonymously, printing a claim code. Needed a
pending-and-unowned state, an expiry sweep, an alphabet a human could
transcribe, and an enrolment endpoint anyone could write rows to — all to
answer a question the current flow never asks, because the owner is signed in
before the machine is touched.
- Dropping it is a reversal worth naming: it had no secret in the owner's command line, where one lands in shell history and in whatever gets pasted into a support thread.
- Provisioning first makes that trade a bad one. The token buys only an allocation that already carries the owner's name, is single-use, and dies within the hour — so what survives in history is a spent secret.
- The host token authorising key changes, with no signature. One call,
no password manager, no manual step.
- The token is readable from the database, used by the proxy, and answered by the agent from the internet. A call that sets the login keys would make it root on the host, and a leak of either holder root on every BYO host.
- Signing moves that authority to a signing key no server holds (M).
- Signing on the server, or in the browser.
admin-operator-keyssigning with the signing key as a Supabase secret, or pasted into/admin, would keep the change to one click.- A Supabase secret is readable by every edge function in the project.
/adminis code Vercel serves, so a tampered page could take a pasted signing key the next time someone uses it.- Signing on the operator's own machine, with the key loaded from the password manager for one signature, keeps it off every server and every Enverge page.
- Delivering key changes by preparing again. It already exists, and
operator-accessalready writesauthorized_keys.- Preparing again waits for an idle host (L), so a rented host would keep a removed operator's key for as long as the rental lasts.
- It is a loop over every host run by hand, where M is one push.
- Retrying a failed push from the heartbeat. It would bring an offline
host up to date without anyone acting.
- A change would reach a host at a moment nobody chose, for a reason
/admindoes not show. - An offline host cannot be logged into, so there is little to protect in
the meantime:
/adminlists it, and sending again is one click.
- A change would reach a host at a moment nobody chose, for a reason
- One Unix account per operator, instead of a shared
spark. Logins would be attributable by user name.- The agent manages one service user on node0 and node1. An account per operator is a home directory and a sudoers file per operator per machine, created and removed on every host with nothing to clean up after them.
- sshd logs the fingerprint of the key it accepted, and
operator_keysmaps that to the operator, so logins are attributable already.