Every part of a BYO host's life, from an owner adding someone else's Spark on the dashboard to the host being handed back — and what keeps it serving in between. This document is the source of truth for how the BYO system behaves: where code, specs or other docs disagree with it, they are the ones to change.
operator-access completes.spark-2), K2 the four-Spark switchless ring (spark-4).Six areas, one continuous numbering across the whole document, so step 44 means one thing and E means one phase.
The Python agent in
platform/agent/, and the hosts that run it, are out of scope. The BYO agent isplatform/byo-agent/.
Everything is provisioned before the box is ever touched. Clicking "add a
GPU" allocates the relay slot, the address, the tunnel and the hosts row — not
the box enrolling, and not an operator afterwards. By the time the owner runs
the one-liner, the only thing missing is the machine itself.
That ordering is what makes the rest fall out:
relay/gcp/bootstrap.sh at first boot,
so nobody SSHes into a relay to bring it up.Everything belongs to one of six places. The diagrams group them, because which side of a boundary a step happens on is usually the point — a credential minted in an edge function and a credential minted on the host are very different facts.
Who belongs to each area — every actor's ID and label, and the verbs they use — is in the BYO glossary, and only there.
1. The owner's Spark host — the machine being onboarded, and the programs on it. The agent is two of them, because preparing needs root and serving must not have it.
2. app.enverge.ai — the Vercel deployment: the owner's and operator's
pages, the static installer assets, and the proxy that redirects to Blob.
3. Supabase edge functions — everything with authority. Nothing here runs on a host.
4. The Supabase database — host_enrollments, host_profiles, hosts,
and vms from E onwards. Written by edge functions under the service role;
read by the pages under RLS. There is no relays table: a relay is named by its
host's slug, so the hosts row is its record, exactly as it is for the
hand-built fleet relays.
5. Relays — one per host, created during provisioning. Includes the GCP API
that builds the instance, reserves the address and points it at the relay. Same
topology and naming as the fleet runs today — <slug>-relay and
<slug>-relay-ip; the difference is that nobody builds it by hand.
6. Cloudflare — the operator tunnel and DNS.
And three people, none of whom belong to an area: the owner, the operator and the customer.
The customer never learns any of this. They land on a spark-1 host and cannot
tell whose hardware it is, which is the tenancy decision spec63 makes
explicitly rather than by omission.
active = false, and nothing writes true except the
operator at the end.host-provision stays an orchestrator, and replacing a dead relay later
does not have to go through host onboarding to get one.hosts
row allocated. Four things contain it:
community_host from the /admin users tab (or a
superadmin) can click at all,host-provision's reap, which hands back everything a lapsed or rejected
enrolment was holding. Nothing calls reap on a schedule yet.wg0 up with the host's
peer, ssh_mapd on 10.8.0.1:8754, Caddy on :80/:443, and the
LOFISPARK_VMSSH chain seeded with the IAP exemption and a blackhole
default.
accessConfigs[].natIP takes the address literal, not a
self-link, so every create was failing 400.As built: the agent's side of enrolment, generated from its tests. Scenario 1 enrols, 2–4 are refused, and 5 finds the host already enrolled.
The agent runs the whole of new_spark.md §4 and §6 as a state machine. Every step has the same shape, and the shape is the point:
As built: the prepare run, generated from the agent's tests. Scenarios 1–3 take the runner down each of its paths, and 4 is the real runbook.
check, so it reports what it would change while
touching nothing.lxd-install must never reach
owner-lockout.| # | Step | Notes |
|---|---|---|
| 1 | preflight | Linux, Ubuntu, uname -m and nvidia-smi match the allocated profile. Fail closed |
| 2 | docker-purge | purge dockerd, iptables -P FORWARD ACCEPT |
| 3 | swap-off | swapoff -a, comment /etc/fstab, mask swap.img.swap |
| 4 | earlyoom | the profile's floor, verbatim (spec28, retuned spec51) |
| 5 | zfs-arc-cap | /etc/modprobe.d/zfs.conf, 4 GiB |
| 6 | headless | multi-user.target, mask gdm |
| 7 | watchdog | RuntimeWatchdogSec + the boot-reset reporter (spec46); skipped when the profile says watchdog_sec: 0 (a guest with no watchdog device), fails closed when it wants one and /dev/watchdog0 is missing |
| 8 | operator-access | spark user, every operator's key, NOPASSWD sudo, the tunnel minted back in A. authorized_keys is written as the whole list, so it holds exactly the operators and nobody else |
| 9 | lxd-install | apt lxd; do not rely on lxd init --storage-backend |
| 10 | lxd-network | lxc network create lxdbr0; lxd init is not run, so nothing else makes it |
| 11 | firewall | if ufw is active: allow in + route on lxdbr0, allow in on wg0, DEFAULT_FORWARD_POLICY=ACCEPT, reload. No ufw or inactive = satisfied |
| 12 | storage-pool | explicit lxc storage create default zfs |
| 13 | nvidia-runtime | nvidia-container-toolkit-base + libnvidia-container-tools, for nvidia.runtime injection |
| 14 | gpu-accounting | NVML per-process accounting. Driver state, so re-asserted |
| 15 | profiles | lofispark-base + lofispark-gpu |
| 16 | base-image | import the prebuilt image, else build it |
| 17 | relay-peer | apply the WireGuard peer collected in B; collision-check first |
| 18 | owner-lockout | gated on hosts.lockout_owner, ships false |
| 19 | self-check | machine-check every verification, and report |
Orderings that are landmines, not preferences:
Named rather than numbered, because a step inserted in the middle renumbers every landmine below it and the numbers are what a reader trusts.
operator-access before owner-lockout, and verified. Prove Enverge's
access before removing the owner's. The reverse strands a box neither party
can administer.docker-purge before base-image. Docker's FORWARD DROP breaks
lxdbr0 NAT, so the image build loses egress to the NVIDIA repo.zfs-arc-cap before lxd-install. modprobe.d binds at module load, and
installing LXD loads zfs even for a non-zfs pool.swap-off is what arms earlyoom. It fires only when MemAvailable
and free swap are low — an AND. A GPU runaway is unswappable, so leaving
swap on disarms the floor entirely.lxd-network before profiles. lofispark-base attaches eth0 to
lxdbr0, and lxd init — which would have made both the pool and the bridge
— is deliberately not run. Creating the pool explicitly without the bridge
gives a host with LXD up, a healthy pool, and no way to attach a NIC; the
failure then surfaces at profiles, several steps after the omission.firewall before base-image and relay-peer. A default-deny ufw
(cloud images and DGX OS alike) drops the bridge's DHCP, forwarded container
egress and the relay's packets, and none of the three failures names the
firewall: the image build dies resolving archive.ubuntu.com and the control
plane just hangs. DEFAULT_FORWARD_POLICY has to change in the file —
docker-purge's iptables -P FORWARD ACCEPT lasts only until the next
ufw reload.base-image builds from a script the agent carries. On the fleet it
arrives with make sparkN-copy; a BYO host has nobody to run that, so the
builder and its login banner are embedded in the binary. It is run with the
profile's cuda_repo_arch — the script defaults to sbsa, which is arm64's
NVIDIA repo and wrong everywhere else.storage-pool must create the pool explicitly. A dir pool costs ~69 s
on every VM create against 0.6 s on zfs, on a host that looks perfectly
healthy.base-image must end with the alias. The lifecycle launches by alias; an
image without one fails every create.relay-peer must check for a subnet collision first. 10.8.0.0/24 is a
common home-router default, and bringing wg0 up into a colliding range makes
the host unreachable the instant it succeeds.spark-1 tier.gpu_type is not decided here. It was fixed when the allocation was made
— BYO accepts one hardware configuration, and the box proved it was that at
enrolment. Activation is the single flag, plus the placement chosen with it.storage-pool
could be activated by anyone who used the other one.admin-host-action and the duplicate button is gone.Five callers, three shapes. What separates them is how the host is chosen, because that decides when the row can be written and where the SSH hostname comes from.
| # | Trigger | Who calls | Drawn in |
|---|---|---|---|
| 1 | the customer asks | the proxy (PXY) | E1 — the picker chooses, and the customer is waiting |
| 2 | capacity frees and the queue drains | drain-waitlist | E2 — no host until one accepts |
| 3 | a paid pre-book window opens | drain-prebook | E3 — the booking pinned its host when it was made |
| 4 | an operator places one on a dedicated host | admin-vm-action | E2's shape, host supplied instead of rotated |
| 5 | an operator launches a booking by hand | admin-prebook-action | E3's shape, one booking by id |
The two operator paths are the manual twins of the two automatic ones — same steps, with the cron's gate replaced by a superadmin and the selection replaced by whatever the operator named. That is why they are not drawn separately:
force that
revives a booking torn down inside its own window.The row is a precondition for dispatch, in all five. On a host whose agent
writes none, a create sent without one produces a container the database has no
record of: unbilled, invisible to the customer, and holding capacity the picker
cannot see. Every caller now refuses rather than dispatches, keyed on the same
pubkey marker — a fleet host is exempt, because its agent inserts the row a
moment later and refusing there would break a path that works.
To the customer every one of these is identical to any other host, which is the
point of BYO joining the spark-1 pool rather than a tier of its own. They are
not identical in who writes the row: see G, where that difference is the
whole story.
A customer's create, once the host is active.
As built: the agent's side of a create, generated from its tests. Scenarios 1–2 create an instance, and 3–5 refuse and revert.
created_at before the host is contacted, so billing runs
from when capacity was committed rather than from whenever a host answered. A
create that is never answered is marked creation_failed_at and bills nothing.ssh_mapd listens on 10.8.0.1:8754, inside the tunnel, and Caddy proxies
only :8080 — there is no route from Vercel. Publishing a second endpoint on
every relay to create one would be more attack surface than the thing it
protects.
The token is the relay's own, minted per relay, and on a per-host relay it
can forward one host's :22 to one port on that same host — which the agent
decides anyway, being what publishes the port. It grants nothing a compromised
host does not already have. A shared relay would break that, since one peer
could forward another's :22 elsewhere, so a pool must move this back to the backend
first.:22 goes nowhere is
not an instance. The fleet's Python agent logs a map failure and carries on,
which is how a host ends up healthy with a blackholed :22 and nobody
noticing.The queue is the only path that creates a VM with nobody present, acting on an intent that may be days old — which is why billing is re-checked live rather than trusted, and why the host is not known when the row is written.
pick_hosts_for_gpu excludes any host holding a live vms row, so a
host_id written before a launch succeeded would make this row occupy the
candidate that just refused it — benching a working host against every
concurrent drain until the row is marked failed.create_vm
stamps the host and the name in one statement. This one cannot, so it stamps
them once a host has actually taken the VM.A booking buys a named box for a date range. The window opening is the trigger, so nobody is present — and the host was decided when the booking was made, not when it launches.
paid_at alone, and no live Stripe check. Payment is
collected off-band and recorded by ops, which is the deliberate difference
from E2 — and the reason a stale paid_at is a known risk rather than a
caught one.launched_vm_id never nulls — without reading the row's liveness, a
customer who deleted their box mid-window would be stranded until end_date.
A non-live row falls through and the booking relaunches.expire-prebook closes it at
end_date and tears the VM down, which is trigger 5 of G.ssh_keys by user_sub
rather than a per-booking id, so it can only ever be a key they hold. No key
on the account leaves the booking open for an operator to follow up.No proxy, no agent, no HTTP. L4 forwarding only.
:22 is a blackhole. Not an error —
nothing is listening for that address yet. The agent:
Six callers, one shape. They differ only in what decided the instance should go. Three of them (3–5) run on a schedule with nobody present, which is why the shape has to be right without anyone reading the result.
| # | Trigger | Who calls |
|---|---|---|
| 1 | the customer asks | DELETE /api/vms/{id} through the proxy (PXY) |
| 2 | an operator, from /admin | admin-vm-action |
| 3 | a payment probe fails mid-dunning | probe-payment |
| 4 | dunning runs out | teardown-vms, the gated reaper |
| 5 | a pre-book window ends | expire-prebook |
| 6 | an operator ends a pre-booking early | admin-prebook-action |
As built: the agent's side of a delete, generated from its tests. Scenario 6, and 7 for an instance that is already gone.
delete_requested_at — the backend accepted a delete. Stops nothing.deleted_at — the container is confirmed gone. Billing stops here.pick_hosts_for_gpu excludes any host holding a live vms row, so
a row closed before its container dies leaves a host that looks free:
launches keep routing to an agent that refuses them, and every user behind it
stalls.
notify-stuck-deletes scans for, and it also sizes the refund: the
customer gave the instance up when they asked, not when an operator noticed.GET /map) and writes it again if it does not match what the host runs, so relay health is the relay as it is this minute, and a stale or cleared target is put right within a beat.pick_hosts_for_gpu excludes a host
holding a live row, so draining before the row closes hides the host that
just freed up from the customer waiting for it.deleted_at from the host, so a caller
that wrote nothing still looked correct there.admin-vm-action was the last caller doing this. It is the same asymmetry
that produced the empty ssh_hostname and the orphaned container: backend
code assuming the agent is a database writer.delete_requested_at=is.null and
deleted_at=is.null — so on a host whose agent does write, the two compose:
whichever lands first wins and the other is a no-op. Two independent paths to
the same value beat either alone.Two different endings. Revoke cuts a host off and keeps it configured; uninstall hands the machine back.
operator-access before owner-lockout, for the same reason.Everything above is the platform driving a machine nobody is sitting in front of. This is the other thing: a person, on a host, because something is wrong.
/admin renders, because "SSH
in and look" does not survive a thousand strangers' boxes.cloudflared, a unit the operator-access step installed with a
connector token minted backend-side.10.8.0.2.:22
into the tenant's container through LOFISPARK_VMSSH, but that chain is on
PREROUTING, so a connection originating on the relay is not caught by
it — ssh [email protected] from the relay lands on the host's sshd, not
the customer's container.spark, with their own operator key and NOPASSWD sudo — so, root. Not a concession: it is
sole administrative control, the same asymmetry that lets owner-lockout remove
the owner's login and never ours. A BYO host has exactly one administrator and
it is us.
spark;
what tells them apart is the operator key. sshd logs the fingerprint of the key it
accepted, and operator_keys records each key's fingerprint against the
operator's user sub, so a login is attributable without a second account.| Depends on | Why it can be missing |
|---|---|
operator-access completed and verified | It is verified precisely because owner-lockout runs after it. An unverified tunnel plus a completed lockout is a box neither party can administer |
| The host has outbound internet | Both paths are outbound-dialled. A host off the network is unreachable by design, and that is what "we detect rather than prevent" means |
| The operator's key is on the host | Delivered at enrolment, then by M. An operator key added while the host was offline arrives when an operator sends again |
relay-peer applied, for path 2 | Before it, the only way in is the tunnel |
| The relay still exists | Releasing an allocation destroys it. Order matters in H: drain the tenant and take what you need from the host before the relay goes, because after that there is no path to that machine at all |
hosts.lockout_owner ships false — the
owner's own login. That is why it ships false: until uninstall exists, their
login is the recovery path if this flow wedges a box.enverge-agent-rollback, and it
moves backwards — once, when serving will not stay up: it consumes
.previous and will not restore the build that is already running (L1).So an upgrade is an operator action, over the same tunnel as I.
GitHub Actions belongs to none of the six areas. It publishes, and publishing is where its involvement ends — which is the point of the section.
The release verifies the path an owner's installer follows, not the
upload. put() returning a URL says the object exists; it says nothing
about whether app.enverge.ai/agent/… resolves.
AGENT_BLOB_BASE_URL sit blank in production — present
in the dashboard, so it read as configured — while every install one-liner
was broken.No token, and that matters more than it sounds. The installer asks for one only when the machine is not enrolled; an enrolled host reuses its credentials.
Detached, deliberately. A dropped SSH session must not land between writing the binary and writing the units.
Preparation starting serving is one line, and it went missing.
Requires=/After= the prepare unit — and that direction only pulls
preparation in when serving starts, never the reverse.Wants=enverge-agent.service closes it. Earlier hosts
hid it because they had been rebooted since installing.Restarting serving is the installer's job, not the Wants='s. On an
upgrade the serving unit is already running, so Wants= is satisfied and
systemd starts nothing — the host would keep answering from the binary the
install just replaced.
Wants= starts serving once credentials exist.An upgrade does not re-check the host. Preparation does not run after go-live (L).
lxd-install gained usermod -aG lxd while keeping its name, so hosts that had completed it skipped the new work and reported "already satisfied".An upgrade does not reach containers that already exist. The login banner is pushed at create, so a running instance keeps what it was given. Recreating it is the only way to change what a customer sees.
One command covers every BYO host, unlike the per-host targets the fleet carries.
One hosts row, one relay, one tier slot, one enrolment, one token, one
command on node0, one prepare run, one heartbeat. Agentless satellites are
reached from node0 — nothing is ever run on them by anyone. One table,
host_nodes, says which machines a host is made of. K1 is two machines;
K2 is four on a switchless ring. Everything from E onwards is
unchanged either way.
A spark-2 host is two Sparks cabled together over ConnectX-7 and sold as one
unit. The second machine is reached from the first.
peer-prep then activates node1's ConnectX-7
ports and pins link-local eui64 before prepare starts, because
NetworkManager often leaves them disconnected while the kernel link is up
— no fe80::, and the cable probe at 125 would see silence forever.
It asks the cable first. A rail is point-to-point, so whatever answers
the multicast on it is node1 by construction; node0 logs in, groups the
two addresses a QSFP shows into the one machine behind them, and prepares
it at its management address — reconfiguring node1's rails through the
rail being spoken over is a way to cut the session. The management LAN
(mDNS _ssh._tcp and ARP on the uplink — not Tailscale) is the fallback,
for exactly the case above: a node1 whose rails cannot yet answer. A rack
of eleven Sparks on one LAN, five accepting the same login, is what that
ordering is for — the scan can only report the ambiguity, and the cable
has already answered. ENVERGE_PEER_ADDR names node1 outright and skips
both.addr_gen_mode=none,
so they have carrier but no fe80:: and the multicast probe sees
silence; NetworkManager's default also rewrites a live sysctl on
reconnect. Before probing, the agent pins each of node0's RDMA ports'
NM connection to ipv6.method=link-local and addr-gen-mode=eui64,
forces the live sysctl, and bounces the iface if an address is still
missing. After a successful login it does the same on node1's rails
over sudo. Later, rail-network (130) keeps eui64 on the planned rails
via netplan. There is no owner/operator runbook for this.rejected, reap releases the relay, the address, the
tunnel and the host row, and the owner is still at the terminal to see
it. Anything found later can only stop a prepare run and leave the
allocation held. So every piece of evidence that can exist before
enrolment is judged here, once: at least one cable and no more than the
plan has rails for, one machine on every cable, a peer that is the profile's
hardware and not node0 itself, and every machine carrying its own
node_index — 0 for the enroller, each other one once. Rows are written
at the index a machine carries, never at its place in the list. nvidia-smi on either box reports GB10 / 1 / aarch64,
byte-identical to a single, which is why the dropdown alone was never
evidence.
host_nodes row per machine, node0 at
index 0 and the peer at index 1, each with its own fingerprint —
hardware identity — and its own rails, the observed state of each
port, with the rail each cable carries and the plan's address for that
node on it stamped in. Every BYO host gets them, so a single Spark is
one row and the node count is a count. unique (hardware_id) is the
"not node0 itself" rule as a constraint, and it also stops one box
being claimed by two hosts. It keys on the GPUs rather than
/etc/machine-id because that file clones with the image it was
captured in; machine_id stays on the row as reported hardware. An
agent older than the field sends none and its machine_id is written
there instead, which is the namespace its own row was already in.constants.rails
says a subnet and an MTU per rail, and a node's address is the subnet's
tenth host plus its index; host_nodes.rails says what these two
machines look like, port by port, and is refreshed at 130–133 as they
change.peer-access creates the spark user, installs the peer key and the operator keys, and deletes the login. Its check still refuses an SSH host key that is not the pinned one. After go-live it runs only in the deliberate run (L).Exec; a remote one over the rail runs the same check, apply
and re-check on node1 and reports each with index: 1, which
host-report writes to node1's own host_nodes row rather than to the
host's prepare_state. The block is node1's subset of the list in the
same order, so every landmine ordering holds there too, and the deliberate
run in L re-checks node1 as it re-checks node0. Node0's track
keeps its shape and its meaning; node0's self-check, after the block,
asserts node1, which is why steps 39–40 and both activation gates are
untouched.host_nodes.rails — per port, how many machines answered — which reads zero on both ports while an instance holds the rails, so it cannot tell a dead node1 from a rented pair.
degraded while every subsystem was green.host_status row for node1, and host_nodes.seen_at is no longer written.spark user
holds the operators' keys, so the tunnel or the relay lands on node0 and one
more hop lands on node1. Node1 has no tunnel of its own.node_count above one,
so 127 judges nothing and 129–132 are skipped, reported as skipped the way
watchdog is on a board without one. The rails still appear in its
fingerprint, which is a few more fields and nothing else.
hardware_id — the machine's GPU UUIDs, digested — not
/etc/machine-id: systemd writes that file once, and only when it is
empty, so an OS image captured with it populated hands every unit the
same value and two genuinely different Sparks present as one machine
that found itself. The agent makes the same comparison before it
enrols, so a real self-pair costs a retry rather than the token.<head>-w1. node0's agent creates and deletes them together over the management network — not over the rail SSH step 129 established, because the dual profile moves the rail netdevs into the containers the moment one starts.
One new table, host_nodes, one row per machine of a host. Everything else
is a table that existed, written at the steps it was always written at, with
the meaning it always had. The database is written at 120, 127, 131 and 133,
plus the approval at 134.
| Step | Reads | Writes |
|---|---|---|
| 118 | host_profiles.staff_only, to decide whether the picker shows the row | — |
| 119 | staff_only for the gate, constants.node_count for the command | — |
| 120 | — | hosts row; host_enrollments row with host_id, profile_id, token_hash, wg_conf. Same as a single |
| 121 | — | Nothing stored. The response carries node_count and --peer --nodes=N on the command (N from the profile) |
| 123–124 | — | /etc/enverge/peer-login and peer-nodes on node0; peer-prep finds node1 on the rail (falling back to the management LAN only when the rail is silent) and activates its RDMA NM connections + link-local eui64 (writes on node1 before enrol, so the cable is visible) |
| 125 | peer-login; node0's RDMA NM connections and sysctls; node1's nvidia-smi (including GPU uuids), uname, /etc/os-release, /etc/machine-id, sysfs ports, and which port holds each address | On node0: NM ipv6.method=link-local + addr-gen-mode=eui64 on every RDMA port (and live addr_gen_mode=0); peer.json with the links the moment their numbering is assigned. On node1, after a successful login only: the same NM link-local pin on its rails (via sudo). No prepare, no spark user, no keys. In memory the fingerprint gains node_index 0, machine_id, hardware_id, rails[].neighbours_discovered, rails[].index, and peers[] each with its own node_index |
| 126 | — | Nothing yet. The request body is the fingerprint from 125 |
| 127 | host_enrollments by token hash for profile_id; host_profiles for gpu_model, arch, gpu_count, constants.node_count, constants.rails | Refused: host_enrollments.state = rejected, token_hash = null, fingerprint, pubkey, source_ip, redeemed_at. Accepted: the same columns with state = ready and node0's identity alone, hosts.pubkey and hosts.profile_id, then a host_nodes row per machine — identity in fingerprint, what detect saw of each port in rails, with the rail index each cable carries and the plan's address for that node stamped in |
| 128 | hosts for the token; constants merged with arch and gpu_model; the tunnel session; operator_keys; wg_conf | host_enrollments.wg_conf = null; credentials.json on node0, carrying node_count, the rails plan and the operator keys |
| 129 | peer-login; every RDMA port, for discovery; node1's ip -6 addr, for which port each cable lands on | On node1: the spark user, authorized_keys with node0's peer key and every operator key, sudoers.d/enverge-spark. On node0: peer.key, peer.json with addr, host_key and the links — index, node0's port, node1's port; peer-login deleted |
| 130 | peer.json for the links; the plan's subnet per rail and node1's index, for node1's addresses | Whatever that step writes on node1: netplan on node1's ends of the links (including ipv6-address-generation: eui64, which keeps the 125 pin on the planned rails), earlyoom, LXD, the base image |
| 131 | — | host_nodes.prepare_state on the row at index 1, replaced with {step, state, attempt, error, final, at}. Node0's own steps keep writing hosts.prepare_state as before. A rail-network transition, from either machine, also carries that machine's ports as observed after it ran — rail index, MTU, active_mtu, the address on the interface — merged into that node's host_nodes.rails |
| 132 | the links; the plan's subnets and both indexes | host_nodes.rails[port].verified_at on both rows for every rail that carried traffic. On failure, hosts.prepare_state with step: rail-verify, state: failed |
| 133 | node0's own ports | host_status for node0; node0's host_nodes.rails refreshed — MTU, active_mtu, neighbours_discovered per port — so a cable that comes loose after approval shows on the port that owns it. Nothing for node1 (L3) |
| 134 | host_enrollments.state must be ready; hosts.prepare_state must be self-check applied or already-satisfied | host_enrollments.state = enrolled, then hosts.active = true with a pool |
Node1 is written at three steps: 125 (NM link-local pin only), 129, and 130. Step 125 is otherwise a look that makes the refusal at 127 possible without preparing the box — the pin is the one pre-enrol exception, applied by the agent on both nodes so the owner never has to.
host_nodes, the new table. One row per machine per host, for every
BYO host, so a single Spark is one row. Identity and state are separate
columns, because one never changes and the other is what an operator reads.
The index is explicit: what the agent, the backend and the plan agree on.
create table host_nodes (
host_id text not null references hosts(id) on delete cascade,
index int not null, -- 0 is node0: the relay, the tunnel, the agent
machine_id text, -- /etc/machine-id, as reported: hardware detail, not identity
hardware_id text not null, -- sha256 over the sorted GPU uuids: what tells identical boxes apart
fingerprint jsonb not null, -- hardware identity: gpu, memory, arch, os, driver
rails jsonb not null default '{}', -- the ports as they are, keyed by interface
prepare_state jsonb, -- that machine's latest {step, state, attempt, error, final, at}
seen_at timestamptz, -- no longer written: node1 is not probed (133)
created_at timestamptz not null default now(),
primary key (host_id, index),
unique (hardware_id)
);
host_profiles.constants, the plan, written once. staff_only is a
column, being policy. The GB10 constants unchanged plus the two things a
pair adds: how many machines, and a rail plan — each rail named by an
explicit index, never by its place in the list, with a subnet and an MTU.
The subnets and hosts are the ones NVIDIA's stacked-Sparks guide assigns:
a node's address on a rail is the subnet's tenth host plus the node's own
index — node 0 is .10, node 1 is .11 — and no interface is named,
because which port carries which rail is the cable's decision.
{
"node_count": 2,
"rails": [
{ "index": 0, "subnet": "192.168.100.0/24", "mtu": 9000 },
{ "index": 1, "subnet": "192.168.101.0/24", "mtu": 9000 }
]
}
host_nodes.rails, the fact beside the plan, keyed by this machine's
own interface. Written at 127 from what detect saw; the rail index and the
address stamped onto the ports that carry a cable; refreshed at 131 after
rail-network ran on that machine; verified_at stamped at 132;
neighbours_discovered kept current at 133 for node0. A port nobody is on
appears with neighbours_discovered: 0 and no index. Here node1's second
cable landed on its other port — recorded, not a fault.
{
"enp1s0f0np0": { "index": 0, "device": "rocep1s0f0", "mtu": 9000, "active_mtu": 4096,
"address": "192.168.100.11/24", "neighbours_discovered": 1, "verified_at": "2026-09-08T14:20:03Z" },
"enP2p1s0f1np1": { "index": 1, "device": "roceP2p1s0f1", "mtu": 9000, "active_mtu": 4096,
"address": "192.168.101.11/24", "neighbours_discovered": 1, "verified_at": "2026-09-08T14:20:03Z" },
"enP2p1s0f0np0": { "device": "roceP2p1s0f0", "mtu": 1500, "neighbours_discovered": 0 }
}
The enrolment request at 126, what the agent sends. This is a wire
shape, not a stored one: host-enroll splits it into the enrolment's own
fingerprint (node0's identity, without rails or peers) and the
host_nodes rows, identity into fingerprint and what detect saw of each
port into rails.
{
"gpu_model": "GB10", "gpu_count": 1, "gpu_mem_mib": 122880, "arch": "aarch64",
"os_id": "ubuntu", "os_version": "24.04", "driver_version": "580.65.06",
"machine_id": "3f2a…", "hardware_id": "b41c…", "node_index": 0,
"rails": [
{ "device": "rocep1s0f0", "iface": "enp1s0f0np0", "mtu": 1500, "active_mtu": 1024, "neighbours_discovered": 1, "index": 0 },
{ "device": "roceP2p1s0f0", "iface": "enP2p1s0f0np0", "mtu": 1500, "active_mtu": 1024, "neighbours_discovered": 1, "index": 1 },
{ "device": "rocep1s0f1", "iface": "enp1s0f1np1", "mtu": 1500, "neighbours_discovered": 0 }
],
"peers": [
{ "gpu_model": "GB10", "gpu_count": 1, "arch": "aarch64", "machine_id": "9c17…", "hardware_id": "7e08…", "os_id": "ubuntu", "node_index": 1,
"rails": [ { "iface": "enp1s0f0np0", "index": 0, "neighbours_discovered": 1 }, { "iface": "enP2p1s0f1np1", "index": 1, "neighbours_discovered": 1 } ] }
]
}
A refused pair lands whole on host_enrollments.fingerprint under
state = rejected, which is how the enrolments page shows what turned up.
hosts.prepare_state, node0's track, unchanged in shape and meaning.
host_nodes.prepare_state on the row at index 1 has the same shape
and is node1's:
{ "step": "base-image", "state": "applied", "attempt": 1, "error": null, "final": false, "at": "2026-09-08T14:02:11Z" }
host_status, node0's row, written at 133. No column for node1: half a pair shows as lxd = degraded (L3).
Node0's files, under /etc/enverge, root-only, never in the database:
peer-login — node1's username and password from 124, deleted at 129peer.key — node0's peer key, Ed25519peer.json — the links from the moment they were numbered, then the pinned node1: {"addr": "fe80::…%enp1s0f0np0", "host_key": "ssh-ed25519 …", "peer_index": 1, "links": [{"index": 0, "local": "enp1s0f0np0", "remote": "enp1s0f0np0", "neighbour_addr": "fe80::…%enp1s0f0np0"}, {"index": 1, "local": "enP2p1s0f0np0", "remote": "enP2p1s0f1np1", "neighbour_addr": "fe80::…%enP2p1s0f0np0"}]}The password never leaves node0, and the backend never learns node1's
link-local address. What it knows about node1 is its host_nodes row:
hardware, the state of its ports, progress, and when it was last seen.
A spark-4 host is four Sparks on a switchless CX7 ring, sold as one unit. Everything in K holds; what differs is how many satellites there are, how node0 finds the one with no cable to it, and how many workers an instance carries.
--nodes=N. host-provision puts --peer --nodes=${node_count} on the command whenever node_count > 1. The installer writes /etc/enverge/peer-nodes, so detect and peer-prep know the target before enrol returns the profile.peer-login as a pair. Wrong password: re-run the one-liner.hardware_id, as in K. Machines on the cables are grouped and deduplicated by GPU identity — never /etc/machine-id, which a ring flashed from one image shares on all four boxes. node_index 1..N−1 is assigned by sorting hardware_id and pinned in peer.json, so later runs do not reshuffle.peer.json lists every satellite: peers: [{index, addr?, iface?, host_key, mgmt_addr?, links}, …], with rail addr/iface only where there is a cable from node0. A satellite's links are its own cables, each carrying peer — the node at the other end — so the diagonal's two cables are recorded on it even though node0 has no end on either. A pair-shaped file (flat fields, one peer) reads too.node{k}/… for each satellite with K1's node1 step code. The diagonal is reached over the management network. rail-verify sends traffic across every cable from both of its ends, the two between the satellites and the diagonal included.<head>-w1…-w3 on the satellites, created and deleted together over the management network. A worker that will not start fails the create, and every worker already launched is reverted.
lofispark-dual, rendered per node. On a ring every node also forwards IP and carries a route to each rail it is not on, through the neighbour that is — so the diagonal reaches the head, and the head the diagonal, as the switchless recipe does it. The head's /etc/enverge/cluster.md names every node, its rail IPs, and which one has no cable to it.profiles rewrites a profile that does not match what this node renders now (a machine re-enrolled from another setup), except while an instance holds the rails.gb10-spark-4) is node_count: 4 and four rails, one per cable (192.168.100–103/24), as the switchless recipe the ring is built from. Each cable is addressed on the PCI-domain-0 port at both ends (enp1s0f*), never on its enP2 twin, and cables are numbered from the cable map by (lower node, higher node). A port that no longer carries a rail loses its rail address, so a re-numbering never leaves the same address on two ports of one wire. Redeem allows fewer cables than rails and refuses more.Everything above ends with a host serving. This is what keeps it serving afterwards, when nobody is onboarding anything: a new binary, either machine restarting, and a reboot of both.
One constraint bounds all of it: node0 is the gateway and satellites have no agent (ARCHITECTURE.md — Dual-host clusters). Whatever happens on a satellite is done by node0's agent over the management network, or by host configuration node0's preparation laid down there — never by a program of ours running on the satellite. On a pair that satellite is node1 (L3); on a ring it is each of node1..node3 the same way.
J could treat the agent as one program. Here its three units are the point: PRP, SRV and RB in the glossary.
Requires= is satisfied without a run, and a host that never completed preparation still cannot serve.agent_unreachable, rather than flipping between two broken builds.boot.autostart: "true" from lofispark-base at the time, yet LXD's boot activation started nothing on either machine of spark-21 across three boots (LXD 5.21.7). LXD starts when something first calls it. Serving waits for snap.lxd.activate, which opens LXD's socket to the lxd group.boot.autostart unset, LXD restores each instance's last state, and the agent starts the head only when that state is running.
boot.autostart: "true" — always start — until make byo-prepare removes it. spark-15 to spark-23 had it removed by hand on 2026-09-14.:22 nowhere until the next restart./etc/nccl.conf and the head's launcher key live inside the container and survive a restart. On spark-21 both containers came back with their rails addressed, RDMA active at MTU 4096, and the head reaching the worker as user.degraded, and de-pooling stays an operator's decision. The signal K.8 once took from the rail is taken from the container, which reads the same during a rental as outside one. It lands on lxd, since host_status has no column for a pair.As built: the agent's side of a reboot, generated from its tests. Scenario 8, with node1 answering.
/admin's are one call, so a pair is reset the same way whoever resets it.The rule is one sentence: the worker exists and runs exactly when the head does. Nothing else needs keeping in step — the rail addresses, /etc/nccl.conf and the launcher key live inside the containers.
| Event | The worker | Built |
|---|---|---|
| Create | launched after the head, trusting its launcher key; failing to start fails the create, and a create refused later reverts it along with the head | yes |
| Delete | removed with the head; both machines' rail addressing restored | yes |
| Start, stop, restart | follows the head; a start that fails is repaired from the heartbeat | yes |
| node0 restarts | started by serving, after the head (L2) | yes |
| node1 restarts | started from the heartbeat (L3) | yes |
| Host reboot | node1 rebooted first, over the management network; back through L2 and L3 (L4) | yes |
| A worker with no head | removed | at create only |
lxd-install gaining usermod -aG lxd under the same name was the first — an operator runs preparation on the host over the operator tunnel (I): make byo-prepare BYO_HOST=…./admin, leaves serving running, and does not fire the rollback, which judges serving only (rule 2).The operator keys are the one thing on a host that changes on Enverge's side rather than the owner's or the agent's. An operator joins, an operator leaves, and every host has to agree — including the ones a customer is renting, which preparing again (L) would not touch until they are idle.
admin-operator-keys will not write a
change that leaves no operator, and the agent will not apply a list with no
keys. Either would leave a box nobody can administer — the same reason
operator-access is verified before owner-lockout (C).spark's login file and nothing else: not the rails, not a container, not a
unit. That is why it is not a preparation step, which on a live host waits
for idle (L).authorized_keys belongs to spark,
the user serving already runs as, and node1's belongs to the spark user
node0 already reaches it as. The serving unit may write spark's .ssh
and nothing else in /home.authorized_keys from what the agent holds, so each push also
updates that copy — the signed list and its version, as received. Otherwise a later run would put back the list from
enrolment, and with it any operator removed since./admin
lists every host that was not updated with the reason, and the attempt stops
there until an operator sends again.
/admin is where that shows.hosts.agent_token is readable from the database and used by the
proxy, and the agent's API answers it from the internet. Until now it could
create and delete containers. If it could also set the login keys, it would
be root on the host, and a leak of the database or the proxy would be root
on every BYO host./admin is code served by
Vercel, so a tampered page could offer an attacker's list for signing. The
signing command prints the operators and key fingerprints from the bytes
themselves, and the operator reads that before approving — what they check
is exactly what they sign, whatever the page showed.admin-operator-keys to the newest list — the one
waiting for a signature, if there is one — so Tudor adding his key and Breno
adding his before either signs end in one version holding both, and one
signature publishes it. Two changes landing at once cannot overwrite each
other: the second is applied again on top of the first. Only the newest
version can be signed, and it carries every change before it.
A host that finished enrolling ends here, for good: its relay, tunnel and DNS go, and nothing it was issued still works. The row stays, and so do its instances and its nodes, so the host's history is still readable. Its machines are free to be enrolled again, as a different setup. This is not H: H revokes an enrolment; N retires a host whose enrolment completed.
active = false) and draining its instances (G) are the operator's decisions, made first. Retire refuses a host that is active or dedicated, has a live instance, or has a pre-booking whose window has not ended, because an inactive host can still be serving a dedicated tenant or be owed to a paid reservation. A failed create is not a live instance.retired_at on a host means everything it held is gone.vms.host_id keeps a host's instances and billing history pointing at it, and the row keeps its slug reserved. A slug that is freed while a teardown is still running is how a new host once lost its relay to the old host's release.*.<slug>.ssh, so a retired host that left it up would keep a name pointing at an address GCP can hand to someone else.pubkey) was not provisioned by an enrolment, and one whose enrolment is still pending or ready is an enrolment to revoke (H), not a host to retire. So when retire runs, the only enrolment it closes is the enrolled one: released, keeping its host_id as the record of the setup, with every enrolment's token and WireGuard config cleared. Nothing is left open for the expiry sweep, so the sweep never tries to delete a retired host.token_hash and SSH-map token are cleared, so an agent left running on the box can no longer report or fetch anything. agent_token stays: it is what the backend presents to the agent, not something the host holds against us. A retired host cannot be reactivated, edited or deleted, and a check on the table holds it inactive and undedicated for every writer, so no placement path can reach it.hardware_id is unique only among nodes that are not retired. The box can join a new host at once, and the retired rows still say which host it was in, and when.| Missing | Phase | Why it matters here |
|---|---|---|
A schedule for host-provision reap | A | The sweep is built; nothing calls it. Until something does, every abandoned click leaks a relay VM and a reserved address |
| Rotating the signing key | M | No call to the agent changes the key a host trusts, so a new signing key is trusted host by hand over the tunnel, as J upgrades are. Losing the password-manager entry before that is done freezes every host's operator keys as they last were |
| Uninstall | H | Revoke is active = false plus the release path, which exists. Handing the machine back does not. It is what gates ever setting hosts.lockout_owner: until it exists, a locked host is one we cannot return |
| Phase | Step | Failure | What happens |
|---|---|---|---|
| A | 5 | the relay will not build | The owner is told to try later; what was allocated is released |
| B | 18 | no token passed | Nothing downloaded; points the owner at the dashboard |
| B | 22 | checksum missing or wrong | Installs nothing, and says so |
| B | 28 | token unknown, spent or expired | One answer for all three. The owner adds the GPU again |
| B | 28 | hardware is not a GB10 Spark | Rejected, and the allocation is released |
| C | 35 | a step will not complete | The run stops there. Visible in /admin, and the host never reaches approval |
| D | 39 | the report shows a GPU that never initialised | The operator does not approve it |
| E | 61 | every idle host refuses the launch | The queue row stays open and the vms row is marked creation_failed_at, so the next trigger retries the same customer |
| E | 60 | a host accepts but its answer cannot be read, and no id was minted | The loop stops rather than rotating: another attempt risks a second container for one queue row. Reported, and left for the next drain to resolve by owner and name |
| E | 62 | the launch worked but this write is lost | The instance is live and reachable — the agent published the port — but its row is missing the host and the customer key. Logged rather than swallowed, because nothing downstream would notice |
| G | 81 | the agent never answers | Nothing is closed. The row keeps billing and keeps its host, which is honest — the capacity is still held — and delete_requested_at is what surfaces it |
| G | 83 | the container is gone but this write is lost | The row outlives the container: still billing, still holding its host, and invisible to every sweep. What the fleet's agent hid, and a BYO host cannot |
| K | 125 | node1 does not answer, or the login is wrong | No enrolment, token unspent. The agent retries every 60 s; the owner re-runs the installer to enter the login again |
| K | 127 | no carrier on any rail | Rejected, token spent, the host released — a lone Spark cannot take a spark-2 allocation |
| K | 127 | more cables than the plan has rails | Rejected: the extra cable has no rail to take |
| K | 127 | the peer is not a GB10 Spark, or is node0 itself | Rejected, and the hardware that turned up is on the enrolment, as it is for node0 |
| K | 132 | the rails carry no traffic once configured | rail-verify fails and retries every 60 s, visible in /admin; approval never opens |
| L | 139 | serving fails to start on a new binary | Rolled back to the previous binary, once. If that fails too, the host stays down and is alerted agent_unreachable |
| L | 145 | node1 cannot be reached when node0 starts | node0 serves; the heartbeat starts the worker when node1 answers (L3) |
| L | 152 | node1 stays unreachable | degraded, half a pair. De-pooling stays an operator's decision |
| L | 156 | node1 cannot be reached for a reboot | node0 reboots anyway, audited as half done; the heartbeat starts the worker when node1 answers (L3) |
| M | 167 | the signature fails the check, or the version is no longer the next one | Nothing is written and nothing is pushed. The operator starts again |
| M | 169 | the host does not answer | Recorded as not updated and listed in /admin with the reason. It stays out of date until an operator sends again (175–177) |
| M | 170 | no signing key trusted, the signature fails the check, or the version is older | Refused, and the keys the host holds stay as they are. The answer says why, and it shows in /admin |
| M | 170 | node1 cannot be reached | node0 is updated and the host answers with node1's version still old, so the host is listed as not updated until an operator sends again. Until then a removed operator's key still opens node1, which has no tunnel of its own: reaching it takes node0, which no longer admits them, or the owner's own network |
| N | 180 | the host never enrolled, its enrolment is still open, it is active or dedicated, has a live instance or a pre-booking to come, or is already retired | Refused with the reason, and nothing changes. An open enrolment is revoked instead (H) |
| N | 184 | the relay instance is still deleting | Nothing is stamped. The address stays reserved until the instance is gone, and retire is run again |
| N | 186 | the customer SSH wildcard cannot be deleted | Nothing is stamped, and retire is run again. The relay and tunnel are already gone, which is safe to repeat |
Set aside, with the reason — because "why didn't you just…" gets asked twice.
nvidia-smi itself, since the GPU claim is billing-relevant.
admin-operator-keys signing
with the signing key as a Supabase secret, or pasted into /admin, would keep the
change to one click.
/admin is code Vercel serves, so a tampered page could take a pasted signing key
the next time someone uses it.operator-access already writes authorized_keys.
/admin does not show./admin lists it, and sending again is one click.spark. Logins would
be attributable by user name.
operator_keys
maps that to the operator, so logins are attributable already.
comments (0)