TensorBoot documentation.
TensorBoot is the single-host daemon that runs each TensorCheck review job inside its own fresh Firecracker microVM. One VM per job, the VM boundary is the tenant boundary, the base image is read-only with a per-job scratch drive, the daemon talks only outbound to the control plane, and it never executes commands inside a running guest. TensorBoot is in Closed beta.
What TensorBoot is.
TensorBoot is a single-host daemon that runs TensorCheck code reviews inside Firecracker microVMs on a worker server. Each review job gets a fresh, fully isolated microVM: the repository under review is attacker-controlled content and the review tooling executes against it, so the VM boundary is the tenant boundary. Every VM runs in its own network namespace with its own veth/TAP and nftables rules, boots from a read-only EROFS base image with a per-job writable scratch drive, and is billed from cgroup counters sampled at the start of Running and again at teardown.
The worker server is not publicly reachable. TensorBoot initiates all connections outbound to the control plane — it receives jobs, boots and supervises the VMs, and reports usage back. It never executes commands inside a running guest; the guest is tensorcheck itself, which pulls the repository, obtains its LLM credentials, posts the review to GitHub, and powers off.
TensorBoot is in Closed beta. The hosted platform already runs every TensorCheck review under it; self-hosting the daemon is available to Closed beta operators only.
What you will notice.
On the hosted platform there is nothing to install or configure, and you do not see or configure TensorBoot from the console.
- Every hosted review runs in its own Firecracker microVM that is destroyed at the end of the job.
- The immutable review guest contains the Codex 0.148.0 and Pi 0.83.0 harnesses only. Claude Code and the Claude Agent SDK are excluded from the Public MVP.
- The repository and the LLM credentials enter the VM only for the duration of the job, and leave with it when the VM is torn down.
- Successful reviews produce usage records that feed billing. (The pricing and billing math live in billing, not here.)
How a review job moves through states.
A dispatched job moves Accepted → Preparing → Prepared → Booting → Running → Draining → Done. At its deadline a running job goes Running → Timeout → Draining. The terminal set is Done and Failed; Timeout and Draining still owe teardown and are therefore non-terminal.
| State | Terminal | Meaning |
|---|---|---|
Accepted | no | A dispatch was accepted and durably recorded; preparation has not started. |
Preparing | no | Allocating the per-job cgroup, uid/gid and jailer chroot; building the per-VM netns and nftables; materialising the read-only EROFS rootfs and the per-job writable scratch drive. |
Prepared | no | All per-job resources are built; the job is ready to boot. |
Booting | no | The Firecracker microVM is starting. |
Running | no | The guest workload is executing. Usage metering from the cgroup counters begins here; a job that reaches Running is billable. |
Timeout | no | The VM was killed at its deadline while Running. Teardown is still owed, so the state is non-terminal (it transitions on to Draining). |
Draining | no | Teardown is in progress — VM and jailer chroot, scratch drive, overlay artifacts, nftables rules, netns/veth and cgroup are destroyed in order. Teardown is still owed, so the state is non-terminal. |
Done | yes | Terminal: teardown is complete. |
Failed(prepare) | yes | Terminal. Preparation failed, or the dispatch was rejected as an identity conflict. |
Failed(boot) | yes | Terminal. Boot failed after the job reached Prepared. |
Failed(deadline) | yes | Terminal. The deadline was reached in a pre-Running state — there is no billable event for a job that never ran. |
Run the daemon yourself (Closed beta operators).
Host prerequisites
- A host Linux kernel ≥ 5.7 (the
clone3(CLONE_INTO_CGROUP)floor). - A usable cgroups v2 hierarchy whose freshly-created delegated child exposes
cpu.stat,memory.currentandmemory.peak. Stock Linux 5.7 through 5.18 lackmemory.peakand do not qualify unless that interface is backported; the metering interfaces are feature-probed. /dev/kvm,firecrackerandjailer,erofs-utils,mkfs.ext4, andnftables(with theinetfamily and named sets).- A pinned guest
vmlinux, andIP forwardingenabled on the transfer subnets only.
1. Build
Build the tensorbootd binary with the pinned toolchain from rust-toolchain.toml (channel 1.90.0).
2. Configure
The daemon takes a single JSON config file. Every struct is deny_unknown_fields, so an unknown key is a startup error. This is the shipped example template (placeholders only — a real deployment replaces the registry, the base URL, the zero digest, and the TLS/key paths):
{
"store_path": "/var/lib/tensorboot/store.redb",
"cgroup_root": "/sys/fs/cgroup/tensorboot",
"kernel": "/var/lib/tensorboot/vmlinux",
"rootfs_dir": "/var/lib/tensorboot/rootfs",
"scratch_dir": "/var/lib/tensorboot/scratch",
"status_dir": "/var/lib/tensorboot/status",
"chroot_dir": "/srv/jailer/firecracker",
"host_egress": "10.100.0.1",
"authorities": ["github.com", "api.github.com"],
"worker_subnets": [{ "net": "10.100.0.0", "prefix": 24 }],
"review_image": {
"reference": "registry.tensorplane.dev/review:2026-08",
"digest": "sha256:0000000000000000000000000000000000000000000000000000000000000000"
},
"review_workload": {
"argv": ["tensorcheck", "review-job", "--from-mmds"],
"env": []
},
"tensorplane_base_url": "https://api.tensorplane.dev",
"jailer": {
"jailer_path": "/usr/bin/jailer",
"firecracker_path": "/usr/bin/firecracker",
"chroot_base_dir": "/srv/jailer",
"uid": 1000,
"gid": 1000,
"cgroup_parent": "tensorboot.slice"
},
"image_prep": {
"work_dir": "/var/lib/tensorboot/convert",
"layer_cache_dir": "/var/lib/tensorboot/layers",
"layer_cache_max_bytes": 34359738368
},
"control": {
"server_addr": "api.tensorplane.dev:8443",
"server_name": "api.tensorplane.dev",
"root_ca_der": ["/etc/tensorboot/tls/root-ca.der"],
"client_cert_der": ["/etc/tensorboot/tls/client.der"],
"client_key_der": "/etc/tensorboot/tls/client.key.der",
"time_authority_keys": ["/etc/tensorboot/keys/time-authority-1.ed25519"],
"backoff_initial_ms": 200,
"backoff_max_ms": 30000,
"backoff_factor": 2,
"tick_interval_ms": 500,
"connect_timeout_ms": 10000,
"io_timeout_ms": 30000,
"listen_timeout_ms": 1000,
"shutdown_deadline_ms": 5000
}
}Three validation rules are easy to miss: chroot_dir must exactly equal jailer.chroot_base_dir joined with the basename of jailer.firecracker_path; review_image.digest must be sha256: followed by exactly 64 lowercase hexadecimal characters; and time_authority_keys is a fail-closed trust store that cannot be empty.
| Field | Type | Meaning (default) |
|---|---|---|
store_path | path | Durable job-state store (redb). Its parent also holds the assignment-watermark replay boundary. |
cgroup_root | path | Root of the delegated cgroups v2 subtree the daemon places per-job cgroups under. |
kernel | path | Pinned guest vmlinux image. |
rootfs_dir | path | Per-job read-only EROFS rootfs and input images. |
scratch_dir | path | Per-job writable scratch drives. |
status_dir | path | Per-job guest status images (see Observability). |
chroot_dir | path | Firecracker jobs directory. Must exactly equal jailer.chroot_base_dir joined with the basename of jailer.firecracker_path; any other value is a config error. |
host_egress | IPv4 | Host egress address for the worker transfer subnets. |
authorities[] | string[] | Egress authorities the guest may reach (for example github.com, api.github.com). |
worker_subnets[] | { net: IPv4, prefix: u8 } | Per-VM transfer subnets. net must parse as IPv4; prefix is the mask length. |
review_image.reference | string | OCI reference of the review image. Non-empty. |
review_image.digest | string | Pinned content digest. Must be `sha256:` followed by exactly 64 lowercase hexadecimal characters. |
review_workload.argv[] | string[] | Guest workload command; argv[0] is the program. Must be non-empty. |
review_workload.env[] | { name, value }[] | Environment passed to the workload. Defaults to empty. |
tensorplane_base_url | string | Control-plane base URL. Non-empty. |
hosted_input.erofs_path | path (optional block) | Pre-supplied hosted input image. Non-empty when hosted_input is present. |
hosted_input.sha256 | string (optional block) | Content hash of the hosted input image. |
jailer.jailer_path | path | Path to the jailer binary. |
jailer.firecracker_path | path | Path to the firecracker binary; its basename derives chroot_dir. |
jailer.chroot_base_dir | path | Base directory under which jailer builds per-job chroots. |
jailer.uid / jailer.gid | u32 | Unprivileged uid/gid the jailer drops the VMM to. |
jailer.cgroup_parent | string (optional) | Optional parent cgroup/slice for the jailer. |
image_prep.registry_auth | { username, password } (optional) | Registry credentials for the OCI pull; absent means anonymous. |
image_prep.registry_ca_der[] | path[] | Extra DER roots trusted only for the configured private registry. |
image_prep.work_dir | path | Converter scratch/work directory. |
image_prep.layer_cache_dir | path | Content-addressed layer cache directory. |
image_prep.layer_cache_max_bytes | u64 | Layer cache ceiling. Default 34359738368 (32 GiB). |
control.server_addr | string | host:port the daemon dials for the single outbound mTLS control channel. Non-empty. |
control.server_name | string | DNS name presented for SNI and validated against the server certificate. Non-empty. |
control.root_ca_der[] | path[] | DER trust root(s) for the server certificate. Non-empty. |
control.client_cert_der[] | path[] | DER client certificate chain (leaf first) the daemon presents. Non-empty. |
control.client_key_der | path | DER PKCS#8 private key for the client certificate. |
control.time_authority_keys[] | path[] | Raw 32-byte Ed25519 time-authority public keys (one file each). Fail-closed: cannot be empty. |
control.connect_timeout_ms | u64 | TCP connect bound. Default 10000. |
control.io_timeout_ms | u64 | Per-operation read/write bound on the control socket. Default 30000. |
control.backoff_initial_ms | u64 | Initial reconnect backoff. Default 200. |
control.backoff_max_ms | u64 | Maximum reconnect backoff. Default 30000. |
control.backoff_factor | u32 | Backoff growth factor. Default 2. |
control.tick_interval_ms | u64 | Steady-state tick interval. Default 500. |
control.listen_timeout_ms | u64 | How long a tick waits for a pushed dispatch before the next supervise/drain/shutdown check. Default 1000. |
control.shutdown_deadline_ms | u64 | Per-operation I/O bound applied during graceful shutdown. Default 5000. |
3. Run
The daemon accepts exactly one argument, the config file path; its usage text is tensorbootd <config.json> (path to the JSON config file). It logs to stderr via tracing, filtered by the RUST_LOG environment variable (default info). Both SIGTERM and SIGINT request a clean shutdown, which exits 0. Run it as root (the jailer, cgroups, nftables and KVM all require it).
The repository ships no systemd unit and no installer; there is no deploy/ directory. Write your own supervisor unit that runs the binary as root under your init system.
| Exit code | Meaning |
|---|---|
0 | Clean shutdown. |
2 | Usage or configuration error, printed as `config error: …` (a usage error shares this code). |
3 | Composition error, printed as `composition error: …`. |
Logs and per-job status.
TensorBoot has no /metrics endpoint and no HTTP surface of any kind. Observability is the structured tracing stream on stderr and the per-job status files. A per-job recovery failure is logged, never fatal.
| stderr line | Meaning |
|---|---|
dispatch accepted and booted | A dispatch was accepted, prepared, and its VM booted. |
dispatch accepted but not booted | Accepted but not yet booted; the cause is logged and the job is left non-terminal for a later restart. |
dispatch re-acked (idempotent duplicate) | A duplicate of a known dispatch; re-acknowledged, nothing re-run. |
dispatch conflict: identity mismatch, existing job untouched | A dispatch reused an id with a different identity; the existing job is left untouched. |
dispatch rejected (fail-closed; channel kept) | A dispatch was refused fail-closed; the control channel is kept open. |
recovered job | A non-terminal job was reconciled on startup. |
job recovery failed; left non-terminal for a later restart | Recovery of one job failed; it is left non-terminal and never aborts the daemon. |
startup orphan sweep clean | The startup sweep reclaimed any leaked named resources and found the host clean. |
The status_dir holds per-job guest status images — <job>.img and <job>.findings.img. The daemon decodes a valid status record purely as an untrusted operator diagnostic, never as authority over the host-observed outcome.
Symptoms, causes, and fixes.
| Symptom | Cause | Fix |
|---|---|---|
tensorbootd: fatal: usage: | The daemon was invoked with the wrong number of arguments. | Pass exactly one argument, the config path. The usage text is tensorbootd <config.json> (path to the JSON config file). |
tensorbootd: fatal: config error: cannot read config file | The config path does not exist or is not readable. | Check the path and permissions of the file you passed. |
tensorbootd: fatal: config error: config is not valid JSON | The file is not valid JSON, or carries an unknown field (every struct is deny_unknown_fields). | Fix the JSON and remove any field not in the reference above. |
required config field `…` is missing or empty | A required field is absent or blank. | Supply the named field. review_workload.argv, control.root_ca_der and control.client_cert_der must be non-empty. |
config field `chroot_dir` is invalid: must exactly equal the derived Firecracker jobs directory | chroot_dir does not match jailer.chroot_base_dir joined with the basename of jailer.firecracker_path. | Set chroot_dir to exactly that derived directory. |
review image is not configured | review_image.reference or review_image.digest is empty, or the digest is not canonical. | Set a reference and a `sha256:` + 64-lowercase-hex digest. |
`time_authority_keys` is empty | The time-authority trust store is empty. | The trust store is fail-closed and cannot be empty; list at least one Ed25519 key file. |
tensorbootd: fatal: composition error: assignment watermark … is present but unreadable/corrupt | The assignment-watermark file next to store_path exists but is unreadable or corrupt. | A corrupt replay boundary could admit a stale assignment, so the daemon refuses to start. Restore or remove the corrupt watermark (a missing file is a legitimate first run). |
dispatch rejected (fail-closed; channel kept) | A dispatch was refused on the fail-closed path. | Expected when a dispatch cannot be admitted; the control channel stays open. Inspect the logged reason. |
job recovery failed; left non-terminal for a later restart | Reconciling one job on startup failed. | Never fatal. The job is left non-terminal and retried on the next restart; inspect the logged per-job error. |
host lacks /dev/kvm or memory.peak | KVM is unavailable, or a freshly-created cgroups v2 child does not expose memory.peak (stock Linux 5.7–5.18). | Run on a host with /dev/kvm and a cgroups v2 hierarchy exposing cpu.stat, memory.current and memory.peak; the metering interfaces are feature-probed. |
What this version does not do.
- Closed beta — behaviour and interfaces can change before general availability.
- Single host — no multi-node or fleet scheduling, no live migration, no Kubernetes, no Cloud Hypervisor.
- Fixed review guest — the Codex and Pi harnesses only; no other harness runs in the guest.
- Managed-key only — the guest obtains its LLM credentials from the control plane under a per-job session; customer-supplied provider keys are not part of this version.
- No metrics endpoint, no installer, and no systemd unit ship in the repository.
- The Closed beta does not claim high availability or an availability SLA.
Get early access to the platform.
Every product is included in one platform subscription. Request early access and we'll onboard your team hands-on — free while we build toward launch.