← All docs
TensorGate — LLM gateway data planeEarly access

TensorGate documentation.

TensorGate is the LLM-gateway data plane that sits between review jobs and the upstream LLM providers. It authenticates a per-job gateway credential, admits each request against your org's entitlement (fail-closed), resolves the provider key from platform custody, enforces request limits, forwards over a closed route table with no retry, and streams the response back — recording usage and denial-audit events to a durable outbox. It is a data plane only, never a management plane.

Overview

What TensorGate is.

TensorGate is the data plane that carries a review job's model calls to the upstream LLM provider. For each request it authenticates a short-lived per-job gateway credential, admits the request against your organization's entitlement over an internal mTLS control-plane seam (fail-closed — if admission cannot be reached, the request is denied), resolves the provider key from platform custody at request time (the gateway never persists it), enforces request limits, and forwards to a single fixed provider host over a closed route table with no gateway-side retry, streaming the response back over SSE under a 10-minute cap. Every call produces usage and denial-audit events that are spooled to a durable outbox. TensorGate is a data plane only — it has no management API and is never a management plane.

There are two audiences for this page:

  • Hosted customers — TensorGate is platform infrastructure behind TensorCheck. There is nothing to install or configure; the sections below describe what you will notice and how to read the status codes when a call is denied.
  • Operators self-hosting the gateway — the repository ships a systemd package. The self-hosted section covers building, configuring, validating, and running the unit. The gateway supports hosted mode only.
Hosted on TensorPlane

What you will notice.

On the hosted platform you do not configure TensorGate. Every TensorCheck review job's provider calls pass through it using a short-lived per-job gateway credential minted for that job. The provider credential you registered in the console is held in platform custody and resolved at request time; the gateway never persists it.

  • Requests outside your org's entitlement, calls on a revoked job, or over-limit traffic are denied with the status codes in the table below and recorded in the denial audit.
  • Successful calls produce signed usage records that feed billing. (The pricing and billing math live in billing, not here.)
  • You never present a provider API key to the gateway yourself; the job credential is the only inbound credential, and the provider key is injected by the gateway from custody.
Client contract

Sending a request.

The gateway exposes a closed, POST-only route table. Each public path maps to one upstream path and one expected model; a body model different from the route's expected model is denied (403). Every other method or path — including /healthz on the public listener — is an audited 404.

Public path (POST)Expected model
/openai/v1/responsesgpt-5-codex
/xai/v1/chat/completionsgrok-build-0.1
/moonshot/v1/chat/completionskimi-k2.7-code-highspeed
/zai/v1/chat/completionsglm-5
  • Send exactly one Authorization: Bearer <job credential> header. A repeated Authorization header is rejected.
  • An x-api-key header is forbidden and rejects the request. Both inbound Authorization and x-api-key are stripped before the request is forwarded, and the provider key is injected outbound.
  • Request headers are capped at max_request_header_bytes (over ⇒ 431) and the body at max_request_body_bytes (over ⇒ 413).
  • The response is streamed back over SSE under a 10-minute cap. The gateway does not retry a forwarded request.
Status codes

Status codes and denial classes.

A denied request returns a bare status code with an empty body. The client sees only the status; the audit class below is what the denial audit records and what support can look up. 503 and 429 carry Retry-After — 1 for limit and admission-unavailable denials, or the exact number of seconds to wait for a budget denial. 401 carries Cache-Control: no-store and no WWW-Authenticate.

StatusAudit classMeaningWhat to do
400body.read_incompleteThe request body ended before its declared length.Resend the complete request body.
401credential.malformed, credential.expired, credential.not_yet_valid, credential.signature_invalid, credential.issuer_mismatch, credential.audience_mismatch, credential.unknown_kidThe per-job gateway credential is missing, malformed, expired, or not valid for this issuer or audience. Carries Cache-Control: no-store and no WWW-Authenticate.On the hosted platform the platform mints job credentials, so a persistent 401 is a platform-side issue. Self-hosting, confirm the credential and the verify-key set.
403admission.org_denied, admission.job_revoked, body.model_not_in_catalogue, g5.live_authority_deniedThe provider or model is outside your org's entitlement, the job was revoked, the body model does not match the route's expected model / the org catalogue, or the live authority policy denied the request.Call a provider and model within your entitlement. A revoked job cannot be retried.
404route.not_found, route.provider_disabledThe method and path are not on the closed route table, or the route's provider is disabled.Use one of the four POST routes below. Anything else — including /healthz on the public listener — is an audited 404.
413limit.body_too_largeThe request body exceeded max_request_body_bytes.Reduce the request body below the configured cap.
429limit.job_concurrency, limit.org_concurrency, limit.org_rate, limit.pre_auth_peer_concurrency, admission.budget_exhaustedPer-job, per-org, or pre-auth peer concurrency, or the per-org request rate, was exceeded; admission.budget_exhausted means the org's credit budget for the period is exhausted. Carries Retry-After.Retry after the Retry-After value; reduce concurrency or request rate. For a budget-exhausted 429 the header is the exact number of seconds to wait — top up or wait (see /docs/billing).
431limit.header_too_largeThe request header fields exceeded max_request_header_bytes.Reduce the size of the request headers.
503limit.global_concurrency, admission.unavailable, admission.no_catalog_entry, admission.budget_integrity_hold, admission.prior_period_liability, limit.peer_identity_unavailable, g5.global_not_ready, g5.live_authority_unavailable, custody.unavailable, custody.invalid_key, credential.keyset_unavailableGlobal concurrency was shed, or admission, peer identity, custody, the verify-key set, or the live authority was momentarily unavailable or not ready (fail-closed); admission.prior_period_liability means an unsettled prior-period liability holds admission. Carries Retry-After.Retry after the Retry-After value.

g5.live_authority_denied is 403 on the production admission path; a 503 variant exists only on a non-production focused-graph branch (src/forward.rs :1522).

Self-hosted

Run the gateway yourself (operators).

Prerequisites

  • The pinned Rust toolchain from rust-toolchain.toml (channel 1.98.1, with rustfmt and clippy).
  • A Linux host with systemd.
  • Operator-issued TLS material — public chains and CA bundles, plus private keys you will encrypt as systemd credentials. Repository fixtures are never production authority.
  • Hosted mode is the only supported mode. A config with mode = "on-prem" parses and then refuses to start with ErrOnPremNotYetSupported: mode = "on-prem" is parsed but not yet supported; the MVP implements hosted mode only.

1. Build

Build the locked release binary.

2. Configure

The single TOML file named by TENSORGATE_CONFIG is the whole deployment configuration. Copy the template to /etc/tensorgate/tensorgate.toml, replace every REPLACE_* placeholder, and set ownership root:tensorgate mode 0640. Install operator-issued public chains and CA bundles under /etc/tensorgate/certs (root:tensorgate, 0640).

# Non-secret operator template. Replace every REPLACE_* placeholder before use.
# Certificate/CA files are public material provisioned by the operator. Private
# keys are delivered only as encrypted systemd credentials and appear at the
# fixed runtime paths below.
mode = "hosted"

[public_listener]
bind = "REPLACE_PUBLIC_IP:8080"
tls_cert = "/etc/tensorgate/certs/public-chain.pem"
tls_key = "/run/credentials/tensorgate.service/public-tls-key"

[review_relay_listener]
bind = "REPLACE_PRIVATE_RELAY_IP:8443"
tls_cert = "/etc/tensorgate/certs/review-relay-chain.pem"
tls_key = "/run/credentials/tensorgate.service/review-relay-tls-key"
guest_relay_roots = "/etc/tensorgate/certs/review-relay-roots.pem"
guest_relay_roots_sha256 = "REPLACE_WITH_64_LOWERCASE_HEX_DIGEST"
guest_relay_roots_len = 0
source_interface = "REPLACE_REVIEW_VETH"
source_admit_cidrs = ["REPLACE_TRANSFER_NETWORK/31"]

[private_listener]
bind = "REPLACE_PRIVATE_OPERATIONAL_IP:9090"
allowed_cidrs = ["REPLACE_PRIVATE_OPERATIONAL_CIDR"]
tls_cert = "/etc/tensorgate/certs/operational-chain.pem"
tls_key = "/run/credentials/tensorgate.service/operational-tls-key"
client_ca = "/etc/tensorgate/certs/operational-client-ca.pem"

[control_plane]
seam_url = "https://REPLACE_SEAM_DNS:8444"
mtls_client_cert = "/etc/tensorgate/certs/control-plane-client-chain.pem"
mtls_client_key = "/run/credentials/tensorgate.service/control-plane-mtls-key"
seam_server_ca = "/etc/tensorgate/certs/control-plane-server-ca.pem"

[store]
path = "/var/lib/tensorgate/state.db"
max_page_count = 262144

[outbox]
max_bytes = 1073741824
audit_reserve_bytes = 134217728
SectionPurpose
modeDeployment mode. Only "hosted" starts; "on-prem" parses then refuses startup.
[public_listener]Public data listener (TLS): bind address and the public chain / key. Serves only the closed POST route table.
[review_relay_listener]Review-relay ingress (TLS): bind, chain / key, pinned guest-relay roots and their digest, source interface, and admitted transfer CIDRs.
[private_listener]Private operational listener (mTLS): bind, allowed_cidrs, server chain / key, and the operational client CA. Serves /healthz, /readyz, /metrics.
[control_plane]mTLS control-plane seam: seam URL and the client cert / key and server CA used for entitlement admission and key custody.
[store]Durable state DB path and its max_page_count.
[outbox]Durable usage and denial-audit outbox: max_bytes total and audit_reserve_bytes reserved for audit events.
[limits]Request limits — see the limits table above. Omit to take the normative defaults.

The [limits] block is optional; omitting a field takes its normative default. Every field except per_org_concurrency (fixed at 32) is tunable.

FieldDefaultOver the limit ⇒
max_request_header_bytes65536 (64 KiB)431
max_request_body_bytes33554432 (32 MiB)413
per_job_concurrency8429
per_org_concurrency32 (fixed — not tunable)429
per_org_rate_per_sec50429
per_org_rate_burst100429
global_concurrency512503
header_read_timeout10 sconnection closed

3. Secrets

Private keys are delivered only as systemd-creds encrypt .cred files. The unit loads exactly four encrypted credentials, each from /etc/tensorgate/credentials/<name>.cred in a root-only directory:

  • public-tls-key — the public data listener's TLS key.
  • review-relay-tls-key — the review-relay listener's TLS key.
  • operational-tls-key — the private operational listener's TLS key.
  • control-plane-mtls-key — the control-plane mTLS client key.

Do not put private keys in the config, argv, environment, shell history, or logs. At runtime the service materializes each credential at the fixed path under /run/credentials/tensorgate.service/, which the template above already references. Encrypt each key into /etc/tensorgate/credentials/ and repeat for all four names:

4. Validate

The preflight parses the full config and loads every public, review-relay, operational, and control-plane TLS input without binding a socket. A preflight failure is terminal.

On a production host the preflight runs as the unit's ExecStartPre=/usr/bin/tensorgate --check-config. systemd materializes the encrypted credentials at the fixed /run/credentials/tensorgate.service/ paths first, so the TLS keys the preflight loads exist only while the unit runs. Validate by starting the unit and reading its journal:

The manual form TENSORGATE_CONFIG=/path/to/tensorgate.toml tensorgate --check-config is only for a config whose TLS key paths are already materialized on disk (a test or staging config). Running it against the production config outside the unit fails, because the /run/credentials/tensorgate.service/ paths the production config names do not exist there — they are materialized only by the unit's ExecStartPre.

5. Install

Run the installer as root, pointing it at the built binary. It reloads systemd, starts the unit only through its fail-closed ExecStartPre preflight, and enables the unit only after start succeeds.

6. Operate

Probe readiness on the private operational listener over mTLS. If preflight fails after a config change, correct the material and restart the unit.

Observability

Health, readiness, metrics, and the outbox.

  • GET /healthz (private operational listener) — liveness, 200 ok.
  • GET /readyz — readiness. Ready: {"ready":true,"reasons":[]} with 200. Not ready: 503 with {"ready":false,"reasons":[…]} naming the unsatisfied dependencies.
  • GET /metrics — Prometheus text exposition. Series include tensorgate_gate_state, tensorgate_gate_reason, tensorgate_gate_epoch, tensorgate_audit_state, tensorgate_audit_refusals, tensorgate_audit_memberships, tensorgate_journal_unwritable, and tensorgate_budget_reservation_exceeded_total.
  • Durable outbox — usage and denial-audit events are spooled to a durable outbox bounded by [outbox] max_bytes with audit_reserve_bytes reserved for audit events, so denial-audit records are retained even under usage-event pressure.

/healthz, /readyz, and /metrics are served only on the private operational listener, restricted to allowed_cidrs over mTLS. They are not exposed on the public data listener.

Troubleshooting

Symptoms and fixes.

SymptomLikely cause / fix
tensorgate: unknown arguments; supported mode is --check-configThe binary was invoked with an argument other than --check-config. No arguments serves; --check-config runs preflight only.
tensorgate: configuration preflight failedPreflight could not load the config or a TLS profile. If you see this running --check-config manually on a production host, the usual cause is that the /run/credentials/tensorgate.service/ key paths are materialized only by the unit (ExecStartPre); rerun through systemctl start tensorgate.service and read journalctl -u tensorgate.service. Otherwise correct the operator-owned config or material and rerun; a preflight failure is terminal.
ErrOnPremNotYetSupportedmode = "on-prem" parses but refuses to start. Set mode = "hosted" — the MVP implements hosted mode only.
401 on every requestThe per-job Authorization: Bearer credential is missing, expired, or from the wrong issuer/audience, or the verify-key set is unavailable (503). No WWW-Authenticate is returned.
403 on a valid credentialThe provider or model is not in the org's scope, the body model does not match the route's expected model, or the job was revoked.
404 on a request you expected to routeThe path or method is not on the closed route table, or the provider is disabled. Only the four POST routes exist.
429 or 503 with Retry-AfterA limit was hit (429), or admission/custody/global capacity was momentarily unavailable (503). Retry after the Retry-After value. A 429 on a credit-exhausted org (admission.budget_exhausted) carries the exact number of seconds to wait in Retry-After.
/readyz returns 503 with reasonsReadiness is not yet satisfied; the JSON body lists the reasons (for example a time-authority or feed dependency). Readiness clears when the listed reasons clear.
Known limitations

What this version does not do.

  • Hosted mode only — on-prem parses then refuses to start.
  • Four fixed routes — a closed, POST-only route table with a single pinned provider host per route. No wildcard routing and no additional provider hosts.
  • No gateway-side retry — a forwarded request is not retried by the gateway.
  • Per-org concurrency is fixed at 32 and cannot be raised by configuration.
  • The MVP does not claim high availability or an availability SLA.

Get early access to the platform.

Every product is included in one platform subscription. Request early access and we'll onboard your team hands-on — free while we build toward launch.