Skip to content

Runners

A fetch runs on a runner, never on the control plane. The control plane receives RunFetch, forwards it to a runner together with the caller’s session and the connection credentials the fetch needs, and streams the runner’s progress back to the cockpit. At startup each process logs which role it has:

🧵 Fetch role: CONTROL-PLANE — forwards fetches to the dedicated runner at <endpoint> (never runs them here).
🧵 Fetch role: RUNNER — executes the fetch pipeline in-process (dedicated compute).

A runner is the FACE backend image started with SERVICE_ROLE=runner. Every boot check on Prerequisites and ordering applies to a runner as well, including the mesh CA, the KEK and the inference endpoint. A runner also mounts the same appfs root as its control plane, because it validates the forwarded session there and writes its fetch results there.

How to connect data sources to runners, and how to choose a runner for a source, is covered in Data sources. This page covers operating the runners themselves.

Managed runners in the cluster

A managed runner is either a FaceInstance (the operator’s face-runner-<instance> Deployment) or the static face-runner Deployment from CORE’s on-host bundle. The control plane dials MANAGED_RUNNER_ENDPOINT (default face-runner:7102) over mutual TLS from the shared mesh CA. It never falls back to plaintext.

TaskHow
Size a runnerSet spec.computeUnits on the FaceInstance. The operator turns it into requests and limits and reports them in status.allocatedCpu and status.allocatedMemory. spec.tier: local uses a fixed, modest footprint instead.
Pin a runner to nodesspec.nodeSelector (and spec.gpu for GPU scheduling).
Idle a tenantspec.suspend: true scales the runner to zero and keeps the resource.
Point the control plane at a runnerspec.managedRunnerEndpoint on the ControlPlane, because the operator names per-tenant runners face-runner-<instance>.
Refuse to run fetch work on the control planeSet REQUIRE_RUNNER_ISOLATION=true on the control plane. An unreachable runner is then a hard failure.
Keep runner state across restartsKeep RUNNER_ID stable. The operator sets it to the Deployment name.

Runners are not consensus voters. They leave RAFT_NODE_ID and RAFT_MTLS_ADDR unset, and a runner’s /readyz reports the raft check as skipped. That is expected: the control plane is the only voting member.

When a runner cannot be reached, the fetch fails with:

Analysis failed: dedicated runner unreachable — fetch does not run on the control plane: <reason>

<reason> explains the dial error. Look for a runner pod that is not Ready, an endpoint that points at the wrong Service, or a mesh certificate that no longer chains to the CA the control plane holds.

Self-hosted runners (outbound enrolment)

A self-hosted runner runs on your own machine and dials out to the control plane. Nothing dials in to it. The runner authenticates with a service-account token issued for your tenant by IdentityService.GenerateServiceAccountKey, holds one enrolment stream open, and runs the fetches the control plane assigns over that stream.

VariablePurpose
SERVICE_ROLE=runnerStarts the process as a runner.
RUNNER_ENROLL_ADDRThe control plane’s gRPC address, as seen from the runner. Setting it turns enrolment on.
RUNNER_ENROLL_TOKENThe service-account token. It is required.
RUNNER_IDA stable id, so that a reconnect takes the same registry slot.
RUNNER_NAME, RUNNER_VERSIONShown in the control plane’s runner registry.
RUNNER_ENROLL_TLSDefaults to TLS with the system roots.

When enrolment starts, the runner logs:

🛰️  Outbound runner enrolment enabled: dialling <addr> (runner id "<id>")

If RUNNER_ENROLL_ADDR is set without a token, the runner refuses to start:

❌ RUNNER_ENROLL_ADDR is set but RUNNER_ENROLL_TOKEN is empty — a runner enrols with a service-account token issued by IdentityService.GenerateServiceAccountKey; there is no anonymous enrolment

When the stream ends, the runner reconnects after 5 seconds.

An enrolled runner handles one kind of work, fetch. It does not persist, queue or retry work. A fetch assigned while the runner is shutting down is lost, and the control plane sees a deadline. Rotate the service-account token the way you rotate any credential, then restart the runner with the new value.

The cockpit’s runner page cannot deploy a FACE runner for you yet. Start self-hosted runners yourself, as above.

Snowflake runners

runink runner create provisions a Snowflake Container Services runner. It reads --compute-pool (or RUNNER_COMPUTE_POOL, which is required for that type), RUNNER_CPU_LIMIT (default 1) and RUNNER_MEMORY_LIMIT (default 1G). It stops if the compute pool is not named, because it cannot invent one.