Operations
Health
GET /health always answers HTTP 200 with:
{"status":"ok","model_ready":false,"port":8080}The status code is deliberately unconditional. PULSE serves every non-AI
feature without a model, so a readiness probe here should catch a wedged or
unbound server, not an AI outage. Read model_ready for the model. It is true
only when the engine is running and has confirmed it can serve a request.
Self-healing
An in-process watchdog probes local resources and the inference and voice
endpoints, and a remediator reacts to what it finds.
SelfHealService/GetStatus shows it working: whether the loop is healthy, the
last event, recent events, and how many remediations were tried and failed.
healthy is false only when the most recent event is critical and was not
remediated. The service only observes. It never triggers remediation itself.
Logs
Logs are OpenTelemetry-shaped JSON on stdout, with service.name
pulse-backend and the version from APP_VERSION. Lines worth watching:
| Log line | Meaning |
|---|---|
refusing to start: ... | A fatal startup condition. The rest of the line names it. |
AUTHENTICATION DISABLED — every request will be rejected as Unauthenticated | The session key is one that authz refuses. Fix PULSE_JWT_SECRET. |
PUBLISHCFG ... | What each outward channel is allowed to do |
grounding ACTIVE / grounding INACTIVE: <reason> | Whether retrieval grounding works |
gRPC reflection ENABLED | Reflection is on. It should not be on a deployed instance. |
security config validated, or a warning listing problems | The result of the production gate |
The production gate
At startup PULSE checks for insecure configuration:
- The session key is missing or refused.
CORS_ALLOWED_ORIGINSis unset or contains a wildcard, while the app is not served same-origin throughPULSE_WEB_DIR.
With PULSE_ENV=production, any problem stops the server with refusing to start in production. Otherwise it is a warning.
PULSE_ENV is not set by any manifest today, so deployments run the warning
path. Confirm the checks pass before you set it, or the pod will refuse to
start. Reflection no longer depends on it: it has its own opt-in,
PULSE_GRPC_REFLECTION.TLS
Pick one:
| Setup | How |
|---|---|
| TLS at the edge (usual) | Nothing to set. The edge terminates TLS and speaks plain HTTP to PULSE. |
| Direct exposure | PULSE_TLS_CERT and PULSE_TLS_KEY. TLS 1.2 minimum, forward-secret AEAD suites only. |
| Service-plane mutual TLS | MESH_MTLS=true, with the mesh CA. Every client needs a mesh certificate. TLS 1.3 only. If the CA is unavailable, the server refuses to start rather than downgrade. |
Rate limiting
PULSE_RATE_RPS and PULSE_RATE_BURST cap calls per peer IP. Over-limit calls
get ResourceExhausted before any other work. Off by default.
Licence gate
With a licence configured (LICENSE or LICENSE_FILE, verified with
LICENSE_PUBLIC_KEY), every non-public call is refused with license expired or invalid when its signature is bad or it has expired. With no licence
configured, nothing is gated.
Backups and the replica
In a cluster, Litestream replicates SQLite to the in-cluster object store on the
schedule in grpc/litestream.yml. The replica is plaintext at rest, including
password hashes; see the comments in that file for the pending fix.
Once encrypted rows in the new envelope format exist, do not roll back to a build from before that format: it cannot read them.