Skip to content

AI safety

FACE’s agents read a customer’s records and propose actions. The controls on this page sit at the points where untrusted text reaches a model, and where a model’s output could turn into an action. None of them makes a language model safe against a determined adversary. Together they make attacks harder and bound what an attack can achieve, and each one says where it stops.

Two kinds of untrusted text

FACE treats text a caller types differently from content the platform fetched, because the two carry different risks.

Caller commands: a phrase filter at entry points

ai.SanitizePrompt rejects inputs that contain common copy-pasted injection phrases, such as ignore previous, system prompt or bypass guardrails. The check is case-insensitive. It runs on user-command entry points (a fetch command, SOP generation, the ReAct loop, a voice utterance). A match is refused with prompt injection detected: blocked by ISO 42001 security guardrail.

This is a ten-phrase substring filter. Paraphrase, another language or a homoglyph gets past it. It is useful because it is cheap, and it is measured not to refuse ordinary operator work. It is not an injection defence against an adversary, and FACE does not rely on it as one.

Fetched content: fenced, not filtered

Connector rows, page text, OCR output and earlier model answers derived from them are wrapped by ai.FenceUntrusted before they reach a prompt:

  • Instruction. The model is told explicitly that the fenced block is untrusted data to analyse, that it carries no authority, and that instructions inside it are part of the data and must never be obeyed.
  • An unguessable delimiter. The markers carry 64 bits of crypto/rand entropy, drawn fresh for each call, and are redrawn if they already occur in the content. Content cannot close its own fence without reading the process’s memory.
  • No bytes changed. The content appears verbatim. Freight records that legitimately say “ignore previous routing advice” are therefore not refused or mangled.

OpenBias guardrails

Each agent service has a judge built from a rule set: the global baseline rules, a shared domain-scope rule, and the service’s own rules. The rule files are in grpc/agents/openbias/rules/, and each one states in its header whether a judge enforces it. A matching rule refuses the request with PermissionDenied, naming the rule ID, and writes a PROMPT_REJECTED audit record. That record carries the rule, severity, category and compliance tags, with the actor ID hashed.

The rules come in three families:

FamilyWhat it coversExamples of rule IDs
OWASPSecurity abuse: prompt injection and jailbreaks, SSRF through agent tools, disclosure of secrets, attempts to tamper with the audit trail, path traversal in uploads.OWASP-LLM01, OWASP-LLM02, OWASP-POSTURE-003, OWASP-TWINS-012
OpenBiasBias and ethics: hate speech, discriminatory claim decisions, retaliatory rules, circumventing ethical-sourcing or governance requirements.OPENBIAS-001, OPENBIAS-RULES-003, OPENBIAS-RULES-006
DomainScopeKeeping each agent inside logistics. This is a boundary hint, not a security control.OPENBIAS-DOMAIN-000

The documentation and the code are checked against each other:

  • Parity. cmd/openbias_rule_parity_test.go compares every documented rule ID and regular expression with what the judge actually compiles. A rule documented with no code behind it fails, and so does code with no documentation.
  • Firing. The firing tests run the real judge. Each rule must fire on its own documented example and must not block its listed allowed examples. The count of rules that cannot fire on their own example is pinned at 0, and so is the count of patterns that fail to compile.

The agent loop fails closed

The ReAct loop (ExecuteLoop) looks up the guardrail mapping for the requested agent. An agent ID with no mapping is refused with a FailedPrecondition that names it as a server-side configuration defect. The loop never runs an unguarded agent. The agent ID is passed through filepath.Base before it is used as a path, so a request cannot choose which template file becomes the system prompt. cmd/ralph_guardrail_failclosed_test.go pins both behaviours.

The voice line

The phone line is the one prompt path a person can reach without an account. Every caller utterance passes three checks, in order, before any model is called:

  1. A length bound. Utterances over 2,000 characters are refused.
  2. The phrase filter, ai.SanitizePrompt.
  3. The voice judge, built from the VoiceService OpenBias rules.

A refusal is spoken to the caller as a sentence. It is not an error, because an error on this path would be silence on the line. The prompt is never built and the model is never called. TestRefusedUtteranceNeverReachesTheModel proves this with a model client that fails the test if anything calls it. Refusals are audited. A caller’s words never go in the process log (TestVoiceTranscriptIsNeverLogged). Only lengths and outcomes are logged there.

The carrier’s webhooks require a valid request signature, and the media socket requires a short-lived token that FACE minted. See Identity & access.

Judging proposed actions

When an agent’s final answer proposes logistics actions, each action is judged before it is accepted. The judge is the shared Runink judging ladder (inference/judgement), running in process. Its one model question goes to the in-cluster plane.

  • Evidence is what the run gathered. Evidence is the data this run fetched and the observations its tools made, passed inside an untrusted fence. The proposer’s own rationale is never treated as evidence for its proposal.
  • Dissent removes the action. The loop is sent back once to revise. An action still dissented from after that is removed from the answer. It is kept only as a record with accepted=false, and never served as a recommendation.
  • “Unable to judge” is never shown as agreement. Actions beyond the per-answer budget are marked unable-to-judge with the reason judgement-budget-exhausted. They are never silently accepted.
  • The verdicts stay in FACE. No verdict is written to another platform’s store.

The judge is on by default. The setting FACE_RALPH_JUDGEMENT controls it, and FACE_RALPH_JUDGE_MAX_ACTIONS and FACE_RALPH_JUDGE_REVISIONS bound its cost.

Bounding what agents may do

A mitigation against injection has to be paired with limits on what model output can cause:

  • Autonomous email. An agent can send email without a human only to recipients on FACE_AGENT_EMAIL_ALLOWLIST. If that setting is unset, no mail is sent autonomously. An agent-proposed email to anyone else is downgraded to a draft for human review, and the downgrade is audited with the recipient hashed. A bare * in the list is ignored, never treated as a wildcard.
  • Tool gating. Agent tools are authorised through the security policy engine, against the verified subject of the request. A service-account token holds the Reader tier and cannot reach write-tier tools.
  • Read-only SQL. Agent-issued and user-issued SQL both go through the same read-only check as every other query. See Network egress controls.

Structured output (TOON)

Model output is parsed as TOON (Token-Oriented Object Notation) into dedicated intermediate structs, then mapped into FACE’s own types. Output is never unmarshalled directly into API messages, and a malformed answer is refused as TOON_PARSE_FAILURE. FACE does not use JSON-schema wrappers or LLM orchestration frameworks.

One exception is documented and bounded: object detection in camera frames asks the vision model for boxes in the model’s native bbox_2d JSON format, because measurement shows no other format localises reliably. Those boxes are parsed into FACE’s own structs immediately. Every verdict about a frame still travels as TOON.

What these controls do not establish

  • Indirect prompt injection is an unsolved problem. Fencing makes injection materially harder. It does not stop a model from being persuaded by text inside the fence. The real limit is what output is allowed to do, which is why the controls above cap agent actions.
  • Regex guardrails stop what they describe and nothing more. A rule blocks the phrasings its pattern matches. Parity and firing tests prove the rules can fire. They do not prove the rules catch every harmful request.
  • The domain-scope rule is not a security control. It keeps agents on topic, and off-topic prompts it does not recognise pass through.
  • A judge verdict is a model’s judgement. It is a check, not proof. Keep a human in the loop for consequential actions.
  • Not every documented rule file is enforced. Some rule files in the repository describe services no judge enforces yet, and each says so in its header. Only services with a judge get guardrail refusals.