Skip to content

Structured output

Every exchange between a model and the FACE backend is TOON (Token-Oriented Object Notation), with one documented exception for vision detection. FACE uses no JSON-schema wrappers and no LLM orchestration framework.

Why TOON

TOON is an indentation-based object notation with compact tabular arrays. For record-shaped answers, such as a list of violations, findings or scenarios, the column names are declared once in a header rather than repeated per row. That takes fewer tokens to generate and gives the model less structure to get wrong. It matters on a sovereign, often CPU-bound model plane, where generation time is the dominant cost.

compliance_status: Warning
violations[2]{category,severity,description,related_rule_id,remediation}:
  PII,HIGH,Unmasked email column in orders,R-12,Apply masking policy
  Retention,MEDIUM,No retention rule on audit table,R-31,Define a retention window
spoken_response: Two compliance findings need attention.

From reply to typed struct

Replies are never unmarshalled straight into generated protobuf messages. The flow has three steps:

  1. Declare an intermediate struct. Each call site defines a Go struct whose fields carry both json and toon tags. For example, the compliance agent parses into ToonComplianceAnalysis, with the fields compliance_status, violations, audit_trail and spoken_response.
  2. Parse. ai.NewToonOutputParser[T]() returns a parser, and Parse returns a value of type T, not a pointer. The parser uses the toon-go library.
  3. Map by hand. The call site maps the struct field by field onto the protobuf response. The model’s output never becomes a wire message without passing through code that names each field.

What the parser tolerates

Small models make predictable formatting mistakes. Before decoding, the parser repairs a fixed set of them without changing any data:

  • Leading prose, markdown fences (```toon, ```json), and reasoning blocks (<think>…</think>) are stripped. So are stray <tool_call> wrappers.
  • Declared array lengths that don’t match the number of values are corrected, and the comma count ignores commas inside quoted strings.
  • A top-level key the model wrote without its colon gets the colon back.
  • Uniformly shifted or four-space indentation is re-levelled to TOON’s two-space unit.
  • Blank lines and malformed rows inside a tabular block are dropped.
  • Object arrays written as indented key/value blocks are rewritten into tabular form.

What the parser refuses

  • An empty completion is a failure, not an empty result. If nothing is left after stripping, which includes a reply that was only a reasoning block, the parser returns TOON_PARSE_FAILURE: empty model completion. A model that said nothing is never read as a model that found nothing.
  • An explicitly empty array is a valid answer. rules[0]: or objects[0]{label,material}: means the model found nothing, and that is how a correctly empty frame or rule set is reported.
  • Ambiguous nesting is refused. If sibling branches disagree about depth, the document is rejected rather than having a structure guessed for it.

Each call site decides what a parse failure means for its own output. Where a failure would otherwise look like a clean result, it surfaces as an error or a stated absence. The compliance agent, for example, doesn’t report “no violations” when the model gave no usable answer.

The one exception: vision detection

One call departs from TOON: cctv.locateInstances, the class-conditioned object detection step of the vision pipeline. It asks the vision model for boxes in the model’s own fine-tuned bbox_2d JSON idiom.

The exception exists because of measurements recorded in the repository, not preference. On a photo of five separable pallet stacks:

  • The grading question, whether asked in TOON or JSON, with corner or width/height coordinates, returned one box covering 79–96% of the frame. That’s a scene-level classification with a rectangle around it.
  • Two changes together produced per-instance boxes. The first was naming the class, in the singular (“detect each individual pallet stack”, not “every distinct object”). The second was asking in bbox_2d JSON. Class-conditioned TOON got part of the way, then degenerated into repeated rows with pinned coordinates.

The scope of the exception is narrow:

  • It covers one detection primitive. The reply is matched by a regular expression and parsed into FACE’s own structs immediately.
  • Every verdict still travels as TOON. Grading, material, damage grade and severity all come from the TOON grading call.
  • It is off by default (VISION_DETECT_INSTANCES=true enables it) because it roughly doubles vision time per frame. The recorded measurement on a CPU inference tier was about 20 s for grading plus about 50 s for detection.

See Vision & CCTV for the full pipeline and coordinate handling.

What this does not establish

  • A successful parse proves shape, not truth. The repairs above keep formatting from discarding an answer. They say nothing about whether the answer is right.
  • TOON limits what a model can express; it doesn’t constrain what it says. Boundaries on content are set elsewhere: the determination/expression split on Evidence documents, and the guardrails and judging ladder on Agents & oversight.