Structured output
Every exchange between a model and the FACE backend is TOON (Token-Oriented Object Notation), with one documented exception for vision detection. FACE uses no JSON-schema wrappers and no LLM orchestration framework.
Why TOON
TOON is an indentation-based object notation with compact tabular arrays. For record-shaped answers, such as a list of violations, findings or scenarios, the column names are declared once in a header rather than repeated per row. That takes fewer tokens to generate and gives the model less structure to get wrong. It matters on a sovereign, often CPU-bound model plane, where generation time is the dominant cost.
compliance_status: Warning
violations[2]{category,severity,description,related_rule_id,remediation}:
PII,HIGH,Unmasked email column in orders,R-12,Apply masking policy
Retention,MEDIUM,No retention rule on audit table,R-31,Define a retention window
spoken_response: Two compliance findings need attention.From reply to typed struct
Replies are never unmarshalled straight into generated protobuf messages. The flow has three steps:
- Declare an intermediate struct. Each call site defines a Go struct whose fields
carry both
jsonandtoontags. For example, the compliance agent parses intoToonComplianceAnalysis, with the fieldscompliance_status,violations,audit_trailandspoken_response. - Parse.
ai.NewToonOutputParser[T]()returns a parser, andParsereturns a value of typeT, not a pointer. The parser uses thetoon-golibrary. - Map by hand. The call site maps the struct field by field onto the protobuf response. The model’s output never becomes a wire message without passing through code that names each field.
What the parser tolerates
Small models make predictable formatting mistakes. Before decoding, the parser repairs a fixed set of them without changing any data:
- Leading prose, markdown fences (
```toon,```json), and reasoning blocks (<think>…</think>) are stripped. So are stray<tool_call>wrappers. - Declared array lengths that don’t match the number of values are corrected, and the comma count ignores commas inside quoted strings.
- A top-level key the model wrote without its colon gets the colon back.
- Uniformly shifted or four-space indentation is re-levelled to TOON’s two-space unit.
- Blank lines and malformed rows inside a tabular block are dropped.
- Object arrays written as indented key/value blocks are rewritten into tabular form.
What the parser refuses
- An empty completion is a failure, not an empty result. If nothing is left after
stripping, which includes a reply that was only a reasoning block, the parser
returns
TOON_PARSE_FAILURE: empty model completion. A model that said nothing is never read as a model that found nothing. - An explicitly empty array is a valid answer.
rules[0]:orobjects[0]{label,material}:means the model found nothing, and that is how a correctly empty frame or rule set is reported. - Ambiguous nesting is refused. If sibling branches disagree about depth, the document is rejected rather than having a structure guessed for it.
Each call site decides what a parse failure means for its own output. Where a failure would otherwise look like a clean result, it surfaces as an error or a stated absence. The compliance agent, for example, doesn’t report “no violations” when the model gave no usable answer.
The one exception: vision detection
One call departs from TOON: cctv.locateInstances, the class-conditioned object
detection step of the vision pipeline. It asks the vision model for boxes in the
model’s own fine-tuned bbox_2d JSON idiom.
The exception exists because of measurements recorded in the repository, not preference. On a photo of five separable pallet stacks:
- The grading question, whether asked in TOON or JSON, with corner or width/height coordinates, returned one box covering 79–96% of the frame. That’s a scene-level classification with a rectangle around it.
- Two changes together produced per-instance boxes. The first was naming the
class, in the singular (“detect each individual pallet stack”, not “every distinct
object”). The second was asking in
bbox_2dJSON. Class-conditioned TOON got part of the way, then degenerated into repeated rows with pinned coordinates.
The scope of the exception is narrow:
- It covers one detection primitive. The reply is matched by a regular expression and parsed into FACE’s own structs immediately.
- Every verdict still travels as TOON. Grading, material, damage grade and severity all come from the TOON grading call.
- It is off by default (
VISION_DETECT_INSTANCES=trueenables it) because it roughly doubles vision time per frame. The recorded measurement on a CPU inference tier was about 20 s for grading plus about 50 s for detection.
See Vision & CCTV for the full pipeline and coordinate handling.
What this does not establish
- A successful parse proves shape, not truth. The repairs above keep formatting from discarding an answer. They say nothing about whether the answer is right.
- TOON limits what a model can express; it doesn’t constrain what it says. Boundaries on content are set elsewhere: the determination/expression split on Evidence documents, and the guardrails and judging ladder on Agents & oversight.