Skip to content

Vision & CCTV

FACE grades what a camera sees with the sovereign vision tier: whether a pallet, crate or container door is damaged, how badly, and what it’s made of. The camera surface is SurveillanceService, with two paths that deliberately answer different questions.

RPCPathModel?
IngestSurveillanceFeedGraded. A pushed frame, or frames sampled from an RTSP/HTTP feed.Yes: the TOON grading turn, plus optional detection.
DetectLiveFrameLive. One validated frame, local detector only.No. It can’t reach the vision model.
GetEvidenceFrameFetches the annotated frame an ingest retained.No.

Camera connections are set up as described in Data sources.

The graded path

    flowchart TD
    F[Frame or sampled feed frames] --> V[Validate + downscale<br/>long edge ≤ VISION_MAX_EDGE]
    V --> R[Region proposals<br/>ml/segment + HDBSCAN]
    R --> G[Grading turn — TOON<br/>subject, material, damage grade, %]
    G --> Q{VISION_DETECT_INSTANCES=true<br/>and a subject found?}
    Q -- no --> O[Response + localisation_note]
    Q -- yes --> D[Class-conditioned detection<br/>bbox_2d JSON, one turn per frame]
    D --> O
  
  1. Frame grading. The vision model is asked, in TOON, for the inspectable subject: the object, its material, a damage_grade (PRISTINE, DAMAGED, DEFECTIVE or DESTROYED) and a model-estimated damage_percentage.
    • A person is not a subject. If the frame holds no inspectable object, subject_object is empty, and every grade field is a no-finding, not a low score. People never appear in objects.
  2. Per-object results. objects (DetectedObject) lists every inspectable object with its own box. Labels are checked against a closed class vocabulary, named in class_vocabulary (for example runink.coldchain.v1). label_in_vocabulary is optional: when it’s absent, the classifier wasn’t asked.
  3. Class-conditioned detection. This step is off by default. With VISION_DETECT_INSTANCES=true, a second, class-conditioned turn asks the model for one box per instance of the graded class, using the model’s bbox_2d JSON idiom. This is the one sanctioned exception to TOON; see Structured output.
    • The detection turn sweeps up to VISION_SWEEP_FRAMES frames (default 2, ceiling 8), with one model turn per frame. frames_detected reports how many it looked at.
    • The recorded cost on a CPU inference tier is about 50 s per detection turn, on top of about 20 s for grading.
    • When a local detector has already boxed the graded objects one instance at a time, the sweep is skipped, and detection_note says so.

Where each box came from

DetectedObject.box_origin (BoxOrigin) names the stage that produced each box:

BoxOriginSource
VISION_GRADINGThe grading turn’s own box columns. The weakest source, because grading answers at scene level.
VISION_DETECTIONThe class-conditioned detection turn.
REGION_CLUSTERINGml/segment region proposals: geometry only, with no label or grade.
LOCAL_DETECTORA detector running inside the FACE process.

A box from a local detector carries detector_label, detector_label_set and an optional detector_confidence. A detector’s class is never written into label, and never mapped onto the cold-chain vocabulary. Material, grade and damage figures come only from the vision model’s grading turn.

Notes that explain an absence

Three response fields explain gaps. An empty object list is also what a correctly empty yard looks like, so these fields are how a reader tells the cases apart.

  • localisation_note: why findings arrived without boxes. The cause might be that detection is off (the note names the switch), or detection’s own reason if it ran.
  • detection_note: what the local detector did. It is never empty. If no detector is wired into the deployment, the note says so, as a fact about the deployment and never about the camera.
  • evidence_note: why no annotated frame was retained. Retention is off unless an operator sets a window.

Coordinate handling and its limits

  • Coordinates are normalised. Every box on the wire is 0 to 1 of the frame’s width and height, so a client can draw it at any size.
  • Frames are downscaled first. Before the model sees a frame, its long edge is reduced to VISION_MAX_EDGE (default 448 px). 0 disables downscaling.
  • Detection boxes are scaled per axis. bbox_2d coordinates may be in the model’s native 0 to 1000 space or in pixels. The parser uses the frame’s real pixel size and scales each axis separately, so boxes on non-square frames aren’t skewed.
  • Above 1000 px, pixels and model space can’t be told apart. Once a frame’s long edge reaches 1000 px, a coordinate like 900 could mean either. The parser resolves it correctly only because the default downscale keeps frames well below that. Raising VISION_MAX_EDGE past 1000 brings the ambiguity back, and a test makes that regime visible.
  • Objects aren’t tracked between frames. Two boxes of the same class in two frames may or may not be the same physical object.

The live path

DetectLiveFrame runs the injected local detector on one frame and returns. It never makes a model turn or a network hop. DetectLiveFrameResponse has no field for a grade, percentage, material or classification, so there’s nowhere to put a verdict. A test walks the call graph and fails if anything reachable from this RPC can reach the vision model, the catalogue resolver or a headless browser.

detection_latency_micros reports the in-process pipeline time per frame. It isn’t a round-trip time or a frame rate.

A live call leaves no trace. It writes no lineage record, evidence frame or sighting. Frames that were only live-detected don’t appear in source activity.

Aggregating sightings

MaterialAnalyticsService.AnalyzeMaterialQuality folds graded ingests into distributions by material, class and damage grade over a time window:

  • Every count is a count of sightings, meaning one object in one frame in one ingest. It is never a count of things.
  • Ungraded is not pristine. Ungraded sightings are counted separately, and averages are optional and carry their denominators.
  • A comparison needs a baseline window. A delta appears only when a baseline window is supplied. It is never invented.
  • Check whether the corpus was read. corpus_scanned and corpus_durable say whether a corpus was actually read and whether it survives a restart.

What this does not establish

  • A grade and a damage percentage are the vision model’s estimate from downscaled pixels. They aren’t a physical inspection.
  • A box drawn where the coordinates say may not be on the right object. detection_note reports which stage ran, not whether it was correct.
  • A retained evidence frame shows the reviewer what was assessed. It isn’t a forensic, chain-of-custody artefact. Nothing signs the bytes or attests the capture time.
  • No detection accuracy is claimed. This page describes the pipeline, not how often its answers are right.