Spring Annuals Trial HarnessHabitat NZ · technical demo by EmerTech

Bake-off

One run is a grid: every selected frame tiled once, that identical tile set sent to every selected model arm, and the counts laid side by side against human truth. Which arms are live and which are drawing a simulated profile is stated on the row, not in a footnote.

16 arms in the registryno provider keys set — all arms simulated

Preselected synth-ladder-2-scattered from the imagery page. Add more frames below to widen the run.

Frames1 of 21 selected

Every selected frame is tiled once and the identical tile set goes to every model — that is the fairness contract. Frames with exact ground truth can be scored; the rest can only be counted.

Model arms6 selected

Grouped by lane, ranked within it. The badge is the honest one: LIVE means this key is set and a client is wired; SIMULATED PROFILE means a seeded error model derived from published benchmarks, never inference on these pixels.

Exemplar / visual prompt

Draw one box, find the rest. Holds recall best at 2–8 px because it matches an appearance instead of regressing a box.

Supervised detectors

Win decisively above ~8 px once fine-tuned — and carry the cross-site transfer risk that fine-tuning brings.

Zero-shot text prompt

No text embedding describes a 4 px yellow fleck. Here to document the honest floor.

Vision-language models

ViT patches are 14–16 px, so watch these fall off a cliff below 8 px and saturate on counts in dense mats.

Tiling, triage and prompt

Chosen once and applied to every arm. A 20 MP frame at 512 px is 200–300 tiles, and the same set must reach every model or the comparison is not a comparison.

Smaller tiles put more pixels on each plant and cost more calls.

Overlap stops a plant being cut in half by a seam; duplicates are merged on intersection-over-smaller at 0.3.

Grounding-DINO-family syntax: phrases separated by . — the exemplar and supervised lanes ignore it.

6 arms are simulated
RF-DETR Small/Medium (+SAHI), SAM 3.1 (exemplar + concept), CountGD, Gemini Flash (tiled), OWLv2 (image-conditioned one-shot), Grounding DINO (open Swin-B) will draw detections from a deterministic, seeded error profile whose recall, false-positive and jitter parameters are read off published benchmarks and applied to the known ground truth. That is a prior, not a measurement: it never becomes this model’s accuracy on Central Otago imagery, and every result carries the badge saying so.

Exemplar lanes use the human boxes on each frame as their visual prompt. No annotations exist on the selected frames yet, so those arms run prompt-only. Draw a few boxes in Annotate first to see the exemplar workflow properly.

1 frame × 6 models = 6 cells.

Previous runs

Every run is kept whole: its config, its per-cell counts, and the reproducibility manifest that says what actually ran. Runs live in this demo's memory and disappear when it restarts.

RunWhenGridStatusDetectionsCost
seed-reference-bakeoff
SIMULATED PROFILE
1 Sept, 12:00 pm10 × 8
80 of 80 cells
done25,067Open →