Filled-region results
Table 1 in full. Every cell reads fixed / parseable-only under IoU² — the solid bar is the fixed score, the faint one behind it is the parseable-only score, and the gap between them is what parse failures cost. IoUp re-scores against the unreduced source part; IoUn divides IoUp by the construction fidelity CN and is a descriptive ratio, not a bounded probability.
▸ Bold = best in the budget block, underline = second. CN = 88.6 / 96.2 / 98.3 / 99.0 / 99.4 for N = 8 / 16 / 24 / 32 / 64.
Where each configuration bends
The same numbers as curves. On the left, one metric across the budget curve; on the right, the threshold profile at one budget. Both switch.
What moves the number
A frozen 13-condition grid on the Qwen endpoint — reasoning, visual evidence, resolution, colour, cue truth, decoding, repeatability — with matched 500-UID cohorts for the image interventions and Gemini replicating them.
Fixed IoU, mean over N = 8 / 16 / 24. Reasoning configuration is the largest input-preserving effect; removing the image is comparably damaging, while halving resolution or removing colour cost far less.
Wrong spatial hints hurt more than wrong colours and combine sub-additively; truthful hints give no reliable gain. This diagnoses susceptibility to conflicting language, not an intrinsic spatial/colour ranking.
Sel is target preference — the prediction overlaps the intended referent more than any same-category distractor. Con is contour IoU conditional on that event. Preference stays high while tracing is the persistent differentiator, so one IoU should not be read as either ability alone.
The complete grid behind §06 — every condition at every budget, plus the parseable-only rows the summary charts leave out. 1,500 UIDs per budget, except matched 500-UID cohorts for the no-image, half-resolution, and greyscale conditions. Fixed IoU ×100.
Diagnostic slices
Best-effort view. Sel records a preference event only when target overlap exceeds every same-category distractor overlap — no absolute threshold, no confidence ranking, no cross-category matching. Con is contour IoU conditional on that event.
Pooled fixed-IoU by composite difficulty, its components, and semantic group. Difficulty, D, Q, and RS use IoU²; semantic groups use best-effort.
Consistent easy → hard decline
Every configuration loses 13–16 points from easy to hard, and matched parseable-only losses implicate conditional geometry rather than format coverage.
Contour complexity dominates
Q = ℓ²/(4πA). The Q1→Q3 drop (−10.2 mean) is the steepest single-axis effect in the table — more than distractor count.
Small targets are the hard case
RS is target-box area over image area. Unlike every other axis, scores rise from RS1 to RS3: small referents cost 8.6 mean points.
Reading these numbers
Every planned question counts
An unparseable response scores zero and stays in the denominator. Conditioning on successful output can make a low-coverage configuration look geometrically strong — which is exactly what happens to Gemini at N = 64 (37.7 fixed, 63.4 parseable).
IoU⁰ / IoU¹ / IoU²
IoU⁰ is the single-call view, IoU¹ adds one parse-failure-triggered stateless retry, IoU² up to two. All headline scores use the uniform IoU² view. Best-effort is reserved for labelled diagnostics.
A snapshot, not a ranking of architectures
A purposive hosted-system snapshot collected 10–17 July 2026 — not a probability sample. Every result is attributed to the complete labelled configuration, endpoint and reasoning setting included.