Full results · every table, sliceable
T1

Filled-region results

Table 1 in full. Every cell reads fixed / parseable-only under IoU² — the solid bar is the fixed score, the faint one behind it is the parseable-only score, and the gap between them is what parse failures cost. IoUp re-scores against the unreduced source part; IoUn divides IoUp by the construction fidelity CN and is a descriptive ratio, not a bounded probability.

Fixed vs parseable-only, by budget
Metric
Fixed denominator Parseable-only
Budget

Bold = best in the budget block, underline = second. CN = 88.6 / 96.2 / 98.3 / 99.0 / 99.4 for N = 8 / 16 / 24 / 32 / 64.

T1b

Where each configuration bends

The same numbers as curves. On the left, one metric across the budget curve; on the right, the threshold profile at one budget. Both switch.

Budget curve
Metric
Pooled threshold profile
Budget
F3 · F4 · T2 · T3

What moves the number

A frozen 13-condition grid on the Qwen endpoint — reasoning, visual evidence, resolution, colour, cue truth, decoding, repeatability — with matched 500-UID cohorts for the image interventions and Gemini replicating them.

F3 · Evidence ≫ resolution ≫ colour

Fixed IoU, mean over N = 8 / 16 / 24. Reasoning configuration is the largest input-preserving effect; removing the image is comparably damaging, while halving resolution or removing colour cost far less.

F4 · Cues act asymmetrically

Wrong spatial hints hurt more than wrong colours and combine sub-additively; truthful hints give no reliable gain. This diagnoses susceptibility to conflicting language, not an intrinsic spatial/colour ranking.

Selection vs tracing

Sel is target preference — the prediction overlaps the intended referent more than any same-category distractor. Con is contour IoU conditional on that event. Preference stays high while tracing is the persistent differentiator, so one IoU should not be read as either ability alone.

Budget

The complete grid behind §06 — every condition at every budget, plus the parseable-only rows the summary charts leave out. 1,500 UIDs per budget, except matched 500-UID cohorts for the no-image, half-resolution, and greyscale conditions. Fixed IoU ×100.

Table 2 · prompt cues and image evidence
Table 3 · thinking, decoding, repeatability
Wrong spatial hints cost 24.1 points against the baseline prompt; wrong colour hints cost 9.2. Truthful hints of either kind give no reliable gain — the asymmetry diagnoses susceptibility to conflicting language, not an intrinsic spatial-over-colour ranking.
T4 · T5

Diagnostic slices

View

Best-effort view. Sel records a preference event only when target overlap exceeds every same-category distractor overlap — no absolute threshold, no confidence ranking, no cross-category matching. Con is contour IoU conditional on that event.

Target preference holds up to N = 32 and then breaks: mean Sel falls from 80.8 to 53.7 at N = 64 while mean Con moves only 2.0 points. Qwen with thinking off collapses to 3.8 — it stops producing parseable answers at all.
§

Reading these numbers

FIXED DENOMINATOR

Every planned question counts

An unparseable response scores zero and stays in the denominator. Conditioning on successful output can make a low-coverage configuration look geometrically strong — which is exactly what happens to Gemini at N = 64 (37.7 fixed, 63.4 parseable).

RECOVERY HISTORY

IoU⁰ / IoU¹ / IoU²

IoU⁰ is the single-call view, IoU¹ adds one parse-failure-triggered stateless retry, IoU² up to two. All headline scores use the uniform IoU² view. Best-effort is reserved for labelled diagnostics.

SCOPE

A snapshot, not a ranking of architectures

A purposive hosted-system snapshot collected 10–17 July 2026 — not a probability sample. Every result is attributed to the complete labelled configuration, endpoint and reasoning setting included.