What closes SigLIP’s predicate gap on BEHAVIOR?

Every published BEHAVIOR baseline reads the world through a frozen SigLIP-family encoder, and the benchmark scores spatial relations, not object identity. A research note with a registered hypothesis, exact simulator labels, and first probe results that sit near chance.

Role: empiricist September 2026 Related: Comet reproduction, first-place review

The problem

The BEHAVIOR head camera renders at 720×720 and the wrist cameras at 480×480, but π0.5 consumes 224×224. That is a 10.3× reduction in area, and at patch 14 one token covers a 45×45 region of the native render. GR00T N1 reads SigLIP-2 through Eagle-2; N1.5 and N1.6 use SigLIP2-So400m/16 at 512 through Eagle 2.5; N1.7 uses SigLIP2-Large/16 through Cosmos-Reason2-2B.

The score does not count object identity. It counts goal predicates, such as whether the plate is inside the cabinet or the radio is toggled_on. Across the 100 tasks of the 2026 challenge there are 359 goal predicates, and 72% are spatial relations: inside appears 123 times across 54 tasks, ontop 81 times across 40, nextto 36 times, open 28 times. Object recognition alone scores zero.

An image-text encoder learns from web captions, which are dominated by nouns and thin on relations, so a gap is plausible. It had not been measured on BEHAVIOR. Two published results measure it elsewhere. Newman et al. (arXiv:2409.10488) report CLIP at 0.918 on objects and 0.614 on states, falling to 0.408 against state-minimal-pair negatives, and their Table 4 shows ground-truth crops do not help. VLM4VLA (arXiv:2601.03309) runs the resolution ablation for policies: resolution buys nothing, and the frozen-encoder penalty grows with resolution.

Neither 2025 BEHAVIOR report supports a resolution effect either. The first-place team lost NGX on their evaluation machines, saw "easily noticeable image-quality degradation" and "very small impact on success rate". Openpi Comet reports that resolution "more than doubled" success, but that doubling is 3 of 10 against 6 of 10, Fisher exact p = 0.37, and the released code pins IMAGE_RESOLUTION = (224, 224).

The hypothesis

H. Frozen SigLIP encodes object categories better than the goal predicates BEHAVIOR scores. Raising the token budget from 256 to 1024 closes less of that gap than swapping to an encoder trained with dense objectives.

BranchHolds if
H1, objectivethe gap survives the budget increase
H2, budgetthe gap closes as tokens rise
H3, readoutthe gap is in the pooled vector but not in the patch tokens

The tests

Frozen forward passes only, no fine-tuning and no rollouts. Labels come from the simulator: the raw 2026 challenge data ships each episode's serialized scene state, OmniGibson rebuilds the task from it and runs the BDDL checking functions, which gives exact per-frame predicate truth. I take the frame before each predicate flip and the frame after, from the same episode. Readouts are mean over patch tokens, max over patch tokens, and the pooled vector; the patch tokens are what the baselines consume, because π0.5 projects last_hidden_state into Gemma and never uses the pooling head. Arms form a 2×2 over objective and budget: SigLIP 2 NaFlex So400m at 256 and 1024 tokens with the weights pinned, which isolates budget from weights; SigLIP 1 So400m/14 at 256; and DINOv2 ViT-L/14 at 256 as a non-contrastive objective.

I test ontop before inside. inside carries the most credit, but an object inside a cabinet is occluded once the door shuts, so a low score there is ambiguous between "not encoded" and "not visible". ontop is visible by construction. Controls still owed: a proprioception-only baseline, a time-shuffled null, arm-present-but-unflipped pairs, and leave-one-scene-out folds.

Preliminary results

Two runs, one per predicate class, with exact labels from simulator replay, matched pairs, chance at 0.50, and episode-grouped cross-validation. No confidence intervals yet.

RunTaskPredicateClassPairsEpisodes
Aturning_on_radiotoggled_onunary state120120
Bputting_up_Christmas_decorationsontopspatial relation14068

Budget, SigLIP 2 NaFlex

ReadoutA: toggled_on, 256 → 1024B: ontop, 256 → 1024
patch mean0.533 → 0.592 (+0.058, p = 0.19)0.518 → 0.568 (+0.050, p = 0.27)
patch max0.525 → 0.571 (+0.046, p = 0.35)0.521 → 0.429 (−0.093, p = 0.033)
pooled0.508 → 0.617 (+0.108, p = 0.019)0.525 → 0.496 (−0.029, p = 0.57)

At 256 tokens, the budget every baseline runs, all six readouts sit between 0.508 and 0.533. The published figure for object categories on comparable encoders is about 0.92. The relation is the flatter of the two classes; all three of its readouts land within 2.5 points of chance. Added tokens do not close the relation gap: on the unary predicate the increase helped, though only the pooled readout reached significance, and on the relation it did not help, with one readout falling significantly. That non-monotone outcome was not anticipated by the design. It agrees with the published work: VLM4VLA found resolution buys nothing for policies, and Newman et al. found crops do not help.

Encoder swap at a matched 256 tokens, run A only

EncoderObjectivepatch meanpatch maxpooled
SigLIP 1 So400m/14image-text contrastive0.5250.5080.554
SigLIP 2 NaFlexthe same, plus dense objectives0.5330.5250.508
DINOv2 ViT-L/14self-supervised, no text0.5620.4750.579

SigLIP 2 adds captioning and self-distillation losses specifically to strengthen local features, yet at a matched budget it gains 0.008 on patch mean over SigLIP 1, against 0.058 from the token increase on the same readout. DINOv2, with no text supervision, scores highest on two readouts of three. So H is contradicted on run A: the budget closed more of the gap than the dense swap. On run B the budget closes nothing, so H cannot be settled there until the encoder arm runs on relations. One caveat covers every number here: they all lie between 0.429 and 0.617, so these are differences between near-chance quantities with no intervals attached.

What needs to run

  1. The encoder arm on relations: SigLIP 1 against SigLIP 2 against DINOv2 at 256 tokens on run B. The frames are already extracted.
  2. A category control. Without the difference between state and category accuracy the numbers have no reference point.
  3. Bootstrap confidence intervals over episodes.
  4. Power: about 310 episodes for the category-by-budget interaction, against 120 and 68 now.
  5. inside, deliberately, reporting occlusion as part of the result.

Method. Every factual claim in the note is covered by 97 scripted checks against the source PDFs, the Comet repository, the challenge page, the released dataset schema and the BDDL goal definitions; each script exits non-zero on failure. The probe numbers were produced on one RTX 5090 on 14 September 2026. An adversarial review pass, four AI-generated reviews with different emphases, changed the design in two places: the probe now reads patch tokens rather than the pooled vector, and the hypothesis is stated as a comparison between two unmeasured levers rather than a ceiling.