Openpi Comet: a reproduction attempt

Team Comet released code, two checkpoints and the rejection-sampling dataset behind their second place in the 2025 BEHAVIOR Challenge. I priced each claim, ran the two I could afford, and found that the free one was the most informative.

Paper: arXiv:2512.10071 Role: empiricist September 2026 my fork with fixes

Comet took second place at NeurIPS 2025 with a Q-score of 0.2514 on the held-out test set, then reported 0.345 on public validation after the challenge, up from 0.192 after pre-training. The recipe is a π0.5 backbone, about 1.5K hours of trajectories, and three rounds of rejection-sampling fine-tuning (RFT). What makes the report worth an empiricist's time is the last line of its abstract: everything was released. Most competition reports are unreproducible by construction. This one shipped its artifacts, so the question becomes which claims survive contact with them.

Price the claims before running anything

Reproduction is not one thing. The report makes five testable claims, and their costs span four orders of magnitude.

ClaimTheir costMy costReachable
Pre-training (pt50)50k steps on about 40 H200sNo. The config is literally named _gpu40
RFT post-training3 rounds of about 8,500 rollouts3,500+ GPU-hoursNo
Per-task evaluation of the released checkpoint30 to 60 GPU-hoursYes
pt12 versus pt5030 to 60 GPU-hoursYes
Table 3 inference ablations10 to 20 GPU-hoursYes, cheapest and sharpest
Characterising the released RFT datasetnoneYes, free

Two things fall out of the table before any GPU is touched. The inference-time ablations are the highest-value target, because they are runtime flags and the released checkpoints make them testable with no training. And the released RFT dataset lets me interrogate the post-training claim, the one I categorically cannot re-run, for free.

The barrier nobody states: RT Cores

BEHAVIOR-1K runs on OmniGibson, which runs on Isaac Sim, which draws the scene with hardware ray tracing. NVIDIA lists the A100 and H100 as unsupported; they have no RT Cores. The V100 is older still. This inverts the usual ranking of hardware: the report trained on H200s, but you cannot evaluate on one.

ArchitectureExampleRT CoresRuns the benchmark
VoltaV100NoNo
Ampere GA100A100NoNo
HopperH100, H200, GH200NoNo
Ampere GA102A40, A10, RTX 3090YesYes
AdaL40S, RTX 4090YesYes
BlackwellRTX 5090YesYes, on driver 580

My two best-specced machines were the two that could not do the job: a DGX with eight V100s and an A100 vGPU allocation. The A40 nodes on NCSA Delta were the only clearly RT-capable GPUs in the ACCESS-CI catalog. The national GPU buildout optimised for tensor-core throughput, and embodied-AI simulation fell outside it.

This explains a design decision in the paper. Section 4.3 rejects online RL partly because it needs GPUs with RT Cores for simulation and GPUs with Tensor Cores for training at the same time. RFT is not only a sample-efficiency choice. It lets training and simulation be decoupled in time, so you never need both GPU types at once. It is arguably the most transferable idea in the report, and it is buried in one sentence.

What the released RFT dataset says

The report attributes the move from 0.224 to 0.345 to a refined task-balancing strategy when selecting 1,469 trajectories from about 25,500 rollouts. If balancing drove the gain, the kept set should be more uniform across tasks than raw rejection sampling would produce. I pulled meta/episodes.jsonl and meta/info.json from delinqu/comet-1.5k and computed per-task counts, per-task mean horizon, and the rank correlation between them. Zero GPU-hours.

  1. Provenance checks out. Exactly 1,469 episodes, matching Section 4.3.
  2. Task balancing is a cap, not a rebalance. Five tasks sit at exactly 120 episodes and three more just below it. Over-represented tasks were capped; rare tasks were not lifted. The set stays skewed 120 to 1, the top five tasks hold 40.8% of the data, and the median task contributes 12 episodes.
  3. RFT yield collapses with task horizon. Spearman correlation between episodes kept and mean episode length is −0.45. The eight tasks with 100 or more episodes average 4.1 minutes; the 31 tasks with fewer than 100 average about 9 minutes.
  4. Eleven of fifty tasks yielded nothing. Thirty-nine tasks appear. About 25,500 rollouts produced no retained success on the other eleven.
  5. Native resolution is 720 for the head camera and 480 for the wrists, which is exactly the "high resolution" row of Table 3.

The report credits task balancing. The artifact shows that balancing did comparatively little. What the data shows instead is a horizon-dependent sampling bias: rejection sampling harvests almost entirely from the short half of the benchmark and is silent exactly where the benchmark is hardest. That sharpens the report's own admission that sampling efficiency is too low. RFT cannot close the gap to the 0.611 theoretical best by construction, because it generates no signal on the eleven tasks that need it most.

Finding 5 also recasts the resolution result in Table 3. The gain from 0.30 to 0.60 is not from adding information; it is from ceasing to discard it. "High-fidelity perception is essential" is more honestly stated as "do not downsample your data."

The radio probe

Figure 4 of the report gives a success rate of 1.00 on turning_on_radio, the task the authors use for every cell of Tables 3 and 4. It is also the shortest task in the benchmark, 934 frames or 31 seconds, and a rate of 1.00 is easy to disprove. I evaluated the released pi05-b1kpt50-cs32 checkpoint with receding-horizon control and native resolution on all 20 public test instances, on two machines with different GPUs, drivers and installs.

MachineGPUReportnOurs95% CI
marklxxxvRTX 30901.00200.00[0.00, 0.16]
b3iqRTX 50901.00200.00[0.00, 0.16]

The policy failed all 40 instances, each running to the 4,300-step limit with a partial-credit Q-score of 0.000. A broken harness also gives zero, so the first question is whether the harness works. On training instances, which the policy saw during training, it succeeded on 1 of 4 with a mean Q of 0.25. The harness can record a success. The policy succeeds about a quarter of the time on data it trained on.

My first run was wrong, and the error was mine. I had changed the wrapper to 224×224 to match Table 3, and the released checkpoint scores zero at that resolution even on training data. The 224×224 numbers in Table 3 come from single-task models that are not public. The positive control caught this. Without it I would have reported a false result about the authors' work.

Three readings of 0 of 20 remain: the claim does not reproduce; the released checkpoint is not the model in Figure 4; or my setup is wrong in a way I did not find. The public record cannot settle the second. The checkpoint's build date and model card point to the 0.345 model, while its directory name and the README's model-zoo text point to the pre-training model, and the README calls the same files both things. The third is bounded by the control: a dead harness cannot succeed on training instances.

What else I found

  • Three faults block the documented setup path, and each stops evaluation: a config pointing at a module path that the install step renames, a dependency step that upgrades numpy past 2.0 and segfaults og.launch(), and a setup script that downloads the 2026 task instances while the evaluator reads the 2025 ones. Fixes for all three are in my fork.
  • The RTX 5090 segfaults on driver 595.84 and runs on 580.173. An earlier version of my notes blamed the GPU architecture. That was wrong; one driver is not a sample.
  • Tables 3 and 4 move in steps of 0.05 to 0.10 at n of 10 to 20. At p = 0.3 the 95% interval is roughly [0.07, 0.65], so several reported effects are not distinguishable from sampling noise. The paper shares this limitation and does not state it.
  • The dataset's info.json frame count is 18% below the sum of episode lengths, probably not rebuilt after the last filter step. Nothing above depends on it.

What I would run next

  • Finish the Table 3 control-mode arm, with the caveat that a policy scoring 0.00 everywhere agrees with every prediction of zero, which means little.
  • Per-task evaluation against Figure 4, which would also settle the checkpoint ambiguity.
  • Test the implication of the dataset finding directly: if RFT yield is horizon-limited, truncating long tasks into sub-episodes before rejection sampling should raise yield on the eleven dead tasks. That is the step the report's own analysis points to but does not take.

Method. The cost table, the dataset analysis and the evaluation runs are mine, carried out in September 2026. Fixes to the evaluation stack are in the fork linked above.